Re: Merging: was Re: Semantic Bleachers: was Re: Afghanistan War Diary as topic map in Maiana

"Andrew S. Townley" <[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
Hi Patrick,

On 2 Nov 2010, at 11:33 PM, Patrick Durusau wrote:
> 
> I don't think of topic maps or subject identity tests as being complex
> but rather that subject identity and handling it adequately as being
> complicated. 


I think we very much agree, and I think that the devil is in the "handling it adequately" part of what you said.

In the implementation of "handling it adequately" you can easily end up needing to process a large number of simple subject identity tests and navigation operations of your topic map.  I don't care what your implementation is, the resulting complexity is going to have to hit you somewhere.  Maybe you can avoid it based on what and how you're modeling, and maybe you can minimize the impact by using clever algorithms and storage representations and/or efficient software tools and fast hardware, but it is still going to be there.

What I actually really love about topic maps is its simplicity, and that's another reason why I'm so enamored with the TMRM.  I find great power in that simplicity.

The richness of the connections within the information I'm trying to manage at the moment means that sometimes resolving lots (and lots and lots) of simple things like individual subject identity tests in a generic fashion quickly can result in a "death by 1,000 cuts" syndrome.  For the particular application that I'm working on at the moment, we're beginning to see some patterns and areas of stability that we can leverage at an application level to minimize this problem, but it doesn't mean that the engine itself can avoid tracking, managing and resolving this complexity.  What we can do is just ask it to do it less frequently in several cases.

I see it really as all part of the learning experience, and that's why I'm very interested in how other people are dealing with these issues both at an application level and at an engine implementation level (where such sharing doesn't get anyone into trouble).  We can't be the only people facing this kind of issue.

At the moment, our data set is extremely small compared to where we intend to go.  We have about 30K proxies and about 500K assertions (properties), of which roughly 80K of them are links.  Going forward, we will have regular and large merging operations of various kinds based on context-specific subject identity, and we don't always have 1:1 mappings between our representations of a subject and the ones we're merging.  Effectively, what we're doing is actually de-merging some of the source information so that it can be re-merged into our ontology based on some well-defined mappings.  That's why I'm so interested in this particular problem.

Thanks for the discussion, and I appreciate the time taken yesterday to continue it.

Cheers,

ast
--
Andrew S. Townley <[email protected]>
http://atownley.org
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.