Merging: was Re: Semantic Bleachers: was Re: Afghanistan War Diary as topic map in Maiana

Patrick Durusau <patrick-Q/[email protected]>
Newsgroups gmane.text.xml.xtm.general
Message-ID <[email protected]>
Andrew,

Two of three I think.

On Tue, 2010-11-02 at 19:59 +0000, Andrew S. Townley wrote:
> On 1 Nov 2010, at 6:55 PM, Patrick Durusau wrote:
> 
> > Andrew,
> > 

<snip>

> > I think the creative fashioning of merging rules, with boundaries for
> > their application is largely unexplored territory. 
> > 
> > We would benefit from a very robust set of comparison operators. 
> 
> That's what I was thinking too, but I wasn't sure.  Again, I don't see how you can deal with Internet-scale datasets if you don't have these type of "overlay" facilities that can be linked to specific contexts.  Simple ones are "I'm user X" and "She's user Y", but others could be a lot more sophisticated.
> 
> However, from personal experience, this stuff gets pretty complicated (and CPU intensive) pretty quickly.  Still working on ways to deal with that issue to my satisfaction...
> 

True but as I have been posting to my blog, there are techniques for
clustering data and then performing additional operations on data that
looks interesting.

In other words, merging isn't a one stage process or even one that
requires all instances be compared. 

> > 
> >> I think you alluded to this the other day with some of your work, but I think that the idea behind what Manina is doing, what I understood was possible with what you have and the way that my implementation handles merging on a conditional rather than absolute basis illustrates that there's still work to be done in this area of the topic maps specifications.  Maybe some implementations take too literal a view on the whole XTM/TMDM merging requirements, but it makes a big deal if you define "equality" on the basis of property values, merge and then one or more of those property values change.
> >> 
> > 
> > Well, those are separate questions:
> > 
> > a. How to merge on something in addition to the standard TMDM basis?
> 
> Or, in our case, "in spite of" or "happily ignoring" the TMDM basis... ;)
> 

Well, maybe yes, maybe no.

The TMDM recognizes that there is likely to be merging beyond what it
prescribes. 

And it would be possible to represent a complex set of merging
conditions with the addition of a URL as subjectIdentifier so that
fairly complex merging could be accomplished with fairly mundane topic
map software. 

That is: Complex merging condition A, B and C, then add URI X to a
topic. 

One of the topics (sorry!) that hasn't been discussed (mostly because it
could not be standardized) is the pre-processing stage or even authoring
stage for topic maps. 

> > 
> > b. What happens if after merging a value changes? (unexplored so far as
> > I know, but good question)
> 
> The way I see it, "merging" by definition is just another context of viewing a particular dataset.  What's available and the legend will define the outcome.  This outcome is totally linked to the context, although the validity of the context can be "until further notice" or "from timestamp T where properties [...] also apply."
> 
> Complexity explosion?  Absolutely, but, again, I think it's essential to work through these scenarios.  "Truth" is relative, and the results of any view of any data set on any given day represent one particular view of "truth" in that domain.  Expand the scope and scale, and the differences between "truth" start to be interesting as artifacts of themselves, but without the traceability and the ability to do "what if" scenario modeling, this can be very difficult to determine.
> 
> Would really be interested in understanding what large TMDM users with real data do in these sorts of scenarios, or if they're just ignored.  Has anyone done any kind of analysis of the way the specs are applied and the corresponding trade-offs that result from going down these paths?
> 
> I'm not a fan of "best practices" because that tends to end up being a fluffy middle-ground of ambiguity that keeps people able to report upwards that "we're following best practices" when results are questioned.  What I think we need are more detailed analysis of the issues people are facing trying to employ topic maps (of any flavor) to solve interesting (and/or business) problems beyond single examples published at topic maps conferences.  Has this been done?  If so (and I hope), I'd love to see it, because I think it would be fascinating reading.
> 

I will have to let others speak for detailed analysis of large user
bases. I don't recall seeing such information in the literature. 

Complexity explosion? Perhaps but then one doesn't have to done all the
merging that is possible. 

When you use a relational database you don't do joins on every table
simply because you want to do a join on two tables.

Much the same should be true for topic maps. Yes, there may be way more
merging possible than we can perform but if it isn't of any interest,
say I am looking for cheap tickets to Bangkok, then merging of
alt.politics posts probably isn't relevant for me. 

Hope you are having a great evening!

Patrick
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.