Re: Typing: was Re: DBpedia as possible project
"Andrew S. Townley" <[email protected]>
| Newsgroups | gmane.text.xml.xtm.general |
|---|---|
| Message-ID | <[email protected]> |
On 3 Nov 2010, at 7:52 PM, Alexander Johannesen wrote: > On Thu, Nov 4, 2010 at 2:09 AM, Andrew S. Townley <[email protected]> wrote: >> Sometimes the classifications are temporal, sometimes they're physical, >> sometimes there based on shared perspectives, communities or other >> forms of context. Saying we must choose one or the other seriously limits >> the things we can do with any information set. I just think that being exact >> requires a bit more than creating a distinction between things like "type" >> and "role type" and that there's a good bit of work to be done (and mistakes >> to be made) trying to nail down exactly what this sort of thing looks like. > > So, you're saying ontology work is hard. :) Nah, as you say below, you just have these things, and they have characteristics. Only one thing, no typing issues. Where's the problem? ;) > > I think parts of this discussion is the difference between *a* model > and *the* model; there simply are too many ways to model anything, and > so in order to actually be productive (vs. being correct) we make lots > of compromises, hoping that we make just enough to be useful without > overdoing it. > > When I look for ontologies, I always go by my gut instinct at several > levels; would this ontology be easy for me to use, easy for my users > to understand, does it "feel" right (ie. match some arbitrary > philosophy in my head about how to model the world), is it much used, > is it made by seemingly smart people, and so on. Every single > evaluation I make is a compromise somewhere, either directly for me or > me doing it for whoever use my systems. It's a lot more organic than > what we like to admit. Totally agree, and this is the same kind of gut feel check I use myself. > But we can take this even further and look at the underlying basics > for all modelling (at least in the TM / RDF world) ; entities and > relationships. There's this thing, and it has properties, and some of > those properties can point to other entities, the recursive > key-value-tree trifecta, the most delicious cake we've had in the last > 30 or so years. But is this the good enough to model *our* world? Maybe not, but I tend to think so. Either way, it sure is tasty! :) [snip] > So the question becomes; can we still rely on our TM way of subject > identification? I'm not so sure. Things change. And here's the catch; > the more you describe that thing, the more you try to pin it down its > definition, the less likely it is for that thing to fit whatever thing > you need in what you're modelling. And the less likely it is that that > model truly represents reality, so there's a whole scale of inherit > dis-ambiguity that you need to have in mind when you knowingly have to > make a million compromises while modelling. Doesn't this make the assumption that you're doing this without any notion of context at all? Of course things change, and they certainly do the more you learn about a domain and all of the intricacies and nuances that describe the essence of that domain. Still, I should be able to "zoom" in and out as necessary to answer particular questions and not be either a) lost in the weeds or b) without sufficient detail to make informed decisions. The only way I can realistically do this is to pull these different perspectives together in meaningful and controlled ways. > To what degree do we need things to be correct vs. useful? And, in the > end, is it useful that things aren't correct? Even useful and incorrect can be useful, depending on what you're trying to learn or say about a particular domain. The thing is that you need to be able to systematically relate your "useful" and my "useful" to each other in meaningful ways without losing either the utility or the assumptions that facilitate that utility. Again, once you start trying to correlate statements about things made by millions of people each with thousands of overlapping but inconsistent assumptions, this stuff matters. In a controlled environment or walled garden, you have a lot more leeway with "useful", but I don't think that's good enough in today's world with over 1 billion addressable pages added to the Web every day. I think it's important to (continue to) talk about these issues now while there's still a chance of influencing how people try to deal with a world with that much data. The less retrofitting and rectifying that needs to be done, the easier it will make things for everyone. Most of the people churning out all that content have no idea these problems exist. After all, they have Google and the magic search box. All they need is just a little bit more link juice and social proof... ;) Great points, Alexander. :) Cheers, ast -- Andrew S. Townley <[email protected]> http://atownley.org