Automatic facet generation

Claudio Gnoli <[email protected]> Mon, 23 Feb 2004 18:03:36 +0100
Newsgroups gmane.comp.infodesign.facetedclassification
Message-ID <[email protected]>
""" [Marcel van M : Jan 13]
1. does it make sense to turn automatically generated 
metadata into faceted metadata? 
2. Does it require humans to determine what is a facet 
and what isn't? 
"""

Thinking again about Marcel's questions, they look as 
very deep and challenging ones...

From the theoretical point of view, I suppose we should
go back to the definition of what is a facet, to see whether
its recognition could automatized in any way. More or less, 
a facet is a category being frequent and relevant within 
a given discipline. So, let's imagine to feed an automatic 
clustering software with a corpus of digital documents 
from a given discipline. AFAIK, such softwares are able 
to extract relevant words from the corpus on the basis 
of their frequency distribution, and to establish relations 
between those and other words -- EG as it can be seen 
in the Vivisimo metasearch engine. 

However, to establish categories we are searching for 
words being in a genus-species relation with other words, 
while I suspect that softwares cannot guess the kinds of 
relations -- EG they can say that "reptile" and "snake" are 
highly associated, but not that snakes _are a kind of_ 
reptiles, which is necessary to decide that the main category 
is "reptiles", not "snakes" (nor "tails", "poisons", etc.).
(Some expert in the statistical analysis of texts could 
correct me on this.)

A second problem is that the supposed facet, in order to be 
correctly used and sorted, should be connected to one of 
the fundamental categories of facet analysis, such as Object, 
Part, Property, Process, Operation, Agent, Space, Time etc.
(though this is not done in most digital applications).
One possible direction towards this could be using 
morphological analysis -- EG guessing that verbal frequent
terms are Process or Operation facets, and so on; but 
we know how irregular is natural language. NPL could play 
some role in this. 

All in all, I'm afraid that such tasks are beyond the current 
capability of machines, hence at least the final stages of 
facet recognition should still be done by humans...



----------
Thanks for playing. 
Yahoo! Groups Links

<*> To visit your group on the web, go to:
     http://groups.yahoo.com/group/facetedclassification/

<*> To unsubscribe from this group, send an email to:
     [email protected]

<*> Your use of Yahoo! Groups is subject to:
     http://docs.yahoo.com/info/terms/