Re: Approximate matches in faceted schema implementations (PRODUCT FEATURES)

"Phil Murray" <[email protected]>
Newsgroups gmane.comp.infodesign.facetedclassification
Message-ID <[email protected]>
Matt --

Matt --

First of all, thanks for introducing yourself and constributing to the
discussion on this list.

I'll let PeterV and others address XFML and Topicmap issues, but let me
voice some opinions on several of the issues you raised.

(1) Full-text search is a vital component of interface to most online
information resources -- especially large collections. You cannot and should
not expect to classify every document for all users. However, although there
have been remarkable improvements in full-text search, you still can't
depend on on full-text search alone. (See "Limitations of full-text search
and retrieval" at http://www.KMconnection.com/DOC100049.htm.)

(2) In general, even a well-defined classification scheme will not
*guarantee* that users will be able to find the information they are looking
for. But the time and effort required by users to find that information will
be reduced substantially by good classification -- especially for retrieval
requirements that occur frequently in the collection. (In some limited
applications -- for example, online product catalogs, a well-conceived and
well-executed scheme should come very close to perfection.)

(3) Faceted classification (FC), IMHO, offers some distinct advantages over
"traditional" approaches to classification of content, especially in
classifying many types of "special collections" -- for example, the
information within an organization:

(a) Placing a word within an appropriately defined facet hierarchy resolves
most issues of polysemy/homographs. If you have defined the facet "armored
vehicles," the topic "tank" in that facet clearly does not refere to a
container for gasoline.

(b) FC further reduces the necessity of knowing the precise name of the
topic you are looking for (that is, the specific term the information-seeker
expects to find), because the different facets can describe the
*characteristics* of the information object catalogued.  For example, in a
company that is selling computer-based fluid dynamics software, one useful
facet is the type of object whose external or internal flow characteristics
customers want to measure.

(c) FC supports adaptability to change and new perspectives. You can add a
new facet at any time without destroying previous work on the classification
schema.

(d) You still need a central arbiter for any classification scheme
(controlled vocabulary), but FC should help contributors or team members
catalog their own information contributions or selections more consistently.
(Ask for feedback but *don't* ask contributors to play a key role in
designing the scheme!) By comparison, the classification provided by such
directories as Yahoo! seems quite arbitrary to me.

(4) You do need to map as many synonyms (and aliases) as possible into a
classification schema, because (as you pointed out) information-seekers may
use different vocabularies.

(5) There are some potential negatives in FC implementations, including
abstractness (especially at higher levels of facet hierarchies), the
presence of empty categories (in part, an implementation/interface issue),
and expectations when returning to the resource (You don't want to have to
repeat your original process of manipulating the facet hierarchies -- an
effort that may have helped you identify the "popular" name of a topic
originally.).   (See my article "Popularity-based categorization — a
complement to faceted classification" at
http://www.KMconnection.com/DOC100124.htm for additional discussion of these
issues.)

A suggestion: Take a look at the Flamenco Search System project
(http://bailando.sims.berkeley.edu/flamenco-interface.html) and some of the
other examples cited in this list. Compare these implementations with
traditional methods that depend heavily on full-text search or traditional
classification approaches.



    Phil


-------------------------------------------
"I have made this letter longer than usual, only because I have
not had the time to make it shorter."
 -- Blaise Pascal, Lettres Provinciales, 1657

Phil Murray -- Chief Knowledge Architect
The Knowledge Management Connection | http://www.KMconnection.com
401-247-7899



----------
Thanks for playing. To unsubscribe from this group, send an email to:
[email protected]

 

Your use of Yahoo! Groups is subject to http://docs.yahoo.com/info/terms/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.