Re: Approximate matches in faceted schema implementations (PRODUCT FEATURES) WAS Introducing myself

"chris_mcmahon2003 <[email protected]>" <[email protected]>
Newsgroups gmane.comp.infodesign.facetedclassification
Message-ID <[email protected]>
Phil, Matt, Pete et al

Sorry for the late reply - examination marking has intervened
I am afraid!  In reponse to your points Phil:

(1) Yes, in our demos at http://nzm.dig.bris.ac.uk "exact" match 
hits only on that concept, but not its children.  In a non-Web
version of the software we handled the issue of the user not
understanding the classification by combining a browsing and text 
earch.  Users could search for a word or phrase either in the 
document text, in the concept names OR (and this was perhaps most 
important) in the words or phrases used in the rules for classifying
documents into concepts (thus giving a sort of thesaurus 
capability).  In the first case the documents are shown, in
the latter two cases those concepts relating to the user's entry are
shown (in a separate window, as in the "show all concepts for the
document" capability in the current demonstrators).  The user could 
then select a concept from the generally small) subset 
interactively.  We hope to have these capabilities in a web-based 
version soon, being developed by Adiuri Systems (see www.adiuri.com).

(2) I partly agree with Matt's point about the need for a controlled
vocabulary.  We have done work looking at classification of documents
used by engineers, and came to the view that there would be a number
of classification schemes used by different specialist groups, and
each would in effect have a "controlled" vocabulary specific to the
group (and the same document set could be classified into multiple
classifications).  We dealt with the thesaurus issues in the automatic
classification algorithms, and did not mind a little imprecision
since the faceted approach meant such rapid homing in on a sub-set
of documents that the final selection could be done easily by the
end user.  For a set of users with 40-50000 poorly organised documents
it was more important to have a scheme to improve the access to the
documents, and correct/improve it as they learnt more about it, 
than to have a completely agreed scheme - but this only works
because the classification is exposed to the user, and would not work
with, for example a form-based interface to a text management system,
which would require a rigidly controlled vocabulary.

(3) Regarding inexact matches, the closest we have got to looking at
this is in another application, for a "best practice" system, in
which we prioritised close hits by the number of steps required to
traverse a hierarchy (which is as Phile suggests in his first
suggestion?).  We were wanting to match advice to a particular
context in an adaptive hypermedia system.  As default we delivered
an exact match to the user's context, (described by multiple facets)
but if an exact match was not available we delivered close matches
ranked on the number of steps up and across the hierarchy from
the given context.

(4) I'm not sure I quite understand the proposal for a percentage-
based weighting scheme.

(5) For the Government Category List example, we used a published
list (updated version now at 
http://195.224.227.150/gcl/content/default.asp), and built our own
constraints for classifying the documents into the list. The List is
part of the emerging UK government e-GMS (Government Metadata
Standard). 

Thanks for Pete and Phil's further comments.

Chris

--- In [email protected], "Phil Murray" 
<pmurray@K...> wrote:
> Chris et al. --
> 
> The NMZ demonstrations are cool.
> 
> The "Exact" match vs. "Default" match distinction is very useful, 
and it
> raises a related question about online implementations of faceted 
schemas.
> (I understand "exact" matches as hits on only that concept, but not 
its
> children. Please correct me if I'm wrong.)
> 
> But how should the inverse case be handled? For example, an 
information
> seeker wants information about "dogs" but the indexer has 
classified a
> document under "canines." This is the classic retrieval problem of 
indexers
> using broader or narrower indexing terms than the information seeker
> expects. It's unavoidable.
> 
> In such cases ...
> 
> -- Is making the schema visible sufficient? (In which case, you see 
the
> broader/parent category.)
> 
> -- Should you prioritize hits by the number of additional steps 
necessary to
> crawl up a facet hierarchy? Display the additional jumps required?
> (MultiCentrix offers a feature along this line.)
> 
> -- Should you use a percentage-based weighting scheme ... with 
lesser weight
> given to those hits that require additional crawling of the facet
> hierarchies? (This may offer the advantage of simplicity for the 
information
> seeker.)
> 
> 
> It would also be interesting to understand your thought processes 
as you
> developed the schema for the acts of parliament example. Did you 
draw on
> existing classification schemas, or did you do the facet analysis 
from the
> ground up?
> 
> Thanks,
> 
>     Phil
> 
> > We built a software demonstrator that is still
> > running at Bristol at http://nzm.dig.bris.ac.uk.  There are three
> > demonstrations at this site – advertisements for cars, legislation
> > from the UK parliament and papers from an academic conference
> > (confidentiality with our collaborators prevents use of 
engineering
> > examples).
> 
> > Chris McMahon
> 
> 
> -------------------------------------------
> "I have made this letter longer than usual, only because I have
> not had the time to make it shorter."
> -- Blaise Pascal, Lettres Provinciales, 1657
> 
> Phil Murray -- Chief Knowledge Architect
> The Knowledge Management Connection | http://www.KMconnection.com
> 401-247-7899


----------
Thanks for playing. To unsubscribe from this group, send an email to:
[email protected]

 

Your use of Yahoo! Groups is subject to http://docs.yahoo.com/info/terms/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.