Re: [Hobbelt b.v.] ...

"Ger Hobbelt" <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
Hi John,


On Thu, Aug 28, 2008 at 3:32 AM, AFP <[email protected]> wrote:
> I'm interested if CRM114 could be used as a tool to group/cluster similar
> content documents / emails / PDFs / etcetera into a database.  Clustering
> would be done by content similarity of an example document submitted as the
> "basis" of the cluster.


Yes, I think this is possible and I believe there are already folks
out there doing exactly that. For additional info I would advise you
subscribe to the crm114 generic mailinglist (see CC address in my
reply to you; you can visit the official crm114 website for info where
and how to subscribe, by the way).

To go into this a little deeper and after trying to read additional
detail from your inquiry I'd like to point out a few technical issues
that may come your way when you choose to do this (and these issues
are are not crm114 specific, though some are
statistical/Bayesian-filter specific):


a)- you may find that 'preprocessing' (or 'munching', or whatever you
like to call it) your documents before submitting them for both
learning or classification may be useful for your purposes: PDF files
may see improved classification when first converted to plain text so
the only thing remaining for classification technology to look at is
the actual content (that is: if document format/container is not
desirable as part of the classification -- I'm thinking along 'search
engine' lines here: google et al also 'munch' PDF and other formats to
get at the actual content therein)

'Preprocess' in such a way that the things you want the filter to
'see' and maybe recognize are readily available (document content
instead of document format? --> preprocess files to make them look
more alike: Word and PDF to unicode text, maybe?)


b)- a statistical (Bayes/Markovian/...) filter needs to be trained.
One sample almost always is not enough -- compare the usual email
training process which basically is this:

1. start with nothing
2. human feeds filter a sample.
3. filter says: 'I dunno!' and dumps it in a 'human, please tell me
what this should be classification-wise' bin
4. human checks document in bin and assigns it a class (good/bad,
red/green/... depending on your system; I use crm114 for particular
signal analysis and there I'm am interested in signal vs. true-noise
and in case of 'signal': rise/fall mostly. Think kinda like SETI but a
different goal and signal sources. Classifying documents is just
another 'signal source' in that regard.)
5. assignment in step 4 triggers 'learn cycle' in the CRM114
(statistical filter): in short: human 'educates' filter what it should
'think'.
6. rinse&repeat: take next sample to classify and goto step 2. This
time around, step 3 will probably produce a 'verdict' from the filter.
Make sure you have a way to 'tell' the filter to train incorrectly
classified samples as well (a.k.a. training false positives and false
negatives)

The above describes -- in a slightly technical manner -- an iterative
classification/learning process. You may find that you will need a
similar iterative process for your purposes to achieve optimum
performance. In control engineering terms, it's a process with TWO
feedback [training] loops: (1) one for the 'I dunno' path, which is
the major feedback/training loop, and (2) second a 'false
pos/neg'-correcting loop.
Quite a few research papers out there concern themselves with the
filter training aspect of the feedback loop and just 'implicitly
assume' those feedback loops are in place one way or another. That's
papers where you'll run into acronyms like TOE, TONE, THTTR, etc.etc.,
by the way.


c)- Most uses seem to concentrate around yes/no binary
classifications, though CRM114 does support multi-way classification.
I myself have run tests with multiway but cannot offer anything
conclusive at this moment as that lab experiment was never really
concluded. From the noises made on the crm114 mailing list, it sounds
to me like others are using that particular feature successfully in
production(?) environments, though.

Rather recently, there was a bit about that on the ML, so when you
like to have a closer look at the technical stuff involved, you might
want to check out a bit of recent mailing list history as well.


d)- best advise I can give today: plan several trial runs to check
your ideas and the actual behavior of your stats classifier (e.g.
CRM114) and most importantly: take the time to evaluate the results
and reconsider your design where necessary. Statistical filters are
not 'drop-in technology'.

Email filtering systems will run [almost] out of the box. If the
number of conversations here and elsewhere is any indication, document
classification comes in second when you look at 'usage numbers' for
CRM114 and others and after that rank 3 and beyond are for what I
would call 'non-standard uses' (i.e. the remaining 20%) where my own
definitely does not rank anywhere above radar in the public available
documentation /anywhere/ on this planet. (Signal analysis using
statistical filters for signal extraction from noisy signals.)
What am I saying here? Depending on your goals, you may find that you
can base your work on that of others (email classification, document
classification?) and the resulting, er, TTM (Time To Market) is quite
short. I cannot read from your email if you are looking for something
'very unusual' like I did/do, but if you do, prepare yourself to find
your way.


This is NOT meant to discourage you in any way -- on the contrary! --
but should be read as a few 'items to keep in mind in the miles
ahead'. Determine for yourself how far you want to take this baby,
both in classification quality (which results in choices to make
regarding training your filter, for example) and quantity (any
software performance criteria lurking around the corner?)


I hope this helps you along the way; feel free to ask for more when
the need arises,

Cheers,

Ger





-- 
Met vriendelijke groeten / Best regards,

Ger Hobbelt

--------------------------------------------------
web: http://www.hobbelt.com/
 http://www.hebbut.net/
mail: [email protected]
mobile: +31-6-11 120 978
--------------------------------------------------

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.