Re: [Hobbelt b.v.] ...
"Ger Hobbelt" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
Hi John, On Thu, Aug 28, 2008 at 3:32 AM, AFP <[email protected]> wrote: > I'm interested if CRM114 could be used as a tool to group/cluster similar > content documents / emails / PDFs / etcetera into a database. Clustering > would be done by content similarity of an example document submitted as the > "basis" of the cluster. Yes, I think this is possible and I believe there are already folks out there doing exactly that. For additional info I would advise you subscribe to the crm114 generic mailinglist (see CC address in my reply to you; you can visit the official crm114 website for info where and how to subscribe, by the way). To go into this a little deeper and after trying to read additional detail from your inquiry I'd like to point out a few technical issues that may come your way when you choose to do this (and these issues are are not crm114 specific, though some are statistical/Bayesian-filter specific): a)- you may find that 'preprocessing' (or 'munching', or whatever you like to call it) your documents before submitting them for both learning or classification may be useful for your purposes: PDF files may see improved classification when first converted to plain text so the only thing remaining for classification technology to look at is the actual content (that is: if document format/container is not desirable as part of the classification -- I'm thinking along 'search engine' lines here: google et al also 'munch' PDF and other formats to get at the actual content therein) 'Preprocess' in such a way that the things you want the filter to 'see' and maybe recognize are readily available (document content instead of document format? --> preprocess files to make them look more alike: Word and PDF to unicode text, maybe?) b)- a statistical (Bayes/Markovian/...) filter needs to be trained. One sample almost always is not enough -- compare the usual email training process which basically is this: 1. start with nothing 2. human feeds filter a sample. 3. filter says: 'I dunno!' and dumps it in a 'human, please tell me what this should be classification-wise' bin 4. human checks document in bin and assigns it a class (good/bad, red/green/... depending on your system; I use crm114 for particular signal analysis and there I'm am interested in signal vs. true-noise and in case of 'signal': rise/fall mostly. Think kinda like SETI but a different goal and signal sources. Classifying documents is just another 'signal source' in that regard.) 5. assignment in step 4 triggers 'learn cycle' in the CRM114 (statistical filter): in short: human 'educates' filter what it should 'think'. 6. rinse&repeat: take next sample to classify and goto step 2. This time around, step 3 will probably produce a 'verdict' from the filter. Make sure you have a way to 'tell' the filter to train incorrectly classified samples as well (a.k.a. training false positives and false negatives) The above describes -- in a slightly technical manner -- an iterative classification/learning process. You may find that you will need a similar iterative process for your purposes to achieve optimum performance. In control engineering terms, it's a process with TWO feedback [training] loops: (1) one for the 'I dunno' path, which is the major feedback/training loop, and (2) second a 'false pos/neg'-correcting loop. Quite a few research papers out there concern themselves with the filter training aspect of the feedback loop and just 'implicitly assume' those feedback loops are in place one way or another. That's papers where you'll run into acronyms like TOE, TONE, THTTR, etc.etc., by the way. c)- Most uses seem to concentrate around yes/no binary classifications, though CRM114 does support multi-way classification. I myself have run tests with multiway but cannot offer anything conclusive at this moment as that lab experiment was never really concluded. From the noises made on the crm114 mailing list, it sounds to me like others are using that particular feature successfully in production(?) environments, though. Rather recently, there was a bit about that on the ML, so when you like to have a closer look at the technical stuff involved, you might want to check out a bit of recent mailing list history as well. d)- best advise I can give today: plan several trial runs to check your ideas and the actual behavior of your stats classifier (e.g. CRM114) and most importantly: take the time to evaluate the results and reconsider your design where necessary. Statistical filters are not 'drop-in technology'. Email filtering systems will run [almost] out of the box. If the number of conversations here and elsewhere is any indication, document classification comes in second when you look at 'usage numbers' for CRM114 and others and after that rank 3 and beyond are for what I would call 'non-standard uses' (i.e. the remaining 20%) where my own definitely does not rank anywhere above radar in the public available documentation /anywhere/ on this planet. (Signal analysis using statistical filters for signal extraction from noisy signals.) What am I saying here? Depending on your goals, you may find that you can base your work on that of others (email classification, document classification?) and the resulting, er, TTM (Time To Market) is quite short. I cannot read from your email if you are looking for something 'very unusual' like I did/do, but if you do, prepare yourself to find your way. This is NOT meant to discourage you in any way -- on the contrary! -- but should be read as a few 'items to keep in mind in the miles ahead'. Determine for yourself how far you want to take this baby, both in classification quality (which results in choices to make regarding training your filter, for example) and quantity (any software performance criteria lurking around the corner?) I hope this helps you along the way; feel free to ask for more when the need arises, Cheers, Ger -- Met vriendelijke groeten / Best regards, Ger Hobbelt -------------------------------------------------- web: http://www.hobbelt.com/ http://www.hebbut.net/ mail: [email protected] mobile: +31-6-11 120 978 -------------------------------------------------- ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/