Re: I need a critique: using crm114 to train on (very) limited word set - or are there better ways?
Paolo <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <20080911135356.GD20788@localhost> |
On Wed, Sep 10, 2008 at 09:30:55PM +0200, Ger Hobbelt wrote: > Now for the fun stuff. The number of parameters (a.k.a. 'words') in > all possible messages is known (requires a bit of work though) AND > severely limited. To talk numbers: I'm pretty sure I've got a > 'vocabulary' of more than 100 words, but I am also pretty sure I'll > never surpass the 5000 words upper limit. In email analogy terms, that > would mean I know the lower and upper limits of the number of > *different* words (no l33t spellings and other tricks for this guy!) > PLUS I know that each message will: > a) only contain each word ONCE, and > b) only contain a limited SUBSET of all the possible words, and > c) all words will appear in a predetermined, FIXED ORDER in each > message. (Did I hear a <unique> there? well, maybe not.) ... > very strong feeling it won't ever get near 95% - contrary to email > claims / SPAMTREC benchmarks - as there's just way too much noise in > the inputs obscuring signal. perhaps a Viterbi decoder is more apt to the task here. ... just a quick 0.02 EUR note - will read up better later/tonight while going through my backlog. -- paolo ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/