Re: I need a critique: using crm114 to train on (very) limited word set - or are there better ways?

Paolo <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <20080911135356.GD20788@localhost>
On Wed, Sep 10, 2008 at 09:30:55PM +0200, Ger Hobbelt wrote:

> Now for the fun stuff. The number of parameters (a.k.a. 'words') in
> all possible messages is known (requires a bit of work though) AND
> severely limited. To talk numbers: I'm pretty sure I've got a
> 'vocabulary' of more than 100 words, but I am also pretty sure I'll
> never surpass the 5000 words upper limit. In email analogy terms, that
> would mean I know the lower and upper limits of the number of
> *different* words (no l33t spellings and other tricks for this guy!)
> PLUS I know that each message will:
> a) only contain each word ONCE, and
> b) only contain a limited SUBSET of all the possible words, and
> c) all words will appear in a predetermined, FIXED ORDER in each
> message. (Did I hear a <unique> there? well, maybe not.)
...
> very strong feeling it won't ever get near 95% - contrary to email
> claims / SPAMTREC benchmarks - as there's just way too much noise in
> the inputs obscuring signal.

perhaps a Viterbi decoder is more apt to the task here.

... just a quick 0.02 EUR note - will read up better later/tonight while 
going through my backlog.

-- 
paolo

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.