Effects of retaining original encoding/headers in mail classification

Jason White <[email protected]> Wed, 6 May 2009 19:42:24 +1000
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
I have been using CRM114 for spam classification, with Mutt as my MUA.

The accuracy hasn't been particularly high. I just discovered a possible
cause: at some point in the long forgotten past, for reasons that I can't
remember, I set "pipe_decode=yes" in my ~/.muttrc configuration file. The
effect of this, due to other configuration settings, is to decode the message
and remove all of the headers (except From, To, Subject and Date) before
piping it to Mailreaver for training. Of course, the original headers and
encoding are intact when messages are processed for classification, since
Mailreaver is called by Procmail in that case.

If this is likely to have significantly degraded accuracy, should I retrain
now from a fresh set of empty CSS files?

Obviously, the cache won't be of much use, since the messages are stored there
in the stripped and decoded form generated by Mutt when pipe_decode was
incorrectly set.

I note that the CRM114 documentation advises users to ensure that the messages
are trained, as far as possible, in the format in which they are received;
thus I assume that the headers and encoding are known to have a significant
impact on accuracy, as is to be expected.


------------------------------------------------------------------------------
The NEW KODAK i700 Series Scanners deliver under ANY circumstances! Your
production scanning environment may not be a perfect world - but thanks to
Kodak, there's a perfect scanner to get the job done! With the NEW KODAK i700
Series Scanner you'll get full speed at 300 dpi even with all image 
processing features enabled. http://p.sf.net/sfu/kodak-com