Effects of retaining original encoding/headers in mail classification
Jason White <[email protected]> Wed, 6 May 2009 19:42:24 +1000
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
I have been using CRM114 for spam classification, with Mutt as my MUA. The accuracy hasn't been particularly high. I just discovered a possible cause: at some point in the long forgotten past, for reasons that I can't remember, I set "pipe_decode=yes" in my ~/.muttrc configuration file. The effect of this, due to other configuration settings, is to decode the message and remove all of the headers (except From, To, Subject and Date) before piping it to Mailreaver for training. Of course, the original headers and encoding are intact when messages are processed for classification, since Mailreaver is called by Procmail in that case. If this is likely to have significantly degraded accuracy, should I retrain now from a fresh set of empty CSS files? Obviously, the cache won't be of much use, since the messages are stored there in the stripped and decoded form generated by Mutt when pipe_decode was incorrectly set. I note that the CRM114 documentation advises users to ensure that the messages are trained, as far as possible, in the format in which they are received; thus I assume that the headers and encoding are known to have a significant impact on accuracy, as is to be expected. ------------------------------------------------------------------------------ The NEW KODAK i700 Series Scanners deliver under ANY circumstances! Your production scanning environment may not be a perfect world - but thanks to Kodak, there's a perfect scanner to get the job done! With the NEW KODAK i700 Series Scanner you'll get full speed at 300 dpi even with all image processing features enabled. http://p.sf.net/sfu/kodak-com