Re: Re: Empty message apparently not tokenized
Brian Burton <[email protected]> Sat, 18 Feb 2006 10:41:38 -0500
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
tBB wrote: > I wouldn't be too sure that mails with empty bodys are caught if this > bug is fixed because many of them originate from zombies so the tokens > SP pulls out might be not that useful either. See the related discussion > on this list a while ago. > > As I mentioned above, I have the "-H all" option enabled since a few > weeks and the processing of all headers noticeably increased the overall > accuracy of SpamProbe and that *not* only for those mails with empty > bodys (of which several hundred reach my server each day). That suggests > that there are rather useful tokens in the headers which are not > processed by default. Besides, there are quite a few Bayesian > classifiers out which do process all headers by default so I wouldn't go > that far to say it is not smart per se. That's a good observation. One caveat though: in my testing -Hall can lead to extra false positives. Not many, but a few. That's why it's not the default setting. Of course this is extremely data dependent. I have some other options I'd like to leave on by default but don't because of increased FPs in testing. One that seemed like a slam dunk was to keep adding tokens to phrases up to a certain length. For example, instead of just taking phrases of 2 words take all phrases of 12 characters or less in length. That's awesome conceptually since "v i a g r a" would be a phrase even though it's 6 "words" long. Seems like an obvious win, right? Well when I compare that option to the default and run my cross validation tests I get only slightly better accuracy and a higher false positive rate. Bummer. :-) The more I experiment with SP the more I'm surprised by test results. Many of the techniques that seem like obvious winners can turn out not to be. I've also thought about scoring messages multiple ways and then combining the scores to make an overall score. I have an experimental mode that does this. The result are, no surprise, mixed. It's hard to know which to believe the most. If you favor the most incriminating score you'll tend to get more false positives. If you favor the least incriminating then your accuracy drops. If you average them you tend to wind up with the worst of both worlds. Like I said though it's all data dependent. Some variations on the data lead to one technique being much more accurate than others. I'm sure some of these techniques could prove to be better for some people but you wouldn't know until after you'd run it for a while. Fun stuff... All the best, ++Brian ------------------------------------------------------- This SF.net email is sponsored by: Splunk Inc. Do you grep through log files for problems? Stop! Download the new AJAX search engine that makes searching your log files as easy as surfing the web. DOWNLOAD SPLUNK! http://sel.as-us.falkag.net/sel?cmd=lnk&kid=103432&bid=230486&dat=121642