Re: Review of bayesian spam filters
Brian Burton <[email protected]> Fri, 24 Feb 2006 09:22:51 -0500
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
Anthony Campbell wrote: > On 23 Feb 2006, Nicolas Duboc wrote: >> Linux Weekly News has just published a review of bayesian spam >>filters. It includes SpamProbe. >> >> http://lwn.net/SubscriberLink/172491/7eaf0d9defca131c/ > > I was surprised at the relatively poor performance of SP here. > Insufficient training? I get NO false positives at all and not more than > one or two false negatives a day, sometimes none. I haven't read the review in detail so take my comments with a grain of salt. I don't take comparative articles all that seriously. "Any given Sunday" (American football cliche) applies to these reviews (maybe better stated as "Any given corpus"). Sometimes one filter will win in one review but lose on another. I don't believe that bogofilter is as inaccurate as this review suggests, for example. It looks as though the reviewer did a good job of trying to be thorough and fair. His description of how he trained on each error throughout the test was a good way to simulate real life use of the filters. One thing I'd point out is that SpamAssassin's performance is probably enhanced by the fact that he was using older mail with a newish version of SpamAssassin (my assumption). Since SA doesn't automatically adapt to errors as a bayesian filter does it really should be judged using a "0-day" version (i.e. one that predates the email being judged) to simulate what would be seen in real use. Also he used the default settings for most other filters but specifically tweaked SA to allow it's bayesian more weight. Note also that since he started counting errors after training on only 1000 messages (1/6) of his corpus he is really evaluating the filters on how quickly they learn, not on what their steady state accuracy is. That's fine but some filters might be more accurate early on while others are more accurate with several thousand emails under their belt. Most testers train on 90% of the corpus and then evaluate the filters on the remaining 10%. That said I think it was a good article and it was obvious that the reviewer went to a lot of effort to make a fair comparison. All the best, ++Brian ------------------------------------------------------- This SF.Net email is sponsored by xPML, a groundbreaking scripting language that extends applications into web and mobile media. Attend the live webcast and join the prime developer group breaking into this new coding territory! http://sel.as-us.falkag.net/sel?cmd=lnk&kid=110944&bid=241720&dat=121642