Re: Re: Empty message apparently not tokenized

Brian Burton <[email protected]> Sat, 18 Feb 2006 10:41:38 -0500
Newsgroups gmane.mail.spam.spamprobe.general
Message-ID <[email protected]>
tBB wrote:
> I wouldn't be too sure that mails with empty bodys are caught if this 
> bug is fixed because many of them originate from zombies so the tokens 
> SP pulls out might be not that useful either. See the related discussion 
> on this list a while ago.
> 
> As I mentioned above, I have the "-H all" option enabled since a few 
> weeks and the processing of all headers noticeably increased the overall 
> accuracy of SpamProbe and that *not* only for those mails with empty 
> bodys (of which several hundred reach my server each day). That suggests 
> that there are rather useful tokens in the headers which are not 
> processed by default. Besides, there are quite a few Bayesian 
> classifiers out which do process all headers by default so I wouldn't go 
> that far to say it is not smart per se.

That's a good observation.  One caveat though: in my testing -Hall can 
lead to extra false positives.  Not many, but a few.  That's why it's 
not the default setting.  Of course this is extremely data dependent.

I have some other options I'd like to leave on by default but don't 
because of increased FPs in testing.  One that seemed like a slam dunk 
was to keep adding tokens to phrases up to a certain length.  For 
example, instead of just taking phrases of 2 words take all phrases of 
12 characters or less in length.  That's awesome conceptually since "v i 
a g r a" would be a phrase even though it's 6 "words" long.

Seems like an obvious win, right?  Well when I compare that option to 
the default and run my cross validation tests I get only slightly better 
accuracy and a higher false positive rate.  Bummer. :-)

The more I experiment with SP the more I'm surprised by test results. 
Many of the techniques that seem like obvious winners can turn out not 
to be.

I've also thought about scoring messages multiple ways and then 
combining the scores to make an overall score.  I have an experimental 
mode that does this.  The result are, no surprise, mixed.  It's hard to 
know which to believe the most.  If you favor the most incriminating 
score you'll tend to get more false positives.  If you favor the least 
incriminating then your accuracy drops.  If you average them you tend to 
wind up with the worst of both worlds.

Like I said though it's all data dependent.  Some variations on the data 
lead to one technique being much more accurate than others.  I'm sure 
some of these techniques could prove to be better for some people but 
you wouldn't know until after you'd run it for a while.

Fun stuff...

All the best,
++Brian


-------------------------------------------------------
This SF.net email is sponsored by: Splunk Inc. Do you grep through log files
for problems?  Stop!  Download the new AJAX search engine that makes
searching your log files as easy as surfing the  web.  DOWNLOAD SPLUNK!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=103432&bid=230486&dat=121642