Re: Problematic e-mails

Brian Burton <[email protected]> Wed, 14 Dec 2005 09:59:32 -0500
Newsgroups gmane.mail.spam.spamprobe.general
Message-ID <[email protected]>
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA256

Patrick Draper wrote:
> Hi, I've been having problems with e-mails that have nothing in them
> but random words and a single http link. They seem to come from different
> machines each time, and also have random subjects. The problem is that
> every time a new http address shows up, they all get through the filter.
> I'm running the very latest version of spamprobe. Anyone know how
> to filter these out?

This sort of spam is highly database dependent.  A given spam might slip 
through for some people and be stopped completely for others.  It all 
depends on the words used by the spammer and their scores in your database.

I can give a very concrete example of how database dependent this sort 
of thing is.  I recently (last weekend) switched back from using a hash 
database to using a PBL database.  The reason was to make it easy for me 
to review the various Igif terms that 1.3x is finding.  I can't do that 
with a hash database since it doesn't store the original terms, just 
their hash codes.  For the same reason when you switch from hash 
database to PBL (or BDB) you can't export/import to bring your database 
along (the other direction does work though) so I had to build a brand 
new database using old emails.

The result was that there were some spams that started leaking through. 
  Mostly they were of exactly this variety.  At first I thought it might 
be a problem with 1.3x but when I scored the exact same emails using my 
old hash database it caught almost all of them.

So on one level the answer is to just keep training and after a while 
those relatively hammy words will become more neutral and the spams will 
be caught.

BUT you do raise an excellent point about the use of zombies by spammers 
and how SP works when headers change frequently for the same spam body. 
  SP uses a loop when training on a message so that it can ensure that 
once it has finished the loop the spam will be identified as spam if it 
arrives again.  The problem in the case of zombies is that the terms in 
the header tend to become spammy really fast (they are unique after all) 
so the words in the body won't become neutral as fast as they should.

I've been thinking of changing SP's training loop to score the message 
both with and without headers when deciding whether or not to stop 
training on the message.  It would keep training on the message until 
both scores were correct (either ham or spam).  That way emails with new 
headers would continue training until their bodies become spammy, not 
just until their headers become spammy.

I'll experiment with that and see if it works well for these spams.  It 
would not eliminate the problem of the first few of these getting 
through for some people but it could make SP learn to recognize and stop 
them much faster.

What do you folks think?

All the best,
++Brian

-----BEGIN PGP SIGNATURE-----
Version: PGP Desktop 9.0.3 (Build 2932)

iQEVAwUBQ6AzYDxRyEoJfXIFAQgtmgf9HF2CpnA2BAORjurDCbZFD/Q5kr9KxRuR
Arjgnf2+DCRN1KWFdKMqtRgJuSAyyH1SrhBkxi5Uy5y4YlQpi02jvZHSbNbZQoBo
IslIajQrfoTjR40W7z3GMqX9wKB3oD1EgfZoXRUDc0A9l9oq9o2Ccw46NAVbxNb5
WOPEGle2dJk2fmvhXGthqQ/nJsc4V6j0K5ZL6Mq4NoaNtfqDErKoiYocYYrQj/fi
Wz5eBWBFARI/3Hf7r+cT9taq117M1GnvJ8dssfIXGC/cMRaAyrrWcnD2RBHMRcfP
mBSyacK9F+zSJLAaiwHe9MY4VbmwIDoRmhizI6iu1O6zRKKK+UZKgg==
=wbo3
-----END PGP SIGNATURE-----


-------------------------------------------------------
This SF.net email is sponsored by: Splunk Inc. Do you grep through log files
for problems?  Stop!  Download the new AJAX search engine that makes
searching your log files as easy as surfing the  web.  DOWNLOAD SPLUNK!
http://ads.osdn.com/?ad_id=7637&alloc_id=16865&op=click