Re: Problematic e-mails
Brian Burton <[email protected]> Wed, 14 Dec 2005 09:59:32 -0500
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA256 Patrick Draper wrote: > Hi, I've been having problems with e-mails that have nothing in them > but random words and a single http link. They seem to come from different > machines each time, and also have random subjects. The problem is that > every time a new http address shows up, they all get through the filter. > I'm running the very latest version of spamprobe. Anyone know how > to filter these out? This sort of spam is highly database dependent. A given spam might slip through for some people and be stopped completely for others. It all depends on the words used by the spammer and their scores in your database. I can give a very concrete example of how database dependent this sort of thing is. I recently (last weekend) switched back from using a hash database to using a PBL database. The reason was to make it easy for me to review the various Igif terms that 1.3x is finding. I can't do that with a hash database since it doesn't store the original terms, just their hash codes. For the same reason when you switch from hash database to PBL (or BDB) you can't export/import to bring your database along (the other direction does work though) so I had to build a brand new database using old emails. The result was that there were some spams that started leaking through. Mostly they were of exactly this variety. At first I thought it might be a problem with 1.3x but when I scored the exact same emails using my old hash database it caught almost all of them. So on one level the answer is to just keep training and after a while those relatively hammy words will become more neutral and the spams will be caught. BUT you do raise an excellent point about the use of zombies by spammers and how SP works when headers change frequently for the same spam body. SP uses a loop when training on a message so that it can ensure that once it has finished the loop the spam will be identified as spam if it arrives again. The problem in the case of zombies is that the terms in the header tend to become spammy really fast (they are unique after all) so the words in the body won't become neutral as fast as they should. I've been thinking of changing SP's training loop to score the message both with and without headers when deciding whether or not to stop training on the message. It would keep training on the message until both scores were correct (either ham or spam). That way emails with new headers would continue training until their bodies become spammy, not just until their headers become spammy. I'll experiment with that and see if it works well for these spams. It would not eliminate the problem of the first few of these getting through for some people but it could make SP learn to recognize and stop them much faster. What do you folks think? All the best, ++Brian -----BEGIN PGP SIGNATURE----- Version: PGP Desktop 9.0.3 (Build 2932) iQEVAwUBQ6AzYDxRyEoJfXIFAQgtmgf9HF2CpnA2BAORjurDCbZFD/Q5kr9KxRuR Arjgnf2+DCRN1KWFdKMqtRgJuSAyyH1SrhBkxi5Uy5y4YlQpi02jvZHSbNbZQoBo IslIajQrfoTjR40W7z3GMqX9wKB3oD1EgfZoXRUDc0A9l9oq9o2Ccw46NAVbxNb5 WOPEGle2dJk2fmvhXGthqQ/nJsc4V6j0K5ZL6Mq4NoaNtfqDErKoiYocYYrQj/fi Wz5eBWBFARI/3Hf7r+cT9taq117M1GnvJ8dssfIXGC/cMRaAyrrWcnD2RBHMRcfP mBSyacK9F+zSJLAaiwHe9MY4VbmwIDoRmhizI6iu1O6zRKKK+UZKgg== =wbo3 -----END PGP SIGNATURE----- ------------------------------------------------------- This SF.net email is sponsored by: Splunk Inc. Do you grep through log files for problems? Stop! Download the new AJAX search engine that makes searching your log files as easy as surfing the web. DOWNLOAD SPLUNK! http://ads.osdn.com/?ad_id=7637&alloc_id=16865&op=click