(unknown)

"Kurt Magnusson" <[email protected]> Sun, 14 Mar 2004 08:57:11 -0300
Newsgroups gmane.ietf.asrg.filtering
Message-ID <[email protected]>
There have earlier in the group been some discussion about comparing spam 
with biological virus/bacteria, an idea I been opposed to, since I believe 
that if we don't look at spam from a economical perspective we loose the 
cause of its existence.

But Yakov sent out a link last week about the biological angel, that gave me 
some thoughts, I would like to see some views about.

Reason is that I have during the last 2 weeks registered a new tactic from 
the spammers.  It is related to others implementation of the same idea I 
have been working with and call Earnest, to filter on the real info in the 
spams. I have at last succeeded in handling some problems with the different 
greps in Solaris and some web char filtering, getting my PoC to work stable, 
so I really catch those spam-domains that exists in the URL's (or phone 
numbers). Once I seen it and caught it, those spams is dead.

Yakov heard that AOL and Yahoo also worked on this type of method, which 
have made some, more solvent, spammers to change domains for "every" 
spamsession send out. From 1-2 passing per day, I have now 3-5 passing of ca 
30 spams. Still better than my ISPs filter, whose filtering I also use, but 
not good enough.

The limitation in my method is just that, I need to have the domain in my 
database, before it works, but how do I catch the first one, if the ISPs 
filter doesn't. I'll think some other methods have the same problem. When 
reading the papers in link Yakov sent, I got a possible attack method, the  
question is how to utilize it without having every MTA changed? For it need 
to be on the MTA level.

In short, apart of their often destructive action, what is the most notable 
feature of a virus in a active stage?

Their amount. One virus doesn't make us sick, 100.000 does. Which is the 
same as for spam.

Some have therefore looked at spam in the sender end, by delaying large 
volume transfers, to make the spammers loose time. And as some in the 
analyze-group did look at, you can view it also on the receiving side. Those 
that worked with it, did look though at the connection state, the number of 
active connections - a lot, spam.

What I am interested in, is more the statistical side. If a site could cache 
all mail, say for 5-10 minutes, and make a analysis on the mails to see if 
there is a number of mails with some equal features, URL's, Phone numbers, 
routing info.

If there exist more than 5 letters with the same features, you have 4 
choices, a group of people with the same contact net (buddy mass mailings), 
several on a same maillist, vendor mailing (which is a sort of spam) and 
spam. That is, with a relatively small amount of mails in a cache, we can 
establish if any mass mailings is at hand.

The cashing is a floating one, that is, incoming mail is on hold for 5-10 
minutes and every minute a scan is done, so every mail is included in least 
4-9 scans. No datafiles need to be matched, we just looking for mails with 
the same active content. If suspect, the letters is quarantined for a check 
with a very thorough check, as [for example] all Spamassassins possible 
criterias. Since it isn't all letters checked, we can afford to do as 
complicated check as possible. If a spam, a [SPAM] tag is attached.

Here I, from a practical perspective, in respect to previous threads think 
we, independent of needs for system headers, should have the subject line 
tagged, since several mail systems don't allows access to all headers and 
makes it harder for the end user to a local end filtering, as example: Notes 
don't show the headers for a end user. If on the subject line, it is very 
simple for the end user to handle it.

If the end user thinks it is spam, he/she should be able to report it and a 
more formal filtering could be done. The two articles also mentioned the 
need of whitelistings and I believe that here it is a need, that people need 
to put mail list their on, on their whitelists.

I think Jose Marcio have worked somewhat in this line, at least looking at 
statistics on the last used URL's and that Brightmail might also do the 
same, so it isn't without precedence and since at a larger site, corporation 
or ISP do receive spam in chunks, you should be able to find more than 5 
spams at a time. Question is if we with for ex. with milter can do such a 
cache. The checking is simple and principle the same as I do in my Earnest 
filter, a simple grep for least common denominator in urls and phone 
nos/routing info.

In order to kill such a filter, the spammers need 10-20 domains for every 
send-out, see to that routing is done through a number of places, before 
arriving to the same target domain, so there is no more that 3-4 in queue at 
the same time. Then it really will cost infrastructure for them and if they 
spread them out, the chance of a domain to be reported gets bigger and can 
be effectively blocked in the normal spam system before there is a volume. 
And the spammers get a lesser ROI.

Regards Kurt

_________________________________________________________________
Add photos to your e-mail with MSN 8. Get 2 months FREE*. 
http://join.msn.com/?page=features/featuredemail