Re: Some last info regarding analysis
Jose Marcio Martins da Cruz <[email protected]> Wed, 11 Feb 2004 17:20:29 +0100
| Newsgroups | gmane.ietf.asrg.analysis |
|---|---|
| Organization | Ecole des Mines de Paris |
| Message-ID | <[email protected]> |
I continue to think that there are something to do about entropy. I don't have enough time now to work on it, but I'll surely do. Spam messages are very short. But when you visually look at it, you can near allways class it in a spam/not spam. So, there are intuitively two possibilities : - you do that based on what you learn before about spam - this means spam isn't volatile - there are too much redundancy on spam - you can identify that based on non-volatility. - both of above options - I think this is true. I'm thinking about how can we think about spam sources and create a model comparing them with information sources. So, in this case, I shall imagine what are the "alphabet" to use to model the source, and how can I measure the source entropy, based on the message which is very short. Probably, a good alphabet to use is some representation of the features a spam filter uses to identify spam. So, that's another way to get the same result you've already got. The next step, is to consider that I can't put an observer on the spam source, as I don't know a "gentleman spammer" who could me let me put an observer on his computer... 8-) - But I can put an observer on my server and consider that the trafic on the server is the merge of many sources - some of them are spammers and some of them are normal users. Jose-Marcio Terry Sullivan wrote: > I remember John Graham-Cummings' presentation (the one mentioned in the > BBC piece) from the MIT conference. It was definitely entertaining. > > One of the questions I got at the end of my talk was along the same > lines: essentially, "So, what exactly are the most time-stable features > in SpamAssassin 2.4?" I deliberately deflected the question, but not > because the answer would be of any real help to spammers. Even if every > spammer on the planet suddenly stopped using the "top 5" most > time-stable features in a given feature set, then time-stable features > 6-10 would simply "bubble up" to the top of the list overnight. > > But I tried to make the point, both in the talk and in answer to the > question, that what would be even *better* is for for everyone to apply > the method independently, and thus to find their own "unique" set of > time-stable features. It would make the whole time-stable approach that > much more robust if, instead of one feature set to circumvent, there > were literally hundreds, or even thousands of them, each slightly > different from the other. > > Here's what I think is the interesting distinction: J. Graham-Cummings > approach to spoofing a Bayesian filter works only because Bayesian > filters are "implicitly" trying to distinguish ham from spam. (Which is > yet another great argument against Bayesian filters, in my mind.) In > contrast, the time-stable spam feature approach is not related in any > way to ham characteristics; time-stable features are affected only by > the spam sent. So, unlike Bayesian classifiers, every attempt to > "probe" a time-stable feature set leaves behind a tiny but measurable > "footprint," sort of making the "probe" into a "tell." (Boy, wouldn't > Heisenberg *love* that.) > > The only countermeasure available to a time-stable feature approach is > to force spam to be truly volatile. But as I pointed out in the talk, > if one believes that spammers have been *trying* for volatility all > along, then achieving it is probably *much* harder than it sounds. (In > fact, though I didn't have time to mention this in the talk, at no time > in the last 2.5 year was spam from any two successive calendar quarters > more different-than-similar. In all 10 cases, spam across successive > quarters was more similar-than-different.) > > It'd require a systematic, sustained, coordinated effort for content, > obfuscation strategies, markup, URLs, etc. to change constantly, with no > spammer doing too much of any one thing, nor any two doing any one thing > simultaneously (at least not to the same domains). As Thomas Juntenen > and I became fond of saying to each other as we examined the volatility > results, "There are only so many ways to munge a message." > > In an off-list correspondence, John Levine recently suggested something > to the effect that "maybe filter builders need better feature analysis." > He's absolutely dead-on right: a concentrated quantitative effort at > feature analysis/optimization can result in a dramatic improvement in > FNR. Properly optimized SpamAssassin can beat Brightmail's FNR by a > hefty margin. > > - Terry > > > On Wed, 04 Feb 2004 16:54:15 -0500, Yakov Shafranovich wrote: > > >>Kurt Magnusson wrote: >> >>>Before the group closes down, here is a interesting BBC article on >>>spamfiltering and beating it. >>> >>>http://news.bbc.co.uk/1/hi/technology/3458457.stm >> >>This raises an interesting point - the same way we can experiment with >>spam to determine whether its volatile or not, and its common features; >>spammers can do the same for anti-spam systems, maybe by using the same >>methods. >> >>Yakov > > > > -- --------------------------------------------------------------- Jose Marcio MARTINS DA CRUZ Tel. :(33) 01.40.51.93.41 Ecole des Mines de Paris http://j-chkmail.ensmp.fr 60, bd Saint Michel http://www.ensmp.fr/~martins 75272 - PARIS CEDEX 06 mailto:[email protected]