Re: Some last info regarding analysis

Jose Marcio Martins da Cruz <[email protected]> Wed, 11 Feb 2004 17:20:29 +0100
Newsgroups gmane.ietf.asrg.analysis
Organization Ecole des Mines de Paris
Message-ID <[email protected]>

I continue to think that there are something to do about entropy. I 
don't have enough time now to work on it, but I'll surely do.

Spam messages are very short. But when you visually look at it, you can 
near allways class it in a spam/not spam.

So, there are intuitively two possibilities :

- you do that based on what you learn before about spam - this means
   spam isn't volatile
- there are too much redundancy on spam - you can identify that
   based on non-volatility.
- both of above options - I think this is true.

I'm thinking about how can we think about spam sources and create
a model comparing them with information sources.

So, in this case, I shall imagine what are the "alphabet" to use
to model the source, and how can I measure the source entropy,
based on the message which is very short.

Probably, a good alphabet to use is some representation of the features 
a spam filter uses to identify spam. So, that's another way to get
the same result you've already got.

The next step, is to consider that I can't put an observer on the
spam source, as I don't know a "gentleman spammer" who could me let
me put an observer on his computer... 8-) - But I can put an
observer on my server and consider that the trafic on the server
is the merge of many sources - some of them are spammers and some
of them are normal users.

Jose-Marcio

Terry Sullivan wrote:
> I remember John Graham-Cummings' presentation (the one mentioned in the 
> BBC piece) from the MIT conference.  It was definitely entertaining.
> 
> One of the questions I got at the end of my talk was along the same 
> lines: essentially, "So, what exactly are the most time-stable features 
> in SpamAssassin 2.4?"  I deliberately deflected the question, but not 
> because the answer would be of any real help to spammers.  Even if every 
> spammer on the planet suddenly stopped using the "top 5" most 
> time-stable features in a given feature set, then time-stable features 
> 6-10 would simply "bubble up" to the top of the list overnight.
> 
> But I tried to make the point, both in the talk and in answer to the 
> question, that what would be even *better* is for for everyone to apply 
> the method independently, and thus to find their own "unique" set of 
> time-stable features.  It would make the whole time-stable approach that 
> much more robust if, instead of one feature set to circumvent, there 
> were literally hundreds, or even thousands of them, each slightly 
> different from the other.
> 
> Here's what I think is the interesting distinction: J. Graham-Cummings 
> approach to spoofing a Bayesian filter works only because Bayesian 
> filters are "implicitly" trying to distinguish ham from spam.  (Which is 
> yet another great argument against Bayesian filters, in my mind.)  In 
> contrast, the time-stable spam feature approach is not related in any 
> way to ham characteristics; time-stable features are affected only by 
> the spam sent.  So, unlike Bayesian classifiers, every attempt to 
> "probe" a time-stable feature set leaves behind a tiny but measurable 
> "footprint," sort of making the "probe" into a "tell."  (Boy, wouldn't 
> Heisenberg *love* that.) 
> 
> The only countermeasure available to a time-stable feature approach is 
> to force spam to be truly volatile.  But as I pointed out in the talk, 
> if one believes that spammers have been *trying* for volatility all 
> along, then achieving it is probably *much* harder than it sounds.  (In 
> fact, though I didn't have time to mention this in the talk, at no time 
> in the last 2.5 year was spam from any two successive calendar quarters 
> more different-than-similar.  In all 10 cases, spam across successive 
> quarters was more similar-than-different.)
> 
> It'd require a systematic, sustained, coordinated effort for content, 
> obfuscation strategies, markup, URLs, etc. to change constantly, with no 
> spammer doing too much of any one thing, nor any two doing any one thing 
> simultaneously (at least not to the same domains).  As Thomas Juntenen 
> and I became fond of saying to each other as we examined the volatility 
> results, "There are only so many ways to munge a message."
> 
> In an off-list correspondence, John Levine recently suggested something 
> to the effect that "maybe filter builders need better feature analysis."  
> He's absolutely dead-on right: a concentrated quantitative effort at 
> feature analysis/optimization can result in a dramatic improvement in 
> FNR.  Properly optimized SpamAssassin can beat Brightmail's FNR by a 
> hefty margin.
> 
> - Terry
> 
> 
> On Wed, 04 Feb 2004 16:54:15 -0500, Yakov Shafranovich wrote:
> 
> 
>>Kurt Magnusson wrote:
>>
>>>Before the group closes down, here is a interesting BBC article on 
>>>spamfiltering and beating it.
>>>
>>>http://news.bbc.co.uk/1/hi/technology/3458457.stm
>>
>>This raises an interesting point - the same way we can experiment with 
>>spam to determine whether its volatile or not, and its common features; 
>>spammers can do the same for anti-spam systems, maybe by using the same 
>>methods.
>>
>>Yakov
> 
> 
> 
> 



-- 
  ---------------------------------------------------------------
  Jose Marcio MARTINS DA CRUZ           Tel. :(33) 01.40.51.93.41
  Ecole des Mines de Paris              http://j-chkmail.ensmp.fr
  60, bd Saint Michel                http://www.ensmp.fr/~martins
  75272 - PARIS CEDEX 06      mailto:[email protected]