Re: Problems with Bayesian filtering

Bill Yerazunis <[email protected]> Fri, 27 Feb 2004 07:15:06 -0500
Newsgroups gmane.ietf.asrg.filtering
Message-ID <[email protected]>
   From: Carl Hutzler <[email protected]>

   Our bayesian filters are being worked around by spammers using "chaff 
   text" which is either in white font on white background, 1 pt font, or 
   in some cases, just normal text a few CR's below the spammer's VERY 
   SHORT message.

   What is the solution to this little trick?

I have no trouble with these "dictionary salad" attacks in CRM114,
which uses a Markovian model of the text.

I also understand that DSPAM (which uses chained tokens) also has
very little problem with them.

     -Bill Yerazunis