Re: Spamprobe -D option

Brian Burton <[email protected]> Sat, 18 Feb 2006 16:55:39 -0500
Newsgroups gmane.mail.spam.spamprobe.general
Message-ID <[email protected]>
Mark Costlow wrote:
> 3. -D shared -d mine
[snip]
> In case 3, it scored them all as GOOD (0.0000).  This was a little suprising.
> I thought it would score them as SPAM, or at least at something higher than
> 0.0000.
> 
> If I train "mine" on just a few of the "mexican pharmacy" samples (I did 10 of
> them), it starts scoring them all as SPAM, most at 1.0000, all over 0.9900.  I
> seem to get the same scores for both case 2 and case 3.  i.e., the -D option
> doesn't seem to affect the scoring.

This is exactly what I'd expect from the README.TXT file:

  -D directory

     Tells SpamProbe to use the database in the specified directory
     (must be different than the one specified with the -d option) as a
     shared database from which to draw terms that are not defined in
     the user's own database.  This can be used to provide a baseline
     database shared by all users on a system (in the -D directory) and
     a private database unique to each user of the system
     ($HOME/.spamprobe or -d directory).

Maybe I didn't word that well but what it's supposed to say is that the 
database in -D is meant to provide an initial value for terms that are 
not yet in the person's own database.  So if "obstacle" is not in your 
database and you specify a shared database using -D that contains 
"obstacle" then SP will use the counts from the shared database for 
"obstacle" and then add "obstacle" to your database.  From that point on 
it will always use the values from your database.

Since you trained your database on those spams I would expect all of the 
terms in those spams to be present in your database.  Thus when SP is 
scoring the emails it will use the records from your database and score 
the emails based on your training.

If, however, you had never trained your own database on these emails and 
if those emails consisted almost entirely of terms not currently in your 
database, then those emails would have been scored based on the training 
of your shared database.

This may not be what some people imagine when they think of a shared 
database but it is how SP works.  I suppose I could have implemented a 
communal database and/or taken an average of both records but that would 
contradict the idea of the supremacy of each user's own database when 
scoring.

All the best,
++Brian


-------------------------------------------------------
This SF.net email is sponsored by: Splunk Inc. Do you grep through log files
for problems?  Stop!  Download the new AJAX search engine that makes
searching your log files as easy as surfing the  web.  DOWNLOAD SPLUNK!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=103432&bid=230486&dat=121642