Re: Spamprobe -D option
Brian Burton <[email protected]> Sat, 18 Feb 2006 16:55:39 -0500
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
Mark Costlow wrote:
> 3. -D shared -d mine
[snip]
> In case 3, it scored them all as GOOD (0.0000). This was a little suprising.
> I thought it would score them as SPAM, or at least at something higher than
> 0.0000.
>
> If I train "mine" on just a few of the "mexican pharmacy" samples (I did 10 of
> them), it starts scoring them all as SPAM, most at 1.0000, all over 0.9900. I
> seem to get the same scores for both case 2 and case 3. i.e., the -D option
> doesn't seem to affect the scoring.
This is exactly what I'd expect from the README.TXT file:
-D directory
Tells SpamProbe to use the database in the specified directory
(must be different than the one specified with the -d option) as a
shared database from which to draw terms that are not defined in
the user's own database. This can be used to provide a baseline
database shared by all users on a system (in the -D directory) and
a private database unique to each user of the system
($HOME/.spamprobe or -d directory).
Maybe I didn't word that well but what it's supposed to say is that the
database in -D is meant to provide an initial value for terms that are
not yet in the person's own database. So if "obstacle" is not in your
database and you specify a shared database using -D that contains
"obstacle" then SP will use the counts from the shared database for
"obstacle" and then add "obstacle" to your database. From that point on
it will always use the values from your database.
Since you trained your database on those spams I would expect all of the
terms in those spams to be present in your database. Thus when SP is
scoring the emails it will use the records from your database and score
the emails based on your training.
If, however, you had never trained your own database on these emails and
if those emails consisted almost entirely of terms not currently in your
database, then those emails would have been scored based on the training
of your shared database.
This may not be what some people imagine when they think of a shared
database but it is how SP works. I suppose I could have implemented a
communal database and/or taken an average of both records but that would
contradict the idea of the supremacy of each user's own database when
scoring.
All the best,
++Brian
-------------------------------------------------------
This SF.net email is sponsored by: Splunk Inc. Do you grep through log files
for problems? Stop! Download the new AJAX search engine that makes
searching your log files as easy as surfing the web. DOWNLOAD SPLUNK!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=103432&bid=230486&dat=121642