Re: Re: spam / train-spam
Ray <[email protected]>
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
On Wed, Jun 22, 2005 at 10:38:12PM -0500, John Chandler wrote: > Should the "good" and "spam" commands _ever_ be used? > If so, when? If you don't mind some disk activity, spam would probably get you higher accuracy. Of course, the difference should be marginal to non-existent and more theoretical than practical. > > Date: Tue, 21 Jun 2005 22:16:55 -0400 > > From: Brian Burton <[email protected]> > > To: k b <[email protected]> > > CC: spamprobe-users-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org > > Subject: Re: [Spamprobe-users] spam / train-spam > > > > k b wrote: > > > I'm sure this has been asked before, > > > but what is the difference between `spam' and `train-spam' ? > > > When should each be used and why? > > > > Not sure I can add much practical information to what the readme says: > > > > FROM THE README: > > > > The train command tries to minimize the number of database updates by > > only updating the database with terms from an incoming message if > > there was insufficient confidence in the message's score. The train > > command always updates the database on the first 1500 of each type > > received. This ensures that sufficient email is classified to allow > > the filter to operate reliably. > > > > As an alternative to using the train command, you can run spamprobe in > > "receive" mode. In that mode SpamProbe scores the email and then > > classifies it as either spam or good based on the score. It always > > automatically adds the word counts for the email to the appropriate > > database. This is essentially like running in score mode followed > > immediately by either spam or good mode. It produces more database > > I/O and a bigger database but ensures that every message has its terms > > reflected in the database. Personally I use train mode. A sample > > procmailrc file using the receive command looks like this: > > > > In more detail: The "spam" command updates the database with terms from > > every message it processes. The "train-spam" command only updates the > > database for a small fraction of messages. The decision of whether or > > not to update the database for a given message is based on how the score > > SP calculated for the message. If the score is dead on (close to 0 for > > ham or 1 for spam) then the database isn't updated for that message > > since there really isn't any need. > > > > > > > what is a `difficult' message to classify? > > > > The threshold is about 0.005 for ham and 0.995 for spam. In my own > > experience that seems to result in around 90% of ham and spam not > > needing to be added to the database. There are some other factors > > involved but that's the idea. > ----- > John Chandler > > > ------------------------------------------------------- > SF.Net email is sponsored by: Discover Easy Linux Migration Strategies > from IBM. Find simple to follow Roadmaps, straightforward articles, > informative Webcasts and more! Get everything you need to get up to > speed, fast. http://ads.osdn.com/?ad_id=7477&alloc_id=16492&op=click > _______________________________________________ > Spamprobe-users mailing list > Spamprobe-users-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org > https://lists.sourceforge.net/lists/listinfo/spamprobe-users ------------------------------------------------------- SF.Net email is sponsored by: Discover Easy Linux Migration Strategies from IBM. Find simple to follow Roadmaps, straightforward articles, informative Webcasts and more! Get everything you need to get up to speed, fast. http://ads.osdn.com/?ad_id=7477&alloc_id=16492&op=click