new feature ideas
Brian Burton <[email protected]>
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
I'm thinking of adding some new features to SP and thought I'd bounce
them off you folks to see if they seem desireable.
* Add aliases for the various *good commands using ham instead of good.
This was actually requested by Karl Fogel. (so train-ham would do the
same thing as train-good for example)
* Add ability to specify different thresholds for classifying spam/ham.
I would allow more than just two levels so that you could put obvious
spam into one folder, less obvious spam into another, etc. I do
something like this now using a procmail recipe but its painful that
way. The option would look something like this:
spamprobe -o thresholds=0.95:SPAM,0.80:SUSPICIOUS,0.0:GOOD train
The decimal number is a score above which the lable that follows it
will be printed. The example above woud have three levels instead of
the normal two that SP uses. These thresholds would only affect the
classification printed by SP. The database related thresholds would
still use the defaults.
* Add ability to set the SPAM/GOOD threshold values for training. I
think this would be pretty dangerous so I'm hesitant to actually do
this one.
* Add ability to operate in "filter" mode with SP printing the email to
stdout with a score and classification added to the Subject: header.
SP could also add its own header instead of modifying the Subject.
I'm on record as not really liking this feature much but several
people have asked for it and it'd be fairly easy to do with the
new parser. This seems fairly dangerous since if SP crashed
(hey, it could happen) the email might be lost. Altering the Subject
of the email could be tricky if the Subject header is encoded in an
alternative charset then the score would mess that up.
* I18N improvements. I'm toying with the idea of making SP more charset
aware. I'd switch to using wchar_t instead of char and add code for
processing each message part in its proper charset. All terms in the
database would be stored in UTF-8. I'd probably use trigrams for
tokens in wide character messages since, from what I've heard, you
can't use whitespace as a token delimiter in asian languages. I'm
not sure if there is a lot of interest in having this feature. It
would be a big code change runtime memory usage would be impacted
significantly. Ideally I'd make this a compile time option so people
who don't ever receive legitimate multi-byte char email would not
have to take the memory/performance hit.
What do you folks think? Do these seem like useful ideas? Anyone else
have an interesting idea? Preferably something that would improve the
user interface without requiring major new code.
All the best,
++Brian
-------------------------------------------------------
SF.Net email is sponsored by: Discover Easy Linux Migration Strategies
from IBM. Find simple to follow Roadmaps, straightforward articles,
informative Webcasts and more! Get everything you need to get up to
speed, fast. http://ads.osdn.com/?ad_id=7477&alloc_id=16492&op=click