Database and diminishing returns
"Jem" <[email protected]>
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <[email protected]> |
Now that I've been using spamprobe since 2003 I'm wondering if there has been a wrong impression about the size of a database you need to have effective filtering. For instance, sure we know there are diminishing returns in effectiveness as your database grows. But perhaps a small database really can be much more capable than we originally thought? Because I was concerned about database growth in the old days (DB etc) I kept using score only, and then trained on corrections. The approx 6 months of recent mail I've used to train now gives me about 99% accuracy, but what I discovered today is I really have very few terms in my database: berkes@mail:~$ spamprobe counts GOOD 1903 SPAM 3438 berkes@mail:~$ spamprobe export | wc -l 139142 With the latest hash database, this 1.6 MB database is enough to give 99% accuracy. That seems pretty amazing! I have two other distinct accounts, and the database size and effectiveness statistics are comparable. All of my accounts get much more spam than non-spam, about 6:1. Perhaps this is the main factor that has helped my effectiveness %? I don't know. I'm trying to think of the possible implications. One might be that a static database (read-only, score) can be nearly as effective as a self updating database (read/write, receive or train). I'd be curious about others' experience in this respect. I can say that spamprobe's performance on my mail doesn't seem to get much better when I use receive/train instead of score, provided the working database is already well tuned of course. ------------------------------------------------------- SF email is sponsored by - The IT Product Guide Read honest & candid reviews on hundreds of IT Products from real users. Discover which products truly live up to the hype. Start reading now. http://ads.osdn.com/?ad_id=6595&alloc_id=14396&op=click