Re: spamprobe-1.1x7 released
"David A. Lee" <[email protected]>
| Newsgroups | gmane.mail.spam.spamprobe.general |
|---|---|
| Message-ID | <002901c534e1$f9ac1ee0$09fd15ac@ENTERPRISE> |
I deleted the email from (?) about the 1.9 MB Hash DB that had 99% effective ... that got me thinking ... My terms are about 10x higher then the referenced ones, but reguardless, I have fluxuated between 98% and 99% for over a year now, I have not done blind side-by-side tests of simply score vs score/train or receive/train ... but I have done some informal studies where I coppied my personal DB to my wife's and then she uses it in read-only mode. Works fairly well ... but her spam count is upping wheras mine is not. I think that is because the nature of the spam we get is slowly changing. Thus using a static DB is good for a while, but eventually gets worse ... of course your spam will vary from mine ... so whiel I think a static DB is a good temporary solution, I think in most cases it will become worse and worse over time, ... but what time ? I dont know ... -1% a month ? anyway that's tangent to my thought lines ... I was thinking "wow" with only 2MB for a good spam DB ... that totally changes the formulas for large groups of users !! especially consider you might be able to maintain, say, a 10MB static DB and a 1MB private DB ... that could scale up really well ... very intersting, I need to try out this new hash DB ! I am wondering, however, what is the failure mode as the hash table gets full ? you mention it starts to 'drop' terms ... sure, the hash buckets get overloaded and/or 2 strings hash to the same value%DB-size ... but whats the failure mode ? What terms get dropped ? Is there a way this could be smart ... say drop terms that are older .. or less weights etc ? so that in time even a 2MB hash DB could be trained and trained maybe way way beyond its capicity but the information it loses is not important information .... Reminds me of an old Married-with-children episode where the blonde is training for a quiz show ... turns out she can memorize quite a lot, except her dad warns her that there's only so much room in her brain ... and once its full, old stuff falls out to make way for new, of course right before the quiz she is exposed to some new (useless) information and out goes all the stuff she memorized ... How do we keep SP from doing this ? In fact how can we *leverage* a fixed size ... there's a coorilary in neural networks, if a neural network is too small it can never train enough for a given dataset, but OTOH if its too large, it never learns to generalize, rather it simply trains on all the specifics ... there's an optimum size range which works best ... big enough to store the data, but not so big that it stores every little bit ... can SP be tuned to do this somehow ? If so, conceviably one could get *better* results from a limited DB then an unlimited one ... ----------------------------------------------------------- David A. Lee [email protected] http://www.calldei.com ------------------------------------------------------- SF email is sponsored by - The IT Product Guide Read honest & candid reviews on hundreds of IT Products from real users. Discover which products truly live up to the hype. Start reading now. http://ads.osdn.com/?ad_id=6595&alloc_id=14396&op=click