Re: spamprobe-1.1x7 released

"David A. Lee" <[email protected]>
Newsgroups gmane.mail.spam.spamprobe.general
Message-ID <002901c534e1$f9ac1ee0$09fd15ac@ENTERPRISE>
I deleted the email from (?) about the 1.9 MB Hash DB that had 99% effective ... that got me thinking ...
My terms are about 10x higher then the referenced ones, but reguardless, I have fluxuated between 98% and 99%  for over a year now,
I have not done blind side-by-side tests of simply score vs  score/train or receive/train ... but I have done some informal studies
where I coppied my personal DB to my wife's and then she uses it in read-only mode.  Works fairly well ... but her spam count is
upping
wheras mine is not. I think that is because the nature of the spam we get is slowly changing. Thus using a static DB is good for a
while, but
eventually gets worse ... of course your spam will vary from mine ... so whiel I think a static DB is a good temporary solution, I
think
in most cases it will become worse and worse over time, ... but what time ? I dont know ...  -1% a month ?

anyway that's tangent to my thought lines ... I was thinking "wow" with only  2MB for a good spam DB ... that totally changes the
formulas for large groups of users !! especially consider you  might be able to maintain, say, a 10MB static DB and a 1MB private DB
... that could scale up really well ... very intersting, I need to try out this new hash DB !

I am wondering, however, what is the failure mode as the hash table gets full ? you mention it starts to 'drop' terms ... sure, the
hash buckets get overloaded and/or 2 strings hash to the same value%DB-size ... but whats the failure mode ?  What terms get dropped
?
Is there a way this could be smart ... say drop terms that are older .. or less weights etc ? so that in time even a 2MB hash DB
could be trained and trained maybe way way beyond its capicity but the information it loses is not important information ....


Reminds me of an old Married-with-children episode where the blonde is training for a quiz show ... turns out she can memorize quite
a lot, except her dad warns her that there's only so much room in her brain ... and once its full, old stuff falls out to make way
for new,
of course right before the quiz she is exposed to some new (useless) information and out goes all the stuff she memorized ...

How do we keep SP from doing this ? In fact how can we *leverage* a fixed size ... there's a coorilary in neural networks,
if a neural network is too small it can never train enough for a given dataset, but OTOH if its too large, it never learns to
generalize,
rather it simply trains on all the specifics ... there's an optimum size range which works best ... big enough to store the data,
but not so big that it stores every little bit ... can SP be tuned to do this somehow ?  If so, conceviably one could get *better*
results from a limited DB then an unlimited one ...
-----------------------------------------------------------
David A. Lee
[email protected]
http://www.calldei.com



-------------------------------------------------------
SF email is sponsored by - The IT Product Guide
Read honest & candid reviews on hundreds of IT Products from real users.
Discover which products truly live up to the hype. Start reading now.
http://ads.osdn.com/?ad_id=6595&alloc_id=14396&op=click
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.