Query/suggestion

John Chandler <jpc-WwYSd80vZ+3bI6/[email protected]>
Newsgroups gmane.mail.spam.spamprobe.general
Message-ID <[email protected]>
spamprobe is working well for me.
It screens out over 90% of incoming spam: 870/954 = 91%+
in a recent eleven-day period.
Handling the seven messages per day that leak through is not a problem; 
handling the 80+ total spam messages per day was a huge problem for me
before I started using spamprobe.

So I am a happy user, 
but am always wondering about results that other users get, 
and about improved algorithms for spam detection.

Query:
Are other users still getting 98% or better efficiency from spamprobe?



A different but related topic
-----------------------------

Here is the word count section from a recent false negative message:

X-SpamProbe: GOOD 0.2436125 ecf89d8f80d6e1583c377c5bb522f6c0
    Spam Prob   Count    Good    Spam  Word
    0.9999967       4       0      93  aamnjgithiifeua8brqtaxihqdykeg
    0.0000032       1      47       0  he might
    0.0000046       1      33       0  says it
    0.0000056       1      27       0  his mind
    0.0000076       1      20       0  follow your
    0.9999938       2       0      49  U_aamnjgithiifeua8brqtaxihqdykeg
    0.9999931       2       0      44  aamnjgithiifeua8brqtaxihqdykeg htm
    0.0000101       1      15       0  what she
    0.0000138       1      11       0  again don
    0.0000138       1      11       0  are cured
    0.0000151       2      10       0  to break
    0.0000168       1       9       0  comforting
    0.0000168       1       9       0  let go
    0.0000168       1       9       0  viewing the
    0.0000189       1       8       0  to concentrate
    0.9999888       1       0      27  U_aamnjgithiifeua8brqtaxihqdykeg jpg
    0.9999888       1       0      27  aamnjgithiifeua8brqtaxihqdykeg jpg
    0.9999862       1       0      22  U_aamnjgithiifeua8brqtaxihqdykeg htm
    0.9999862       1       0      22  aamnjgithiifeua8brqtaxihqdykeg html
    0.0000216       1       7       0  causes some
    0.9999724       1       0      11  Hfrom_i
    0.0000379       1       4       0  body for
    0.9999696       1       0      10  Hfrom_ayala


Now the first line,

    0.9999967       4       0      93  aamnjgithiifeua8brqtaxihqdykeg

seems to me to be highly relevant, 
whereas later lines such as

    0.0000032       1      47       0  he might

would seem to be much less relevant.

(It is true that this message is a little atypical 
in containing such a long string
as "aamnjgithiifeua8brqtaxihqdykeg"
that had many appearances in previous spam messages,
but that should only have made it easier to classify.)

Ignoring of common words has long been done 
in information indexing and retrieval systems.
It would seem to me that counting occurrences of such common words as
"he", "it", "his", "your", "what", "she", "to", etc.
may be counterproductive, 
and even "might", "says", "mind", "follow", etc.
may be of questionable utility.
Not considering these words would remove 
almost all of the "good" words from the list above.
Of course other "good" words would then appear.

Would it be reasonable to add a feature to spamprobe
(or other Bayesian spam detectors) to allow the user
to specify a file of common words 
that are not to be counted by the detector,
either singly or in pairs, 
or possibly specify them as being counted in pairs but not singly,
or some such feature?
Or the detector itself could count all words 
and not use those with a relative frequency above some threshold
in computing the spam score.

No doubt this has occurred to authors of spam detectors...
--------------
John Chandler


-------------------------------------------------------
SF email is sponsored by - The IT Product Guide
Read honest & candid reviews on hundreds of IT Products from real users.
Discover which products truly live up to the hype. Start reading now.
http://ads.osdn.com/?ad_id=6595&alloc_id=14396&op=click
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.