Re: score graph

"Ger Hobbelt" <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
Eric,

I don't have any good explanation for this behaviour, apart from some
handwaving regarding the way OSB determines the pR from the number of
matched features, and since no-one else seems able to shed more light
on this issue, I'll take you up on your challenge stated over private
email: if you can give me an environment to test this in, I might be
able to find the reason why.

When I had time for crm114 I've been working with mailtrainer.crm the
last few weeks and I ran into some surprises there. It's probably
yours truly the Nutty Professor, but I'd like to have a peek at yours,
if you don't mind.


Alternatively, you might want to have a look at the feature hitcounts
& totalcounts produced by crm114 - that's what I intend to look at at
least - as those are used to calculate the pR value.
It might be that:

- as these hitcounts are 'weighted' (i.e. reduced) according to the
total number of features stored in the spam and good CSS databases
respectively, the resulting pR may be adversily affected due to this
weighting. (spam does not have much in common, and there's a lot of
it, so spam hitcounts will be rather low. When weighted linearly
against CSS database size, those numbers will be reduced further:
that's my guess #1.

- mailtrainer seems to act up overhere and refused to retrain 'mails'
more than twice. To force spam charactistics to 'pop out' despite the
low hitcount and large 'totalcount', you can learn a mail several
times. This helps because a single feature match is not counted as '1'
but as a 1..N number, depending on how many times that particular
feature was learned.

In other words: once you've learned an email, does a recheck with
classify actually report it as spam/good as desired? (I.e. does the
learn process actually push the document beyond the decision
threshold? It might not, when mailtrainer is not configured 'properly'
or acts up otherwise.)


Another thought - but that would mean the pR calculus in OSB mode will
have to be changed: my spam/good collection tends to have a huge load
of spams with a LOT of features, with only VERY FEW of those common
between the 'spams'. Meanwhile, that 'spam' carries a lot of
obfuscation, meaning features which also appear in the 'good'
messages.
Since the pR weighting for both good and spam are linear against the
total number of learned features, learning new spams 'degrades' the pR
for previous ones.
Given the fact that spam these days has a lot of fake 'good' and only
a bit of 'spam' in it feature-wise, I would suggest the spam hitcounts
should NOT be weighted linearly against the total number learned, but
weighting spam should maybe be a power curve. In other words: if I
have a few spam feature hits, say, 8 or more, I would like the pR to
be STRONGLY influenced by such a number, while a count of 1 or 2 spam
hits should be considered minimal. This can be done by weighting the
spam hitcount using a power curve (hitcount ** power) so medium spam
hitcounts weight in stronger than a bunch of 'good'.
I haven't tested this idea for lack of time and most of all: a good
evaluation setup.


Anyway, Eric, I have got a new 64-bit Windows box operational now
(thanks to blown fuzes, fizzing power supplies and all that: Finally I
have some x64 Windows hardware - it was waiting for over a year
already, but now was the time. I'll install VMware Player on it and
see what happens.

Cheers,

Ger



On Fri, May 16, 2008 at 5:18 AM, Eric S. Johansson <[email protected]> wrote:
> here is a graph of the numbers I dumped on the list earlier this week.  the
> light blue area is the "uncertain" region as defined by 2penny blue.  the
> training cutoffs are +-3 as shown in the code also mailed earlier.
>
> as I expect, I get a peak on the green side up around 20 with 200 messages.
>  the red side is very unpleasant.  it peaks right in the prime viewing
> region at just over 800 messages.  why?  help please.
>
> -------------------------------------------------------------------------
> This SF.net email is sponsored by: Microsoft
> Defy all challenges. Microsoft(R) Visual Studio 2008.
> http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/
> _______________________________________________
> Crm114-general mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/crm114-general
>
>



-- 
Met vriendelijke groeten / Best regards,

Ger Hobbelt

--------------------------------------------------
web: http://www.hobbelt.com/
 http://www.hebbut.net/
mail: [email protected]
mobile: +31-6-11 120 978
--------------------------------------------------

-------------------------------------------------------------------------
This SF.net email is sponsored by: Microsoft
Defy all challenges. Microsoft(R) Visual Studio 2008.
http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.