I don't get it anymore (if ever): OSB pR increases linearly with message size. Is that by design?

"Ger Hobbelt" <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
While I have been working on mailreaver et al I ran into a - to me -
weird phenomenon, which I had not seen with my own data - because I
never tried it.

train a few (small) messages in good and bad.css. One of those trained
messages (in good) is 'A B C', just to make matters _Extremely_
simple.

Now classify that same message 'A B C' and you'll get a certain pR.
Should be large as we should have an exact match on the good side and
only one or two words of those three appearing in 'bad.css' as well
(pR ~ 4 over here).

(Side track: pR will be very low (~ 0.4) when 'A B C' was also trained
into 'bad.css', but more often in 'good.css'; I won't discuss that
'collision' scenario right now.)

So far, so good. (The test was meant to see what ONE word of
difference would do.)



Now here's the weird bit:

classify 'A B C' --> pR = 4

classify message 'A B C A B C' (copy concatenated to it, so new
message is double size, same content) --> pR = 8

In l33t speak, my response was 'WTF?'

now classify 'A B C A B C A B C'... pR = 12

hairs start to raise on their volition here...

classify 'A B C A B C A B C A B C' (message x 4, concatenated) --> pR = 16



Now can **ANYBODY** explain to me why this should be so?


Because my common sense was that 'pR' represents a rating of how much
a message is 'akin to' one of two sides: good versus spam. I *never*
trained 'A B C A B C', but I *did* train 'A B C'. I also trained a
message 'A B D' as *BAD*, so 'A B C' should be found to have 66% in
common with BAD and 100% in common with GOOD. NO MATTER HOW MANY 'A B
C's YOU'VE GOT CONCATENATED IN THAT MESSAGE: duplicate, triplicate,
quadruple, I don't care: the *ratios* will always be 66% bad hits and
100% good hits, which should land it in the 'good' space with a
probability of 33% thereabouts.

Of course, given multiple training rounds and the cpcorr[] balancing
act - which has much bigger impact than I ever thought possible: I
simply CANNOT 'push' 'A B C' into a higher pR by repetitive training!
- the '33% good' should produce slightly different pRs, but WITHOUT
INTERMEDIATE TRAINING, simply feeding the OSB classifier a 'massage *
2' or *3 or *4 dupe, will amplify your pR values by the same factor.
Now I wasn't top-of-class when it came to probability and statistics,
so I may simply have my math screwed up utterly, but that doesn't take
away the fact that this result is still very, ah, 'counterintuitive'
to me.

I expect this perceived 'linearity of scaling' to subside once we go
nearer the 300+ mark as it's a logarithm really, but this amazes me no
end.



So where's my error?








PS: Same phenomenon for bigger messages. I found it because the
mailstripper in mailreaver is _quite_ different from the one in
mailtrainer, resulting in mailtrainer working on basically
'duplicated' messages, resulting in doubled pR scores used to
determine THTTR training regime. Which ended up in mailtrainer NOT
training messages (or stopping the repeat cycle for them) while
mailreaver itself would still incorrectly classify them when using the
same threshold values. Gives me the heebie-jeebies for sure.


-- 
Met vriendelijke groeten / Best regards,

Ger Hobbelt

--------------------------------------------------
web:    http://www.hobbelt.com/
        http://www.hebbut.net/
mail:   [email protected]
mobile: +31-6-11 120 978
--------------------------------------------------

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.