Re: score graph
"Ger Hobbelt" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
Eric, I don't have any good explanation for this behaviour, apart from some handwaving regarding the way OSB determines the pR from the number of matched features, and since no-one else seems able to shed more light on this issue, I'll take you up on your challenge stated over private email: if you can give me an environment to test this in, I might be able to find the reason why. When I had time for crm114 I've been working with mailtrainer.crm the last few weeks and I ran into some surprises there. It's probably yours truly the Nutty Professor, but I'd like to have a peek at yours, if you don't mind. Alternatively, you might want to have a look at the feature hitcounts & totalcounts produced by crm114 - that's what I intend to look at at least - as those are used to calculate the pR value. It might be that: - as these hitcounts are 'weighted' (i.e. reduced) according to the total number of features stored in the spam and good CSS databases respectively, the resulting pR may be adversily affected due to this weighting. (spam does not have much in common, and there's a lot of it, so spam hitcounts will be rather low. When weighted linearly against CSS database size, those numbers will be reduced further: that's my guess #1. - mailtrainer seems to act up overhere and refused to retrain 'mails' more than twice. To force spam charactistics to 'pop out' despite the low hitcount and large 'totalcount', you can learn a mail several times. This helps because a single feature match is not counted as '1' but as a 1..N number, depending on how many times that particular feature was learned. In other words: once you've learned an email, does a recheck with classify actually report it as spam/good as desired? (I.e. does the learn process actually push the document beyond the decision threshold? It might not, when mailtrainer is not configured 'properly' or acts up otherwise.) Another thought - but that would mean the pR calculus in OSB mode will have to be changed: my spam/good collection tends to have a huge load of spams with a LOT of features, with only VERY FEW of those common between the 'spams'. Meanwhile, that 'spam' carries a lot of obfuscation, meaning features which also appear in the 'good' messages. Since the pR weighting for both good and spam are linear against the total number of learned features, learning new spams 'degrades' the pR for previous ones. Given the fact that spam these days has a lot of fake 'good' and only a bit of 'spam' in it feature-wise, I would suggest the spam hitcounts should NOT be weighted linearly against the total number learned, but weighting spam should maybe be a power curve. In other words: if I have a few spam feature hits, say, 8 or more, I would like the pR to be STRONGLY influenced by such a number, while a count of 1 or 2 spam hits should be considered minimal. This can be done by weighting the spam hitcount using a power curve (hitcount ** power) so medium spam hitcounts weight in stronger than a bunch of 'good'. I haven't tested this idea for lack of time and most of all: a good evaluation setup. Anyway, Eric, I have got a new 64-bit Windows box operational now (thanks to blown fuzes, fizzing power supplies and all that: Finally I have some x64 Windows hardware - it was waiting for over a year already, but now was the time. I'll install VMware Player on it and see what happens. Cheers, Ger On Fri, May 16, 2008 at 5:18 AM, Eric S. Johansson <[email protected]> wrote: > here is a graph of the numbers I dumped on the list earlier this week. the > light blue area is the "uncertain" region as defined by 2penny blue. the > training cutoffs are +-3 as shown in the code also mailed earlier. > > as I expect, I get a peak on the green side up around 20 with 200 messages. > the red side is very unpleasant. it peaks right in the prime viewing > region at just over 800 messages. why? help please. > > ------------------------------------------------------------------------- > This SF.net email is sponsored by: Microsoft > Defy all challenges. Microsoft(R) Visual Studio 2008. > http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/ > _______________________________________________ > Crm114-general mailing list > [email protected] > https://lists.sourceforge.net/lists/listinfo/crm114-general > > -- Met vriendelijke groeten / Best regards, Ger Hobbelt -------------------------------------------------- web: http://www.hobbelt.com/ http://www.hebbut.net/ mail: [email protected] mobile: +31-6-11 120 978 -------------------------------------------------- ------------------------------------------------------------------------- This SF.net email is sponsored by: Microsoft Defy all challenges. Microsoft(R) Visual Studio 2008. http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/