Re: I don't get it anymore (if ever): OSB pR increases linearly with message size. Is that by design?
"Ger Hobbelt" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
On Mon, Oct 6, 2008 at 1:40 AM, Bill Yerazunis <[email protected]> wrote: > Well, it gets rid of some of it. It does not get rid of all of it; Yup, that I understood. <unique> will only take out exact matches, so features which correlate fully (as they are identical). At least that removes issues with doubled messages. (Nevertheless, have a look at mailreaver and mailtrainer for non-<unique>-d :clf: : as mailtrainer doesn't use the exact same 'email preprocessing' code as mailreaver, the thick_threshold in there is compared against pR produced by such message=original*2 dupes, thus resulting in <waves hands here> an actual thick_threshold of about half the configured value (try with clf=osb (without unique) and make sure you print all classify output; that was also what confused me regarding that 'missing |': classify pR's in mailtrainer will report much higher pRs than the classify's on the same input message in mailreaver (before any training is done, mind you) > that would require knowledge of the cross-correlation of every phrase > fragment versus every other phrase fragment in the entire language. > > That's why the bit-entropy and compressive classifiers hold such > promise; they are built on the expectation of that correlation. But > experimental reality shows that they are still in the same ballpark > accuracy-wise, and that only for very large test sets. uh-huh. Got it. > I fear this is another error of mine due to (over)simplification of my > thought model and that really bothers me. > > What cpcorr is for is to renormalize the total number of learns done > in the current file versus the total number of learns in all of the > files. The idea is that a feature in a file with only one LEARN done > on it will have each feature weighted three times as heavily as > a file with three LEARNs done in it. > > And yes, it does tend to marginalize repeated learnings of the > same single text. That's supposedly a good thing, in that > it makes one-shot learning more of a reality. Yeah. Well, when I look at my numbers, it surely takes all the air out THTTR, as that is simply TOE with a repeat loop and this behaviour is counter-weighting that learn loop almost completely. > If you want to take it out, go to line 1232++ in crm_osb_bayes.c > and change the cpcorr setting to 1.00 always. Seen that. ;-) If I read the code correctly, osb chi2-based classify runs would also neglect cpcorr[], but that idea is only based on glossing over the code, not through testing that particular config. Besides, chi2 does other things that I might like less, so I'll just have to see... Thanks for taking the time to explain. -- Met vriendelijke groeten / Best regards, Ger Hobbelt -------------------------------------------------- web: http://www.hobbelt.com/ http://www.hebbut.net/ mail: [email protected] mobile: +31-6-11 120 978 -------------------------------------------------- ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/