Re: I don't get it anymore (if ever): OSB pR increases linearly with message size. Is that by design?

"Ger Hobbelt" <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
On Mon, Oct 6, 2008 at 1:40 AM, Bill Yerazunis <[email protected]> wrote:
> Well, it gets rid of some of it.  It does not get rid of all of it;

Yup, that I understood. <unique> will only take out exact matches, so
features which correlate fully (as they are identical).

At least that removes issues with doubled messages. (Nevertheless,
have a look at mailreaver and mailtrainer for non-<unique>-d :clf: :
as mailtrainer doesn't use the exact same 'email preprocessing' code
as mailreaver, the thick_threshold in there is compared against pR
produced by such message=original*2 dupes, thus resulting in <waves
hands here> an actual thick_threshold of about half the configured
value (try with clf=osb (without unique) and make sure you print all
classify output; that was also what confused me regarding that
'missing |': classify pR's in mailtrainer will report much higher pRs
than the classify's on the same input message in mailreaver (before
any training is done, mind you)



> that would require knowledge of the cross-correlation of every phrase
> fragment versus every other phrase fragment in the entire language.
>
> That's why the bit-entropy and compressive classifiers hold such
> promise; they are built on the expectation of that correlation.  But
> experimental reality shows that they are still in the same ballpark
> accuracy-wise, and that only for very large test sets.

uh-huh. Got it.


>   I fear this is another error of mine due to (over)simplification of my
>   thought model and that really bothers me.
>
> What cpcorr is for is to renormalize the total number of learns done
> in the current file versus the total number of learns in all of the
> files.  The idea is that a feature in a file with only one LEARN done
> on it will have each feature weighted three times as heavily as
> a file with three LEARNs done in it.
>
> And yes, it does tend to marginalize repeated learnings of the
> same single text.  That's supposedly a good thing, in that
> it makes one-shot learning more of a reality.

Yeah. Well, when I look at my numbers, it surely takes all the air out
THTTR, as that is simply TOE with a repeat loop and this behaviour is
counter-weighting that learn loop almost completely.


> If you want to take it out, go to line 1232++ in crm_osb_bayes.c
> and change the cpcorr setting to 1.00 always.

Seen that. ;-) If I read the code correctly, osb chi2-based classify
runs would also neglect cpcorr[], but that idea is only based on
glossing over the code, not through testing that particular config.
Besides, chi2 does other things that I might like less, so I'll just
have to see...


Thanks for taking the time to explain.




-- 
Met vriendelijke groeten / Best regards,

Ger Hobbelt

--------------------------------------------------
web:    http://www.hobbelt.com/
        http://www.hebbut.net/
mail:   [email protected]
mobile: +31-6-11 120 978
--------------------------------------------------

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.