Re: Why is my pR always the same?
"Robert Swirsky-Warner" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
> Hmmmm... yeah, that's not right. Your suspicious nature is correct.
>
> (although 304 as a pR is "special" in the sense that it's the maximum
> pR you can get in an OSB situation.)
>
> Let's start from the beginning:
>
> 1) What OS are you running?
> 2) What version of CRM114 are you running?
> 3) Are you using mailreaver, mailfilter, minifilter, or something else?
>
>
Thanks for the reply, Bill!
Here are my answers
1. Mac OS 10.5
swirsky-tppllc:~ swirsky$ uname -a
Darwin swirsky-tppllc.local 9.5.0 Darwin Kernel Version 9.5.0: Wed Sep 3
11:29:43 PDT 2008; root:xnu-1228.7.58~1/RELEASE_I386 i386*
*2*.*swirsky-tppllc:~ swirsky$ crm -v
This is CRM114, version 20070810-BlameTheSegfault (0.6.8)
(It seems that the Mac came with one pre-installed, but I downloaded and
build what I presume to be the latest version)
3. No*
*
We're trying to use CRM114 to distinguish between "good" chat conversations
and "bad" chat conversations. Since I had very promising results training a
Paul Graham style Bayesean system to distinguish these results, I thought
I'd give CRM 114 a whirl.
I took the same "good" and "bad" sets that I trained the Paul Graham/Baysean
system with and trained CRM 114, like this:
*{
learn <osb unique microgroom> (:*:_arg2:)
}
*
(with arguments like)
crm learn.crm BAD < bad.txt
Then I tried to see how it did, like this:
#
{
isolate (:stats:)
classify <osb unique microgroom> (GOOD BAD) (:stats:)
output /:*:stats:/
}
and the result lines always start with a line that says
CLASSIFY succeeds; success probability: 1.0000 pR: 304.6527
Best match to file #1 (BAD) prob: 0.7139 pR: 0.3972
Total features in input file: 704
#0 (GOOD): features: 438256, hits: 1498, prob: 2.86e-01, pR: -0.40
#1 (BAD): features: 446250, hits: 1687, prob: 7.14e-01, pR: 0.40
Note that the "Best match" line has a different pR. From TFM, it seems that
a good pR should be 10 or more units away from zero. The ones on the "Best
match" line are always fractions.
Curiously, in the document CRM114_REVEALED, his examples also show pR always
at 304.6527 For example, from page 175 of CRM114_Revealed
CLASSIFY succeeds; success probability: 1.0000 pR: 304.6527
Best match to file #1 (BAD) prob: 0.7139 pR: 0.3972
Total features in input file: 704
#0 (GOOD): features: 438256, hits: 1498, prob: 2.86e-01, pR: -0.40
#1 (BAD): features: 446250, hits: 1687, prob: 7.14e-01, pR: 0.40
What's the real story here?
*
*
-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
_______________________________________________
Crm114-general mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/crm114-general