Re: Why is my pR always the same?
Bill Yerazunis <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
well, only different printout and rounding.
> #0 (GOOD): features: 438256, hits: 1498, prob: 2.86e-01, pR: -0.40
> #1 (BAD): features: 446250, hits: 1687, prob: 7.14e-01, pR: 0.40
> What's the real story here?
dunno ... might be perfectly right, depends on training sets and text you
have feeded it here. We can only see it's a short one (704 features).
Actual examples would be useful (though likely not avail for a public
ML). But for testing any good/bad short texts collection would be fine.
I realized something else on this: if it's IM rather than email,
a lot of the conventions are different.
You _might_ want to try <unigram> instead of <osb unique> and see if that
gives you better results.
Also try <hyperspace> and <hyperspace unigram>.
Note that both of these formats are _incompatible_ with your current
training set (and with each other) so before trying either save your
statistics files first to another directory and remove them from your
work directory (fresh new ones will be created by the first LEARN
you execute).
- Bill Yerazunis
-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/