Re: Can some one enlight me regarding the results of megatest?
Bill Yerazunis <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
From: "Ger Hobbelt" <[email protected]> Unfortunately, the diff doesn't tell you the Classifier this way so it's kinda hard to see which classifier we're talking about here, but given the NN b0rk just before, I'm heavily betting on this being NN classify going pearshaped as a result. Apropos that - I have been using the subterfuge of putting classifier specific terms into the individual file output lines. What about this change: instead of CLASSIFY succeeds or CLASSIFY fails we would have [Classifier_ID] CLASSIFY succeeds and [Classifier_ID] CLASSIFY fails where [Classifier_ID] is one of "Markov", "OSB", "OSBF", I don't think this breaks any known script, and it would make debugging megatest output much easier. feature series for particular classifiers); other count differences are due to the fact that I rewrote the VT code a long time ago and it will produce more feature hashes for boundary cases. The latest mainline is getting much closer to my numbers, though. Another bit Bill & I need to go through when time allows. As to which is 'better' then? That depends on the bugfix being actually a fix or a tweak due to a misconception. Of course, I am opinionated, but that's not what this is about. It about this: if you see any differences in feature counts or documents (integer) numbers reported, that is enough reason to ask around on the ML if it's okay. It probably is (there are types of bugs, which can be tolerated some times; those bugs are about accuracy instead of operation -- that's also where most of the Blame is shifting around for. :-) ) Yeah. That too. That's one problem with machine learning code; unlike ( for example ) calculating the value of pi or a symbolic manipulation, you can't say this is "correct" or "incorrect", all you get is "better" or "worse". > 1022,1023c1032,1033 [...] >> #1 (q_test.css):documents: 175, features: 37599, prob: 1.53e-02, pR: -18.10 > ------------------------------------ Bill can probably dream these stats, but I need a diff -u where the classifier print lines show up as well (-C5 ? maybe even -C10 ? or better yet: -B10 for the diff command?); you might benefit from that too. No, I look it up. :) There's usually enough info in the actual .log file to figure out what went wrong. - Bill ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/