Re: Can some one enlight me regarding the results of megatest?
"Steve" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
-------- Original-Nachricht -------- > Datum: Thu, 4 Sep 2008 08:22:24 -0400 > Von: Bill Yerazunis <[email protected]> > An: "Ger Hobbelt" <[email protected]> > CC: [email protected], [email protected] > Betreff: Re: [Crm114-general] Can some one enlight me regarding the results of megatest? > > From: "Ger Hobbelt" <[email protected]> > > Unfortunately, the diff doesn't tell you the Classifier this way so > it's kinda hard to see which classifier we're talking about here, but > given the NN b0rk just before, I'm heavily betting on this being NN > classify going pearshaped as a result. > > Apropos that - I have been using the subterfuge of putting classifier > specific terms into the individual file output lines. > > What about this change: instead of > > CLASSIFY succeeds > or > CLASSIFY fails > > we would have > > [Classifier_ID] CLASSIFY succeeds > and > [Classifier_ID] CLASSIFY fails > > where [Classifier_ID] is one of "Markov", "OSB", "OSBF", > I don't think this breaks any known script, and it would > make debugging megatest output much easier. > I would welcome that. Would be a very good addition. > feature series for particular classifiers); other count differences > are due to the fact that I rewrote the VT code a long time ago and it > will produce more feature hashes for boundary cases. The latest > mainline is getting much closer to my numbers, though. Another bit > Bill & I need to go through when time allows. As to which is 'better' > then? That depends on the bugfix being actually a fix or a tweak due > to a misconception. Of course, I am opinionated, but that's not what > this is about. It about this: > > if you see any differences in feature counts or documents (integer) > numbers reported, that is enough reason to ask around on the ML if > it's okay. It probably is (there are types of bugs, which can be > tolerated some times; those bugs are about accuracy instead of > operation -- that's also where most of the Blame is shifting around > for. :-) ) > > Yeah. That too. > > That's one problem with machine learning code; unlike ( for example ) > calculating the value of pi or a symbolic manipulation, you can't say > this is "correct" or "incorrect", all you get is "better" or "worse". > > > > 1022,1023c1032,1033 > [...] > >> #1 (q_test.css):documents: 175, features: 37599, prob: 1.53e-02, > pR: -18.10 > > ------------------------------------ > > Bill can probably dream these stats, but I need a diff -u where the > classifier print lines show up as well (-C5 ? maybe even -C10 ? or > better yet: -B10 for the diff command?); you might benefit from that > too. > > No, I look it up. :) There's usually enough info in the > actual .log file to figure out what went wrong. > > - Bill -- GMX Kostenlose Spiele: Einfach online spielen und Spaß haben mit Pastry Passion! http://games.entertainment.gmx.net/de/entertainment/games/free/puzzle/6169196 ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/