Re: Can some one enlight me regarding the results of megatest?

Bill Yerazunis <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
   From: "Ger Hobbelt" <[email protected]>

   Unfortunately, the diff doesn't tell you the Classifier this way so
   it's kinda hard to see which classifier we're talking about here, but
   given the NN b0rk just before, I'm heavily betting on this being NN
   classify going pearshaped as a result.

Apropos that - I have been using the subterfuge of putting classifier
specific terms into the individual file output lines.  

What about this change: instead of

  CLASSIFY succeeds
or
  CLASSIFY fails

we would have 

  [Classifier_ID] CLASSIFY succeeds
and
  [Classifier_ID] CLASSIFY fails

where [Classifier_ID] is one of "Markov", "OSB", "OSBF", 
I don't think this breaks any known script, and it would
make debugging megatest output much easier.

   feature series for particular classifiers); other count differences
   are due to the fact that I rewrote the VT code a long time ago and it
   will produce more feature hashes for boundary cases. The latest
   mainline is getting much closer to my numbers, though.  Another bit
   Bill & I need to go through when time allows. As to which is 'better'
   then? That depends on the bugfix being actually a fix or a tweak due
   to a misconception. Of course, I am opinionated, but that's not what
   this is about. It about this:

   if you see any differences in feature counts or documents (integer)
   numbers reported, that is enough reason to ask around on the ML if
   it's okay. It probably is (there are types of bugs, which can be
   tolerated some times; those bugs are about accuracy instead of
   operation -- that's also where most of the Blame is shifting around
   for. :-) )

Yeah.  That too.

That's one problem with machine learning code; unlike ( for example )
calculating the value of pi or a symbolic manipulation, you can't say
this is "correct" or "incorrect", all you get is "better" or "worse".


   > 1022,1023c1032,1033
   [...]
   >> #1 (q_test.css):documents: 175, features: 37599,  prob: 1.53e-02, pR: -18.10
   > ------------------------------------

   Bill can probably dream these stats, but I need a diff -u where the
   classifier print lines show up as well (-C5 ? maybe even -C10 ? or
   better yet: -B10 for the diff command?); you might benefit from that
   too.

No, I look it up.  :)  There's usually enough info in the 
actual .log file to figure out what went wrong.

    - Bill

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.