Re: second attempt at dstttr training
Thomas Michael Hagen <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
thanks for an extremely thorough and helpful answer, ger! i've been struggling with what you bring up as your big point as well, but this is how i understood the description of dstttr in the book (by bill). i have a different version of my classify, where the user is asked for input on every classification, but i was hoping that this one could run for a while before i got to that point, and somehow improve the accuracy. however, as you point out, it is a utopic illusion to think that the classifier is able to improve itself - at least by improving its own confidence in its weaker conclusions! i guess my next version will have to ask for user help whenever it is unsure. i'll need to look more closely at your code before i can begin to incorporate; just one quick question: i'm assuming ist-wert means current value, and soll-wert means future value. is this about right? so, i'll get to work again, well helped, and you'll hear from me next week :) ps: i've had complaints that my mails contain html-encoding. i think i have fixed that now. please let me know if i haven't. On Thu, Mar 5, 2009 at 1:34 PM, Ger Hobbelt <[email protected]> wrote: > On Thu, Mar 5, 2009 at 10:28 AM, Thomas Michael Hagen > <[email protected]> wrote: >> here's my second attempt, after valuable input from ger: >> >> comments and harsh criticism is very welcome :) > > Had to stare at it for a little while... anyway, harsh criticism is my forte ;-) > > What I see happening in your classify code is this, written in > 'Hobbelt'[*] pseudocode: > > ------------------------------------ > function classify(article) > { > crm114.classify(article, CFC_set) -> [pR, class_conc] > if (failed = no) > { > # CRM114 selected a cfc file; the 'winner' > # lands in :class_conc: > if (pR > 30) > { > print "Classify succeeded as $class_conc" > return > } > > # pR wasn't high enough; accumulating statistics: > conc_file <- class_conc ".cfc" > crm114.learn(article, conc_file) > # GHO: are you absolutely positively sure you want that? > > print "Classify succeeded, but threshold low. ..." > > # reclassifying to check improvement: > > crm114.classify(article, CFC_set) -> [pR2, ---] > imp <- pR2 - pR > if (pR > 10 && imp > 3) > { > print "Improvement high enough, ..." > return > } > > # imp wasn't high enough, and new threshold wasn't met > > refute(article, conc_file) > return > } > else > { > # classify failed (category is unc or out) > learn (article) > if (failed = yes) > { > # did learn fail?: > interaction(...) > return > } > else > { > print "Learn succeeded ... after classify FAILED ..." > return > } > } > } > ------------------------------------ > > I find the pseudocode a little easier to grok than crm114 script, but > that's probably just me. > > Anyway, here's the harsh criticism 'n all: > > From a code flow perspective, nothing to nag about. > > From a functional flow perspective, the big one remains: are you > absolutely positively 100% sure you want to crm114.learn() an article > into a category (CFC) based on an automated classification? NOT > especially when that classification didn't make it in the match > olympics already, i.e. is definitely on the shady side. > > > Okay, let me chop that one up for you: > > 1) do you want to train an article into the category GUESTIMATED by an > automaton, when that automaton already indicated, according to your > own standards, that 'it isn't really sure about this'? > Try to rephrase this with human actors (roleplay), where the librarian > isn't exactly sure where to put that item he's got. Exercise question > to finish this roleplay: how do you rate your chances at retrieving > articles managed/archived by a librarian who 'isn't exactly sure about > it'? Hmmmm. > > 2) do you want to train an article into a category GUESTIMATED by an > automaton, period? > A classifier can be viewed as a retard ^H^H^H^H^H^H^H mentally > challenged human down on his luck, which you've put in charge of the > library. We've done our MBA, read our Dilbert so we enjoy the A/P > skyrocketing towards black numbers again, thanks to this initiative, > while we tolerate the slight reduction in performance. That is: the > automaton can do retrieval for us fine. But we ARE aware it WILL make > [a few] mistakes. Which is perfectly acceptable to us. > Now ask ourselves the #2 here: do we want such a process/individual to > steer our archival process as well? One mistake in archival will > produce a PERMANENT error as such a misplaced article will never be > found. Until we invest a lot of $$$ to sift through the entire library > (reorganization). > > > Of course, #2 is disputable as some members may argue that the > classifiers are _smarter_ than the human type of same. Unfortunately, > I don't attend that church. The only thing in my book that makes > automated classifiers perform better than humans is because: > a) the don't fall asleep / don't strike for higher pay > b) at least attempt to be long term consistent (read: are unemotional) > > > Given 1) + 2) I'd rather expect a flow like this: > > ------------------------------------------ > function da_works(article) > { > // ist-wert > classify(article) -> [sure, pR, category] > if (sure) > { > print "assigned category:", category > return category > } > # else: not ok > go_ask_human_where_this_belongs(article) -> soll_wert > > # training SSTTTR/DSTTTR now: > rounds = 5 > # ^^^ example: upper bound to keep things sane in the > # bayes box > do > { > crm114.train(article, soll_wert.as_category) > > # check 'thick' threshold: > classify(article) -> [sure, pR, category] > > # in case of threshold identical to regular > # classify margin: > --> if (sure) exit loop > > # otherwise, when you want to use a different > # 'thick threshold', e.g. like mailtrainer, > # which has/had the regular threshold at > # pR = +10/-5.0 (assym!) while the 'thick' > # threshold is +/- 5.0: > if (pR > 5.0) exit loop > > # and another round! > decrement rounds > } while rounds > 0 > > if (sure) > { > print "trained, and we're sure now" > return category > } > else > { > print "trained, yet the automaton still didn't 'get' it" > return soll_wert.as_category > } > } > > > function classify(article) > { > crm114.classify(article, cfc_set) -> [pR, category] > if (failed) > { > return NOT SURE, 0.0, unknown_category > } > else if (pR > 30) > { > return SURE, pR, category > } > else > { > return NOT SURE, pR, category > } > } > ------------------------------------------ > > > Of course, didn't get to see your own :learn: and other bits, but that > doesn't really matter: the Q1 and Q2 stand and I only wrote the whole > process (da_works()) as pseudocode so you can see what I'd come up > with. In any language / software process. > > The key part is that, as soon as the automaton starts to make unsure > noises, I defer to a human. No training is done, unless the human > tells the machine what to do. > It's the 'I don't know where this goes' librarian roleplay again: > there's the human [supervisor] to tell the librarian where to put that > bit he just got and the librarian is not permitted to just put it > anywhere; if this dilemma would block the flow, than countermeasures > (SOPs aka Standard Operating Procedures) are in place, I'm sure, so > the librarian can go on doing other work he IS sure about, until this > item is resolved by his super. > > Basic library management out in a classifier/software environment. So > we all can find our articles and books in the future. ;-) > > > > In closing: I am sure you've designed your process like this and made > those important design decisions before you started working on that > crm114 script. It's just that sometimes design decisions can get > obscured during the implementation phase. Hence me asking them again > here. > And as that 'sollwert' (and the implicit classified 'istwert') may > have hinted to you already, this is control engineering (it's just > that software engineers don't realize it). Which means there's still > one important question unasked. > > Nice work so far, keep it up! > > > > -- > Met vriendelijke groeten / Best regards, > > Ger Hobbelt > > -------------------------------------------------- > web: http://www.hobbelt.com/ > http://www.hebbut.net/ > mail: [email protected] > mobile: +31-6-11 120 978 > -------------------------------------------------- > ------------------------------------------------------------------------------ Open Source Business Conference (OSBC), March 24-25, 2009, San Francisco, CA -OSBC tackles the biggest issue in open source: Open Sourcing the Enterprise -Strategies to boost innovation and cut costs with open source participation -Receive a $600 discount off the registration fee with the source code: SFAD http://p.sf.net/sfu/XcvMzF8H