Re: Twitter
Chris Babcock <[email protected]> Thu, 27 Sep 2012 15:36:02 -0700
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <CAFpQuatGWOu4iw4vv2Z-ZjS18o8in_yoyxGyjwnsB4fD6VDTbQ@mail.gmail.com> |
On Thu, Sep 27, 2012 at 2:41 PM, Chris Babcock <[email protected]> wrote: > On Wed, Sep 19, 2012 at 5:44 AM, <[email protected]> wrote: >> Chris Babcock <[email protected]> writes: >>> I tried correlate on the bios only. It was predictably a failure. Then >>> I found that I could request an arbitrary number of tweets with the >>> bio. The default of 30 tweets trained the correlator almost >>> satisfactorily, but I decided to try hyperspace. >>> >>> I'm getting reasonable results with the hyperspace classifier using >>> TOE, but I'm wondering how it would do with SSTTT. Are there any good >>> training examples that someone could put up on Pastebin? I tried >>> Google, but all I got was mailing list articles. The one post >>> containing code was too complex to serve as a general example. >>> >>> Here's my TOE code: >>> >>> http://pastebin.com/jCDshUn5 >>> >>> I'm probably just one loop away from SSTTT... >> >> Less than that, actually. Right now, if I read it correctly out of >> pastebin, the built-in test in CLASSIFY is thresholding at a pR of 0.0 >> and you do this test at lines 35 and 51 >> >> The difference is to grab the output of the classify in line 35 and 51, >> and then check it to see whether the pR listed is sufficiently high >> and in the right direction - if insufficient, then train. >> >> It takes about two or three lines in each place, which you can >> crib from the mailing list articles. >> >> Here's the relevant part of one that looks nice: >> >> { do your classify here, capture the result as :stats: } >> match (:: :pR:) [:stats:] /Best match.*: (.*) \nTotal/ >> { >> { >> eval /:@::*:pR: > 0.5:/ #SSTTT >> { >> [[[ do your learn here... ]]] >> } >> } >> } >> >> >> >> >> (note: we wrap the CLASSIFY in {} so that the built-in FAIL doesn't >> cause a skip if it fails) >> >> - Bill Yerazunis > > I had some difficulty bootstrapping the code with your simple > recommendation. When there's an error, like a missing CSS file, > nothing is placed in the stats variable. My solution is probably more > complicated than it needs to be: > > http://pastebin.com/Bnqwfkb5 > > I'd be grateful if you'd take a glance at my code and make sure my > logic is correct, please. The training code looks at the bio to see > whether I'm following that user and, depending on the result, checks > the tweet history to see whether the user is already correctly > classified. If the classification fails then the case is trained. If > the classification succeeds then we check the pR and train if it's > lower than 0.5. If the stats variable isn't set due to an error then > we assume the pR is 0 and train that case. > > The advantage of training on the tweets without the bio attached is > that I can use the same CSS files for other purposes, like following > hash tags. > > Chris > -- > "Nothing says, 'I love you,' like 200 feet of parachute cord and a cargo net." > - Fred, Scooby-Doo! Mystery Incorporated Here... this is better: http://pastebin.com/wmdq9bLW Chris -- "Nothing says, 'I love you,' like 200 feet of parachute cord and a cargo net." - Fred, Scooby-Doo! Mystery Incorporated ------------------------------------------------------------------------------ Everyone hates slow websites. So do we. Make your web apps faster with AppDynamics Download AppDynamics Lite for free today: http://ad.doubleclick.net/clk;258768047;13503038;j? http://info.appdynamics.com/FreeJavaPerformanceDownload.html