Re: (no subject)
"Eric S. Johansson" <[email protected]> Tue, 12 May 2009 10:34:08 -0400
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
Bill Yerazunis wrote: > It seems like the optimal ratio between the smaller corpora (like the > SA corpus and a couple other 4000ish element corpora) and a > near-infinite corpus like the 100,000+ element TREC corpus is about > 2:1, so if your threshold on a 4000 element corpus is 10, then use 5 > for the infinite-corpus "I'm gonna run this on my mail for years" > situation. based on your previous descriptions, I would say that the thickness changes as the training history grows. Also, don't forget that in the real world (TM), the "corpus", shifts based on your last X days of training and implies a changing thickness. > The easiest way is to use the program "tenfold_validate.crm", with > various values of --clf="Your_Classifier_Flags_Here" and > --thickness=0.5 (or whatever), and --results=Results_File_Here.txt. > You do need to create an "index" file, which is the type and location > of each file in your corpus- it looks like this: ... > The tenfold_validate program will then run a ten-fold validation > (basically, divide the incoming knowns into ten roughly equal chunks, > train on nine, test the tenth, and repeat until each of the ten chunks > is tested once. Then, scan the results file with "grep" and "wc" > to count errors: > > grep WRONG my_results.txt | wc > > For reasonably small corpora (4000 or so, like the SA corpus), this > will take a few minutes per --thickness value. in my world, I'm seeing classification times on the order of roughly a half a second per message ranging up to four or five seconds as the CSS database slowly corrupts itself. I have typically 5000 to 7000 messages per five-day window which means I'm looking at roughly 83 minutes to process all these messages. Obviously, I will run some tests to validate this model. I seem to remember with a brand-new CSS file, I'm looking at closer to 15 to 30 minutes. Again, from my experience, it sounds like I will need to regenerate CSS files and recalculate thickness about once every 14 days. I could probably get by with stretching out the time between reconstructions to something like 30 or even 45 days. this shifting threshold problem might explain the creeping crud phenomenon I'm seeing within my training window (the thickness). by creeping crud, I referred to the increasing intrusion of negative scored messages into the thickness wide region around zero. They tend to pile up just short of the negative boundary (i.e. minus five in this case). It's quite consistent, quite repeatable and, quite annoying. I'm going to have to play with the 10 full training a bit. ---eric ------------------------------------------------------------------------------ The NEW KODAK i700 Series Scanners deliver under ANY circumstances! Your production scanning environment may not be a perfect world - but thanks to Kodak, there's a perfect scanner to get the job done! With the NEW KODAK i700 Series Scanner you'll get full speed at 300 dpi even with all image processing features enabled. http://p.sf.net/sfu/kodak-com