Re: Sorting many small data sets
Paolo <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <20081215090244.GA26964@localhost> |
On Sun, Dec 14, 2008 at 02:00:37PM -0700, Chris Babcock wrote: ... > Each record is only a about a dozen fields (~250 bytes each). There are > about 15,000 records distributed over a dozen or so flat file ... > for this Pr or better.'" Intuitively, hyperspace seems to be the most > likely candidate because of the speed of learning and the small disk > foot print. Each corpus will be very small, the documents small and > fairly similar to begin with, and there will be 15,000+ classification > files floating around at the end of it, which would be nice to keep for > reducing future duplications. Is there a reasonable chance of this > approach working given the obvious weaknesses in the quality of the > data? choice of HS seems fine - works better with small amount of data, _growing_ cssfile size - I'd rather say the other way around: you've got a nice multiclass case, pls go ahead, try and report back how it goes ;) -- paolo ------------------------------------------------------------------------------ SF.Net email is Sponsored by MIX09, March 18-20, 2009 in Las Vegas, Nevada. The future of the web can't happen without you. Join us at MIX09 to help pave the way to the Next Web now. Learn more and register at http://ad.doubleclick.net/clk;208669438;13503038;i?http://2009.visitmix.com/