Re: Help needed with a "wedged" CRM114 installation
[email protected] Thu, 22 Mar 2012 10:32:52 -0400
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
Martin Lucina <[email protected]> writes: > Hi Bill, > > Should I also upgrade to the latest codebase? If so, which one? I'm > currently using 20090807-BlameThorstenAndJenny (TRE 0.7.5 (LGPL)). Only if you want the SVM classifier. The base classifer (Markovian) is unchanged. > Assuming the various different Korean character sets all have what renders > as "space" on my screen in the same place as "ascii space", which is a > pretty safe assumption, then yes. Hangul is alphabet-based rather than > glyph-based, the only reason it looks superficially similar to Japanese is > that they use the trick of combining multiple charaters to form a single > glyph. Go into EMACS or another byte-accurate editor (or a hex editor) and check! It can be quite important. (i.e. is there really a hex 0x20 between each word? Or does the Hangul representation you get have a "space" in the glyphset that is NOT 0x20, and it gets used because that way you don't have to change glyphsets twice for every word. > Alternatively, given that the volume of mail I get in Korean is quite low, > is there a way to tell the system to just pass through mail from certain > senders, completely ignoring it for training purposes? Yes, there is. That's what the whitelist is for. Although one might regard it as "cheating", I am a firm believer in "whatever gets you through the night." - Bill Yerazunis ------------------------------------------------------------------------------ This SF email is sponsosred by: Try Windows Azure free for 90 days Click Here http://p.sf.net/sfu/sfd2d-msazure