Re: Request for Comment: Callable CRM114 Classifiers (libcrm114)
"Trever L. Adams" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
Bill Yerazunis wrote: > From: "Trever L. Adams" <[email protected]> > > If possible, please, please, please allow a command line or library call > option to use UTF-8. > > I only need regex matching and breaking to be UTF-8 (breaking meaning > knowing when it is a space or punctuation that would normally cause word > breaks, not just assuming that ascii is the way to go). > > Do you mean in the regex engine? > > If so, then it's not something libcrm114 itself handles. We accept > whatever the regex engine itself accepts (and that includes locality > flags). > > Or... I am missing your point? > > -Bill Yerazunis > Hello Bill, Yes, you are missing the point. It isn't just regex where UTF-8 would be good. The various "classifiers" all have their tokenization rules. I am not asking that they be made completely generic. The problem is that they assume various 8 bit sequences are ALWAYS a given mark. I would like to see a command line/library call that would enable UTF-8 tokenization. If everything is in UTF-8 (should be if they are using the call/command line), then it doesn't matter what the given text string means, it just matters where it is broken. Obviously, what I say only applies to those designed for text. The various binary learners would not or should not be affected. Trever ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/