Re: Universal Charset Detector
Simon Montagu <[email protected]> Fri, 24 Jun 2005 01:23:16 +0300
| Newsgroups | gmane.comp.mozilla.internationalization |
|---|---|
| Organization | Another Netscape Collabra Server User |
| Message-ID | <[email protected]> |
Tay, William wrote: > Hi, > > > > I read about the composite approach to language/encoding detection at > http://www.mozilla.org/projects/intl/UniversalCharsetDetection.html and > would like to know if the component (Universal Charset Detector) > supports the detection of Bi-Di encodings/languages (e.g. Hebrew and > Arabic). > > There is a work-in-progress patch to support Hebrew at https://bugzilla.mozilla.org/show_bug.cgi?id=86999 Doing the same for Arabic would be feasible, given a large quantity of text to use in building the language model. (https://bugzilla.mozilla.org/show_bug.cgi?id=265030) > Is there a JAVA version for the component? > See http://www.i18nfaq.com/chardet.html