Re: Universal Charset Detector

Simon Montagu <[email protected]> Fri, 24 Jun 2005 01:23:16 +0300
Newsgroups gmane.comp.mozilla.internationalization
Organization Another Netscape Collabra Server User
Message-ID <[email protected]>
Tay, William wrote:
> Hi,
> 
>  
> 
> I read about the composite approach to language/encoding detection at 
> http://www.mozilla.org/projects/intl/UniversalCharsetDetection.html and 
> would like to know if the component (Universal Charset Detector) 
> supports the detection of Bi-Di encodings/languages (e.g. Hebrew and 
> Arabic).
> 
>  

There is a work-in-progress patch to support Hebrew at 
https://bugzilla.mozilla.org/show_bug.cgi?id=86999

Doing the same for Arabic would be feasible, given a large quantity of 
text to use in building the language model. 
(https://bugzilla.mozilla.org/show_bug.cgi?id=265030)

> Is there a JAVA version for the component?
> 

See http://www.i18nfaq.com/chardet.html