GuessUnicodeCharset

Nerijus Baliunas <[email protected]>
Newsgroups gmane.mail.mahogany.devel
Message-ID <[email protected]>
Hello,

GuessUnicodeCharset() is not working sometimes with Lithuanian
texts, when the first non ASCII character (ð, þ for example) is in
ISO-8859-1 or -2 (in addition to -4 or -13, i.e. Baltic encodings).
Then ISO-8859-1 or -2 is chosen for conversion from UTF-8 and
the text is garbled, as it usually has characters which are not in -1
or -2. I thought of 2 solutions:
* find all non ASCII characters instead of the first only (or at least
3-5), and analyze all of them.
* If found encoding is ISO-8859-1 or -2, continue searching for
non ASCII chars until some other encoding is matched, then stop.

Which is better? Do you have other ideas?

Regards,
Nerijus


-------------------------------------------------------
This SF.Net email is sponsored by: YOU BE THE JUDGE. Be one of 170
Project Admins to receive an Apple iPod Mini FREE for your judgement on
who ports your project to Linux PPC the best. Sponsored by IBM.
Deadline: Sept. 24. Go here: http://sf.net/ppc_contest.php
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.