[poedit-users] Problem with utf-8 encoding
David Bolen <db3l-H7rzD4QtxWhWk0Htik3J/[email protected]>
| Newsgroups | gmane.editors.poedit.user |
|---|---|
| Organization | Fitlinxx, Inc. - Stamford, CT |
| Message-ID | <[email protected]> |
I'm running into problems with poEdit versions greater than 1.3.1
(running on a Windows 2K SP3 system) trying to read in PO (or POT)
files marked to use a utf-8 encoding, which were being generated by an
automatic tool I want to use. I simplified things to a pure metadata
entry which is just in ASCII and it still has the problem.
Attempting to open an existing POT/PO (or use the New Catalog from POT
file... functionality) yields errors on each line of the input file
indicating that:
"Line # of file 'XXXX' is corrupted (not valid utf-8 data)."
where XXXX is the filename, and # is all lines from 0 to n-1 of the
file. However, in each case the line quoted is pure ASCII, which
should also be valid UTF-8, so I'm at a loss to determine what is
corrupted about it.
It fails so completely I can only assume that I'm doing something
stupid, or being dense. In searching the list I've seen some prior
posts for similar items, but none with a resolution that would seem to
explain failing to handle even a pure-ASCII case. One reference
commented about the lack of a BOM, but to my understanding a BOM is
not required for UTF-8 files (although I did try manually adding one
to no effect - a UTF-8 bom as the 8-bit characters 0xEF,0xBB,0xBF).
Another person had a reference to upgrading their Win2K system to SP4
and the problem going away - I happen to be on SP3 but can't do a full
system upgrade at the moment to test that theory.
I also find that attempting to create a brand new catalog from within
poEdit itself (and indicating that it should be UTF-8) generates errors
on any attempt to save the catalog of (filename shortened to <file>):
19:38:10: msgfmt: <file>.po: warning: PO file header missing or invalid
19:38:10: warning: charset conversion will not work
19:38:10: msgfmt: found 1 fatal error
and the PO file itself is always 0 bytes long at this point.
I'm able to run the msgfmt binary from the poEdit installation on the
same PO files (although the empty one results in no output) without
these errors showing up, but I'm not that familiar with the toolchain,
so assume there's something environmentally different when run under
poEdit.
As a sample, here's one of the POT files (automatically created) that
I'm trying to work with. I'm including it inline as it contains no
special characters, although in the filesystem it does use CRLF line
endings (e.g., it was written natively under Windows).
- - - - -
msgid ""
msgstr ""
"POT-Creation-Date: Tue Jan 17 18:32:40 2006\n"
"PO-Revision-Date: YEAR-MO-DA HO:MI+ZONE\n"
"Last-Translator: FULL NAME <EMAIL@ADDRESS>\n"
"Language-Team: Silva i18n team <[email protected]>\n"
"MIME-Version: 1.0\n"
"Content-Type: text/plain; charset=utf-8\n"
"Content-Transfer-Encoding: 8bit\n"
"Generated-By: zope/app/translation_files/extract.py\n"
- - - - -
I'm willing to find out I'm just making a fundamental error about how
the charset in a POT/PO file works, but shouldn't the above be able to
load properly? At least I don't have a problem reading the same file
using UTF-8 decoders in various other tools/languages.
It seems that whatever changed happened between 1.3.1 and 1.3.2 since
1.3.1 will at least load the above without error (and create new PO
files successfully) whereas 1.3.2, 1.3.3 and 1.3.4 all yield the above
errors.
Thanks for any help.
-- David
-------------------------------------------------------
This SF.net email is sponsored by: Splunk Inc. Do you grep through log files
for problems? Stop! Download the new AJAX search engine that makes
searching your log files as easy as surfing the web. DOWNLOAD SPLUNK!
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=103432&bid=230486&dat=121642