Re: Serbian Latin 30% finished
"Philippe Verdy" <[email protected]>
| Newsgroups | gmane.network.gnutella.limewire.translate |
|---|---|
| Organization | Ordinateur Personnel |
| Message-ID | <00a101c690b8$5836b3c0$0b01a8c0@harnon> |
Hmmm. Hello Sasha! There does not seem to exist any translation attachment in your email. May be you forgot it? ---- The following section is quite technical, if you can't understand it, don't worry, we'll address this issue. I just want to inform you about some dificulties with the current codification of this language. I propose a solution, that will work in LimeWire, but there may be alternatives. Note that for LimeWire, the language selection does not offer a clean way to select the script. We've code the following codes, assigned according to RFC 3066 and based on iSO 639 language codes: * sr : Serbian (Cyrillic) * hr : Croatian (Latin) * bs : Bosnian (Latin) Unfortunately, Serbian Latin uses the same ISO 639 code "sr" for the language, and the distinction of script (ISO 15924) is still not standardized in RFC 3066. So the only way to support it will be to create a language code that does not conflict with RFC 3066 codes. So we would use the code "srLatn" for Serbian Latin (note the use of a capital in the middle of the code, to separate the language code from the script code). The alternative would be using a dash separator would make "Latn" a country/region code to get "sr_Latn" in the filename extension recognized by the Java class loader. This will work simply because there's no ISO 3166 code using 4 letters codes, and also because these codes are normally fully written with capitals, for example "pt_BR" for Portuguese in brasil). Note that according to RFC 3066, Serbian spoken in Serbia and in Monenegro would have the code "sr_CS" (but this "CS" country code may soon become deprecated due to the independance of Montenegro, unless Serbia chooses to keep the "CS" code). It used to be "sr_YU" when the country was still named "Yougoslavia"; It's impossible for now to determine a reliable country code, so it's best to use a script code. If we code the Serbian latin resource as "sr_Latn", it will be impossible to make a clear distinction of country between Serbian Latin written in Serbia, in Montenegro and in North-East Bosnia, if there are such differences (but Serbian also has a known distinct language variant named "Zlatiborian" spoken in Central Serbia, near the border of Bosnia-Herzegovina, around the city of Zlatibor, this language variant currently has no codification, and was typically written in Cyrillic, but now also in Latin... How to encode these variants? "srLatn_CS_zl" for the Latin version, and "sr_CS_zl" for the Cyrillic version... Hmmm, unfortunately there's no clean way to codify all these 4 fields, and two fields need to be merged. That's why I tend to think that merging the language code and script code into "srLatn" seems the best solution for now, but there will be no inheritance of missing resources between Serbian Latin and Serbian Cyrillic! Anyway, such inheritance would be problematic because mixing scripts is a bad idea for a user interface, so this should not be a problem. So you can name your file like this: "MessagesBundle_srLatn.UTF-8.txt" (for a plain-text file format saved in Unicode UTF-8), or you can edit it using Windows codepage 1250 (based on ISO 8859-2, i.e. Latin 2 for Central European languages) and we'll make the conversion to UTF-8 (if so, name your file "MessagesBundle_srLatn.txt") >From this UTF-8 file, we'll generate automatically the actual properties file used in LimeWire "MessagesBundle_srLatn.properties"which will represent specially the non ASCII characters into hexadecimal sequences. There will be NO automatic selection of the language, because there's still no correct support to map the "srLatn" combined code from OS-level locale identifiers: the selection of language may have to be performed manually at LimeWire installation or at runtime. --- For now, Serbian Latin was best supported by selecting either Croatian or Bosnian which are written with the same Latin script, despite there are known differences of vocabulary and grammatical rules (and I don't know if there's a way to create a version that will be acceptable by both speakers communities.) I suggest then to start translating by taking the Croatian or Bosnian versions as a model. But consider also looking into the existing Serbian Cyrillic version (for the choice of vocabulary): this will be helpful for you to get a consistant translation and may give you ideas for some concepts difficult to translate. We're waiting for your file and comments. Thanks. Philippe. ----- Original Message ----- From: Sasha Erceg To: LimeWire translation Sent: Thursday, June 15, 2006 1:24 AM Subject: [trans] Serbian Latin 30% finished Hi guys, I just want to announce that Serbian (Latin) translation is about 30% finished, and I expect it will be done in a next few days. I am doing this alone, because there was no Serbian Latin language pack until now, and current Bosnian & Serbian version needs many improvements to be really usefull. I will need help with conversion from plain text (Notepad) to Unicode UTF-8, please contact me if you know how to do this. I found some "explanation" but I do not have a JDK and native2ascii to do that. Sasha, Belgrade, Serbia. _______________________________________________ translate mailing list [email protected] http://lists.limewire.org/mailman/listinfo/translate