Re: Global and national e-mail address
"Mark Davis" <[email protected]> Wed, 10 Dec 2003 07:33:08 -0800
| Newsgroups | gmane.ietf.imaa |
|---|---|
| Message-ID | <002601c3bf32$e9a65d10$7900a8c0@DAVIS1> |
It is easy to fall into a trap of being Eurocentric: to overestimate the ease of Latin and underestimate the difficulty of transliteration. For many languages there are no good transliteration standards; or rather, there are many conflicting ones. And many of these are transcriptions, and not transliterations. (The difference is that transliteration to Latin is reversable; one can recover precisely the original text; transcription is not reversable -- but is more pronouncable). For example, here is are some sample transliterations (from http://oss.software.ibm.com/cgi-bin/icu/tr). If you asked an average user of each of these languages to do a transliteration, even of those who know Latin, the odds of their coming up with exactly these results are very small. По своей природе компьютеры могут работать лишь с числами. И для того, чтобы они могли хранить в памяти буквы или другие символы, каждому такому символу должно быть поставлено в соответствие число. => Po svoej prirode kompʹûtery mogut rabotatʹ lišʹ s čislami. I dlâ togo, čtoby oni mogli hranitʹ v pamâti bukvy ili drugie simvoly, každomu takomu simvolu dolžno bytʹ postavleno v sootvetstvie čislo. (ISO 9) Οι ηλεκτρονικοί υπολογιστές, σε τελική ανάλυση, χειρίζονται απλώς αριθμούς. Αποθηκεύουν γράμματα και άλλους χαρακτήρες αντιστοιχώντας στο καθένα τους από έναν αριθμό (ονομάζουμε μία τέτοια αντιστοιχία κωδικοσελίδα). => Oi i̱lektronikoí ypologistés, se telikí̱ análysi̱, cheirízontai apló̱s arithmoús. Apothi̱kév̱oun grámmata kai állous charaktí̱res antistoichó̱ntas sto kathéna tous apó énan arithmó (onomázoume mía tétoia antistoichía ko̱dikoselída). (ISO 843) 기본적으로 컴퓨터는 숫자만 처리합니다. 글자나 다른 문자에도 숫자를 지정하여 저장합니다. => gibonjeog'euro keompyuteoneun susjaman ceorihabnida. geuljana dareun munja'edo susjareul jijeonghayeo jeojanghabnida. (Korean Ministry of Culture & Tourism Transliteration regulations with the clause 8 variant) कम्प्यूटर, मूल रूप से, नंबरों से सम्बंध रखते हैं। ये प्रत्येक अक्षर और वर्ण के लिए एक नंबर निर्धारित करके अक्षर और वर्ण संग्रहित करते हैं। => kampyūṭara, mūla rūpa sē, nambarōṁ sē sambandha rakhatē haiṁ. yē pratyēka akṣara aura varṇa kē li'ē ēka nambara nirdhārita karakē akṣara aura varṇa saṅgrahita karatē haiṁ. (ISO 15919) And there is no known transliteration method for Chinese or Japanese -- there are (many) transcription standards, but none that preserve the original characters. John's point about IPA is particularly apt. Think about the example: if an English user had to use IPA for web addresses, how would he do it? For "bort.com" would he use "bɔrt.kɑm", "bɔːt.kɒm", "bɔrʔ.kɑm", ... It would be very difficult to predict exactly what the spelling in IPA would be, even if he were conversant with IPA. The same problem is faced by someone using a language normally written in a non-Latin script when trying to transliterate into Latin. Mark __________________________________ http://www.macchiato.com ► शिष्यादिच्छेत्पराजयम् ◄ ----- Original Message ----- From: "John C Klensin" <[email protected]> To: "Dan Oscarsson" <[email protected]>; <[email protected]> Sent: Wed, 2003 Dec 10 06:11 Subject: Re: Global and national e-mail address > > Dan, > > Three observations (short this time)... > > (i) Computer geeks and their possible preferences aside, people > tend to not like transliterations (writing of a name that would > normally be written in one character set in the characters of > another). Whether they like "better" transliterations more than > "worse" transliterations is a cultural issue. > > (ii) To reasonably transliterate names and languages, one needs > not only a collection of the right phonemes, but an appropriate > and accurate notation for tones. You can't get those out of a > small extension to Latin letters. I'm told that one can get a > reasonable approximation of all of the relevant phonemes and > tones with IPA, but IPA not only uses some characters that are > distinctly non-Latin-based, but also uses a rather complex > collection of combining diacriticals. And, at least unless one > is a professional phonologist, learning IPA and how to use it > accurately is _hard_ (having had people attempt to teach it to > me twice, once when I was young enough to learn these things). > For some hints in a reference that is easily accessible to most > of us, see the discussion of IPA Characters in the Unicode > definition (3.0 or 4.0, take your pick). > > (iii) If one wants even an approximation to accurate > transliteration, the symbol-overloading in Latin scripts is bad > news. E.g., the sound of "ö" (o with diaresis, U+00F6) is > different in, e.g., Swedish and German. I.e., they are > different characters, even if they look the same and even if > Unicode "unified" them. If one is trying to transliterate, > e.g., Arabic into Roman characters, does one pick a character on > the basis of the Swedish phoneme or the German ones? (Hint, as > soon as you start down that path, you end up sliding toward IPA.) > > It appears to me that you are proposing a very Euro-centric view > of things, and it won't work all that well even for Europe. > > regards, > john > > > --On Wednesday, 10 December, 2003 10:32 +0100 Dan Oscarsson > <[email protected]> wrote: > > > > > I will here comment several comments frpm John C Klensin, Adan > > M Costello and others. > > > > With this topic I did not want to talk about replacing ASCII > > with ISO 8859-1 or about what character encoding to use. > > Instead I wanted to discuss what names could be suitable to > > use as a global fallback name. > > > > To be able to write a name you need, at least, to be able to > > have letters so you can write all phonemes used. ASCII only > > contains 26 letters and they cannot represent all phonemes > > very well. For example, Swedish have three vowals in addition > > to the ones available in ASCII. They are represented by the > > letters "åäö". These are letters, not an "a" or "o" with an > > accent above. Without those three letters you cannot write all > > Swedish names. Accents I can live without, but not the letters > > for our additional phonemes. > > > > To be able to write most names in the world I think you need > > to be able to write about 60 phonems. Not everybody need their > > own letter (English have about 45 phonems but the 26 letters > > are enough). So I would expect by adding not that many more > > letters to the ones in ASCII we could get a quite faire > > representation of all names in the world. That would be more > > acceptable to have on a business card. > > > > And just like Adam my Swedish keyboard do not have any > > accented or diacritic letters. But I can still type quite a > > lot of them by using "compose" or "alt graph". > > > > ASCII will never be good enough for use as a "global" name for > > Swedish. But with a few more letters added it would be > > possible. I expect the same would work for most other > > languages. I would not be surprised if 35-45 letters would be > > enough to give quite good representation of the native names. > > > > Dan > > > > > > > > > >