Re: Global and national e-mail address

"Mark Davis" <[email protected]> Wed, 10 Dec 2003 07:33:08 -0800
Newsgroups gmane.ietf.imaa
Message-ID <002601c3bf32$e9a65d10$7900a8c0@DAVIS1>
It is easy to fall into a trap of being Eurocentric: to overestimate the ease of
Latin and underestimate the difficulty of transliteration. For many languages
there are no good transliteration standards; or rather, there are many
conflicting ones. And many of these are transcriptions, and not
transliterations. (The difference is that transliteration to Latin is
reversable; one can recover precisely the original text; transcription is not
reversable -- but is more pronouncable).

For example, here is are some sample transliterations (from
http://oss.software.ibm.com/cgi-bin/icu/tr). If you asked an average user of
each of these languages to do a transliteration, even of those who know Latin,
the odds of their coming up with exactly these results are very small.

По своей природе компьютеры могут работать лишь с числами. И для того, чтобы они
могли хранить в памяти буквы или другие символы, каждому такому символу должно
быть поставлено в соответствие число.
=>
Po svoej prirode kompʹûtery mogut rabotatʹ lišʹ s čislami. I dlâ togo, čtoby oni
mogli hranitʹ v pamâti bukvy ili drugie simvoly, každomu takomu simvolu dolžno
bytʹ postavleno v sootvetstvie čislo. (ISO 9)

Οι ηλεκτρονικοί υπολογιστές, σε τελική ανάλυση, χειρίζονται απλώς αριθμούς.
Αποθηκεύουν γράμματα και άλλους χαρακτήρες αντιστοιχώντας στο καθένα τους από
έναν αριθμό (ονομάζουμε μία τέτοια αντιστοιχία κωδικοσελίδα).
=>
Oi i̱lektronikoí ypologistés, se telikí̱ análysi̱, cheirízontai apló̱s
arithmoús. Apothi̱kév̱oun grámmata kai állous charaktí̱res antistoichó̱ntas sto
kathéna tous apó énan arithmó (onomázoume mía tétoia antistoichía
ko̱dikoselída). (ISO 843)

기본적으로 컴퓨터는 숫자만 처리합니다. 글자나 다른 문자에도 숫자를 지정하여 저장합니다.
=>
gibonjeog'euro keompyuteoneun susjaman ceorihabnida. geuljana dareun munja'edo
susjareul jijeonghayeo jeojanghabnida.
(Korean Ministry of Culture & Tourism Transliteration regulations with the
clause 8 variant)

कम्प्यूटर, मूल रूप से, नंबरों से सम्बंध रखते हैं। ये प्रत्येक अक्षर और वर्ण के
लिए एक नंबर निर्धारित करके अक्षर और वर्ण संग्रहित करते हैं।
=>
kampyūṭara, mūla rūpa sē, nambarōṁ sē sambandha rakhatē haiṁ. yē pratyēka akṣara
aura varṇa kē li'ē ēka nambara nirdhārita karakē akṣara aura varṇa saṅgrahita
karatē haiṁ. (ISO 15919)

And there is no known transliteration method for Chinese or Japanese -- there
are (many) transcription standards, but none that preserve the original
characters.

John's point about IPA is particularly apt. Think about the example: if an
English user had to use IPA for web addresses, how would he do it? For
"bort.com" would he use "bɔrt.kɑm", "bɔːt.kɒm", "bɔrʔ.kɑm", ... It would be very
difficult to predict exactly what the spelling in IPA would be, even if he were
conversant with IPA. The same problem is faced by someone using a language
normally written in a non-Latin script when trying to transliterate into Latin.

Mark
__________________________________
http://www.macchiato.com
► शिष्यादिच्छेत्पराजयम् ◄

----- Original Message ----- 
From: "John C Klensin" <[email protected]>
To: "Dan Oscarsson" <[email protected]>; <[email protected]>
Sent: Wed, 2003 Dec 10 06:11
Subject: Re: Global and national e-mail address


>
> Dan,
>
> Three observations (short this time)...
>
> (i) Computer geeks and their possible preferences aside, people
> tend to not like transliterations (writing of a name that would
> normally be written in one character set in the characters of
> another).  Whether they like "better" transliterations more than
> "worse" transliterations is a cultural issue.
>
> (ii) To reasonably transliterate names and languages, one needs
> not only a collection of the right phonemes, but an appropriate
> and accurate notation for tones.   You can't get those out of a
> small extension to Latin letters.  I'm told that one can get a
> reasonable approximation of all of the relevant phonemes and
> tones with IPA, but IPA not only uses some characters that are
> distinctly non-Latin-based, but also uses a rather complex
> collection of combining diacriticals.  And, at least unless one
> is a professional phonologist, learning IPA and how to use it
> accurately is _hard_ (having had people attempt to teach it to
> me twice, once when I was young enough to learn these things).
> For some hints in a reference that is easily accessible to most
> of us, see the discussion of IPA Characters in the Unicode
> definition (3.0 or 4.0, take your pick).
>
> (iii) If one wants even an approximation to accurate
> transliteration, the symbol-overloading in Latin scripts is bad
> news.  E.g., the sound of "ö" (o with diaresis, U+00F6) is
> different in, e.g., Swedish and German.  I.e., they are
> different characters, even if they look the same and even if
> Unicode "unified" them.  If one is trying to transliterate,
> e.g., Arabic into Roman characters, does one pick a character on
> the basis of the Swedish phoneme or the German ones?  (Hint, as
> soon as you start down that path, you end up sliding toward IPA.)
>
> It appears to me that you are proposing a very Euro-centric view
> of things, and it won't work all that well even for Europe.
>
> regards,
>     john
>
>
> --On Wednesday, 10 December, 2003 10:32 +0100 Dan Oscarsson
> <[email protected]> wrote:
>
> >
> > I will here comment several comments frpm John C Klensin, Adan
> > M Costello and others.
> >
> > With this topic I did not want to talk about replacing ASCII
> > with ISO 8859-1 or about what character encoding to use.
> > Instead I wanted to discuss what names could be suitable to
> > use as a global fallback name.
> >
> > To be able to write a name you need, at least, to be able to
> > have letters so you can write all phonemes used. ASCII only
> > contains 26 letters and they cannot represent all phonemes
> > very well. For example, Swedish have three vowals in addition
> > to the ones available in ASCII. They are represented by the
> > letters "åäö". These are letters, not an "a" or "o" with an
> > accent above. Without those three letters you cannot write all
> > Swedish names. Accents I can live without, but not the letters
> > for our additional phonemes.
> >
> > To be able to write most names in the world I think you need
> > to be able to write about 60 phonems. Not everybody need their
> > own letter (English have about 45 phonems but the 26 letters
> > are enough). So I would expect by adding not that many more
> > letters to the ones in ASCII we could get a quite faire
> > representation of all names in the world. That would be more
> > acceptable  to have on a business card.
> >
> > And just like Adam my Swedish keyboard do not have any
> > accented or diacritic letters. But I can still type quite a
> > lot of them by using "compose" or "alt graph".
> >
> > ASCII will never be good enough for use as a "global" name for
> > Swedish. But with a few more letters added it would be
> > possible. I expect the same would work for most other
> > languages. I would not be surprised if 35-45 letters would be
> > enough to give quite good representation of the native names.
> >
> >    Dan
> >
> >
>
>
>
>
>
>