Re: language (or script?) tagging in draft-obispo-epp-idn-00.txt

Keith Gaughan <[email protected]>
Newsgroups gmane.ietf.provreg
Organization Blacknight Internet Solutions
Message-ID <[email protected]>
Apologies in advance for all the footnotes.

On 22/12/11 16:38, Jim Reid wrote:

> On 22 Dec 2011, at 16:09, Keith Gaughan wrote:
> 
>> Essentially, what Patrik (and I) are asking is, if an IDN fits into a common
>> subset of the character tables of multiple languages, should it not be
>> possible to allow all those languages to be specified when registering the
>> IDN?
> 
> It might not be that simple. Should it be "all languages" or "all scripts"?
> I'm thinking of Japanese where there could be as many as four different
> alphabets (scripts) used for representing one word.

The point is primarily not to allow homograph attacks. 'Script' mightn't be
the best way to think of things, especially with the likes of Japanese.
Instead, 'writing system'[1] would be better: so long as the writing system
contains no homographs[2], there's no issue.

> I'm also having flashbacks
> to the confusion between scripts and languages when ICANN was discussing IDN
> TLDs 2-3 years ago. ISTR there are differences between the Chinese characters
> used in Taiwan and those in the People's Republic.

Well, that goes back to the whole headache of Han unification in Unicode, and
it doesn't just apply to Simplified vs Traditional Hanzi, but also Kanji
(Japanese) and Hanja (Korean). All may be rendered differently, even if they
share codepoints[3], which I'm sure made plenty of headaches for the people
behind .asia. RFC 3743 sorted that mess out for the most part, didn't it?

K.

[1] By which I mean, all the languages that use the Latin alphabet could be
    said to share a common script and writing system, which are both the same
    thing. It's not an ideal term, I know. Japanese and Chinese could be said
    to have separate writing systems even if they both have variants of a
    common script as part of them. Japanese has a writing system that
    consists of four--six if you count Hentaigana and Man'yōgana--different
    scripts.

[2] And between Hiragana, Katakana, Romaji, and Kanji, there are none that I
    know of, though variant Han characters are problematic, which is a
    known problem that's been dealt with by the various registries already.
    Sometimes I wonder why the Japanese haven't gone down the same road as the
    Koreans are going down (gradually abandoning Hanja for exclusive use of
    Hangeul) and Vietnamese (which no longer uses Han characters at all),
    especially when it has two excellent syllabaries.

[3] I can't recall whether Traditional and Simplified Han characters occupy
    different code points in general, or the same ones. It's years since I
    last looked that up, so I'd ask leniency about any inaccurate remarks I
    may have made on that issue.

-- 
Keith Gaughan, Senior Developer
PGP/GPG key ID: 3E896381
Blacknight Internet Solutions Ltd. <http://blacknight.com/>
12A Barrowside Business Park, Carlow, Ireland
Registered in Ireland, Company No.: 370845
_______________________________________________
provreg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/provreg
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.