Re: New Version Notification for draft-obispo-epp-idn-00.txt
Keith Gaughan <[email protected]>
| Newsgroups | gmane.ietf.provreg |
|---|---|
| Organization | Blacknight Internet Solutions |
| Message-ID | <[email protected]> |
On 22/12/11 20:25, Michael Young wrote:
> First of all Janusz, who submitted that Afilias IDN I-D was working for me
> at the time, I thought it seemed like a good idea at the time to encourage
> standardization.
It was, very much so.
> We were thinking primarily languages versus scripts due guidance we had from
> John Klensin at the time.
>
>> Though more specifically, the point of even asking for the language in the
> first place is to prevent homograph attacks. The use of languages rather
> than scripts leads to slightly odd situations where a person wouldn't be
> able to register a domain >like æçöÿ.info (a silly example, I know, but
> valid), as though there's no language that uses all those characters in
> their alphabet, it's in no way a vector for homograph attacks.
> So specifically we thought about the above case and decided that a language
> was a construct of sorts, usually indicated by a group or community
> declaring and codifying its uniqueness.
Even that's not entirely correct, as the linguist Max Weinreich wrote about
the social plight of Yiddish, "a shprakh iz a dialekt mit an armey un flot".
That's why EURid went down the route of specifying codepoint tables for
scripts rather than languages: it's *relatively* easy to deliminate scripts
and thus avoid homograph attacks within scripts[1] than it is to decide what
language ought to be allowed, especially when that means that going the
language route can end up inadvertently discriminating against minority
languages for no good reason.
> We figured any policy in
> registration should be made against a language table that was in turn
> defined by a recognized language authority. Now "language authority" is a
> wide range, but so are the number of defined languages.
Not all recognised languages have language academies though, especially
minority languages, and in some parts of the world (India, for example),
there's a common folk belief that to be a real language without having its
own distinct script, even if those languages with the common script are
members of totally different language families, like Indo-Aryan and Dravidian.
IIRC, the Grantha script was an historical example of this. And then there's
the interesting situation in Chinese that bungs 'dialects' as different as
Welsh and Greek are from one another under a single banner as 'Chinese'.
> We decided against defining the policies (at that time in 2004, I am not
> speaking to anyone's current practise/policy) by script because we didn’t
> see a graceful way to address different spoken languages that shared the
> same script.
The graceful way to do this is to have multiple tables per script to check
labels against. If all the characters in the label are whitelisted in any
one of those tables, then it's valid. That solves the homograph problem
without discriminating against minority languages (and Europe still has a lot
of them), and copes just fine with situations like Serbian, which uses both
the Latin and Cyrillic scripts. All that's saved by requiring a language code
is CPU cycles.
K.
[1] Odd cases like s-comma in Romanian versus s-cedilla in Turkish aside,
though that's an example of a near-homograph rather than a true homograph.
This kind of thing is the best argument for going down the language rather
than the script.
--
Keith Gaughan, Senior Developer
PGP/GPG key ID: 3E896381
Blacknight Internet Solutions Ltd. <http://blacknight.com/>
12A Barrowside Business Park, Carlow, Ireland
Registered in Ireland, Company No.: 370845
_______________________________________________
provreg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/provreg