Re: New Version Notification for draft-obispo-epp-idn-00.txt

Andrew Sullivan <[email protected]>
Newsgroups gmane.ietf.provreg
Message-ID <[email protected]>
Hi all,

Doing some catch-up.  Sorry for coming in late.

On Wed, Dec 21, 2011 at 10:10:52PM -0800, Francisco Obispo wrote:

> Looking at the IDN implementation guidelines, item #5 states:
> 
> 5.All code points in a single label will be taken from the same
>   script as determined by the Unicode Standard Annex #24: Script
>   Names <http://www.unicode.org/reports/tr24>. Exceptions to this
>   guideline are permissible for languages with established 
>   orthographies and conventions that require the commingled use of
>   multiple scripts.

i.e. "every language in the world".  It turns out that there is a
major problem with this guideline in general use, even if it turns out
to work for some languages for TLDs: most languages need some things
from Common or Inherited.  That is true even of LDH labels. 

You might get around this by hand-waving Common and Inherited, but
then you have a different problem.  For instance, suppose you wanted
to permit the traditional digits 0-9; they're Common.  But if you said
that Latin implicitly included everything in Common, you'd implicitly
include 0660..0669, which are ARABIC-INDIC DIGIT ZERO..ARABIC-INDIC
DIGIT NINE.  Which is presumably not what you wanted.

Worse,

> So it would not be possible to have multiple languages associated to a label.

you're conflating "language" and "script" here, since the restriction
above is about scripts and not languages.  

> XML Schema "language" type[1]:
> 
> [Definition:]   language represents natural language identifiers as defined by by [RFC 3066][2] . The ·value space· of language is the set of all strings that

I guess this reference ought to be changed to 4646?

Anyway, that doesn't wholly help you, becuase just because you have an
identifier doesn't mean you have a reasonable repertoire of
characters.

More broadly, I'd like it a lot if people went and looked at the
discussion of zone repertoires and the label generation rules in the
recent ICANN Variant Issues Project (announcement here:
http://www.icann.org/en/announcements/announcement-2-23dec11-en.htm).
I came to believe over the last six or so months that neither
"language" nor "script" is what we want here, and I think it would be
extremely helpful to have some other eyes on the discussion of this
issue in that report.  While the report is actually aimed only at the
top level, I think this part of the report is broadly applicable to
any zone that has to serve a linguistically diverse population
(i.e. pretty much every gTLD and some ccTLDs too).

Best regards,

A

-- 
Andrew Sullivan
<[email protected]>
_______________________________________________
provreg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/provreg
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.