Re: New Version Notification for draft-obispo-epp-idn-00.txt
Andrew Sullivan <[email protected]>
| Newsgroups | gmane.ietf.provreg |
|---|---|
| Message-ID | <[email protected]> |
Hi all, Doing some catch-up. Sorry for coming in late. On Wed, Dec 21, 2011 at 10:10:52PM -0800, Francisco Obispo wrote: > Looking at the IDN implementation guidelines, item #5 states: > > 5.All code points in a single label will be taken from the same > script as determined by the Unicode Standard Annex #24: Script > Names <http://www.unicode.org/reports/tr24>. Exceptions to this > guideline are permissible for languages with established > orthographies and conventions that require the commingled use of > multiple scripts. i.e. "every language in the world". It turns out that there is a major problem with this guideline in general use, even if it turns out to work for some languages for TLDs: most languages need some things from Common or Inherited. That is true even of LDH labels. You might get around this by hand-waving Common and Inherited, but then you have a different problem. For instance, suppose you wanted to permit the traditional digits 0-9; they're Common. But if you said that Latin implicitly included everything in Common, you'd implicitly include 0660..0669, which are ARABIC-INDIC DIGIT ZERO..ARABIC-INDIC DIGIT NINE. Which is presumably not what you wanted. Worse, > So it would not be possible to have multiple languages associated to a label. you're conflating "language" and "script" here, since the restriction above is about scripts and not languages. > XML Schema "language" type[1]: > > [Definition:] language represents natural language identifiers as defined by by [RFC 3066][2] . The ·value space· of language is the set of all strings that I guess this reference ought to be changed to 4646? Anyway, that doesn't wholly help you, becuase just because you have an identifier doesn't mean you have a reasonable repertoire of characters. More broadly, I'd like it a lot if people went and looked at the discussion of zone repertoires and the label generation rules in the recent ICANN Variant Issues Project (announcement here: http://www.icann.org/en/announcements/announcement-2-23dec11-en.htm). I came to believe over the last six or so months that neither "language" nor "script" is what we want here, and I think it would be extremely helpful to have some other eyes on the discussion of this issue in that report. While the report is actually aimed only at the top level, I think this part of the report is broadly applicable to any zone that has to serve a linguistically diverse population (i.e. pretty much every gTLD and some ccTLDs too). Best regards, A -- Andrew Sullivan <[email protected]> _______________________________________________ provreg mailing list [email protected] https://www.ietf.org/mailman/listinfo/provreg