BCP 47 and subdivisions (was: RE: Montenegrin)
"Doug Ewell" <[email protected]>
| Newsgroups | gmane.ietf.languages |
|---|---|
| Message-ID | <20171010113252.665a7a7059d7ee80bb4d670165c8327d.0a2fa18a75.wbe@email03.godaddy.com> |
Mark Davis wrote on the Montenegrin thread: > And such action further muddies the line for what constitutes a > language: one step further along a slippery slope to an "American" > language code, or a "Texan" language code, etc. Arthur Reutenauer replied to Mark: > If the latter is ever requested, we may consider extending the set of > region subtags to include ISO 3166-2 ;-) BCP 47 extension U allows this, except that the subtags for subdivisions are CLDR identifiers, a superset of ISO 3166-2 in which deleted 3166-2 code elements are deprecated but still present in CLDR. CLDR subdivision codes and Extension U may not have been specifically invented for one another, but I don't see what violation would result from using them in this way. John Cowan responded to Arthur: > We considered that when writing the RFC. But it falls down on a > number of counts: > > 1) 3166-2 codes aren't freely and publicly available direct from ISO > (although Wikipedia is pretty good about keeping its copies up to date). Extension U solves this, by making the IP CLDR's problem instead of ours. > 2) They often represent merely administrative units that keep changing > all the time: after centuries of stability in county names and > boundaries until 1974, the UK has rearranged its administrative units > every few years, most recently in 2015. > > 3) Partly because of (2), dialect boundaries don't actually follow > ISO-3166-2 unit boundaries. These arguments could be, and have often been, made against region subtags in general: they are variably too broad, too narrow, and just right. > 4) The codes are variable-length (1-3 alphanumerics) , which does not > fit into our rather rigid structure. Solved by Extension U, as shown below. > If necessary they could be added using a reserved singleton, like > en-us-2-tx or en-gb-2-sct. IIRC, with Extension U these would be "en-US-u-sd-ustx" and "en-GB-u-sd-gbsct" respectively. -- Doug Ewell | Thornton, CO, US | ewellic.org