Re: UTS #35: Unicode Locale Data Markup Language (LDML)

Mark Davis ☕️ <[email protected]> Mon, 22 Jul 2019 18:42:00 -0700
Newsgroups gmane.ietf.languages
Message-ID <CAJ2xs_H7Ly1YL2QDvZ+g77yq_NVycO6jMtGi03ckx3_hy_0J_w@mail.gmail.com>
--===============4425200038544018470==
Content-Type: multipart/alternative; boundary="00000000000030f7b7058e4f4d22"

--00000000000030f7b7058e4f4d22
Content-Type: text/plain; charset="UTF-8"

See https://www.unicode.org/reports/tr35/#BCP_47_Conformance

Use of the *Unicode BCP 47 locale identifier* is preferred: it is valid BCP
47 language tag. (The term *Unicode CLDR locale identifier* applies where a
backwards compatibility syntax is used.)

Mark


On Mon, Jul 22, 2019 at 4:08 PM Gordon P. Hemsley <[email protected]> wrote:

> [I wouldn't expect this to be news to many of the people who are active
> on this mailing list, given that they were apparently involved in its
> creation, but it was news to me, so I wanted to share it.]
>
> It appears that Unicode has defined what is effectively a superset of
> BCP 47 for use in the CLDR:
>
> https://www.unicode.org/reports/tr35/
>
> This goes beyond maintaining the 't' and 'u' extensions to defining
> their own rules for parsing and canonicalizing BCP 47 language tags.
>
> Can someone who is familiar with the matter explain the relationship
> between Unicode's UTS #35 and IETF's BCP 47 standards?
>
>
> --
> Gordon P. Hemsley
> [email protected]
> http://gphemsley.org/
>
> _______________________________________________
> Ietf-languages mailing list
> [email protected]
> https://www.ietf.org/mailman/listinfo/ietf-languages
>

--00000000000030f7b7058e4f4d22
Content-Type: text/html; charset="UTF-8"
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div class=3D"gmail_default" style=3D"font-family:times ne=
w roman,serif">See=C2=A0<a href=3D"https://www.unicode.org/reports/tr35/#BC=
P_47_Conformance" style=3D"font-family:Arial,Helvetica,sans-serif">https://=
www.unicode.org/reports/tr35/#BCP_47_Conformance</a></div><div class=3D"gma=
il_default" style=3D"font-family:times new roman,serif"><br></div><div clas=
s=3D"gmail_default" style=3D"font-family:times new roman,serif">Use of the=
=C2=A0<i>Unicode BCP 47 locale identifier</i> is preferred: it is valid BCP=
 47 language tag.=C2=A0(The term=C2=A0<i>Unicode CLDR locale identifier</i>=
 applies where a backwards compatibility syntax is used.)</div><div class=
=3D"gmail_default" style=3D"font-family:times new roman,serif"><br></div><d=
iv><div dir=3D"ltr" class=3D"gmail_signature" data-smartmail=3D"gmail_signa=
ture"><div dir=3D"ltr"><div><div dir=3D"ltr"><div><div dir=3D"ltr"><div dir=
=3D"ltr"><div dir=3D"ltr"><div dir=3D"ltr"><div dir=3D"ltr"><font face=3D"&=
#39;times new roman&#39;, serif"><div style=3D"background-color:transparent=
;margin-top:0px;margin-left:0px;margin-bottom:0px;margin-right:0px"><div></=
div></div><div style=3D"background-color:transparent;margin-top:0px;margin-=
left:0px;margin-bottom:0px;margin-right:0px">Mark</div></font><div><div><fo=
nt face=3D"&#39;times new roman&#39;, serif"><i><span style=3D"font-style:n=
ormal"><i></i></span><i></i></i></font></div></div></div></div></div></div>=
</div></div></div></div></div></div></div><br></div><br><div class=3D"gmail=
_quote"><div dir=3D"ltr" class=3D"gmail_attr">On Mon, Jul 22, 2019 at 4:08 =
PM Gordon P. Hemsley &lt;<a href=3D"mailto:[email protected]">[email protected]=
rg</a>&gt; wrote:<br></div><blockquote class=3D"gmail_quote" style=3D"margi=
n:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex=
">[I wouldn&#39;t expect this to be news to many of the people who are acti=
ve<br>
on this mailing list, given that they were apparently involved in its<br>
creation, but it was news to me, so I wanted to share it.]<br>
<br>
It appears that Unicode has defined what is effectively a superset of<br>
BCP 47 for use in the CLDR:<br>
<br>
<a href=3D"https://www.unicode.org/reports/tr35/" rel=3D"noreferrer" target=
=3D"_blank">https://www.unicode.org/reports/tr35/</a><br>
<br>
This goes beyond maintaining the &#39;t&#39; and &#39;u&#39; extensions to =
defining<br>
their own rules for parsing and canonicalizing BCP 47 language tags.<br>
<br>
Can someone who is familiar with the matter explain the relationship<br>
between Unicode&#39;s UTS #35 and IETF&#39;s BCP 47 standards?<br>
<br>
<br>
-- <br>
Gordon P. Hemsley<br>
<a href=3D"mailto:[email protected]" target=3D"_blank">[email protected]</a><=
br>
<a href=3D"http://gphemsley.org/" rel=3D"noreferrer" target=3D"_blank">http=
://gphemsley.org/</a><br>
<br>
_______________________________________________<br>
Ietf-languages mailing list<br>
<a href=3D"mailto:[email protected]" target=3D"_blank">Ietf-languages=
@ietf.org</a><br>
<a href=3D"https://www.ietf.org/mailman/listinfo/ietf-languages" rel=3D"nor=
eferrer" target=3D"_blank">https://www.ietf.org/mailman/listinfo/ietf-langu=
ages</a><br>
</blockquote></div>

--00000000000030f7b7058e4f4d22--


--===============4425200038544018470==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Ietf-languages mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/ietf-languages

--===============4425200038544018470==--