Re: Fwd: I-D Action: draft-msporny-d-langtag-ext-00.txt
Mark Davis ☕️ <[email protected]> Mon, 27 May 2019 15:52:23 +0200
| Newsgroups | gmane.ietf.languages |
|---|---|
| Message-ID | <CAJ2xs_EwKg3Tu5etk-ELXXd0u2Go-6TZbGm3QsBxV1upKTa8_g@mail.gmail.com> |
--===============4583254440210378145== Content-Type: multipart/alternative; boundary="0000000000005265510589dedc51" --0000000000005265510589dedc51 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Doug, I agree most of the points you are making, especially #1. I think what they are trying to do is shoehorn in a parameter that lets them set the paragraph embedding level (https://unicode.org/reports/tr9/#BD= 4) for the Bidi Algorithm. But instead of that hack, as you point out, one can deduce the direction from the language tag. The best way to do this is to get the ordering from CLDR for an ordinary language tag like "ar" or "ar-Arab". 1. So from the tag "ar-Arab", we get the script "Arab". Then use https://github.com/unicode-org/cldr/blob/master/common/properties/scriptMet= adata.txt, which has a mapping from script to direction (RTL=3DYES). (I'm pointing to trunk, just so people can read the file easily; one would use the latest release.) 2. But let's suppose that you have just "ar". Since the script is not explicit, the best way to get it is also CLDR. You can use https://github.com/unicode-org/cldr/blob/master/common/supplemental/likelyS= ubtags.xml, which has a mapping from language or language+region to default language-script-region. So "ar" =3D> "ar_Arab_EG", from which we get the script "Arab", and then use step 1. Or from "fr" you'd get "Latn" and map it to RTL=3DNO. A few more comments below. On Mon, May 27, 2019 at 2:47 AM Doug Ewell <[email protected]> wrote: > Manu Sporny wrote: > > > There is a time pressure here. Our i198n concerns have been hanging > > out there for more than 9 months and our WG charter is up in a couple > > of months. We need to wrap this up in 3 weeks. Or to put it another > > way, if we don't wrap this up in 3 weeks, we won't be addressing this > > issue, which would be a shame. > > I know it is flippant to say "that's not our problem," and I apologize in > advance for that, but trying to push through this extension quickly, > without consulting or even notifying the language-tagging community, does > not seem to me an appropriate way to compensate for this lapse. It was on= ly > by chance that Martin happened to spot this I-D and was able to bring it = to > our attention. > > Apparently Addison did know about this effort, and is credited in the > Acknowledgements section of the I-D, but it would be nice if the author(s= ) > of an extension proposal would check in with ietf-languages as part of > their effort. RFC 5646 does not require this; I wish it did. The IETF at > large and W3C are not experts in this field, and probably will not be abl= e > to detect significant operational problems in such a proposal. > > > In any case, if you're going to engage in this discussion, the issue > > #3 above is probably the place to do it. > > I believe THIS LIST is the place to discuss this I-D. (Definitely not on > some GitHub account.) > > I have other questions and/or concerns, some of which overlap with > Martin's: > > 1. In the proposal's lone example, the Arabic script is a right-to-left > script. How does "ar-d-rtl" indicate right-to-left directionality in a wa= y > that "ar-Arab" does not? > 2. Given #1, and given that the script subtag 'Arab' is a Suppress-Script > for the language subtag 'ar' (which means "ar" is equivalent to "ar-Arab" > for almost all purposes), how is "ar" not sufficient? I agree with Martin= 's > comment here: what rendering process is likely to display Arabic > left-to-right? > It isn't that Arabic would be displayed left to right, it is what establishes the paragraph ordering. The problem arises when you have mixed text. Look at the following example, using the convention that lowercase = =3D English and uppercase=3DArabic. The majority of the text and the first stro= ng character are both English, but the sentence is meant to be used in an Arabic environment, so the default paragraph embedding level needs to be RTL. rindfleischetikettierungs=C3=BCberwachungsaufgaben=C3=BCbertragungsgesetz I= S A LONG WORD. > > 3. I also agree with Martin that the definition "automatically detected" > for subtag 'auto' is not adequate. How does it differ from leaving off th= e > D extension altogether? > Agreed, not well specified. But -d- is not needed in the first place, so moot. > > 4. Scripts exist in other directionalities besides LTR and RTL. Chinese, > Japanese, and Korean can be written top-to-bottom, right-to-left. Mongoli= an > in Mongolian script is properly written top-to-bottom, left-to-right, but > is sometimes (although incorrectly) rendered LTR as well. Some languages > have been written boustrophedon, either with or without reversing the > glyphs when transitioning from LTR to RTL. None of these scenarios are > covered in the proposal, but some of them seem much more in need of > explicit marking than the Arabic example. > While this is true, for the fast majority of cases, LTR and RTL are the important issues. Most computer systems don't really handle vertical natively; one needs to have more specialized text processing systems, and that is not, I imagine, the target for this syntax. > 5. Given #4, the lack of a registry for the proposed extension, or even > the mention of one, is a significant problem. The set of exactly 3 values > associated with this extension ('ltr', 'rtl', and 'auto') would be fixed; > adding to it would require updating the RFC, which is much more work than > updating a registry. > Agreed, that would be a major drawback. But -d- is not needed in the first place, so moot. > Without these issues being addressed in a satisfactory way, I would lobby > IETF not to approve this I-D. > I don't see that there is any reason to approve it, given that it is, as far as I can tell, completely unnecessary and would just complicate implementer's lives to no good end. > > -- > Doug Ewell | Thornton, CO, US | ewellic.org > > > _______________________________________________ > Ietf-languages mailing list > [email protected] > https://www.ietf.org/mailman/listinfo/ietf-languages > --0000000000005265510589dedc51 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div dir=3D"ltr"><div class=3D"gmail_default" style=3D"fon= t-family:times new roman,serif">Doug, I agree most of the points you are ma= king, especially #1.</div><div class=3D"gmail_default" style=3D"font-family= :times new roman,serif"><br></div><div class=3D"gmail_default" style=3D"fon= t-family:times new roman,serif">I think what they are trying to do is shoeh= orn in a parameter that lets them set the paragraph embedding level (<a hre= f=3D"https://unicode.org/reports/tr9/#BD4" style=3D"font-family:Arial,Helve= tica,sans-serif">https://unicode.org/reports/tr9/#BD4</a>) for the Bidi Alg= orithm. But instead of that hack, as you point out, one can deduce the dire= ction from the language tag. The best way to do this is to get the ordering= from CLDR for an ordinary language tag like "ar" or "ar-Ara= b".</div><div class=3D"gmail_default" style=3D"font-family:times new r= oman,serif"><br></div><div class=3D"gmail_default" style=3D"font-family:tim= es new roman,serif">1. So from the tag "ar-Arab", we get the scri= pt "Arab". Then use=C2=A0<a href=3D"https://github.com/unicode-or= g/cldr/blob/master/common/properties/scriptMetadata.txt">https://github.com= /unicode-org/cldr/blob/master/common/properties/scriptMetadata.txt</a>, whi= ch has a mapping from script to direction (RTL=3DYES). (I'm pointing to= trunk, just so people can read the file easily; one would use the latest r= elease.)</div><div class=3D"gmail_default" style=3D"font-family:times new r= oman,serif"><br></div><div class=3D"gmail_default" style=3D"font-family:tim= es new roman,serif">2. But let's suppose that you have just "ar&qu= ot;. Since the script is not explicit, the best way to get it is also CLDR.= You can use=C2=A0<a href=3D"https://github.com/unicode-org/cldr/blob/maste= r/common/supplemental/likelySubtags.xml">https://github.com/unicode-org/cld= r/blob/master/common/supplemental/likelySubtags.xml</a>, which has a mappin= g from language or language+region to default language-script-region. So &q= uot;ar" =3D>=C2=A0"ar_Arab_EG", from which we get the scr= ipt "Arab", and then use step 1. Or from "fr" you'd= get "Latn" and map it to RTL=3DNO.</div><div class=3D"gmail_defa= ult" style=3D"font-family:times new roman,serif"><br></div><div class=3D"gm= ail_default" style=3D"font-family:times new roman,serif">A few more comment= s below.</div></div><br><div class=3D"gmail_quote"><div dir=3D"ltr" class= =3D"gmail_attr">On Mon, May 27, 2019 at 2:47 AM Doug Ewell <<a href=3D"m= ailto:[email protected]">[email protected]</a>> wrote:<br></div><blockquot= e class=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px s= olid rgb(204,204,204);padding-left:1ex">Manu Sporny wrote:<br> <br> > There is a time pressure here. Our i198n concerns have been hanging<br= > > out there for more than 9 months and our WG charter is up in a couple<= br> > of months. We need to wrap this up in 3 weeks. Or to put it another<br= > > way, if we don't wrap this up in 3 weeks, we won't be addressi= ng this<br> > issue, which would be a shame.<br> <br> I know it is flippant to say "that's not our problem," and I = apologize in advance for that, but trying to push through this extension qu= ickly, without consulting or even notifying the language-tagging community,= does not seem to me an appropriate way to compensate for this lapse. It wa= s only by chance that Martin happened to spot this I-D and was able to brin= g it to our attention.<br> <br> Apparently Addison did know about this effort, and is credited in the Ackno= wledgements section of the I-D, but it would be nice if the author(s) of an= extension proposal would check in with ietf-languages as part of their eff= ort. RFC 5646 does not require this; I wish it did. The IETF at large and W= 3C are not experts in this field, and probably will not be able to detect s= ignificant operational problems in such a proposal.<br> <br> > In any case, if you're going to engage in this discussion, the iss= ue<br> > #3 above is probably the place to do it.<br> <br> I believe THIS LIST is the place to discuss this I-D. (Definitely not on so= me GitHub account.)<br> <br> I have other questions and/or concerns, some of which overlap with Martin&#= 39;s:<br> <br> 1. In the proposal's lone example, the Arabic script is a right-to-left= script. How does "ar-d-rtl" indicate right-to-left directionalit= y in a way that "ar-Arab" does not?</blockquote><blockquote class= =3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rg= b(204,204,204);padding-left:1ex"> <br> 2. Given #1, and given that the script subtag 'Arab' is a Suppress-= Script for the language subtag 'ar' (which means "ar" is = equivalent to "ar-Arab" for almost all purposes), how is "ar= " not sufficient? I agree with Martin's comment here: what renderi= ng process is likely to display Arabic left-to-right?<br></blockquote><div>= <br></div><div class=3D"gmail_default" style=3D"font-family:"times new= roman",serif">It isn't that Arabic would be displayed left to rig= ht, it is what establishes the paragraph ordering. The problem arises when = you have mixed text. Look at the following example, using the convention th= at lowercase =3D English and uppercase=3DArabic. The majority of the text a= nd the first strong character are both English, but the sentence is meant t= o be used in an Arabic environment, so the default paragraph embedding leve= l needs to be RTL.</div><div class=3D"gmail_default" style=3D"font-family:&= quot;times new roman",serif"><br></div><div class=3D"gmail_default" st= yle=3D"font-family:"times new roman",serif">rindfleischetikettier= ungs=C3=BCberwachungsaufgaben=C3=BCbertragungsgesetz IS A LONG WORD.</div><= div class=3D"gmail_default" style=3D"font-family:"times new roman"= ;,serif"></div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0p= x 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"> <br> 3. I also agree with Martin that the definition "automatically detecte= d" for subtag 'auto' is not adequate. How does it differ from = leaving off the D extension altogether?<br></blockquote><div><br></div><div= class=3D"gmail_default" style=3D"font-family:"times new roman",s= erif">Agreed, not well specified. But -d- is not needed in the first place,= so moot.</div><blockquote class=3D"gmail_quote" style=3D"margin:0px 0px 0p= x 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"> <br> 4. Scripts exist in other directionalities besides LTR and RTL. Chinese, Ja= panese, and Korean can be written top-to-bottom, right-to-left. Mongolian i= n Mongolian script is properly written top-to-bottom, left-to-right, but is= sometimes (although incorrectly) rendered LTR as well. Some languages have= been written boustrophedon, either with or without reversing the glyphs wh= en transitioning from LTR to RTL. None of these scenarios are covered in th= e proposal, but some of them seem much more in need of explicit marking tha= n the Arabic example.<br></blockquote><div><br></div><div class=3D"gmail_de= fault" style=3D"font-family:"times new roman",serif">While this i= s true, for the fast majority of cases, LTR and RTL are the important issue= s. Most computer systems don't really handle vertical natively; one nee= ds to have more specialized text processing systems, and that is not, I ima= gine, the target for this syntax.</div><div class=3D"gmail_default" style= =3D"font-family:"times new roman",serif"><br></div><blockquote cl= ass=3D"gmail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid= rgb(204,204,204);padding-left:1ex"> <br> 5. Given #4, the lack of a registry for the proposed extension, or even the= mention of one, is a significant problem. The set of exactly 3 values asso= ciated with this extension ('ltr', 'rtl', and 'auto'= ;) would be fixed; adding to it would require updating the RFC, which is mu= ch more work than updating a registry.<br></blockquote><div><br></div><div = class=3D"gmail_default" style=3D"font-family:"times new roman",se= rif">Agreed, that would be a major drawback.=C2=A0 But -d- is not needed in= the first place, so moot.</div><div class=3D"gmail_default" style=3D"font-= family:"times new roman",serif"><br></div><blockquote class=3D"gm= ail_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,= 204,204);padding-left:1ex"> <br> Without these issues being addressed in a satisfactory way, I would lobby I= ETF not to approve this I-D.<br></blockquote><div><br></div><div class=3D"g= mail_default" style=3D"font-family:"times new roman",serif">I don= 't see that there is any reason to approve it, given that it is, as far= as I can tell, completely unnecessary and would just complicate implemente= r's lives to no good end.</div><div class=3D"gmail_default" style=3D"fo= nt-family:"times new roman",serif"></div><blockquote class=3D"gma= il_quote" style=3D"margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,2= 04,204);padding-left:1ex"> <br> --<br> Doug Ewell | Thornton, CO, US | <a href=3D"http://ewellic.org" rel=3D"noref= errer" target=3D"_blank">ewellic.org</a><br> <br> <br> _______________________________________________<br> Ietf-languages mailing list<br> <a href=3D"mailto:[email protected]" target=3D"_blank">Ietf-languages= @ietf.org</a><br> <a href=3D"https://www.ietf.org/mailman/listinfo/ietf-languages" rel=3D"nor= eferrer" target=3D"_blank">https://www.ietf.org/mailman/listinfo/ietf-langu= ages</a><br> </blockquote></div></div> --0000000000005265510589dedc51-- --===============4583254440210378145== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ Ietf-languages mailing list [email protected] https://www.ietf.org/mailman/listinfo/ietf-languages --===============4583254440210378145==--