Re: Directionality Standard
pemuro <[email protected]> Tue, 18 Dec 2007 18:44:25 +0100
| Newsgroups | gmane.text.unicode.devel |
|---|---|
| Message-ID | <[email protected]> |
This is a multi-part message in MIME format. --------------060403050804030503000606 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable X-MIME-Autoconverted: from 8bit to quoted-printable by unicode.org id lBIHsM8t030805 <!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN"> <html> <head> <meta content=3D"text/html;charset=3DUTF-8" http-equiv=3D"Content-Type"= > <title></title> </head> <body bgcolor=3D"#ffffff" text=3D"#000000"> Interesting Mark that the bidi-Algorithm defines the default direction of a paragraph.=C2=A0=C2=A0 I took bidi to mean text which uses both directionalities (in some textpassage, NOT necessarily changing only from paragraph to paragraph).=C2=A0 <br> I like to collect etymological information of one word in one paragraph switching often several time per line between R2L and L2R and had no problems with corrrect linebreaking e.g. with a semitic phrase embedded in an English text). I was not missing the RTL-paragraph-icon in Open-Office (exits now). In Word switching such a multi-language paragraph's directionality made it unrecognizable.=C2=A0 I thought these icons a relic of pre-bidi/ pre-Unicode times when some fonts did not contain automatic directionality. <br> Only one irritant exists: Entering the directionally neutral SPace after an RTL-word makes the cursor jump to the right (when the paragraph is LTR) before i enter the next RTL-letter when it jumps back to the left.=C2=A0=C2=A0 I consider this jumping premature!=C2=A0 It woul= d be more logical to decide only after the next symbol with directionality is entered.<br> That is of course not regulated by Unicode, but I would much prefer it to be the default behavior. <br> <br> For my mnemonic multi-Latin++ (IPA, Sanskrit, ) keyboard layout and textexamples look at the directory pemuro.funpic.de <br> <br> pemuro<br> <br> <br> <br> Mark Davis wrote: <blockquote cite=3D"[email protected]" type=3D"cite">There may be some misunderstanding. Unicode does define the default direction of a paragraph for use with the bidi algorithm (which determines the ordering of characters containing bidirectional scripts like Arabic or Hebrew). <br> <br> See <a href=3D"http://unicode.org/reports/tr9/">http://unicode.org/report= s/tr9/</a><br> <br> Mark<br> <br> <div class=3D"gmail_quote">On Dec 17, 2007 4:23 PM, Behnam <<a href=3D"mailto:[email protected]">[email protected] </a>> wrote:<br> <blockquote class=3D"gmail_quote" style=3D"border-left: 1px solid rgb(204, 204, 204); margin: 0pt 0pt 0pt = 0.8ex; padding-left: 1ex;"> <div style=3D"">Thank you. <div>So the answer is no. Unicode does not define the directionality of a paragraph. Then I guess my next question should be why?</div> <div>I think I have some explaining to do.</div> <div>Unicode defines a very complex bidi behaviour of characters, and it defines the beginning and ending of a paragraph (I assume). Yet, it doesn't define what directionality this paragraph should take to arrange these characters within the paragraph. </div> <div>Defining the directionality of a paragraph is more important than defining the language of a text. Yes, language tag can help language aware devices and=C2=A0applications=C2=A0behave accordingly. But directionality definition is not about ' user friendly' behaviour of a text, it is about reproducing the raw text, as intended by its Unicode encoding. </div> <div>Understanding this issue I suppose, may be very easy or very difficult, depending on to the extend you were exposed to rtl experience. In the next paragraph, I write a Persian line, throwing a couple of English words within, and in left to right directionality to give you an idea about what right to left users are experiencing in everyday basis. </div> <div>=D9=BE=D8=B1=D8=B3=D8=B4 =D9=85=D9=86 =D8=A7=D8=B2 Unicode =D8=A7= =DB=8C=D9=86 =D8=A7=D8=B3=D8=AA =DA=A9=D9=87 =DA=86=D8=B1=D8=A7 =D8=A8=D8= =B1=D8=A7=DB=8C =D9=BE=D8=A7=D8=B1=D8=A7=DA=AF=D8=B1=D8=A7=D9=81 directio= nality =D8=AA=D8=A8=DB=8C=DB=8C=D9=86 =D9=86=DA=A9=D8=B1=D8=AF=D9=87 =D8=A7=D8=B3= =D8=AA.</div> <div>In order to read the above phrase correctly in Persian, the order of words should be as I numbered below (from right to left): </div> <div>=D9=BE=D8=B1=D8=B3=D8=B41 =D9=85=D9=862 =D8=A7=D8=B23 Unicode4 =D8= =A7=DB=8C=D9=865 =D8=A7=D8=B3=D8=AA6 =DA=A9=D9=877 =DA=86=D8=B1=D8=A78 =D8= =A8=D8=B1=D8=A7=DB=8C9 =D9=BE=D8=A7=D8=B1=D8=A7=DA=AF=D8=B1=D8=A7=D9=8110 directionality11 =D8=AA=D8=A8=DB=8C=DB=8C=D9=8612 =D9=86=DA=A9=D8=B1=D8=AF= =D9=8713 =D8=A7=D8=B3=D8=AA14.</div> <div><br> </div> <div>Of-course I can set this paragraph in my application to "rtl" and thanks to wonders of bidi behaviour of characters, everything will be put in place: </div> <div style=3D"direction: rtl;"><br> </div> <div> <div style=3D"direction: rtl;">=D9=BE=D8=B1=D8=B3=D8=B4 =D9=85=D9=86 = =D8=A7=D8=B2 Unicode =D8=A7=DB=8C=D9=86 =D8=A7=D8=B3=D8=AA =DA=A9=D9=87 =DA= =86=D8=B1=D8=A7 =D8=A8=D8=B1=D8=A7=DB=8C =D9=BE=D8=A7=D8=B1=D8=A7=DA=AF=D8=B1=D8=A7=D9=81 directionality =D8=AA=D8= =A8=DB=8C=DB=8C=D9=86 =D9=86=DA=A9=D8=B1=D8=AF=D9=87 =D8=A7=D8=B3=D8=AA.<= /div> <div style=3D"direction: rtl;"><br> </div> <div style=3D"direction: ltr;">But I have absolutely no guarantee that my rtl text in an email, in a text message, in an online forum posting... will be received in rtl setting. This perfectly Unicode encoded text is at the mercy of applications, devices, mediums and platforms. And more likely than not, my rtl paragraph will be received in ltr and in the order that I numbered above! Even in a more controlled situations such as word processors, as a friend of mine has experienced, this Persian phrase written in rtl setting of Nisus on a Mac, exported in a .doc format, and opened on a Windows platform will produce an rtl, but 'Arabic' document! not only an Arabic script document which is, but an Arabic language document! </div> <div style=3D"direction: ltr;"><br> </div> <div style=3D"direction: ltr;">You can experiment this dilemma yourself. Set your application to rtl (which can be done in many applications), write something in English or any Roman language. As long as the whole phrase is Roman, you only get a misplaced final period in far left. But if you throw a couple of Hebrew words within the phrase, then you'll see what a wrong directionality setting can do to your English. Of-course you are not exposed to this dilemma because the default directionality of all computerized devices and applications is left to right. But it gives you an idea what rtl users are going through in everyday basis. </div> <div style=3D"direction: ltr;"><br> </div> <div style=3D"direction: ltr;">Again, this is not about requesting a convenience. It is about requesting Unicode to do what it is set to do. Unicode encodes bidi behaviour of characters, the beginning of a paragraph, the end of a paragraph. It must encode its directionality too. </div> <div style=3D"direction: ltr;"><br> </div> <div style=3D"direction: ltr;">Behnam</div> <div style=3D"direction: rtl;"><br> </div> </div> <div><br> <div> <div>On 17-Dec-07, at 4:20 AM, Stephane Bortzmeyer wrote:</div> <br> <blockquote type=3D"cite"> <div>On Sat, Dec 15, 2007 at 11:08:40AM -0500,</div> <div>=C2=A0Behnam <<a href=3D"mailto:[email protected]" target=3D"_blank">[email protected]</a>> wrote=C2=A0</div> <div>=C2=A0a message of 78 lines which said:</div> <div><br> </div> <blockquote type=3D"cite"> <div>Is there any Unicode standard to identify a text? i.e. primary</div> <div>script>directionality>language?</div> </blockquote> <div><br> </div> <div>Not an Unicode standard but, yes, there is a standard to tag texts to </div> <div>indicate language, script, etc. It's RFC 4646. See</div> <div><a href=3D"http://www.langtag.net" target=3D"_blank">http://ww= w.langtag.net</a>/ for a start.</div> </blockquote> </div> <br> </div> </div> </blockquote> </div> <br> <br clear=3D"all"> <br> -- <br> Mark </blockquote> <br> </body> </html> --------------060403050804030503000606 Content-Type: text/x-vcard; charset=utf-8; name="pmr.vcf" Content-Disposition: attachment; filename="pmr.vcf" Content-Transfer-Encoding: base64 YmVnaW46dmNhcmQNCmZuOkRyLiBQZXRlciBSLiAgTXVlbGxlci1Sb2VtZXINCm46TXVlbGxl ci1Sb2VtZXI7RHIuIFBldGVyIFIuIA0KZW1haWw7aW50ZXJuZXQ6cG1yQGNzLnVuaS1mcmFu a2Z1cnQuZGUsIHBldGVtdWVsQHN0dWQudW5pLWZyYW5rZnVydC5kZQ0KeC1tb3ppbGxhLWh0 bWw6VFJVRQ0KdmVyc2lvbjoyLjENCmVuZDp2Y2FyZA0KDQo= --------------060403050804030503000606--