Re: Bidi issues

Paul Hoffman / IMC <[email protected]>
Newsgroups gmane.ietf.imaa
Message-ID <p0600204bbb9d08eae133@[63.202.92.152]>
At 10:31 PM +0100 9/28/03, Roy Badami wrote:
>Unfortuately, as far as I can see, the additional check on single
>components is probably too complex to go into any protocol spec
>(essentially it's a regular expression).

That's sort of what we came to in earlier discussions. The Unicode 
BIDI algorithm makes things complicated enough that conformance in 
free text is hard enough, but trying to make it foolproof in 
structured text (like fully-qualified domain names or structured 
email local parts) is probably too far out there.

>   It's possible to simplify
>it, of course, at the expense of disallowing perfectly safe labels.
>For instance, if we restrict the discussion to labels that follow
>hostname rules, then it is sufficient, I think, to disallow ARABIC
>COMMA (being the only character of class ES or CS that survives
>nameprep).  But this would be a little unfair on users of arabic
>languages, given ARABIC COMMA is in fact perfectly safe when preceded
>by arabic characters.

There are probably other tradeoff rules like this. None are simple.

>I hope to post my full analysis in a week or two.

Actually, we'd like to proceed earlier than that. I think we'll turn 
in our "final" draft for consideration as an IETF standard some time 
this week. If we find a willing Area Director, it will go to a 
four-week IETF-wide last call. At the same time, Adam and Patrik and 
I will probably make our first round or editorial changes for the 
IDNA specs soon. Your analysis will certainly be useful in both of 
those discussions.

--Paul Hoffman, Director
--Internet Mail Consortium
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.