Re: Bidi issues
Paul Hoffman / IMC <[email protected]>
| Newsgroups | gmane.ietf.imaa |
|---|---|
| Message-ID | <p0600204bbb9d08eae133@[63.202.92.152]> |
At 10:31 PM +0100 9/28/03, Roy Badami wrote: >Unfortuately, as far as I can see, the additional check on single >components is probably too complex to go into any protocol spec >(essentially it's a regular expression). That's sort of what we came to in earlier discussions. The Unicode BIDI algorithm makes things complicated enough that conformance in free text is hard enough, but trying to make it foolproof in structured text (like fully-qualified domain names or structured email local parts) is probably too far out there. > It's possible to simplify >it, of course, at the expense of disallowing perfectly safe labels. >For instance, if we restrict the discussion to labels that follow >hostname rules, then it is sufficient, I think, to disallow ARABIC >COMMA (being the only character of class ES or CS that survives >nameprep). But this would be a little unfair on users of arabic >languages, given ARABIC COMMA is in fact perfectly safe when preceded >by arabic characters. There are probably other tradeoff rules like this. None are simple. >I hope to post my full analysis in a week or two. Actually, we'd like to proceed earlier than that. I think we'll turn in our "final" draft for consideration as an IETF standard some time this week. If we find a willing Area Director, it will go to a four-week IETF-wide last call. At the same time, Adam and Patrik and I will probably make our first round or editorial changes for the IDNA specs soon. Your analysis will certainly be useful in both of those discussions. --Paul Hoffman, Director --Internet Mail Consortium