Re: [idn] Re: FYI: BOF on Internationalized Email Addresses (IEA)

Dave Crocker <[email protected]> Wed, 29 Oct 2003 12:18:36 -0800
Newsgroups gmane.ietf.imaa
Organization Brandenburg InternetWorking
Message-ID <[email protected]>
Michel,


MS> Could you all read the Unicode spec as pointed by Mark instead of trying
MS> to recreate it (http://www.unicode.org/versions/Unicode4.0.0/ch02.pdf,
MS> section 2-4 to 2-6). There is no such a term sequences as a Unicode
MS> native representation.

Unicode documentation is not the only place that defines and discusses
computer science constructs involving data representation and encoding. For
example, since the work is being done in the IETF I suggest you look at RFC
1521. It has a nice review of the difference between native form and encoded
form. Obviously, that text is tailored to MIME, but it communicates the
general concepts adequately.

So, I apologize for trying to keep the discussion generic. Silly me. I thought
we were having an architectural discussion, rather than debating precise
Unicode terminology. In fact I was trying to be careful to express the issues
using generic computer science terminology, just to avoid a linguistic pissing
contest.

So I'm sorry that it has proven difficult for some folks to translate the
typical term "native representation" into the Unicode term "abstract
character" that is used at the beginning of section 2.4 that you cite.


MS> Abstract characters use a code point part of a
MS> set called codespace and can be referred as an encoded character within
MS> that context (paraphrasing text in 2.4).

The third paragraph of 2.4 cites the range of characters, in base 16. Note
that that range consumes 24 bits. That's the native representation I was
referring to. (When I wrote my original note, I was not sure that 24-bits was
the right number and it was not important that I get it exactly right.)


MS> Note also that Unicode favors three encoding form: UTF-8, UTF-16 and
MS> UTF-32. While ACE is as well another encoding, it does not have the same
MS> software libary support that any of the three above.

I used the word "efficient" to focus on the relevant difference between ACE
and UTF-8.  My point was bit-encoding efficiency.  You want to focus on
software development ease.  That's fine too.  However neither of these has to
do with inherent goodness or purity.  They are all encodings, so that debating
one versus another is only debating trade-offs.


MS> And now if we could get back to the subject instead of debating Unicode
MS> principle and terminology which belongs to a another list.

Ahh.  I see that you entirely missed the point I was raising.

Here's a reminder:

JCK>> If one is going to consider internationalization of email
JCK>> addresses in a way that permits them to move through the mail
JCK>> protocol in some traditional Unicode encoding  (e.g., UTF-8),
JCK>> then

DC> ...then we get to repeat the mime/esmtp debates all over again.  After all,
DC> why should we even try to learn anything from 10 years of experience.  (And
DC> no, John, I'm not directing my comment at you.)
DC>
DC> To be specific: I am not suggesting that pure utf-8 is a bad goal -- although
DC> the fact that utf-8 is, itself, a condensed representation of unicode should
DC> strike folks as a just a tad ironic, with respect to these discussions.
DC>
DC> Rather, I suggest that it be a _separate_ goal from near-term support of an
DC> edge-only enhancement for Unicode support, the same as we did for mime and
DC> IDN.

Please note that John was careful to say "traditional" and that I was careful
to respect that goal.

So I heartily agree with your suggestion that we get back to the subject.

The subject was about pursuing an infrastructure-based solution separately
from an end-point solution.

d/
--
 Dave Crocker <dcrocker-at-brandenburg-dot-com>
 Brandenburg InternetWorking <www.brandenburg.com>
 Sunnyvale, CA  USA <tel:+1.408.246.8253>