RE: [idn] Re: FYI: BOF on Internationalized Email Addresses (IEA)

"Michel Suignard" <[email protected]> Wed, 29 Oct 2003 10:48:20 -0800
Newsgroups gmane.ietf.imaa
Message-ID <84DD35E3DD87D5489AC42A59926DABE9046E6B8B@WIN-MSG-10.wingroup.windeploy.ntdev.microsoft.com>
Could you all read the Unicode spec as pointed by Mark instead of trying
to recreate it (http://www.unicode.org/versions/Unicode4.0.0/ch02.pdf,
section 2-4 to 2-6). There is no such a term sequences as a Unicode
native representation. Abstract characters use a code point part of a
set called codespace and can be referred as an encoded character within
that context (paraphrasing text in 2.4).

It just happens that one of the encoding form defined by Unicode
(UTF-32) is a simpler mapping to and from the original Unicode
codespace.

Note also that Unicode favors three encoding form: UTF-8, UTF-16 and
UTF-32. While ACE is as well another encoding, it does not have the same
software libary support that any of the three above.

And now if we could get back to the subject instead of debating Unicode
principle and terminology which belongs to a another list.

Michel Suignard 

-----Original Message-----
From: [email protected] [mailto:[email protected]]
On Behalf Of Dave Crocker

John,

JC> That's only true if you take the position that there are no 
JC> native/direct/raw encodings of Unicode.

Oh?  You mean that Unicode does not fit directly -- ie, with no special
encoding rules -- into 32 bits, or 24 bits, or somesuch.

You mean that Unicode does not need special rules to stuff it into 8
bits, and another set of rules to stuff it into 16 bits?

Because if the answer is that yes it does -- and the answer _is_ yes it
does
-- then my point stands.

That's the difference between native representation, versus "encoding".

d/
--
 Dave Crocker <dcrocker-at-brandenburg-dot-com>  Brandenburg
InternetWorking <www.brandenburg.com>  Sunnyvale, CA  USA
<tel:+1.408.246.8253>