Re: 10646 & Unicode

John C Klensin <[email protected]> Tue, 06 Jan 2004 08:29:11 -0500
Newsgroups gmane.ietf.imaa
Message-ID <[email protected]>
Keld,

As you know, I took the same "reference 10646, not Unicode"=20
position for many years.  I have given up, and suggest you do so=20
also.  For better or worse, the fight is lost:

	* The Unicode mapping and normalization tables, which
	ISO has not incorporated, are critical to key IETF
	standards, especially those that are connected to, or
	depend on, Stringprep.
=09
	* Changes in the structure and management of ISO/IEC
	JTC1/SC2, and the transfer of formal responsibility for
	key work from SC22 to SC2, pretty clearly put the
	ISO-based management process and direction for this work
	in the hands of people who will ensure that nothing of
	significance is or will be approved in the ISO context
	that is inconsistent with Unicode (and, probably, that
	has not already reached consensus in UTC).  So, for
	better or worse, the concern about important divergence
	is no longer important: the only difference of substance
	between 10646 and friends and Unicode, in the areas of
	overlap, is likely to be lag time.
=09
	* Where UTF-8 differs between the two, important
	computer and software vendors, and IETF standards work,
	seem to be tracking Unicode.  I suspect, but have not
	been following the progress of the work (largely because
	I no longer consider it worth the effort), that the
	differences are also a matter of lag time, i.e., that
	SC2 will, sooner or later, catch up.
=09
	* I wish UTC would be more orderly about handling and
	identifying revisions, etc., especially to materials
	that are critical to the practice of the standard like
	some of the TRs.  However, being in a situation of
	having to reference one or two standards from one
	organization and some technical reports from another is
	worse than a situation in which we at least can be
	reasonably assured of synchrony in the references.  In
	other words, those critical TRs refer back to [versions
	of] the base Unicode standard, not to 10646.  If we
	reference 10646 as the base, and they reference the TRs,
	we risk serious confusion about what is being specified
	if there are any differences, even temporarily.  And,
	while an identification process that does not provide
	concise and precise identification of versions of TRs
	(or keep earlier versions readily accessible) adds
	complication, the fact that ISO doesn't have those
	documents as standards is a showstopper on the ISO path.

This, IMO, just isn't worth going around any more.  ISO and its=20
member bodies have decided -- in this area and others -- to=20
yield its authority, development, and review process to an=20
outside body, limiting its role to formal final review=20
(typically now primarily by the same people).  I don't like that=20
outcome in this area, and like it still less in others.  But=20
insisting on recognition of documents just because they have=20
gone through that weakened process doesn't seem to me to enhance=20
anything other than ISO's dwindling reputation.   So, in this=20
area and others, perhaps the best thing other standards bodies=20
can do is to say "Ok, the key standard is really being developed=20
in another place, and ISO is rubber-stamping and reprinting=20
(usually at high price and long delay) its key content.  We=20
should reference the other documents unless there is obvious=20
value-added (and, in this case, comprehensive coverage) in the=20
ISO work.   And assignment of a number and reprinting in another=20
format is _not_, at least in my opinion, sufficient value-added=20
to justify the time lag and risk of non-synchrony between, e.g.,=20
the coding standard and the normalization rules.

I don't like this outcome but the wounds in ISO's feet are=20
self-inflicted and we can't ignore them any more in this area.

regards,
      john


--On Tuesday, 06 January, 2004 11:07 +0100 Keld J=F8rn Simonsen=20
<[email protected]> wrote:

>
> On Mon, Jan 05, 2004 at 06:56:37PM -0800, Mark Davis wrote:
>>
>> Yes, Keld's statement is completely inaccurate.
>
> True, there are a lot of information on character properties
> etc in unicode. Some of that info can be found in other ISO
> standards such as  ISO 14651 on sorting and ISO TR 14652 with
> an enhanced POSIX like  locales.
>
> What I meant is that the character table that defines the
> charset in IETF terms, is almost the same, and the UTF-8 spec
> is almost the same, and these are the specs that we probably
> will be using in  our enhanced mail spec.
>
> Best regards
> keld
>
>> Mark
>> __________________________________
>> http://www.macchiato.com
>> ???
>> ?????????????????????????????????????????????????????????????
>> ?? ???
>>
>> ----- Original Message -----
>> From: "Adam M. Costello"
>> <[email protected]> To:
>> <[email protected]>
>> Sent: Mon, 2004 Jan 05 17:15
>> Subject: 10646 & Unicode
>>
>>
>> >
>> > Keld J=F8rn Simonsen <[email protected]> wrote:
>> >
>> > > 10646, not unicode, ... (as the technical differences are
>> > > almost negligible).
>> >
>> > I'm under the impression that the technical differences are
>> > quite great, because Unicode includes a great deal of
>> > technical information that is absent from 10646 (but
>> > whatever does exist in 10646 is identical with or at least
>> > compatible with Unicode).
>> >
>> > I haven't seen 10646 (it's too expensive), but does it
>> > include much more than a code chart and specs for UTF-8 and
>> > UTF-16?  Unicode includes a lot of character properties and
>> > algorithms for doing useful things based on those
>> > properties.  See sections C.4 and C.7 of the Unicode
>> > standard:
>> >
>> > http://www.unicode.org/versions/Unicode4.0.0/appC.pdf
>> >
>> > Does 10646 include normalization forms (NFC, etc)?  Those
>> > are likely to be important in any effort to use the UCS in
>> > identifiers (like email addresses and web links).  Does
>> > 10646 include case folding data?  That is needed in any
>> > effort to use the UCS in case-insensitive identifiers.
>> >
>> > AMC
>> >