Re: 10646 & Unicode
"Mark Davis" <[email protected]> Tue, 6 Jan 2004 07:27:59 -0800
| Newsgroups | gmane.ietf.imaa |
|---|---|
| Message-ID | <008401c3d469$aafc1e50$6901a8c0@DAVIS1> |
Now I am one of the first to say that referencing the Unicode standards i= s the right direction, in no small part because of the wealth of specifications= and data over and above the character repertoire itself, items that are simpl= y not found in ISO standards. However, your message goes a bit too far in that direction. ISO does not = merely rubber-stamp the Unicode standard; there is an extremely successful degre= e of cooperation between the Unicode consortium and the ISO subcommittees. Thi= s has allowed the resulting character repertoire to be more thoroughly vetted t= han if either organization had done it alone. In terms of collation, cooperation between the consortium and ISO has pro= duced synchronization between the UCA and ISO 14651 on a base level (although t= he UCA does add substantial capabilities). While we have not been happy at the s= peed with which ISO/IEC SC22 was able to address the issues involved in collat= ion, the main reason for the transfer of authority of 14651 to ISO/IEC SC2 is = that at this point the bulk of the work is in refining the approaches to more specialized scripts, and SC2 has many, many times the representation in t= hese areas. As to UTF-8, there is no effective difference between ISO, Unicode, or http://www.ietf.org/rfc/rfc3629.txt remaining. >* I wish UTC would be more orderly about handling and identifying revisions, etc., especially to materials that are critical to the practice of the standard like some of the TRs. The UTC has a well-defined process for producing and identifying all revi= sions to all specifications. You may find the FAQ at http://www.unicode.org/faq/reports_process.html useful. If you have any c= oncerns about the processes that it uses, or suggestions for improvements, please= let me know and I will communicate those to the committee. Mark __________________________________ http://www.macchiato.com =E2=96=BA =E0=A4=B6=E0=A4=BF=E0=A4=B7=E0=A5=8D=E0=A4=AF=E0=A4=BE=E0=A4=A6= =E0=A4=BF=E0=A4=9A=E0=A5=8D=E0=A4=9B=E0=A5=87=E0=A4=A4=E0=A5=8D=E0=A4=AA=E0= =A4=B0=E0=A4=BE=E0=A4=9C=E0=A4=AF=E0=A4=AE=E0=A5=8D =E2=97=84 ----- Original Message -----=20 From: "John C Klensin" <[email protected]> To: "Keld J=C3=B8rn Simonsen" <[email protected]>; "Mark Davis" <mark.davis@j= tcsv.com> Cc: "IETF IMAA list" <[email protected]> Sent: Tue, 2004 Jan 06 05:29 Subject: Re: 10646 & Unicode Keld, As you know, I took the same "reference 10646, not Unicode" position for many years. I have given up, and suggest you do so also. For better or worse, the fight is lost: * The Unicode mapping and normalization tables, which ISO has not incorporated, are critical to key IETF standards, especially those that are connected to, or depend on, Stringprep. * Changes in the structure and management of ISO/IEC JTC1/SC2, and the transfer of formal responsibility for key work from SC22 to SC2, pretty clearly put the ISO-based management process and direction for this work in the hands of people who will ensure that nothing of significance is or will be approved in the ISO context that is inconsistent with Unicode (and, probably, that has not already reached consensus in UTC). So, for better or worse, the concern about important divergence is no longer important: the only difference of substance between 10646 and friends and Unicode, in the areas of overlap, is likely to be lag time. * Where UTF-8 differs between the two, important computer and software vendors, and IETF standards work, seem to be tracking Unicode. I suspect, but have not been following the progress of the work (largely because I no longer consider it worth the effort), that the differences are also a matter of lag time, i.e., that SC2 will, sooner or later, catch up. * I wish UTC would be more orderly about handling and identifying revisions, etc., especially to materials that are critical to the practice of the standard like some of the TRs. However, being in a situation of having to reference one or two standards from one organization and some technical reports from another is worse than a situation in which we at least can be reasonably assured of synchrony in the references. In other words, those critical TRs refer back to [versions of] the base Unicode standard, not to 10646. If we reference 10646 as the base, and they reference the TRs, we risk serious confusion about what is being specified if there are any differences, even temporarily. And, while an identification process that does not provide concise and precise identification of versions of TRs (or keep earlier versions readily accessible) adds complication, the fact that ISO doesn't have those documents as standards is a showstopper on the ISO path. This, IMO, just isn't worth going around any more. ISO and its member bodies have decided -- in this area and others -- to yield its authority, development, and review process to an outside body, limiting its role to formal final review (typically now primarily by the same people). I don't like that outcome in this area, and like it still less in others. But insisting on recognition of documents just because they have gone through that weakened process doesn't seem to me to enhance anything other than ISO's dwindling reputation. So, in this area and others, perhaps the best thing other standards bodies can do is to say "Ok, the key standard is really being developed in another place, and ISO is rubber-stamping and reprinting (usually at high price and long delay) its key content. We should reference the other documents unless there is obvious value-added (and, in this case, comprehensive coverage) in the ISO work. And assignment of a number and reprinting in another format is _not_, at least in my opinion, sufficient value-added to justify the time lag and risk of non-synchrony between, e.g., the coding standard and the normalization rules. I don't like this outcome but the wounds in ISO's feet are self-inflicted and we can't ignore them any more in this area. regards, john --On Tuesday, 06 January, 2004 11:07 +0100 Keld J=C3=B8rn Simonsen <[email protected]> wrote: > > On Mon, Jan 05, 2004 at 06:56:37PM -0800, Mark Davis wrote: >> >> Yes, Keld's statement is completely inaccurate. > > True, there are a lot of information on character properties > etc in unicode. Some of that info can be found in other ISO > standards such as ISO 14651 on sorting and ISO TR 14652 with > an enhanced POSIX like locales. > > What I meant is that the character table that defines the > charset in IETF terms, is almost the same, and the UTF-8 spec > is almost the same, and these are the specs that we probably > will be using in our enhanced mail spec. > > Best regards > keld > >> Mark >> __________________________________ >> http://www.macchiato.com >> ??? >> ????????????????????????????????????????????????????????????? >> ?? ??? >> >> ----- Original Message ----- >> From: "Adam M. Costello" >> <[email protected]> To: >> <[email protected]> >> Sent: Mon, 2004 Jan 05 17:15 >> Subject: 10646 & Unicode >> >> >> > >> > Keld J=C3=B8rn Simonsen <[email protected]> wrote: >> > >> > > 10646, not unicode, ... (as the technical differences are >> > > almost negligible). >> > >> > I'm under the impression that the technical differences are >> > quite great, because Unicode includes a great deal of >> > technical information that is absent from 10646 (but >> > whatever does exist in 10646 is identical with or at least >> > compatible with Unicode). >> > >> > I haven't seen 10646 (it's too expensive), but does it >> > include much more than a code chart and specs for UTF-8 and >> > UTF-16? Unicode includes a lot of character properties and >> > algorithms for doing useful things based on those >> > properties. See sections C.4 and C.7 of the Unicode >> > standard: >> > >> > http://www.unicode.org/versions/Unicode4.0.0/appC.pdf >> > >> > Does 10646 include normalization forms (NFC, etc)? Those >> > are likely to be important in any effort to use the UCS in >> > identifiers (like email addresses and web links). Does >> > 10646 include case folding data? That is needed in any >> > effort to use the UCS in case-insensitive identifiers. >> > >> > AMC >> >