Re: 10646 & Unicode
Keld Jørn Simonsen <[email protected]> Wed, 7 Jan 2004 22:35:38 +0100
| Newsgroups | gmane.ietf.imaa |
|---|---|
| Message-ID | <[email protected]> |
On Tue, Jan 06, 2004 at 08:29:11AM -0500, John C Klensin wrote: > Keld, > > As you know, I took the same "reference 10646, not Unicode" > position for many years. I have given up, and suggest you do so > also. For better or worse, the fight is lost: Yes, I know. I don't think the fight is lost, but I am probably a die-hard. The last defender of the free and open world:-) Anyway the dominance of American multinational companies in the cultural specifications for IT in the 1970'es was what motivated me into all my i18n work, and I actually thinh we are better off now with some ways to be in charge of what are our own culture (eg Danish) now, than then. I am more involved than you as I am the editor of some of the ISO specifications that are not directly controlled by Unicode, but I do cooperate with the Unicode people, and although they are quite negative to this work at large, they are some of the biggest contributers to these specs too, and the ISO specs, including mine, are well aligned with the Unicode specs. > * The Unicode mapping and normalization tables, which > ISO has not incorporated, are critical to key IETF > standards, especially those that are connected to, or > depend on, Stringprep. I know, this is a hard one. I think it was the wrong decision from IETF, and I hope we can avoid that decision for the specs we are discussing here. There are viable ISO alternatives (IMHO). > * Changes in the structure and management of ISO/IEC > JTC1/SC2, and the transfer of formal responsibility for > key work from SC22 to SC2, pretty clearly put the > ISO-based management process and direction for this work > in the hands of people who will ensure that nothing of > significance is or will be approved in the ISO context > that is inconsistent with Unicode (and, probably, that > has not already reached consensus in UTC). So, for > better or worse, the concern about important divergence > is no longer important: the only difference of substance > between 10646 and friends and Unicode, in the areas of > overlap, is likely to be lag time. That is probably right, although SC2 has some independence from Unicode. Anyway it is very normal in the ISO process that the specifications have been originated somewhere else, and that ISO then standardizes it. Think of POSIX, which was made by AT&T. And many character sets were originated in ECMA, eg iso-8859-1. > * Where UTF-8 differs between the two, important > computer and software vendors, and IETF standards work, > seem to be tracking Unicode. I suspect, but have not > been following the progress of the work (largely because > I no longer consider it worth the effort), that the > differences are also a matter of lag time, i.e., that > SC2 will, sooner or later, catch up. There are in practice not much difference between ISO and Unicode UTF-8. ISO is still 31 bit, while unicode is 21 bits only. The IETF spec was thus changed from 31 bits to 21 bits under the hood, when it said that the Unicode version was the reference version. I am not so happy about that. Unicode UTF-8 then also have some more restrictions on coding of some characters like the ASCII range. I believe that ISO UTF-8 would catch up on the more restrictive spec, while I am not sure whether ISO will restrict it to 21 bits. In reality there is not allocated any characters beyond the 21 bits, neither in ISO nor in Unicode. But some clever software could use some of the extra defined user space in 10646, I remember a proposal frm Marcus Kuhn using this to represent all other charsets. > > * I wish UTC would be more orderly about handling and > identifying revisions, etc., especially to materials > that are critical to the practice of the standard like > some of the TRs. However, being in a situation of > having to reference one or two standards from one > organization and some technical reports from another is > worse than a situation in which we at least can be > reasonably assured of synchrony in the references. In > other words, those critical TRs refer back to [versions > of] the base Unicode standard, not to 10646. If we > reference 10646 as the base, and they reference the TRs, > we risk serious confusion about what is being specified > if there are any differences, even temporarily. And, > while an identification process that does not provide > concise and precise identification of versions of TRs > (or keep earlier versions readily accessible) adds > complication, the fact that ISO doesn't have those > documents as standards is a showstopper on the ISO path. Yes, that is why I advocate using the ISO references only. ISO do have specs that can do the work, AFAICT. > > This, IMO, just isn't worth going around any more. ISO and its > member bodies have decided -- in this area and others -- to > yield its authority, development, and review process to an > outside body, limiting its role to formal final review > (typically now primarily by the same people). As others have said, this is not truely the case. ISO has not yeild itd authority in these matters. Anyway it is natural in the ISO process that work is done outside ISO and then brought to ISO, as described earlier. best regards keld