Re: 10646 & Unicode
James Seng <[email protected]> Sat, 10 Jan 2004 01:15:44 +0800
| Newsgroups | gmane.ietf.imaa |
|---|---|
| Message-ID | <[email protected]> |
These discussion about 10646 vs. Unicode is interesting but shouldn't we go back to the topic of IMAA? -James Seng Martin Duerst wrote: > > At 22:35 04/01/07 +0100, Keld J甚n Simonsen wrote: > >> On Tue, Jan 06, 2004 at 08:29:11AM -0500, John C Klensin wrote: > > >> > * The Unicode mapping and normalization tables, which >> > ISO has not incorporated, are critical to key IETF >> > standards, especially those that are connected to, or >> > depend on, Stringprep. >> >> I know, this is a hard one. I think it was the wrong decision from IETF, >> and I hope we can avoid that decision for the specs we are discussing >> here. There are viable ISO alternatives (IMHO). > > > I would be interested to know which alternatives you are thinking of. > > >> > * Where UTF-8 differs between the two, important >> > computer and software vendors, and IETF standards work, >> > seem to be tracking Unicode. I suspect, but have not >> > been following the progress of the work (largely because >> > I no longer consider it worth the effort), that the >> > differences are also a matter of lag time, i.e., that >> > SC2 will, sooner or later, catch up. >> >> There are in practice not much difference between ISO and Unicode UTF-8. >> ISO is still 31 bit, while unicode is 21 bits only. The IETF spec was >> thus changed from 31 bits to 21 bits under the hood, when it said that >> the Unicode version was the reference version. I am not so happy >> about that. Unicode UTF-8 then also have some more restrictions on >> coding of some characters like the ASCII range. I believe that ISO UTF-8 >> would catch up on the more restrictive spec, > > > That would be great. It is important for the IETF for security > reasons. > > >> while I am not sure whether >> ISO will restrict it to 21 bits. In reality there is not allocated any >> characters beyond the 21 bits, neither in ISO nor in Unicode. >> But some clever software could use some of the extra defined user space >> in 10646, I remember a proposal frm Marcus Kuhn using this to represent >> all other charsets. > > > That's the main problem with huge amounts of empty space: > Somebody will come up with an idea of how to use it that they > think is very clever, but is bothering everybody else. > Restricting the amount of space available, even if artificially, > has the beneficial effect to guide people to efficient use of > space; too much space available just leads to 'sprawl'. > In this sense, I think the limitation of the original Unicode > approach to 16 bits was a good thing, and the current limitation > to 21 bits hopefully at least provides a bit of pressure in the > right direction, even if things look still really very empty > at the moment, and it will take at the very least this century > to fill up this space. > > As for the IETF, I don't think that a proposal to reflect > all legacy charsets inside the UCS would meet the IETF's > interoperability goals in any way. > > > Regards, Martin. > >