Re: URL internationalization!

"Martin J. Duerst" <[email protected]> Wed, 26 Feb 1997 16:11:07 +0100 (MET)
Newsgroups gmane.ietf.url
Message-ID <Pine.SUN.3.95q.970226155156.245J-100000@enoshima>
Masataka,

On Tue, 25 Feb 1997, you wrote:

> Martin;
> 
> > > And, ISO 10646 can't handle multiple scripts of Hanzi and Kanji.
> > 
> > This issue has been mentionned before. Of course, ISO 10646
> > can handle CJK(V) ideographs
> 
> What we, Japanese, daily use is not "CJK(V) ideograph" script
> but Kanji-Kana-majiri script.

I never spoke about "CJK(V) ideograph *script*", I only spoke
of individual characters. Some people use the term "script" for
CJK(V) ideographs, others, such as you, prefer not to do so.
For the discussion here, it's largely irrelevant, because
it is the individual characters that matter.

You started the discussion with the terms "Hanzi script" and
"Kanji script", but please note that the majority of Japanese
will use the word "Kanji" also for the characters used in
China, and the average Chinese will use the word "Hanzi" also
for the ideographic characters in Japan, and so on. If you
want to speak about Kanji-Kana-majiri-bun, then please say
so up front.


> As ISO 2022 based encoding already supports Kanji-Kana-majiri
> script in fully internationalized way but ISO 10646 can't, there
> is no further discussion possible that ISO 10646 is
> internationalized.

How come that you claim ISO 10646 can't? It can represent all
the characters in the Kanji-Kana-majiri "script". What else
would be needed?


> BTW, JIS X 0208 itself already contains too many similar
> characters that exact match of code points is useless for real
> world search.

Agreed. Search engines have to take this into account.
But current URLs allow case distinction, and search engines
also have to take this into account, so it's not a problem.


Regards,	Martin.