Re: Some thoughts on Unicode.

Aidan Kehoe <[email protected]> Wed, 1 Sep 2004 09:15:37 +0100
Newsgroups gmane.emacs.xemacs.mule,gmane.emacs.xemacs.beta
Message-ID <[email protected]>
 Ar an ch=E9ad l=E1 de m=ED M=E9an F=F3mhair, scr=EDobh Stephen J. Turnbu=
ll:=20

 >     Aidan> I know [latin-unity is] a kludge--I still think its API
 >     Aidan> should be preserved. An efficient way to ask "can this
 >     Aidan> buffer be encoded in iso-8859-1 without losing data" would
 >     Aidan> be, and is, worthwhile.
 >=20
 > What's wrong with Just Doing It, and dealing with the error if data
 > loss would occur?=20

If you're writing lisp code that works on current XEmacs, you don't see a=
n
error when data loss occurs. If lisp code is going to work on 21.4 (sure,
you don't ever use it--but I was up until the other day, because this
21.5.17 dropped characters for me when I was accessing it via Apple's
Terminal. I've switched over to XTerm (changed machine and OS) and this
latter is sufficiently slow that XEmacs has time to catch up. Mostly.), a=
nd
every version of 21.5 up to now, then it has to assume that data loss wil=
l
occur and it won't know about it.

 > [...] Note that the problem that latin-unity addresses is not efficien=
cy,
 > but rather that our current ISO-8859-X, X !=3D 1, coding systems simpl=
y use
 > ISO 2022 extensions to encode non-ISO-8859-X characters (typically not
 > what is wanted), while ISO-8859-1 (aka binary) just throws away anythi=
ng
 > that isn't ISO-8859-1 and replaces it with ~.
 >=20
 > With that context, do you still see a need for a testing API?

_If_ encode-coding-region signalled error when it saw something it couldn=
't
encode, then sure, the approach you describe would absolutely make sense.=
 It
doesn't at the moment, and I don't see any projections for when it will.

For people in the Lisp world, writing portable code, testing is what they
have to do. (Or, like many of them do, ignore it.) FSF Emacs does the
degrade-to-iso-2022 thing too; which _is_ an error outside East Asia. _No=
_
other program (besides the web browser) understands it. Users, in general=
,
look at the escape sequences and have no idea what to do with them, besid=
es
delete them.

 > Most politically important, though, is Han disunity ;-).  In general,
 > different languages will prefer different fonts, and we should cater
 > to that, not by changing faces, but by providing faces that map
 > languages (not charsets) to fonts.=20

Right. Sounds cool. (Though I don't imagine there will be much metadata y=
ou
can trust to differentiate between French and German--the differing Han
character sets provide that metadata, there's no obligation to do provide=
 it
in any form in Western Europe. Or in the Americas, for English and Spanis=
h.
And heuristic algorithms going wrong are very, very sucky. )

--=20
Like the early Christians, Marx expected the millennium very soon; like
their successors, his have been disappointed--once more, the world has sh=
own
itself recalcitrant to a tidy formula embodying the hopes of some section=
 of
mankind. (Russell)