Re: Invalid UTF-8
Oren Ben-Kiki <[email protected]>
| Newsgroups | gmane.text.yaml.general |
|---|---|
| Message-ID | <1251807978.10156.279.camel@nero> |
On Mon, 2009-08-31 at 13:17 -0700, William Spitzak wrote:
> I think you are AGAIN proposing the "solution" of double-encoding the
> UTF-8. At first appearance you may think you have just made all
> non-ASCII unreadable, but if you examine the "unreadable" Unicode you
> will realize that you have just defined the encoding as ISO-8859-1. This
> is EXACTLY what I am trying to prevent!
I don't see where you get the idea I am promoting ISO-8859-1 encoding.
You keep saying that but it makes no sense and I am completely baffled
by it. For the record, and for the last one, I do not suggest YAML uses
any encoding other than Unicode (UTF-*), under any circumstance, at any
place, in any library, file, API, anywhere, *ever*. Please never
interpret anything I say as implying use of ISO-8859-1, and if it seems
to you this is what I am saying, you can take it as a given that you
misunderstood what I am trying to say. OK?
All I said was that you can use "\xNN" for any value of NN. Yes, some of
these become two-byte sequences in UTF-8. Nothing more, nothing less.
> That said, I would favor a redefinition of yaml so that \xNN means the
> raw byte (to write what \xNN does currently for 0x80-0xFF, you must
> instead write \u00NN). The main reason for this is to be compatible with
> C strings. My current plan is to patch libyaml to do exactly this.
Raw bytes are not portable (portable = we can load them in *all* YAML
libraries, even if they use UTF-16 or UTF-32 for the in-memory
representation).
> > However (1) *Not* all 16-bit
> > "\uNNNN" or 32-bit "\UNNNNNNNN" are valid and (2) this doesn't solve
> > William's problem.
>
> I believe you are referring to the idea that the "surrogate halves"
> codes of 0xD800..0xDFFF are somehow invalid and should not be allowed.
Not only that. "\uFFFF" is explicitly forbidden, for example.
> This is entirely due to a defect in UTF-16's design.
I disagree but even if this were the case, your problem is with Unicode,
not YAML. It isn't YAML's job to fix any (perceived of actual) flaws in
Unicode. The next thing you know, we'd have pressure to fix the Han
unification problems :-)
Have fun,
Oren Ben-Kiki
------------------------------------------------------------------------------
Let Crystal Reports handle the reporting - Free Crystal Reports 2008 30-Day
trial. Simplify your report design, integration and deployment - and focus on
what you do best, core application coding. Discover what's new with
Crystal Reports now. http://p.sf.net/sfu/bobj-july