Re: Unicode strings and invalid codepoints

Henrik Grubbström <[email protected]>
Newsgroups gmane.comp.lang.pike.user
Organization Roxen Internet Software AB
Message-ID <[email protected]>
On Tue, 17 Mar 2015, Chris Angelico wrote:

> On Tue, Mar 17, 2015 at 4:30 AM, Henrik Grubbström <[email protected]> wrote:
>> On Mon, 9 Mar 2015, Chris Angelico wrote:
>>
>>> There's a bit of a discussion happening on python-list about whether
>>> or not it should be legal to have codepoints like U+D800 in Unicode
>>> strings. Currently, both Python and Pike permit them, but reject them
>>> if you try to, for example, convert to UTF-8. But a suggestion has
>>> been made that the mere presence of \uD800 in a string literal should
>>> be a syntax error, and I'm wondering: Has anyone considered and
>>> rejected this, or is it simply something that nobody's thought to
>>> disallow?
>>
>>
>> Well, as U+D800 is legal in UTF-16 strings (which BTW is why it is
>> reserved), I don't see a reason to prohibit it.
>>
>>> Are there situations in which it's necessary to be able to store these
>>> kinds of noncharacters in a string?
>>
>>
>> Yes.
>
> UTF-16 is represented with individual bytes, though:
>
>> string_to_unicode("\U00012345");
> (1) Result: "\330\b\337E"
>> String.string2hex(string_to_unicode("\U00012345"));
> (2) Result: "d808df45"

Well, take UCS-2 then.

-- 
Henrik Grubbström					[email protected]
Roxen Internet Software AB
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.