Fwd: mbrtowc(3) state after an invalid sequence "undefined" or "unspecified"?

Kang-Che Sung <[email protected]>
Newsgroups org.kernel.vger.linux-man
Message-ID <CADDzAfPRptY_yTxVAL-5vmqf01UMdmnMLL6f3ocTm1kTXC+keg@mail.gmail.com>
---------- Forwarded message ---------
From: Kang-Che Sung <[email protected]>
Date: Thu, May 21, 2026 at 11:08 PM
Subject: mbrtowc(3) state after an invalid sequence "undefined" or
"unspecified"?
To: Alejandro Colomar <[email protected]>
Cc: <[email protected]>, <[email protected]>


Hi, Alejandro (or anyone else interested),

There's a discrepancy in the wording of the mbrtowc(3) function (and
similarly, mbsrtowcs(3) function) between in POSIX and ISO C. It could
be reported as an issue to POSIX (the Austin Group), and I am not sure
if you can do that.

In ISO C (I checked in both C99 and C23, in particular the N3220
draft), there's a statement that if mbrtowc() returns a (size_t)(-1)
as an encoding error occurs, "the conversion state is unspecified".

POSIX (see <https://pubs.opengroup.org/onlinepubs/9799919799/functions/mbrtowc.html>),
for the same part it says "the conversion state is undefined".

This wording difference matters when the "unspecified behavior" and
"undefined behavior" are technically different. An example is how the
mbstate_t object can be reused after an invalid sequence is
encountered. When the state is said to be "undefined" it's implied to
be not usable again (unless it is reset, e.g., by an `mbrtowc(NULL,
"", 1, ps)` call). When it's "unspecified" then implementations can
allow the state to be reused for certain encodings (possible for
UTF-8, for example).

This is something I discovered accidentally when researching the
multibyte functions in the C standard library and how they work with
an encoding like UTF-8.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.