Fwd: mbrtowc(3) state after an invalid sequence "undefined" or "unspecified"?
Kang-Che Sung <[email protected]>
| Newsgroups | org.kernel.vger.linux-man |
|---|---|
| Message-ID | <CADDzAfPRptY_yTxVAL-5vmqf01UMdmnMLL6f3ocTm1kTXC+keg@mail.gmail.com> |
---------- Forwarded message --------- From: Kang-Che Sung <[email protected]> Date: Thu, May 21, 2026 at 11:08 PM Subject: mbrtowc(3) state after an invalid sequence "undefined" or "unspecified"? To: Alejandro Colomar <[email protected]> Cc: <[email protected]>, <[email protected]> Hi, Alejandro (or anyone else interested), There's a discrepancy in the wording of the mbrtowc(3) function (and similarly, mbsrtowcs(3) function) between in POSIX and ISO C. It could be reported as an issue to POSIX (the Austin Group), and I am not sure if you can do that. In ISO C (I checked in both C99 and C23, in particular the N3220 draft), there's a statement that if mbrtowc() returns a (size_t)(-1) as an encoding error occurs, "the conversion state is unspecified". POSIX (see <https://pubs.opengroup.org/onlinepubs/9799919799/functions/mbrtowc.html>), for the same part it says "the conversion state is undefined". This wording difference matters when the "unspecified behavior" and "undefined behavior" are technically different. An example is how the mbstate_t object can be reused after an invalid sequence is encountered. When the state is said to be "undefined" it's implied to be not usable again (unless it is reset, e.g., by an `mbrtowc(NULL, "", 1, ps)` call). When it's "unspecified" then implementations can allow the state to be reused for certain encodings (possible for UTF-8, for example). This is something I discovered accidentally when researching the multibyte functions in the C standard library and how they work with an encoding like UTF-8.