Re: on the irresponsibility of pursuing C language reform (was: [PATCH 1/2] man/man3/{mem,strn}*(): SYNOPSIS, STANDARDS: Document these as provided by <memory.h>)

Douglas McIlroy <[email protected]> Sat, 1 Aug 2026 08:39:08 -0400
Newsgroups org.kernel.vger.linux-man
Message-ID <CAKH6PiWTmDsuFOOLGqMjEBY7+GafEmnfwoy-8-ZDxoJ-14JPkA@mail.gmail.com>
Branden wrote
> I'm sure I don't need to bring to your attention what a mine field
> string/`char` sequence/memory buffer handling has been in C since the
> language's inception.

Yes, this is a property of the language, not a peculiar deficiency of the
<string.h> functions. As I see it, patching up perceived deficiencies of th=
e
functions adds complexity to the language definition and to the task of
code-reading, with little real benefit.

Doug

On Fri, Jul 31, 2026 at 7:45=E2=80=AFPM Alejandro Colomar <[email protected]> =
wrote:
>
> Hi Joseph,
>
> > Date: 2026-07-31 23:11:56+0000
> > From: Joseph Myers <[email protected]>
> >
> > On Sat, 1 Aug 2026, Alejandro Colomar wrote:
> >
> > > > In other words, they were in that header for 6 years, and it's been
> > > > implicitly obsolescent by virtue of the standard choice for the 37 =
years
> > > > since then.
> > >
> > > Yes.  And I'm trying to revert the implicit obsolescence.  Obsolescen=
ce
> > > isn't a one-way process.  Sometimes, evidence shows up, and the obsol=
ete
> > > feature must come back for $reasons.
> >
> > When it's been obsolescent for 37 years, bringing back a header with th=
e
> > same name is just going to confuse people with 37 years of past
> > information saying it's obsolescent (normally if one source says "use X=
"
> > and another says "X is obsolescent", you can reliably assume that "X is
> > obsolescent" is the more recent information, even without 37 years of
> > history involved).
>
> This is highly speculative.
>
> If the official documentation says use <memory.h>, use <memory.h>.
> Moreover, I've never seen anything saying <memory.h> is obsolescent.
> It's been only implicitly obsolescent, and most programmers don't even
> know <memory.h> exists at all.  FWIW, I've sometimes seen it in the
> system headers, and eventually wondered why it was there, but nothing
> else.  Thus, there's no contradicting information about it.
>
> > Any reasonable change there would involve a new
> > header, say <strnpad.h> for strncpy and strncat,
>
> I think that misses that strn*() are still relatively misunderstood.
> It's only because they don't write that they aren't as dangerous.
>
> But we would be much better with strprefix/suffix() and other related
> functions in <string.h> to even further remove uses of strn*()
> functions.
>
> Moving the entirety of strn*() might be less urgent/important than
> moving strncpy/cat(), but it'd still be good to move them.  They are all
> about handling [[gnu::nonstring]]s.
>
> > rather than resurrecting
> > a very old one.
>
> I think the very old one is much better, because it can be used by
> virtually all programs already, since it already is there in all systems
> that matter, and already provides the functions were moving there.
> Apart from documentation, there's no real change.
>
> On the other hand, if we were adding a new header, it'd have to be
> wrapped in #if __has_include(), and wouldn't be used until 10 years from
> now or so.
>
> > > The solution of moving both mem*() and strn*() to <memory.h> and leav=
ing
> > > just str*() in <string.h> is a consistent one, because <string.h> the=
n
> > > remains strictly for string APIs, and <memory.h> is for the rest of b=
yte
> > > handling.
> >
> > It's inconsistent with how people have understood C ever since it was
> > standardized.  Changing the header memcpy is in is just as ridiculous a=
t
> > this point as the proposal there once was to obsolesce NULL.
>
> I disagree.  Obsolescing NULL is a deeply breaking change.
> Standardizing an existing header file is just like standardizing an
> existing function.  Nobody is saying <string.h> won't provide memcpy(3).
> My proposal to C2y says that <string.h> includes <memory.h>, so nothing
> changes.  All of the existing understanding remains valid (although less
> preferred).
>
> > > It wouldn't be reasonable to move strn*() to <memory.h>, and then lea=
ve
> > > mem*() in <string.h>, of course.
> > >
> > > Similarly, it wouldn't be reasonable to more strncpy/cat() to <memory=
.h>
> > > and leave the rest of strn*() and all of mem*() in <string.h>.
> >
> > On the contrary, it's only the functions for null-padded fixed-width
> > buffers that are niche functions causing confusion, compared to all the
> > rest of the functions in <string.h> for which it's a very well-establis=
hed
> > and well-understood location.  Some others like memccpy are *obscure*, =
but
> > not confusing in the same way.
>
> Did you read my alx-0096 paper?  It shows how n2349 --which introduced
> memccpy(3) in C23-- is full of UB in examples introduced as
>
>         To avoid the risk of buffer overflow, the appropriate bound
>         needs to be determined for each call and provided as an
>         argument.
>
> If memccpy(3) was standardized for reducing buffer overflows (which is
> implied, but not very explicitly said in the paper), it's ironic that
> the examples of how it's supposed to be used invoke UB.  If the author
> of the paper can't avoid UB, you can guess it's a bomb as bad as
> strncpy(3) once was.  And I'm not saying that memccpy(3) is bad, not
> that strncpy(3) is bad.  I use strncpy(3) and know it's fine, and have
> been recently shown that memccpy(3) is actually good for implementing
> fgets(3).  But it's not a function for copying strings, and n2349 is the
> ultimate proof.  It is just as bad as strncpy(3) if misused for copying
> strings.
>
> I'll paste here part of alx-0096, which clarifies how n2349 invokes UB.
>
>         Reading N2349 further, one finds an example of copying with
>         truncation:
>
>                 char *p =3D memccpy (d, s1, '\0', dsize);
>                 dsize -=3D (p - d - 1);
>                 memccpy (p - 1, s2, '\0', dsize);
>
>         This is more prone to bugs than the case above, and more than
>         strncpy(3).  Anyone suggesting to use this to improve safety
>         compared to strncpy(3), please, please, explain to me how they
>         think this can be safe in any way.
>
>         In fact, the code above is completely bogus, because if the
>         string is truncated, p will be NULL, and it invokes UB in
>         line 2.  See how it was predictably prone to bugs?  :)
>
>         At the bottom of the N2349 paper, there's a more correct --and
>         also more worrying-- example of how memccpy(3) could be used for
>         copying strings with truncation.  This shows how terrible
>         memccpy(3) is for copying strings:
>
>                 char *p =3D memccpy (d, s1, '\0', dsize);
>                 if (p) {
>                   --p;
>                   p =3D memccpy (p, "/", '\0', dsize - (p - d));
>                   if (p) {
>                     --p;
>                     p =3D memccpy (p, s2, '\0', dsize - (p - d));
>                   }
>                 }
>                 if (!p)
>                   d[dsize - 1] =3D '\0';
>
>         I think I don't need to explain what can go wrong in such
>         unreadable, brittle, and complex code.
>
>         ...
>
>         Using a more suitable function --similar to POSIX's stpcpy(3)--,
>         this could be written much more safely:
>
>                 char *p =3D d;
>                 char *e =3D d + dsize;
>                 p =3D stpecpy(p, e, s1);
>                 p =3D stpecpy(p, e, "/");
>                 p =3D stpecpy(p, e, s2);
>                 if (p =3D=3D NULL)
>                         goto trunc;  // The string was truncated
>
>         Here's how stpecpy() can be implemented for this:
>
>                 char *
>                 stpecpy(char *dst, const char *end, const char *restrict =
src)
>                 {
>                         ssize_t  dlen;
>
>                         if (dst =3D=3D NULL)
>                                 return NULL;
>
>                         dlen =3D strtcpy(dst, src, end - dst);
>                         if (dlen =3D=3D -1)
>                                 return NULL;
>
>                         return dst + dlen;
>                 }
>
>
> Cheers,
> Alex
>
> --
> <https://www.alejandro-colomar.es>