bug#81615: [PATCH] cut: speed up -w in UTF-8 locales
Pádraig Brady <[email protected]>
| Newsgroups | gmane.comp.gnu.core-utils.bugs |
|---|---|
| Message-ID | <[email protected]> |
On 13/08/2026 19:22, Iván Ezequiel Rodriguez wrote:
> In UTF-8 locales, -w (without --whitespace-delimited=trimmed) used the
> slow cut_fields_mb_any path. Route it through cut_fields_bytesearch
> with memchr2 for ASCII blanks plus a linear scan for c32issep
> multibyte blanks, with the same semantics as mcel_isblank.
>
> Preserve only truly incomplete UTF-8 tails via
> mbrtoc32 == (size_t) -2, so complete characters are not held across
> refill (cf. mbbuf_fill responsiveness). Invalid UTF-8 and
> continuation-only input remain field data and always make progress.
>
> Timings on this host (LC_ALL=C.UTF-8, cut -w -f1):
>
> mostly-ASCII ~20MiB: ~0.073s -> ~0.018s
> heavy UTF-8 ~136MiB: ~0.67s -> ~0.41s
Hmm, the implementation looks potentially quadratic to me,
as memchr2() is done for each multi-byte separator found.
This is OK for -f1 as you benchmarked but then gets worse.
I suppose you could cache uni-byte and multi-byte separators
to avoid the rescan?
Testing a pathological case shows a 70x slow down:
$ { yes $(printf 'a\xe2\x80\x83') | head -n40000 | tr -d '\n'; echo b; } > mb.in
$ time src/cut -w -f999999 mb.in
real 0m0.541s
$ git show src | patch -p1 -R
patching file src/cut.c
$ make src/cut
$ time src/cut -w -f999999 mb.in
real 0m0.007s
cheers,
Padraig