Re: should we use PCRE2_MATCH_INVALID_UTF ?

Stephane Chazelas <[email protected]>
Newsgroups gmane.comp.shells.zsh.devel
Message-ID <[email protected]>
2026-06-08 19:33:07 +0100, Stephane Chazelas:
[...]
> PCRE2 have a:
> 
> > PCRE2_MATCH_INVALID_UTF  Enable support for matching invalid UTF
> 
> flag.
> 
> That would not make "." match that $'\x80' byte but would align
> the behaviour with that of GNU's ERE's at least.
> 
> Would it be worth adding? Patch below.
[...]

Argh! "make test" fails with it with:

Testing PCRE multibyte with locale en_US.UTF-8
Test ./V07pcre.ztst failed: bad status 1, expected 0 from:
  pcre_compile 'cat(er(pillar)?)?'
  pcre_match -d 'the caterpillar catchment' && print $match
Error output:
(eval):pcre_match:2: error in pcre matching for the caterpillar catchment: PCRE2_MATCH_INVALID_UTF is not supported for DFA matching
Was testing: pcre_match -d

So maybe not an option if we care for that "-d"/DFA matching.

-- 
Stephane
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.