Re: [docs] [PATCH v2 5/5] tools: Add check-confusables pre-commit hook

"Antonin Godard" <[email protected]> Tue, 28 Jul 2026 09:53:12 +0200
Newsgroups org.yoctoproject.lists.docs
Message-ID <[email protected]>
Hi,

On Tue Jul 28, 2026 at 9:23 AM CEST, Niko Mauno via lists.yoctoproject.org =
wrote:
> From: Niko Mauno <[email protected]>
>
> Add a check-confusables script, in the same fashion as
> check-glossaries, that scans the documentation .rst sources for
> non-ASCII "confusable" characters (curly quotes, en/em dashes,
> horizontal ellipsis, non-breaking and zero-width spaces, etc.) and
> reports each occurrence with its location and suggested ASCII
> replacement, exiting non-zero if any are found. This guards against
> the class of breakage fixed in the preceding commit, e.g. curly quotes
> causing recipe ParseErrors.
>
> Legitimate non-ASCII such as box-drawing characters used in directory
> trees, accented letters in contributor names and CJK characters are
> intentionally left untouched.
>
> Wire it up both as a local pre-commit hook and in the Makefile "checks"
> target, alongside check-glossaries.
>
> Suggested-by: Quentin Schulz <[email protected]>
> Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
> Signed-off-by: Niko Mauno <[email protected]>
> ---
>  .pre-commit-config.yaml               |  5 ++
>  documentation/Makefile                |  1 +
>  documentation/tools/check-confusables | 77 +++++++++++++++++++++++++++
>  3 files changed, 83 insertions(+)
>  create mode 100755 documentation/tools/check-confusables
>
> diff --git a/.pre-commit-config.yaml b/.pre-commit-config.yaml
> index f2b73a481..d6008b609 100644
> --- a/.pre-commit-config.yaml
> +++ b/.pre-commit-config.yaml
> @@ -6,3 +6,8 @@ repos:
>          entry: ./documentation/tools/check-glossaries
>          language: python
>          pass_filenames: false
> +      - id: check-confusables
> +        name: Check for non-ASCII confusable characters
> +        entry: ./documentation/tools/check-confusables
> +        language: python
> +        pass_filenames: false

I think we should pass filenames and check only the modified files, no need=
 to
checking the entire tree each time?

> diff --git a/documentation/Makefile b/documentation/Makefile
> index fe0574537..87a6f8a8b 100644
> --- a/documentation/Makefile
> +++ b/documentation/Makefile
> @@ -37,6 +37,7 @@ clean:
> =20
>  checks:
>  	$(SOURCEDIR)/tools/check-glossaries --docs-dir $(SOURCEDIR)
> +	$(SOURCEDIR)/tools/check-confusables --docs-dir $(SOURCEDIR)
> =20
>  stylecheck:
>  	vale sync
> diff --git a/documentation/tools/check-confusables b/documentation/tools/=
check-confusables
> new file mode 100755
> index 000000000..f79ee046c
> --- /dev/null
> +++ b/documentation/tools/check-confusables
> @@ -0,0 +1,77 @@
> +#!/usr/bin/env python3
> +
> +import argparse
> +import sys
> +
> +from pathlib import Path
> +
> +
> +def parse_arguments() -> argparse.Namespace:
> +    parser =3D argparse.ArgumentParser(
> +        description=3D"Check documentation sources for non-ASCII typogra=
phic "
> +                    "characters that should be plain ASCII")
> +
> +    parser.add_argument("-d", "--docs-dir",
> +                        type=3DPath,
> +                        default=3DPath(__file__).resolve().parent.parent=
,
> +                        help=3D"Path to documentation/ directory in yoct=
o-docs")
> +
> +    return parser.parse_args()
> +
> +
> +# Map of "confusable" characters that are frequently introduced by edito=
rs,
> +# word processors or copy-pasting, to their plain ASCII replacement. The=
se
> +# look almost identical to regular ASCII but break tooling, e.g. a curly
> +# quote in a recipe example causes:
> +#
> +#   ERROR: ParseError ...: unparsed line: 'RDEPENDS:${PN} =3D "foo"'

I don't see the curly quote here

> +#
> +# Only these characters are flagged; legitimate non-ASCII such as box-dr=
awing
> +# characters used in directory trees, accented letters in contributor na=
mes
> +# and CJK characters are intentionally left alone.
> +confusables =3D {
> +    "=E2=80=98": "'",      # LEFT SINGLE QUOTATION MARK
> +    "=E2=80=99": "'",      # RIGHT SINGLE QUOTATION MARK
> +    "=E2=80=9C": '"',      # LEFT DOUBLE QUOTATION MARK
> +    "=E2=80=9D": '"',      # RIGHT DOUBLE QUOTATION MARK
> +    "=E2=80=B2": "'",      # PRIME
> +    "=E2=80=B3": '"',      # DOUBLE PRIME
> +    "=E2=80=93": "-",      # EN DASH
> +    "=E2=80=94": "--",     # EM DASH
> +    "=E2=80=90": "-",      # HYPHEN
> +    "=E2=80=91": "-",      # NON-BREAKING HYPHEN
> +    "=E2=88=92": "-",      # MINUS SIGN
> +    "=E2=80=A6": "...",    # HORIZONTAL ELLIPSIS
> +    "=C2=A0": " ",      # NO-BREAK SPACE

Actually, since the `tree` command uses no-break spaces, I think it's OK to=
 keep
them otherwise we'd have to convert each tree output, which also might be
confusing.

> +    "=E2=80=AF": " ",      # NARROW NO-BREAK SPACE
> +    "=E2=80=8B": "",       # ZERO WIDTH SPACE
> +    "=EF=BB=BF": "",       # ZERO WIDTH NO-BREAK SPACE / BOM
> +    "=C2=AD": "",       # SOFT HYPHEN
> +}
> +
> +
> +def main():
> +
> +    args =3D parse_arguments()
> +    exit_code =3D 0
> +
> +    for rst_path in sorted(Path(args.docs_dir).rglob("*.rst")):
> +        rel =3D rst_path.relative_to(args.docs_dir)
> +
> +        with open(rst_path, "r", encoding=3D"utf-8") as f:
> +            for lineno, line in enumerate(f, start=3D1):
> +                for col, char in enumerate(line, start=3D1):
> +                    if char in confusables:
> +                        replacement =3D confusables[char]
> +                        hint =3D (f"'{replacement}'" if replacement
> +                                else "(remove)")
> +                        print(f"WARNING: {rel}:{lineno}:{col}: non-ASCII=
 "
> +                              f"character U+{ord(char):04X} should be "
> +                              f"replaced with {hint}")
> +                        exit_code =3D 1
> +
> +    sys.exit(exit_code)
> +
> +
> +if __name__ =3D=3D "__main__":
> +    main()


Antonin