Re: need macOS/Darwin help (was: preconv test failure on Darwin with nixpkgs 25.11)

"G. Branden Robinson" <[email protected]>
Newsgroups gmane.comp.printing.groff.general
Message-ID <20260227172533.z5sh4nrlfs5jxfgx@illithid>
Hi Alexis,

At 2026-02-27T11:13:07-0600, G. Branden Robinson wrote:
> At 2026-02-27T10:05:30+0100, Alexis wrote:
> > I've done as you've asked and replaced
> > src/preproc/preconv/tests/smoke-test.sh from the 1.24.0.rc4
> > distribution archive with the smoke-test.sh attached to your mail
> > in the build pipeline I have for building groff on macOS Darwin
> > using nixpkgs.
> > Unfortunately the smoke test still fails. Please find attached the
> > nixpkgs build log, hopefully it provides valuable and helpful
> > insights.
> 
> Thanks!  Yeah, it's not great that the test failed, but given that I
> heavily refactored the logic at the end of the script, it appears to
> have worked just fine to tell me I made an incorrect assumption about
> what macOS/Darwin's libc uses as a character encoding for the "C"
> locale, which is fine.  The script failed in a controlled way.
> 
> An incorrect assumption should be easy to correct.  :)
> 
> I expect to have a fresh cut of the test script for you shortly.

Here you go.  Same procedure as before.  Or just replace the
src/preproc/preconv/tests/smoke-test.sh you already have with this one.

The difference is that we now assume that preconv will fall back
"ISO-8859-1" on macOS/Darwin, just as we expect from GNU libc.

Regards,
Branden
smoke-test.sh (application/x-sh, 4.4 KB)
#!/bin/sh
#
# Copyright 2020-2024 G. Branden Robinson
#
# This file is part of groff, the GNU roff typesetting system.
#
# groff is free software; you can redistribute it and/or modify it under
# the terms of the GNU General Public License as published by the Free
# Software Foundation, either version 3 of the License, or (at your
# option) any later version.
#
# groff is distributed in the hope that it will be useful, but WITHOUT
# ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or
# FITNESS FOR A PARTICULAR PURPOSE.  See the GNU General Public License
# for more details.
#
# You should have received a copy of the GNU General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.

preconv="${abs_top_builddir:-.}/preconv"

fail=

wail () {
    echo ...FAILED >&2
    fail=YES
}

# Ensure a predictable character encoding.
export LC_ALL=C

echo "testing -e flag override of BOM detection" >&2
printf '\376\377\0\100\0\n' \
    | "$preconv" -d -e euc-kr 2>&1 > /dev/null \
    | grep -q "no search for coding tag" || wail

echo "testing detection of UTF-32BE BOM" >&2
printf '\0\0\376\377\0\0\0\100\0\0\0\n' \
    | "$preconv" -d 2>&1 > /dev/null \
    | grep -q "found BOM" || wail

echo "testing detection of UTF-32LE BOM" >&2
printf '\377\376\0\0\100\0\0\0\n\0\0\0' \
    | "$preconv" -d 2>&1 > /dev/null \
    | grep -q "found BOM" || wail

echo "testing detection of UTF-16BE BOM" >&2
printf '\376\377\0\100\0\n' \
    | "$preconv" -d 2>&1 > /dev/null \
    | grep -q "found BOM" || wail

echo "testing detection of UTF-16LE BOM" >&2
printf '\377\376\100\0\n\0' \
    | "$preconv" -d 2>&1 > /dev/null \
    | grep -q "found BOM" || wail

echo "testing detection of UTF-8 BOM" >&2
printf '\357\273\277@\n' \
    | "$preconv" -d 2>&1 > /dev/null \
    | grep -q "found BOM" || wail

# We do not find a coding tag on piped input because it isn't seekable.
echo "testing detection of Emacs coding tag in piped input" >&2
printf '.\\" -*- coding: euc-kr; -*-\\n' \
    | "$preconv" -d 2>&1 >/dev/null \
    | grep -q "no coding tag" || wail

test -z "$fail" || exit

# We need uchardet to work to get past this point.
if ! "$preconv" -v | grep -q 'with uchardet support'
then
    echo "$0: preconv lacks uchardet support; skipping" >&2
    exit 77 # skip
fi

# Instead of using temporary files, which in all fastidiousness means
# cleaning them up even if we're interrupted, which in turn means
# setting up signal handlers, we use files in the build tree.

# TODO: groff_mmse(7) is no longer UTF-8-encoded; find another.
#doc=contrib/mm/groff_mmse.7
#echo "testing uchardet detection on UTF-8 document $doc" >&2
#"$preconv" -d -D us-ascii 2>&1 >/dev/null $doc \
#    | grep -q 'charset: UTF-8' || wail

# uchardet can't seek on a pipe either.
echo "testing uchardet detection on pipe (expect fallback to -D)" >&2
printf 'Eat at the caf\351.\n' \
    | "$preconv" -d -D euc-kr 2>&1 > /dev/null \
    | grep -q "encoding used: 'EUC-KR'" || wail

test -z "$fail" || exit

# Fall back to the locale.
#
# It's hard to determine the character encoding of the 'C' locale
# because the only POSIX-standard way to do so is to build a C program
# to call `nl_langinfo(CODESET)`.  There's also no POSIX-standard way
# to ask a system to report the byte sequence it uses to encode, say,
# "lowercase e with acute accent".
#
# (I think Perl can do that, though.)
#
# We're just a shell script, so on non-glibc systems, we guess at it.
#
# On glibc systems, the 'C' locale uses "ANSI_X3.4-1968" for the
# character set, and `locale charmap` tells us as much, but preconv
# assumes Latin-1 instead of US-ASCII, so we override that.
#
# On Darwin (macOS) systems, we do the same.  See
# <https://lists.gnu.org/archive/html/groff/2026-02/msg00129.html>.
#
# For everything else, we assume UTF-8.

libc_vendor=

if command -v locale > /dev/null
then
    libc_vendor=gnu
    charset=ISO-8859-1
elif [ "$(uname -s)" = "Darwin" ]
then
    libc_vendor=apple
    charset=ISO-8859-1
else
    libc_vendor=unknown
    charset=UTF-8
fi

printf "standard C library vendor: %s;" $libc_vendor >&2
printf " expecting preconv character encoding %s\n" $charset >&2

echo "testing fallback to locale setting in environment" >&2
printf 'Eat at the caf\351.\n' \
    | "$preconv" -d 2>&1 > /dev/null \
    | grep -q "encoding used: '$charset'" || wail

test -z "$fail"

# vim:set autoindent expandtab shiftwidth=4 tabstop=4 textwidth=72:
signature.asc (application/pgp-signature, 833 B)
-----BEGIN PGP SIGNATURE-----

iQIzBAABCAAdFiEEh3PWHWjjDgcrENwa0Z6cfXEmbc4FAmmh04QACgkQ0Z6cfXEm
bc5Akw//ROMtm02ZKtTxjt1y4DVwhdyIldt5CtfwlnByp25TtJs07sHcIZhRwcRF
oySsdopX/8FvJw+gLTMIZLjeK40ZxUq9hVeipzRE7lywLZ1+vyCPzGa5HsYrmJ+X
+E65AvxT2v2Sx2/keWYb4IUt1kM/yn6A2oFmUgp4LItIxLyagHJidl9pdevTbZMS
q0qrulNwQjqqyeqZwA23NDfiSoN+SMLk0TtOa4149nEqIydi9DiJ2e6AaRU9ObA0
cmUmZSLtlb2UWKLjAK3ZHHrldeN/rXeVGR89B9CPr+Au4T4T3Uo3EsOksjbB8lTa
3lLoJaoc13rXYe1Pdaw1856Y2clj57PC/CyNHbr8jgSFVc/IFVVys+kvtobGnnBS
ZPyDY+sEvAd9CTt3EnUF7Dqblj/zUrox/+Bvyujg+7iZS7mucDCuP3pS95Y7CMfM
V2W5/Fky1NuQ2DnqQWSwVnP3nTvlLxvkX2JJlMBDEIRWlzV2fQJLh3gjO0XWOT7W
ExN7SgaMC0ymOwzQYZwRfYC8i8wlp6dBqwsSxeULsbtDDC3+mwBRVC5pCC7mSvAs
efmVhejm6YBMM87IAcXSS4QdVQA5y5/qhHaJmSkfqfusll+soPx3Uq67XIT6qu+Q
qbxZV6G0q7VoGzD6xOxD8tQ6CrOiJ2FzUQ1vfNq6IbIgq8Aa3Ag=
=Vj5X
-----END PGP SIGNATURE-----
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.