Re: SBCL 2.6.1 SIGILL during make-target-2 warm init on riscv64/linux/musl
Stas Boukarev <[email protected]> Tue, 24 Feb 2026 04:54:49 +0300
| Newsgroups | gmane.lisp.steel-bank.devel |
|---|---|
| Message-ID | <CAF63=11WHOW9Xo0R83OyQfepNiT4MZ7KPY652kQJ=DUbCzMSsQ@mail.gmail.com> |
I noticed that this breaks backtraces for invalid-arg-count (besides being overall bad). I added some code to force it to reside in RA/CS0. It's still not perfect, it gets moved back and forth between RA/CS0 if different environments have different conflicts, but that was the case before too. (That actually didn't entirely fix backtraces) On Sun, Feb 22, 2026 at 8:29 AM Stas Boukarev <[email protected]> wrote: > > Finally, I figured out the original riscv problem. > $RA gets moved into a non-descriptor register, never spilled onto the > control stack, and the function does not get pinned. > I made riscv/loongarch always spill RA. Whoever wants to make RA > always stay in RA until it's spilled to the CS is welcome to optimize > it. > > Here's a test function: > (defun c (n h) (values n h)) > (defun n (m) (declare (fixnum m)) (c (make-array 10000) (> m 20))) > (defun l () (declare (optimize speed (debug 0) (safety 0))) (n 0)) > (loop (l))) > > On Fri, Feb 20, 2026 at 3:47 AM Stas Boukarev <[email protected]> wrote: > > > > I have made some changes to the register pinning process. And it seems > > to have fixed a --gc-stress crash on riscv (and on loongarch). > > But the riscv machine I have access to is so slow I can't really make > > --gc-stress tell anything. > > Maybe someone who has a fast riscv machine can try --gc-stress. > > > > On Sat, Feb 14, 2026 at 1:44 AM Charles Zhang <[email protected]> wrote: > > > > > > I agree with Stas, it’s most likely a GC thing. Threaded builds were always a bit flaky because of that, and I was never sure after removing LRA whether on-stack code really does get pinned correctly or not. > > > > > > Although this seems like a regression - I think I had the test suite at least flushing through at one point. > > > > > > On Friday, February 13, 2026, 10:26 PM, Stas Boukarev <[email protected]> wrote: > > > > > > I'm willing to wager that memory just moves away from under it. > > > > > > E.g. I run ./run-tests.sh debug.* --gc-stress on riscv and get > > > > > > Unhandled TYPE-ERROR in thread #<SB-THREAD:THREAD tid=69302 "main > > > thread" RUNNING > > > {4F96AC73}>: > > > The value > > > 0 > > > is not of type > > > SB-C:TN > > > when binding SB-C::Y > > > > > > On Sat, Feb 14, 2026 at 12:19 AM Christophe Rhodes via Sbcl-devel > > > <[email protected]> wrote: > > > > > > > > SIGILL can mean a number of different things. One thing it can mean is > > > > that we have managed to encode an illegal instruction for your hardware. > > > > If that's the case, you could probably at least verify that by attaching > > > > gdb to the process and disassembling the memory around the faulting > > > > instruction: does most of it look like plausible Lisp code and there's > > > > just one bogus instruction? Then that's probably it. > > > > > > > > Another thing it can mean is that our assembler is mostly fine, but > > > > somehow the Lisp process has decided to jump into things that aren't > > > > meant to be machine code (header data, error encoding, that kind of > > > > thing). That's probably quite painful to start debugging: where have we > > > > jumped from? If this looks like it might be the case, I would probably > > > > start by checking to see if *any* riscv/sbcl builds on your hardware: if > > > > none does, then it's something more fundamental, but if there is an SBCL > > > > that builds, you could bisect to find the first version that doesn't, > > > > and stare suspciously at that change. If no version of SBCL runs on > > > > your hardware, I'd perhaps try swapping out the C library, because I'm > > > > not very familiar with musl, to see if that makes a difference. > > > > > > > > A third thing it might mean is that your hardware requires more > > > > instruction cache flushing (I think the SBCL term is > > > > "sanctify-for-execution") than we are currently doing. (Some hardware > > > > has separate and non-coherent data and instruction caches to memory, so > > > > writing data to memory that is intended later to be executed, as SBCL > > > > does, requires us to make sure that the instruction cache will pick up > > > > the new stuff). > > > > > > > > Best wishes, > > > > > > > > Christophe > > > > > > > > Will Sinatra <[email protected]> writes: > > > > > > > > > Building SBCL 2.6.1 from source on native riscv64 hardware using ECL > > > > > 24.5.10 as the cross-compilation host results in a SIGILL (signal 4) > > > > > crash during the make-target-2 warm init phase. The cross-compilation > > > > > (make-host-1) and genesis complete successfully, and the freshly built > > > > > SBCL boots and begins warm compilation, but crashes after compiling > > > > > src/code/room.lisp. > > > > > > > > > > Build progression > > > > > > > > > > make-host-1 (cross-compilation via ECL): Completes successfully. > > > > > make-host-2 / genesis: Completes successfully. Cold core is generated. > > > > > Genesis passes are consistent. > > > > > make-target-2(warm init): The freshly built SBCL boots from the cold > > > > > core and begins compiling. During warm init, three missing foreign > > > > > symbols are reported at startup: > > > > > > > > > > Missing required foreign symbol 'fun_end_breakpoint_trap'Missing > > > > > required foreign symbol 'fun_end_breakpoint_end'Missing required > > > > > foreign symbol 'fun_end_breakpoint_guts' > > > > > > > > > > Despite these warnings, cold init proceeds and warm compilation > > > > > begins. The build successfully compiles many files (including the > > > > > disassembler) but crashes immediately after compiling > > > > > src/code/room.lisp. > > > > > > > > > > Crash output > > > > > > > > > > ; compiling file "src/code/room.lisp" (written 26 JAN 2026 09:10:11 PM): > > > > > > > > > > ; wrote /builds/.../obj/from-self/src/code/room.fasl > > > > > ; compilation finished in 0:00:01.999 > > > > > CORRUPTION WARNING in SBCL pid 2678 tid 2678: > > > > > Signal 4 received (PC: 0x518e6f14) > > > > > Exiting. > > > > > Error opening /dev/tty: No such device or address > > > > > ldb> Welcome to LDB, a low-level debugger for the Lisp runtime environment. > > > > > Command exited with non-zero status 1 > > > > > > > > > > Signal 4 is SIGILL. The crash PC (0x518e6f14) falls within the dynamic > > > > > space (base 0x4F000000), indicating the illegal instruction is in > > > > > compiled Lisp code rather than the C runtime. > > > > > > > > > > The full build log is attached, and any suggestions on how to proceed > > > > > are sincerely appreciated! > > > > > > > > > > > > _______________________________________________ > > > > Sbcl-devel mailing list > > > > [email protected] > > > > https://lists.sourceforge.net/lists/listinfo/sbcl-devel > > > > > > > > > > > > _______________________________________________ > > > Sbcl-devel mailing list > > > [email protected] > > > https://lists.sourceforge.net/lists/listinfo/sbcl-devel _______________________________________________ Sbcl-devel mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/sbcl-devel