Re: Unaligned access trade-offs for SFrame FRE layout
Indu Bhagat <[email protected]> Sun, 14 Sep 2025 22:42:46 -0700
| Newsgroups | org.kernel.vger.linux-toolchains |
|---|---|
| Message-ID | <[email protected]> |
On 9/14/25 11:10 AM, Jan Beulich wrote: > On 14.09.2025 18:18, Rainer Orth wrote: >> Jan Beulich <[email protected]> writes: >> >>> On 14.09.2025 16:39, Rainer Orth wrote: >>>>> On 12.09.2025 19:34, Indu Bhagat via Binutils wrote: >>>>>> TL;DR: Thinking and experimenting a bit on the possible approaches for >>>>>> avoiding unaligned accesses in the SFrame FRE layout (in SFrame V3), I >>>>>> am not convinced that avoiding unaligned accesses for performance is >>>>>> worth it. IMO, forsaking compactness for avoiding unaligned accesses is >>>>>> not a good trade off for SFrame. >>>>>> >>>>>> Problem Statement >>>>>> On architectures such as x86_64, AArch64, and s390x, unaligned memory >>>>>> accesses are handled transparently by the hardware but incur a >>>>>> performance penalty. >>>>> >>>>> As you say in a reply, may incur. However, shouldn't we also consider >>>>> possible ports of SFrame to architectures which don't handle this as >>>>> transparently? Off the top of my head I don't, for example, recall >>>>> whether RISC-V requires unaligned accesses to be handled transparently >>>>> by the hardware. >>>> >>>> look for STRICT_ALIGNMENT in the GCC sources in gcc/config. While >>>> several are embedded targets, there's also sparc in that list. >>> >>> But is this setting a good reference for the purpose here. For RISC-V it's >>> command line (?) controlled (TARGET_STRICT_ALIGN), despite the spec saying >>> >>> "An EEI may not guarantee misaligned loads and stores are handled invisibly. >>> In this case, loads and stores that are not naturally aligned may either >>> complete execution successfully or raise an exception. The exception raised >>> can be either an address-misaligned exception or an access-fault exception." >>> >>> It's okay for gcc to make assumptions (assuming they're properly documented), >>> but I don't think such assumptions can be extended to a discussion like the >>> one here. >> >> please look at the actual code: >> >> gcc/config/riscv/riscv.cc: TARGET_STRICT_ALIGN ? 0 : 1); >> >> on RISC-V, the setting is controlled by -mstrict-align. On most others, >> it's just 1. > > Precisely my point: By (not or wrongly) using the command line option, you > can break things. Whereas such breakage wants avoiding here. > >> I'm just pointing out an easy way to answer the question: there's a >> considerable number of strict-alignment targets. If they is relevant to >> the discussion at hand is for the SFrame developers to decide: if they >> come to the conclusion that none of the affected CPUs is of interest, >> that's certainly fine. However, from my recently experience porting >> LLVMs openmp to SPARC, fixing this as an afterthought takes some time >> and analysis, so you need to decide if you want to support such targets >> or not. > > Yes, that's exactly why I brought up the point. > In such cases, the routines reading the SFrame data under consideration here (SFrame FRE start addr, and SFrame FRE stack offsets) from memory will need to use a memcpy to copy out the data to an aligned location. In GNU Binutils libsframe (used by ld), we do the above. Such a "SFrame FRE decoding" routine could be provided in a arch-specific manner in SFrame stack tracers.