Re: Unaligned access trade-offs for SFrame FRE layout
Indu Bhagat <[email protected]> Tue, 16 Sep 2025 10:33:52 -0700
| Newsgroups | org.kernel.vger.linux-toolchains |
|---|---|
| Message-ID | <[email protected]> |
On 9/16/25 9:32 AM, Fangrui Song wrote: > On Tue, Sep 16, 2025 at 9:03 AM Indu Bhagat <[email protected]> wrote: >> >> On 9/15/25 11:05 PM, Fangrui Song wrote: >>> On Mon, Sep 15, 2025 at 9:12 AM Steven Rostedt <[email protected]> wrote: >>>> >>>> On Sun, 14 Sep 2025 22:42:46 -0700 >>>> Indu Bhagat <[email protected]> wrote: >>>> >>>>> In such cases, the routines reading the SFrame data under consideration >>>>> here (SFrame FRE start addr, and SFrame FRE stack offsets) from memory >>>>> will need to use a memcpy to copy out the data to an aligned location. >>>>> >>>>> In GNU Binutils libsframe (used by ld), we do the above. Such a "SFrame >>>>> FRE decoding" routine could be provided in a arch-specific manner in >>>>> SFrame stack tracers. >>>> >>>> I'm perfectly fine with making it a requirement for the reader of the >>>> SFrame section having to use memcpy into an aligned structure for reading >>>> if the architecture requires it. Let only the architectures that have >>>> issues with unaligned access take the performance hit. >>>> >>>> -- Steve >>> >>> I agree. Unaligned access has nearly zero performance impact on modern >>> architectures, provided the access doesn't span additional cache >>> lines. >>> The padding required for alignment would increase the size, likely >>> creating more overhead than any alignment benefit would justify. >>> >>> ( >>> From a linker and binary utilities perspective, I'd even suggest >>> adopting a universal little-endian format regardless of the target >>> system's native endianness. >>> This would eliminate the need for endianness templates in the C++ code >>> and simplify toolchain implementation across platforms. >>> >> >> (Perhaps I am missing something) Wouldnt a toolchain implementation need >> endianness handling anyway to support cross toolchains? >> >>> On the big-endian z/Architecture, this is efficient: the LOAD REVERSED >>> instructions are used by the bswap versions in the following program, >>> not even requiring extra instructions. >>> #define WIDTH(x) \ >>> typedef __UINT##x##_TYPE__ [[gnu::aligned(1)]] uint##x; \ >>> uint##x load_inc##x(uint##x *p) { return *p+1; } \ >>> uint##x load_bswap_inc##x(uint##x *p) { return __builtin_bswap##x(*p)+1; }; \ >>> uint##x load_eq##x(uint##x *p) { return *p==3; } \ >>> uint##x load_bswap_eq##x(uint##x *p) { return __builtin_bswap##x(*p)==3; }; \ >>> >>> WIDTH(16); >>> WIDTH(32); >>> WIDTH(64); >>> ) >> >> For AArch64 which SFrame supports too, this is not true. AArch64 has >> both LE and BE. > > While runtime consumers typically handle a single endianness, other > tools like linkers and binary utilities must support both. They have > to support cross compilation, producing a big-endian executable from a > little-endian host. > Right. Sorry, I am still missing the link between "complexity of endianness templates in the C++ code" vs what you say in the next paragraph: endian aware read/write is anyway necessary. > A universal little-endian approach simplifies code. Instead of using a > function like read32le(config, p), where config->endian specifies the > object file's endianness, or read32(p) with an internal endianness > check, the code can simply use read32le(p). > > The read32le(p) function is either a standard read or a byte-swapped > read. This byte-swapping is fast on aarch64be (thanks to REV16 and > REV32 instructions) and s390x (byte-swap load). The rev* instruction is in the data dependency chain. This means that using little-endian for AArch64 BE then defers the task of endian swap on to the stack tracers. E.g., aarch64 (added insn in dependency chain): load_inc16: ldrh w0, [x0] add w0, w0, 1 ret load_bswap_inc16: ldrh w0, [x0] rev16 w0, w0 add w0, w0, 1 ret s390x (same height of dependency chain): load_inc16: lh %r2,0(%r2) ahi %r2,1 llghr %r2,%r2 br %r14 load_bswap_inc16: lrvh %r2,0(%r2) ahi %r2,1 llghr %r2,%r2 br %r14