Re: Unaligned access trade-offs for SFrame FRE layout
Fangrui Song <[email protected]> Tue, 16 Sep 2025 09:32:30 -0700
| Newsgroups | org.kernel.vger.linux-toolchains |
|---|---|
| Message-ID | <CAN30aBFW1T7WhBn9QBDig6i1Nh23XTJ1eRFHeNLNU2nfahv_7Q@mail.gmail.com> |
On Tue, Sep 16, 2025 at 9:03 AM Indu Bhagat <[email protected]> wrote: > > On 9/15/25 11:05 PM, Fangrui Song wrote: > > On Mon, Sep 15, 2025 at 9:12 AM Steven Rostedt <[email protected]> wrote: > >> > >> On Sun, 14 Sep 2025 22:42:46 -0700 > >> Indu Bhagat <[email protected]> wrote: > >> > >>> In such cases, the routines reading the SFrame data under consideration > >>> here (SFrame FRE start addr, and SFrame FRE stack offsets) from memory > >>> will need to use a memcpy to copy out the data to an aligned location. > >>> > >>> In GNU Binutils libsframe (used by ld), we do the above. Such a "SFrame > >>> FRE decoding" routine could be provided in a arch-specific manner in > >>> SFrame stack tracers. > >> > >> I'm perfectly fine with making it a requirement for the reader of the > >> SFrame section having to use memcpy into an aligned structure for reading > >> if the architecture requires it. Let only the architectures that have > >> issues with unaligned access take the performance hit. > >> > >> -- Steve > > > > I agree. Unaligned access has nearly zero performance impact on modern > > architectures, provided the access doesn't span additional cache > > lines. > > The padding required for alignment would increase the size, likely > > creating more overhead than any alignment benefit would justify. > > > > ( > > From a linker and binary utilities perspective, I'd even suggest > > adopting a universal little-endian format regardless of the target > > system's native endianness. > > This would eliminate the need for endianness templates in the C++ code > > and simplify toolchain implementation across platforms. > > > > (Perhaps I am missing something) Wouldnt a toolchain implementation need > endianness handling anyway to support cross toolchains? > > > On the big-endian z/Architecture, this is efficient: the LOAD REVERSED > > instructions are used by the bswap versions in the following program, > > not even requiring extra instructions. > > #define WIDTH(x) \ > > typedef __UINT##x##_TYPE__ [[gnu::aligned(1)]] uint##x; \ > > uint##x load_inc##x(uint##x *p) { return *p+1; } \ > > uint##x load_bswap_inc##x(uint##x *p) { return __builtin_bswap##x(*p)+1; }; \ > > uint##x load_eq##x(uint##x *p) { return *p==3; } \ > > uint##x load_bswap_eq##x(uint##x *p) { return __builtin_bswap##x(*p)==3; }; \ > > > > WIDTH(16); > > WIDTH(32); > > WIDTH(64); > > ) > > For AArch64 which SFrame supports too, this is not true. AArch64 has > both LE and BE. While runtime consumers typically handle a single endianness, other tools like linkers and binary utilities must support both. They have to support cross compilation, producing a big-endian executable from a little-endian host. A universal little-endian approach simplifies code. Instead of using a function like read32le(config, p), where config->endian specifies the object file's endianness, or read32(p) with an internal endianness check, the code can simply use read32le(p). The read32le(p) function is either a standard read or a byte-swapped read. This byte-swapping is fast on aarch64be (thanks to REV16 and REV32 instructions) and s390x (byte-swap load).