Re: Unaligned access trade-offs for SFrame FRE layout

Fangrui Song <[email protected]> Tue, 16 Sep 2025 09:32:30 -0700
Newsgroups org.kernel.vger.linux-toolchains
Message-ID <CAN30aBFW1T7WhBn9QBDig6i1Nh23XTJ1eRFHeNLNU2nfahv_7Q@mail.gmail.com>
On Tue, Sep 16, 2025 at 9:03 AM Indu Bhagat <[email protected]> wrote:
>
> On 9/15/25 11:05 PM, Fangrui Song wrote:
> > On Mon, Sep 15, 2025 at 9:12 AM Steven Rostedt <[email protected]> wrote:
> >>
> >> On Sun, 14 Sep 2025 22:42:46 -0700
> >> Indu Bhagat <[email protected]> wrote:
> >>
> >>> In such cases, the routines reading the SFrame data under consideration
> >>> here (SFrame FRE start addr, and SFrame FRE stack offsets) from memory
> >>> will need to use a memcpy to copy out the data to an aligned location.
> >>>
> >>> In GNU Binutils libsframe (used by ld), we do the above. Such a "SFrame
> >>> FRE decoding" routine could be provided in a arch-specific manner in
> >>> SFrame stack tracers.
> >>
> >> I'm perfectly fine with making it a requirement for the reader of the
> >> SFrame section having to use memcpy into an aligned structure for reading
> >> if the architecture requires it. Let only the architectures that have
> >> issues with unaligned access take the performance hit.
> >>
> >> -- Steve
> >
> > I agree. Unaligned access has nearly zero performance impact on modern
> > architectures, provided the access doesn't span additional cache
> > lines.
> > The padding required for alignment would increase the size, likely
> > creating more overhead than any alignment benefit would justify.
> >
> > (
> >  From a linker and binary utilities perspective, I'd even suggest
> > adopting a universal little-endian format regardless of the target
> > system's native endianness.
> > This would eliminate the need for endianness templates in the C++ code
> > and simplify toolchain implementation across platforms.
> >
>
> (Perhaps I am missing something) Wouldnt a toolchain implementation need
> endianness handling anyway to support cross toolchains?
>
> > On the big-endian z/Architecture, this is efficient: the LOAD REVERSED
> > instructions are used by the bswap versions in the following program,
> > not even requiring extra instructions.
> > #define WIDTH(x) \
> > typedef __UINT##x##_TYPE__ [[gnu::aligned(1)]] uint##x; \
> > uint##x load_inc##x(uint##x *p) { return *p+1; } \
> > uint##x load_bswap_inc##x(uint##x *p) { return __builtin_bswap##x(*p)+1; }; \
> > uint##x load_eq##x(uint##x *p) { return *p==3; } \
> > uint##x load_bswap_eq##x(uint##x *p) { return __builtin_bswap##x(*p)==3; }; \
> >
> > WIDTH(16);
> > WIDTH(32);
> > WIDTH(64);
> > )
>
> For AArch64 which SFrame supports too, this is not true. AArch64 has
> both LE and BE.

While runtime consumers typically handle a single endianness, other
tools like linkers and binary utilities must support both. They have
to support cross compilation, producing a big-endian executable from a
little-endian host.

A universal little-endian approach simplifies code. Instead of using a
function like read32le(config, p), where config->endian specifies the
object file's endianness, or read32(p) with an internal endianness
check, the code can simply use read32le(p).

The read32le(p) function is either a standard read or a byte-swapped
read. This byte-swapping is fast on aarch64be (thanks to REV16 and
REV32 instructions) and s390x (byte-swap load).