Re: Unaligned access trade-offs for SFrame FRE layout

Indu Bhagat <[email protected]> Tue, 16 Sep 2025 09:03:24 -0700
Newsgroups org.kernel.vger.linux-toolchains
Message-ID <[email protected]>
On 9/15/25 11:05 PM, Fangrui Song wrote:
> On Mon, Sep 15, 2025 at 9:12 AM Steven Rostedt <[email protected]> wrote:
>>
>> On Sun, 14 Sep 2025 22:42:46 -0700
>> Indu Bhagat <[email protected]> wrote:
>>
>>> In such cases, the routines reading the SFrame data under consideration
>>> here (SFrame FRE start addr, and SFrame FRE stack offsets) from memory
>>> will need to use a memcpy to copy out the data to an aligned location.
>>>
>>> In GNU Binutils libsframe (used by ld), we do the above. Such a "SFrame
>>> FRE decoding" routine could be provided in a arch-specific manner in
>>> SFrame stack tracers.
>>
>> I'm perfectly fine with making it a requirement for the reader of the
>> SFrame section having to use memcpy into an aligned structure for reading
>> if the architecture requires it. Let only the architectures that have
>> issues with unaligned access take the performance hit.
>>
>> -- Steve
> 
> I agree. Unaligned access has nearly zero performance impact on modern
> architectures, provided the access doesn't span additional cache
> lines.
> The padding required for alignment would increase the size, likely
> creating more overhead than any alignment benefit would justify.
> 
> (
>  From a linker and binary utilities perspective, I'd even suggest
> adopting a universal little-endian format regardless of the target
> system's native endianness.
> This would eliminate the need for endianness templates in the C++ code
> and simplify toolchain implementation across platforms.
> 

(Perhaps I am missing something) Wouldnt a toolchain implementation need 
endianness handling anyway to support cross toolchains?

> On the big-endian z/Architecture, this is efficient: the LOAD REVERSED
> instructions are used by the bswap versions in the following program,
> not even requiring extra instructions.
> #define WIDTH(x) \
> typedef __UINT##x##_TYPE__ [[gnu::aligned(1)]] uint##x; \
> uint##x load_inc##x(uint##x *p) { return *p+1; } \
> uint##x load_bswap_inc##x(uint##x *p) { return __builtin_bswap##x(*p)+1; }; \
> uint##x load_eq##x(uint##x *p) { return *p==3; } \
> uint##x load_bswap_eq##x(uint##x *p) { return __builtin_bswap##x(*p)==3; }; \
> 
> WIDTH(16);
> WIDTH(32);
> WIDTH(64);
> )

For AArch64 which SFrame supports too, this is not true. AArch64 has 
both LE and BE.