Re: [PATCH 1/2] arm64: ftrace: enable single ftrace_ops for direct calls

Leon Hwang <[email protected]>
Newsgroups org.kernel.vger.linux-kselftest,org.infradead.lists.linux-arm-kernel,org.kernel.vger.bpf,org.kernel.vger.linux-kernel,org.kernel.vger.linux-trace-kernel
Message-ID <[email protected]>
On 31/7/26 07:03, Ihor Solodrai wrote:
> On 7/27/26 7:28 AM, Leon Hwang wrote:
>> The BPF tracing multi link updates several direct-call sites through one
>> ftrace_ops. Its implementation is therefore gated by
>> HAVE_SINGLE_FTRACE_DIRECT_OPS in addition to
>> DYNAMIC_FTRACE_WITH_DIRECT_CALLS.
>>
>> Select HAVE_SINGLE_FTRACE_DIRECT_OPS whenever arm64 enables dynamic ftrace
>> direct calls. This enables BPF tracing multi links on arm64. Also
>> generalize the unreachable-trampoline comment because the single-ops path
>> does not use ops->direct_call.
> 
> Hi Leon,
> 
> I don't think this change can land as is yet. The series doesn't even
> apply cleanly to bpf-next, but that's minor.
> 
> More importantly, it depends on Jose's series [1], which is not in the
> mainline yet. And there Mark has raised performance concerns [2] and
> the discussion still seems to be open.
> 
> [1] https://lore.kernel.org/all/[email protected]/
> [2] https://lore.kernel.org/all/amjnf5gz0xP5PTSB@J2N7QTR9R3/
> 
>>
>> Assisted-by: Codex:gpt-5.6-sol
>> Signed-off-by: Leon Hwang <[email protected]>
>> ---
>>  arch/arm64/Kconfig         | 2 ++
>>  arch/arm64/kernel/ftrace.c | 3 +--
>>  2 files changed, 3 insertions(+), 2 deletions(-)
>>
>> diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
>> index 0de419ed780f..c98dca76859b 100644
>> --- a/arch/arm64/Kconfig
>> +++ b/arch/arm64/Kconfig
>> @@ -188,6 +188,8 @@ config ARM64
>>  		    CLANG_SUPPORTS_DYNAMIC_FTRACE_WITH_ARGS)
>>  	select HAVE_DYNAMIC_FTRACE_WITH_DIRECT_CALLS \
>>  		if DYNAMIC_FTRACE_WITH_ARGS
>> +	select HAVE_SINGLE_FTRACE_DIRECT_OPS \
>> +		if DYNAMIC_FTRACE_WITH_DIRECT_CALLS\
> 
> The select is only conditional on DYNAMIC_FTRACE_WITH_DIRECT_CALLS, so
> it can be set along with HAVE_DYNAMIC_FTRACE_WITH_CALL_OPS. And AFAIU
> this would make the fast path effectively dead: every BPF direct call
> routed through ftrace_caller now goes through the slow path.
> 
> As Mark noted in the other thread, on arm64 trampolines come from
> EXECMEM_BPF, so they always land out of BL range (chance of landing in
> range is 256M/terabytes).
> 
> I vibe-slop-coded a benchmark and ran it on a Neoverse V2 machine, and
> toggling your config change seems to be causing a 1.3x regression in
> the tracing overhead:
> 
>   do-nothing fentry (r0=0; exit) on __arm64_sys_getpid, 20 M calls, min-of-N, several boots. Results:
> 
>               untraced (base)   traced        overhead
>   Kernel A    ~128 ns/call      ~145.5 ns     ~17.6 ns   (fast path: br x17)
>   Kernel B    ~127 ns/call      ~150.4 ns     ~23.4 ns   (slow path: save regs + call_direct_funcs + hash)


Thanks for your testing.

> 
> This confirms Jiri's suspicion.

True.

> 
> However my understanding is the regression should mostly disappear in
> case some version of in-range trampoline allocation on arm64 lands.
> 
> So, I think the landing sequence should be something like follows:
>   * in-range BPF-trampoline allocation that Jose proposed [3]
>   * then HAVE_SINGLE_FTRACE_DIRECT_OPS selection
> 
> After all of that reaches mainline, then a selftest patch can go
> through the bpf-next.

Sounds reasonable.

This series is based on Jose's series and is intended for the arm64
tree, rather than bpf-next. Like Jose's series, this series aims to
enable BPF tracing_multi link on arm64. And yes, with in-range
BPF-trampoline allocation, the regression should mostly disappear.

I'll follow the sequence and repost the patches afterward.

Thanks,
Leon

> 
> [3] https://lore.kernel.org/all/[email protected]/
> 
>>  	select HAVE_DYNAMIC_FTRACE_WITH_CALL_OPS \
>>  		if (DYNAMIC_FTRACE_WITH_ARGS && !CFI && \
>>  		    (CC_IS_CLANG || !CC_OPTIMIZE_FOR_SIZE))
>> diff --git a/arch/arm64/kernel/ftrace.c b/arch/arm64/kernel/ftrace.c
>> index e1a3c0b3a051..56ba72a87dfa 100644
>> --- a/arch/arm64/kernel/ftrace.c
>> +++ b/arch/arm64/kernel/ftrace.c
>> @@ -301,8 +301,7 @@ static bool ftrace_find_callable_addr(struct dyn_ftrace *rec,
>>  
>>  	/*
>>  	 * If a custom trampoline is unreachable, rely on the ftrace_caller
>> -	 * trampoline which knows how to indirectly reach that trampoline
>> -	 * through ops->direct_call.
>> +	 * trampoline which knows how to indirectly reach that trampoline.
>>  	 */
>>  	if (*addr != FTRACE_ADDR && !reachable_by_bl(*addr, pc))
>>  		*addr = FTRACE_ADDR;
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.