Re: [RFC PATCH net-next v2] net: gro: coalesce short IPv4 packets padded to the minimum frame size

Eric Dumazet <[email protected]>
Newsgroups dev.linux.lists.llvm,org.kernel.vger.linux-kernel,org.kernel.vger.netdev
Message-ID <CANn89iKFHCA9B1+0FkML=BxZ+0ytu3s5YBLjbt5-+X+CCox09A@mail.gmail.com>
On Fri, Aug 14, 2026 at 12:26 AM Glenn Judd <[email protected]> wrote:
>
> Software GRO fails to coalesce a small IPv4 segment that was
> padded up to the 60-byte minimum Ethernet frame.
>
> The selftest tools/testing/selftests/drivers/net/gro.py subtest
> sw_ipv4_data_lrg_1byte sends {100, 1} expecting to receive {101}.
> In current code, it receives {100, 1} (no coalescing) instead.
>
> Cause: inet_gro_receive() computes its flush term from
> tot_len ^ skb_gro_len() before skb_gro_pull(), while skb_gro_len()
> still includes trailing Ethernet padding. A small IPv4 segment
> padded up to the 60-byte minimum frame has tot_len != skb_gro_len(),
> so flush is set and the runt never coalesces.
>
> v1 detected the padding with an added iph->tot_len read and
> skb_gro_len() comparison on every IPv4 GRO packet. Instead, split
> inet_gro_receive() so that everything after the header validation
> takes the flush term as a parameter, and pass a literal 0 on the
> common path. That folds away both flush updates and lets the
> transport dispatch become a tail call, leaving the common path
> shorter than before this patch rather than merely unchanged.
>
> Assisted-by: Claude:claude-opus-5
> Assisted-by: Codex:gpt-5.6
> Assisted-by: Meta:internal-AI-tooling
> Signed-off-by: Glenn Judd <[email protected]>
> ---

I am pretty sure I met NIC where the extra bytes were not zero.
Unfortunately I no longer have access to them.

Your patch seems to target a specific NIC.

Look at my prior commit : I made it generic. Checking if the bytes are
zero has the same cost as computing their partial checksum.
(real cost is the cache line miss)

commit 88078d98d1bb085d72af8437707279e203524fa5
Author: Eric Dumazet <[email protected]>
Date:   Wed Apr 18 11:43:15 2018 -0700

    net: pskb_trim_rcsum() and CHECKSUM_COMPLETE are friends

    After working on IP defragmentation lately, I found that some large
    packets defeat CHECKSUM_COMPLETE optimization because of NIC adding
    zero paddings on the last (small) fragment.

    While removing the padding with pskb_trim_rcsum(), we set skb->ip_summed
    to CHECKSUM_NONE, forcing a full csum validation, even if all prior
    fragments had CHECKSUM_COMPLETE set.

    We can instead compute the checksum of the part we are trimming,
    usually smaller than the part we keep.

    Signed-off-by: Eric Dumazet <[email protected]>
    Signed-off-by: David S. Miller <[email protected]>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.