Re: [PATCH net v2 1/2] net: Fix UDP length overflow with PMTU discover and big MTU

"Alice Mikityanska" <[email protected]>
Newsgroups org.kernel.vger.linux-kselftest,org.kernel.vger.netdev
Message-ID <[email protected]>
On Thu, Aug 13, 2026, at 23:11, Willem de Bruijn wrote:
> Alice Mikityanska wrote:
>> From: Alice Mikityanska <[email protected]>
>> 
>> This commit bounds cork->base.fragsize to IP(6)_MAX_MTU to avoid a
>> possible overflow of UDP length that triggers a WARN in
>> udp_set_len_short when setsockopt IP(V6)_MTU_DISCOVER is set to
>> IPV6_PMTUDISC_DO or IP(V6)_PMTUDISC_PROBE, and a large packet is sent
>> over a netdev with an unusually large MTU.
>> 
>> Steps to reproduce (included in the new selftest):
>> 
>> 1. Set device MTU bigger than IP6_MAX_MTU (or IP_MAX_MTU + 20).
>>    cork->base.fragsize will be set to that MTU in ip(6)_setup_cork.
>> 2. Set IP(V6)_MTU_DISCOVER to IP(V6)_PMTUDISC_PROBE or IPV6_PMTUDISC_DO.
>>    It lets maxnonfragsize be set to device MTU (cork->fragsize) in
>>    __ip(6)_append_data, rather than to IP(6)_MAX_MTU.
>> 3. Send 65528 bytes of payload (+8 bytes of UDP header, +20/40 bytes of
>>    IPv4/IPv6 header). Device MTU allows it (it's only one byte bigger
>>    than IP6_MAX_MTU or IP_MAX_MTU + IPv4 header, and the device MTU is
>>    bigger than that).
>> 4. The UDP length in the built packet is 65536, which overflows the
>>    16-bit length field and triggers the WARN in udp_set_len_short.
>> 
>> Note: IP_PMTUDISC_DO with IPv4 is safe, because ip_dst_mtu_maybe_forward
>> always clamps at IP_MAX_MTU, unlike ip6_dst_mtu_maybe_forward.
>> 
>> The Fixes tag points at the first commit where I could reproduce the
>> overflow with IPv4 and IP_PMTUDISC_PROBE.
>> 
>> Fixes: daba287b299e ("ipv4: fix DO and PROBE pmtu mode regarding local fragmentation with UFO/CORK")
>> Reported-by: [email protected]
>> Closes: https://lore.kernel.org/netdev/[email protected]/
>> Signed-off-by: Alice Mikityanska <[email protected]>
>> Assisted-by: Claude:claude-sonnet-4.6
>> Cc: Willem de Bruijn <[email protected]>
>
> Not sure whether IPv4 and IPv6 should be separate patches.
> IPv6 will have a different Fixes tag.
>
> Per maintainer-netdev.rst, mixing code and test changes in the same
> patch is also discouraged.
>
>> ---
>>  net/ipv4/ip_output.c                         |  2 +
>>  net/ipv6/ip6_output.c                        |  2 +-
>>  tools/testing/selftests/net/Makefile         |  1 +
>>  tools/testing/selftests/net/cork_fragsize.py | 77 ++++++++++++++++++++
>>  4 files changed, 81 insertions(+), 1 deletion(-)
>>  create mode 100755 tools/testing/selftests/net/cork_fragsize.py
>> 
>> diff --git a/net/ipv4/ip_output.c b/net/ipv4/ip_output.c
>> index e6dd1e5b8c32..2bc997c4fa7b 100644
>> --- a/net/ipv4/ip_output.c
>> +++ b/net/ipv4/ip_output.c
>> @@ -1299,6 +1299,8 @@ static int ip_setup_cork(struct sock *sk, struct inet_cork *cork,
>>  
>>  	cork->fragsize = ip_sk_use_pmtu(sk) ?
>>  			 dst4_mtu(&rt->dst) : READ_ONCE(rt->dst.dev->mtu);
>> +	if (cork->fragsize > IP_MAX_MTU)
>> +		cork->fragsize = IP_MAX_MTU;
>
> if respinning: use min(), like IPv6? It's even more concise and easier to read.
>
>>  
>>  	if (!inetdev_valid_mtu(cork->fragsize))
>>  		return -ENETUNREACH;
>> diff --git a/net/ipv6/ip6_output.c b/net/ipv6/ip6_output.c
>> index 2c44e5ed6171..81fb61b94759 100644
>> --- a/net/ipv6/ip6_output.c
>> +++ b/net/ipv6/ip6_output.c
>> @@ -1430,7 +1430,7 @@ static int ip6_setup_cork(struct sock *sk, struct inet_cork_full *cork,
>>  	if (frag_size && frag_size < mtu)
>>  		mtu = frag_size;
>>  
>> -	cork->base.fragsize = mtu;
>> +	cork->base.fragsize = min(mtu, IP6_MAX_MTU);
>
> Not sure whether a limitation that we seriously have to care about, but:
>
> This restricts all IPv6 protocols, while only UDP has the length field
> limitation.
>
> With IPv6 jumbo frames and RAW sockets, could there be a valid case
> for larger fragsize?

True; if a valid jumbo UDP (length=0, with IPv6 HBH exthdr) packet is
crafted manually with SOCK_RAW, it's dropped after my change. I'll
respin, limiting the clamp to UDP sockets only.

I'll also split IPv4/IPv6/selftest into separate patches, also adding a
selftest for SOCK_RAW.

>>  	cork->base.gso_size = ipc6->gso_size;
>>  	cork->base.tx_flags = 0;
>>  	cork->base.mark = ipc6->sockc.mark;
>> diff --git a/tools/testing/selftests/net/Makefile b/tools/testing/selftests/net/Makefile
>> index 708d960ae07d..82a1407f6a69 100644
>> --- a/tools/testing/selftests/net/Makefile
>> +++ b/tools/testing/selftests/net/Makefile
>> @@ -24,6 +24,7 @@ TEST_PROGS := \
>>  	cmsg_so_mark.sh \
>>  	cmsg_so_priority.sh \
>>  	cmsg_time.sh \
>> +	cork_fragsize.py \
>>  	double_udp_encap.sh \
>>  	drop_monitor_tests.sh \
>>  	ecmp_rehash.sh \
>> diff --git a/tools/testing/selftests/net/cork_fragsize.py b/tools/testing/selftests/net/cork_fragsize.py
>> new file mode 100755
>> index 000000000000..7a1bb069227a
>> --- /dev/null
>> +++ b/tools/testing/selftests/net/cork_fragsize.py
>> @@ -0,0 +1,77 @@
>> +#!/usr/bin/env python3
>> +# SPDX-License-Identifier: GPL-2.0
>> +
>> +# Test possible UDP length overflow in udp_send_skb/udp_v6_send_skb.
>> +
>> +from lib.py import ksft_run, ksft_exit, ksft_true
>> +from lib.py import ip, NetNS, NetNSEnter
>> +import errno
>> +import socket
>> +import subprocess
>> +
>> +
>> +IP_MTU_DISCOVER = 10
>> +IP_PMTUDISC_PROBE = 3
>> +IPV6_MTU_DISCOVER = 23
>> +IPV6_PMTUDISC_DO = 2
>> +
>> +
>> +def check_dmesg_clean(func) -> bool:
>> +    dmesg = subprocess.Popen(['dmesg'], stdout=subprocess.PIPE)
>> +    result = subprocess.run(['grep', '-q', f'WARNING:.*{func}'], stdin=dmesg.stdout)
>> +    dmesg.wait()
>> +    return result.returncode != 0 and dmesg.returncode == 0
>> +
>> +
>> +def test_ipv6() -> None:
>> +    with NetNS() as ns:
>> +        ip('link add dummy type dummy', ns=ns)
>> +        ip('link set dummy mtu 65576', ns=ns)
>> +        ip('link set dummy up', ns=ns)
>> +        ip('-6 addr add fd00::1/64 dev dummy nodad', ns=ns)
>> +        ip('-6 neigh add fd00::2 lladdr 02:00:00:00:00:02 dev dummy nud permanent', ns=ns)
>> +
>> +        with NetNSEnter(ns):
>> +            with socket.socket(socket.AF_INET6, socket.SOCK_DGRAM) as fd:
>> +                fd.setsockopt(socket.IPPROTO_IPV6, IPV6_MTU_DISCOVER, IPV6_PMTUDISC_DO)
>> +                try:
>> +                    fd.sendto(b' ' * 65528, ('fd00::2', 1234))
>> +                except OSError as e:
>> +                    # Ignore EMSGSIZE: it happens on kernels with the fix.
>> +                    if e.errno != errno.EMSGSIZE:
>> +                        raise
>> +
>> +        ip('link del dummy', ns=ns)
>> +
>> +    ksft_true(check_dmesg_clean('udp_v6_send_skb'), 'WARNING detected in dmesg')
>> +
>> +
>> +def test_ipv4() -> None:
>> +    with NetNS() as ns:
>> +        ip('link add dummy type dummy', ns=ns)
>> +        ip('link set dummy mtu 65556', ns=ns)
>> +        ip('link set dummy up', ns=ns)
>> +        ip('addr add 10.0.0.1/24 dev dummy', ns=ns)
>> +        ip('neigh add 10.0.0.2 lladdr 02:00:00:00:00:02 dev dummy nud permanent', ns=ns)
>> +
>> +        with NetNSEnter(ns):
>> +            with socket.socket(socket.AF_INET, socket.SOCK_DGRAM) as fd:
>> +                fd.setsockopt(socket.IPPROTO_IP, IP_MTU_DISCOVER, IP_PMTUDISC_PROBE)
>> +                try:
>> +                    fd.sendto(b' ' * 65528, ('10.0.0.2', 1234))
>> +                except OSError as e:
>> +                    # Ignore EMSGSIZE: the check happens after the WARN is printed.
>> +                    if e.errno != errno.EMSGSIZE:
>> +                        raise
>> +
>> +        ip('link del dummy', ns=ns)
>> +
>> +    ksft_true(check_dmesg_clean('udp_send_skb'), 'WARNING detected in dmesg')
>> +
>> +
>> +if __name__ == "__main__":
>> +    ksft_run([
>> +        test_ipv6,
>> +        test_ipv4,
>> +    ])
>> +    ksft_exit()
>> -- 
>> 2.55.0
>>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.