[PATCH net v2 0/1] ipv4: fix a non-progressing fragmentation loop from undersized effective MTUs

Ren Wei <[email protected]>
Newsgroups gmane.linux.network
Message-ID <[email protected]>
From: Yong Wang <[email protected]>

Hi Linux kernel maintainers,

This series fixes a non-progressing IPv4 fragmentation loop in
net/ipv4/ip_output.c. The bug is reachable by a non-root user with
CAP_NET_ADMIN in a user-created user and network namespace.

---- details below ----

Bug details:

ip_do_fragment() subtracts the IPv4 header length from the effective
MTU and passes the resulting payload MTU to ip_frag_next().

If the effective MTU is smaller than hlen + 8, ip_frag_next() rounds
the fragment payload length down to zero. The fragmentation state then
never makes forward progress: state->left, state->ptr and state->offset
stay unchanged while ip_do_fragment() keeps allocating and transmitting
header-only fragments until the softlockup detector fires.

This is reproducible with a route installed using "mtu lock 20", but it
is also reproducible without route MTU lock, for example by forwarding a
packet to a device whose MTU is 20.

Because the failure happens in the fragmentation path itself, this
series fixes it in ip_do_fragment() by rejecting mtu < hlen + 8 with
-EMSGSIZE, matching the existing IPv6 fragmentation check.

The bug is reachable through more than one userspace-controlled path. A
user with CAP_NET_ADMIN in a user-created network namespace can trigger
it either by installing an undersized locked route MTU or by forwarding
through an egress device whose MTU is too small.

Tested in a 2 vCPU, 2 GB RAM x86 QEMU guest:

- the old kernel reproduces the softlockup both with "mtu lock 20" and
  with a dummy egress device whose MTU is 20
- the new kernel returns packet loss in the no-route-mtu-lock reproducer
  and the guest remains alive

Reproducer:

Run inside the guest as root:

    bash poc.sh

The PoC sets up forwarding through an egress dummy device whose MTU is
20, without installing any route MTU lock, and injects one forwarded
non-DF ping from a child netns.

We run the PoC in a 2 vCPU, 2 GB RAM x86 QEMU environment.

v2:
  - move the fix from ip_metrics_convert() to ip_do_fragment()
  - update the commit message and cover letter per reviewer and Sashiko
    feedback
  - switch the cover-letter reproducer to the no-route-mtu-lock
    dummy-device case

v1: https://lore.kernel.org/netdev/[email protected]/

------BEGIN poc.sh------

#!/bin/bash

sysctl -w net.ipv4.ip_forward=1

ip link add name dummy1 up mtu 20 type dummy
ip address add 192.0.2.1/24 dev dummy1

ip link add veth0 type veth peer name veth1
ip addr add 198.51.100.1/24 dev veth0
ip link set veth0 up

ip netns add ns1
ip link set veth1 netns ns1
ip -n ns1 address add 198.51.100.2/24 dev veth1
ip -n ns1 link set veth1 up
ip -n ns1 route add default via 198.51.100.1
ip netns exec ns1 ping -M dont -s 1000 -c 1 192.0.2.2

------END poc.sh--------

Best regards,
Yong Wang

Yong Wang (1):
  ipv4: reject undersized MTUs in ip_do_fragment()

 net/ipv4/ip_output.c | 4 ++++
 1 file changed, 4 insertions(+)

-- 
2.53.0
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.