Re: [PATCH net] netfilter: nf_dup_netdev: scrub duplicates to preserve the direct path
Alexandre Ferrieux <[email protected]> Wed, 5 Aug 2026 11:02:48 +0200
| Newsgroups | gmane.comp.security.firewalls.netfilter.devel,gmane.linux.network |
|---|---|
| Organization | Orange |
| Message-ID | <[email protected]> |
On 8/5/26 10:24 AM, Pablo Neira Ayuso wrote:
> Hi,
>
> On Tue, Aug 04, 2026 at 10:11:50PM +0200, Alexandre Ferrieux wrote:
>> The nftables 'dup' action clones the skb with its full glory of
>> metadata, including references to its destination and conntrack
>> information. As a consequence, a link failure on the duplicate's
>> egress path ends up doing the same as it would for the direct path,
>> for example invalidating the original packet's destination, which
>> typically breaks all TCP connections to that address.
>>
>> In other words, the "dup" path has the potential to wreak havoc
>> in the direct path as a consequence of secondary link failures. This
>> is very bad behavior for a monitoring tool, which is the most
>> obvious application of 'dup'.
>
> Can you describe your use-case a bit and how it breaks?
Sure:
- assume hosts A.eth0 and B.eth0 have active production traffic (say TCP)
- assume we have a monitoring tool on A that dups eth0's egress to some other
interface $MON
nft add chain netdev ta ch '{type filter hook egress device "eth0" priority
filter ; policy accept ; }'
nft add rule netdev ta ch dup to $MON
- assume something goes wrong on $MON generating link failures. In my case it
was a GRETAP with L3 destination suddenly unreachable.
- The next A->B packet goes through normally, but its duplicate hits
ipv4_link_failure(), hence the dst (which is B) is expired.
- As a result, (say) TCP disruptions occur. The thermometer killed the patient :)
Note: as a straightforward repro, you can simply witness "noise" in simple ping
sessions, with ghost unreach reports muxed with normal measurement:
ip link add gre1 type gretap remote 192.168.1.99 ;# on the LAN, nonexistent
IP => will generate link failures
ip link set dev gre1 up
nft add table netdev ta
nft add chain netdev ta ch '{type filter hook egress device "eth0" priority
filter ; policy accept ; }'
nft add rule netdev ta ch counter dup to gre1
ping -n 8.8.8.8
=>
PING 8.8.8.8 (8.8.8.8) 56(84) bytes of data.
64 bytes from 8.8.8.8: icmp_seq=1 ttl=115 time=19.9 ms
64 bytes from 8.8.8.8: icmp_seq=2 ttl=115 time=12.8 ms
64 bytes from 8.8.8.8: icmp_seq=3 ttl=115 time=36.0 ms
64 bytes from 8.8.8.8: icmp_seq=4 ttl=115 time=30.3 ms
From 192.168.1.13 icmp_seq=5 Destination Host Unreachable
64 bytes from 8.8.8.8: icmp_seq=5 ttl=115 time=31.0 ms
From 192.168.1.13 icmp_seq=6 Destination Host Unreachable
64 bytes from 8.8.8.8: icmp_seq=6 ttl=115 time=38.0 ms
From 192.168.1.13 icmp_seq=7 Destination Host Unreachable
64 bytes from 8.8.8.8: icmp_seq=7 ttl=115 time=36.8 ms
From 192.168.1.13 icmp_seq=8 Destination Host Unreachable
64 bytes from 8.8.8.8: icmp_seq=8 ttl=115 time=23.5 ms
^C
>
>> This patch fixes all similar scenarii by calling skb_scrub_pkt()
>> on the clone, severing its link to precious direct-path state.
>
> This patch is targetted at the net tree, but nf.git is preferred.
Okay, will retarget :)
> As for the conntrack and dst, you have to explain what it breaks on
> your end.
Dst as shown above. Conntrack is more speculation, but my take is that in any
case the dup path should *never* have any kind of retroaction on the observed
path, so any "complex state" attached to the direct path should be absolutely
isolated from "whatever happens on the dup path". Am I mistaken ?
-Alex