Re: Asymmetric forwarded TCP TX collapse over one GRE path on FreeBSD 14.3-RELEASE-p8 (PF independent, offloads disabled)
Michael Tuexen <[email protected]>
| Newsgroups | gmane.os.freebsd.devel.net |
|---|---|
| Message-ID | <[email protected]> |
> On 25. Feb 2026, at 10:08, Ivan K. <[email protected]> wrote: > > The problem is that I've tried everything I know/found on the internet/was recommended by Codex and GPT, and I (and LLVMs) have run out of ideas :( Could you capture the traffic of the connection experiencing bad throughput on the sender and on the receiver side in a .pcap file and post them or prove a link to them? You can send it privately, if preferred. What TCP stack are you using? Best regards Michael > >> 25 февр. 2026 г., в 10:55, Ivan K. <[email protected]> написал(а): >> >> I don't quite understand your question. >> iperf3 shows 200-300 Mbit through VPS0 and 100 kbit through VPS2 at the same time when I run them from my laptop. >> If I run them from the router, the speeds are good on both tunnels. >> >>> 25 февр. 2026 г., в 01:12, Mark Saad <[email protected]> написал(а): >>> >>> >>> >>> On Tue, Feb 24, 2026 at 3:35 PM Ivan K. <[email protected]> wrote: >>> Yes, I wrote about this in my original letter. It is an Intel Gigabit server network card. I disabled everything possible on it with the command: >>> >>> ifconfig igb0 -mextpg -vlanhwtag -vlanhwfilter -vlanhwtso -vlanhwcsum -rxcsum -txcsum -rxcsum6 -txcsum6 -tso -lro >>> >>> The options that remained are: VLAN_MTU, JUMBO_MTU, WOL_MAGIC, HWSTATS >>> >>>> 24 февр. 2026 г., в 23:23, Mark Saad <[email protected]> написал(а): >>>> >>>> >>>> >>>> On Tue, Feb 24, 2026 at 3:02 PM Ivan K. <[email protected]> wrote: >>>> Each real life - 1500, gre 1472. >>>> I checked this. I tried setting mtu 1200 on the client device. I have mss clamping at 1400 in pf. >>>> If it's mtu, it doesn't explain why the identical tunnel on the same machine functions normally. >>>> >>>> >>>>> 24 февр. 2026 г., в 21:10, Mark Saad <[email protected]> написал(а): >>>>> >>>>> >>>>> >>>>> On Tue, Feb 24, 2026 at 12:19 PM Ivan K. <[email protected]> wrote: >>>>> Hello, everyone. The text below was written by me and “polished” by a neural network, which I have been using to research this problem for three days now. >>>>> >>>>> I am investigating a reproducible GRE issue and would appreciate advice on further debugging. >>>>> >>>>> Versions / hosts >>>>> - Router: FreeBSD 14.3-RELEASE-p8 >>>>> - VPS0: FreeBSD 14.3-RELEASE-p8 >>>>> - VPS2: FreeBSD 14.3-RELEASE-p8 >>>>> >>>>> Topology >>>>> - Router has WAN on VLAN interface `onp` (parent `igb0`) and LAN on `igb0`. >>>>> - Two GRE tunnels from the same router: >>>>> - GRE to VPS0: healthy >>>>> - GRE to VPS2: problematic >>>>> - There is also an ng tunnel to VPS2 (works better than bad GRE, but not full speed). >>>>> - Clients behind router tested via both Wi-Fi and wired Ethernet (same problem). >>>>> - Also tested traffic from a VM on the router: VM path does not reproduce the severe collapse. >>>>> >>>>> Core symptom (asymmetric) >>>>> - Only TX direction is affected (client uploads through GRE->VPS2). >>>>> - RX/download direction is not affected. >>>>> - Through GRE->VPS2, TCP upload collapses to about ~100 Kbit/s per flow. >>>>> - 5 parallel TCP flows produce roughly 5 x ~100 Kbit/s. >>>>> - iperf3 TCP shows many retransmissions and very small cwnd (~2.7 KB). >>>>> - Same client via GRE->VPS0 is fast (normal high throughput). >>>>> - Traffic generated locally on router through GRE->VPS2 is also fast. >>>>> >>>>> Firewall / policy checks >>>>> - PF disabled completely (`pfctl -d`) for tests: problem remains. >>>>> - Therefore this is not PF rule/policy dependent in current testing. >>>>> - Also checked for ipfw/dummynet: >>>>> - `kldstat | egrep 'ipfw|dummynet|pf'` shows only `pf.ko`. >>>>> - `ipfw show` returns “Protocol not available”. >>>>> - `net.inet.ip.fw.enable` OID not present. >>>>> >>>>> Offload / NIC checks performed >>>>> - Explicitly disabled offload features (multiple attempts), including: >>>>> - `-lro` >>>>> - `-tso` / tso-related toggles >>>>> - `-rxcsum` / `-txcsum` / v6 checksum toggles >>>>> - VLAN hardware offload related toggles where supported >>>>> - No improvement. >>>>> - Current interface snapshot during testing: >>>>> - `ifconfig igb0` shows options: `VLAN_MTU,JUMBO_MTU,WOL_MAGIC,HWSTATS` (no LRO/TSO flags active) >>>>> - `ifconfig onp` shows `options=0` >>>>> - Reboots of router and VPS2 done after changes: no improvement. >>>>> - Also tested with minimal manual GRE setup (without tunnel-management scripts, only one GRE up): no improvement. >>>>> >>>>> Queue/counters/sysctl checks >>>>> - `netstat -Q` shows no queue drops (`QDrops=0`). >>>>> - `netstat -iW -I ...` during tests shows no obvious interface error growth. >>>>> - `sysctl net.inet.ip.fastforwarding` does not exist on this release. >>>>> >>>>> UDP vs TCP behavior >>>>> - UDP over problematic GRE path can run at high throughput (similar order from client and from router), i.e., no analogous dramatic collapse. >>>>> - TCP over same path collapses hard in TX direction from forwarded clients. >>>>> >>>>> Packet capture findings >>>>> - Collected captures at multiple points (router and VPS side). >>>>> - In bad runs, client-side TCP packets are seen, but many segments are retransmitted repeatedly. >>>>> - Comparing capture points suggests segments often appear on VPS GRE side only much later / after retransmission, consistent with loss on the GRE path for this forwarded TCP case. >>>>> - In control runs (local router-generated traffic), packet timing/count alignment is normal. >>>>> >>>>> What makes this puzzling >>>>> - Same router, same OS version, same general config: >>>>> - GRE to VPS0 works well. >>>>> - GRE to VPS2 fails for forwarded TCP TX only. >>>>> - PF disabled does not change behavior. >>>>> - Offloads disabled does not change behavior. >>>>> - No obvious drops in standard queue/interface counters. >>>>> >>>>> Request >>>>> Any suggestions on what to inspect next for FreeBSD 14.3-p8 in this scenario: >>>>> - netisr / forwarding path diagnostics >>>>> - GRE-specific kernel instrumentation >>>>> - known igb/VLAN/GRE forwarding edge cases >>>>> - any tunables or debug knobs that could explain “forwarded TCP TX collapse only on one GRE peer/path” >>>>> >>>>> Thanks. >>>>> >>>>> >>>>> Ivan >>>>> Just a guess but it sounds like you may have a MTU issue . What is the MTU settings on each of the real interfaces and the gre tunnels ? >>>>> -- >>>>> mark saad | [email protected] >>>> >>>> Ivan >>>> What sort of network cards are you using? This is another long shot but do you have anything that does GRE offload? Some of the intel >>>> and mellanox cards support some sort of checksum offloads. >>>> >>>> -- >>>> mark saad | [email protected] >>> >>> Ivan >>> So back to how you are testing. Do you have iperf / iperf3 blasting traffic from router to target VPS0 and VPS2 at the same time ? >>> >>> >>> -- >>> mark saad | [email protected] >> >