Re: Asymmetric forwarded TCP TX collapse over one GRE path on FreeBSD 14.3-RELEASE-p8 (PF independent, offloads disabled)
"Ivan K." <[email protected]>
| Newsgroups | gmane.os.freebsd.devel.net |
|---|---|
| Message-ID | <[email protected]> |
The problem is that I've tried everything I know/found on the internet/was recommended by Codex and GPT, and I (and LLVMs) have run out of ideas :( > 25 февр. 2026 г., в 10:55, Ivan K. <[email protected]> написал(а): > > I don't quite understand your question. > iperf3 shows 200-300 Mbit through VPS0 and 100 kbit through VPS2 at the same time when I run them from my laptop. > If I run them from the router, the speeds are good on both tunnels. > >> 25 февр. 2026 г., в 01:12, Mark Saad <[email protected]> написал(а): >> >> >> >> On Tue, Feb 24, 2026 at 3:35 PM Ivan K. <[email protected] <mailto:[email protected]>> wrote: >>> Yes, I wrote about this in my original letter. It is an Intel Gigabit server network card. I disabled everything possible on it with the command: >>> >>> ifconfig igb0 -mextpg -vlanhwtag -vlanhwfilter -vlanhwtso -vlanhwcsum -rxcsum -txcsum -rxcsum6 -txcsum6 -tso -lro >>> >>> The options that remained are: VLAN_MTU, JUMBO_MTU, WOL_MAGIC, HWSTATS >>> >>>> 24 февр. 2026 г., в 23:23, Mark Saad <[email protected] <mailto:[email protected]>> написал(а): >>>> >>>> >>>> >>>> On Tue, Feb 24, 2026 at 3:02 PM Ivan K. <[email protected] <mailto:[email protected]>> wrote: >>>>> Each real life - 1500, gre 1472. >>>>> I checked this. I tried setting mtu 1200 on the client device. I have mss clamping at 1400 in pf. >>>>> If it's mtu, it doesn't explain why the identical tunnel on the same machine functions normally. >>>>> >>>>> >>>>>> 24 февр. 2026 г., в 21:10, Mark Saad <[email protected] <mailto:[email protected]>> написал(а): >>>>>> >>>>>> >>>>>> >>>>>> On Tue, Feb 24, 2026 at 12:19 PM Ivan K. <[email protected] <mailto:[email protected]>> wrote: >>>>>>> Hello, everyone. The text below was written by me and “polished” by a neural network, which I have been using to research this problem for three days now. >>>>>>> >>>>>>> I am investigating a reproducible GRE issue and would appreciate advice on further debugging. >>>>>>> >>>>>>> Versions / hosts >>>>>>> - Router: FreeBSD 14.3-RELEASE-p8 >>>>>>> - VPS0: FreeBSD 14.3-RELEASE-p8 >>>>>>> - VPS2: FreeBSD 14.3-RELEASE-p8 >>>>>>> >>>>>>> Topology >>>>>>> - Router has WAN on VLAN interface `onp` (parent `igb0`) and LAN on `igb0`. >>>>>>> - Two GRE tunnels from the same router: >>>>>>> - GRE to VPS0: healthy >>>>>>> - GRE to VPS2: problematic >>>>>>> - There is also an ng tunnel to VPS2 (works better than bad GRE, but not full speed). >>>>>>> - Clients behind router tested via both Wi-Fi and wired Ethernet (same problem). >>>>>>> - Also tested traffic from a VM on the router: VM path does not reproduce the severe collapse. >>>>>>> >>>>>>> Core symptom (asymmetric) >>>>>>> - Only TX direction is affected (client uploads through GRE->VPS2). >>>>>>> - RX/download direction is not affected. >>>>>>> - Through GRE->VPS2, TCP upload collapses to about ~100 Kbit/s per flow. >>>>>>> - 5 parallel TCP flows produce roughly 5 x ~100 Kbit/s. >>>>>>> - iperf3 TCP shows many retransmissions and very small cwnd (~2.7 KB). >>>>>>> - Same client via GRE->VPS0 is fast (normal high throughput). >>>>>>> - Traffic generated locally on router through GRE->VPS2 is also fast. >>>>>>> >>>>>>> Firewall / policy checks >>>>>>> - PF disabled completely (`pfctl -d`) for tests: problem remains. >>>>>>> - Therefore this is not PF rule/policy dependent in current testing. >>>>>>> - Also checked for ipfw/dummynet: >>>>>>> - `kldstat | egrep 'ipfw|dummynet|pf'` shows only `pf.ko`. >>>>>>> - `ipfw show` returns “Protocol not available”. >>>>>>> - `net.inet.ip.fw.enable` OID not present. >>>>>>> >>>>>>> Offload / NIC checks performed >>>>>>> - Explicitly disabled offload features (multiple attempts), including: >>>>>>> - `-lro` >>>>>>> - `-tso` / tso-related toggles >>>>>>> - `-rxcsum` / `-txcsum` / v6 checksum toggles >>>>>>> - VLAN hardware offload related toggles where supported >>>>>>> - No improvement. >>>>>>> - Current interface snapshot during testing: >>>>>>> - `ifconfig igb0` shows options: `VLAN_MTU,JUMBO_MTU,WOL_MAGIC,HWSTATS` (no LRO/TSO flags active) >>>>>>> - `ifconfig onp` shows `options=0` >>>>>>> - Reboots of router and VPS2 done after changes: no improvement. >>>>>>> - Also tested with minimal manual GRE setup (without tunnel-management scripts, only one GRE up): no improvement. >>>>>>> >>>>>>> Queue/counters/sysctl checks >>>>>>> - `netstat -Q` shows no queue drops (`QDrops=0`). >>>>>>> - `netstat -iW -I ...` during tests shows no obvious interface error growth. >>>>>>> - `sysctl net.inet.ip.fastforwarding` does not exist on this release. >>>>>>> >>>>>>> UDP vs TCP behavior >>>>>>> - UDP over problematic GRE path can run at high throughput (similar order from client and from router), i.e., no analogous dramatic collapse. >>>>>>> - TCP over same path collapses hard in TX direction from forwarded clients. >>>>>>> >>>>>>> Packet capture findings >>>>>>> - Collected captures at multiple points (router and VPS side). >>>>>>> - In bad runs, client-side TCP packets are seen, but many segments are retransmitted repeatedly. >>>>>>> - Comparing capture points suggests segments often appear on VPS GRE side only much later / after retransmission, consistent with loss on the GRE path for this forwarded TCP case. >>>>>>> - In control runs (local router-generated traffic), packet timing/count alignment is normal. >>>>>>> >>>>>>> What makes this puzzling >>>>>>> - Same router, same OS version, same general config: >>>>>>> - GRE to VPS0 works well. >>>>>>> - GRE to VPS2 fails for forwarded TCP TX only. >>>>>>> - PF disabled does not change behavior. >>>>>>> - Offloads disabled does not change behavior. >>>>>>> - No obvious drops in standard queue/interface counters. >>>>>>> >>>>>>> Request >>>>>>> Any suggestions on what to inspect next for FreeBSD 14.3-p8 in this scenario: >>>>>>> - netisr / forwarding path diagnostics >>>>>>> - GRE-specific kernel instrumentation >>>>>>> - known igb/VLAN/GRE forwarding edge cases >>>>>>> - any tunables or debug knobs that could explain “forwarded TCP TX collapse only on one GRE peer/path” >>>>>>> >>>>>>> Thanks. >>>>>>> >>>>>> >>>>>> >>>>>> Ivan >>>>>> Just a guess but it sounds like you may have a MTU issue . What is the MTU settings on each of the real interfaces and the gre tunnels ? >>>>>> -- >>>>>> mark saad | [email protected] <mailto:[email protected]> >>>>> >>>> >>>> Ivan >>>> What sort of network cards are you using? This is another long shot but do you have anything that does GRE offload? Some of the intel >>>> and mellanox cards support some sort of checksum offloads. >>>> >>>> -- >>>> mark saad | [email protected] <mailto:[email protected]> >>> >> >> Ivan >> So back to how you are testing. Do you have iperf / iperf3 blasting traffic from router to target VPS0 and VPS2 at the same time ? >> >> >> -- >> mark saad | [email protected] <mailto:[email protected]> >