Re: Asymmetric forwarded TCP TX collapse over one GRE path on FreeBSD 14.3-RELEASE-p8 (PF independent, offloads disabled)
Ronald Klop <[email protected]>
| Newsgroups | gmane.os.freebsd.devel.net |
|---|---|
| Message-ID | <2067978522.504.1772018196806@localhost> |
A next step could be:
Post more actual data from your system. Your original mail contains interpreted information which might miss the cause. For example post the output of ifconfig, netstat -r -n, dmesg, tcpdump of the slow traffic vs the fast traffic. You can use something like https://pastebin.com/ if the output is very big.
Make an image of how the systems are connected:
Host X --- 192.168.1.* --- Router A -------tunnel-------- Router B --- 192.168.2.* --- Host Y
\ /
\ /
+------ the Internet ------+
This helps people understand what is talking to what.
Try to test individual systems, like
iperf Host X to Router A
iperf router A to Router B (directyly as well as via the tunnel)
iperf Router B to Host Y.
iperf Host X to Router B.
iperf Host X to Host Y.
This could help understanding where the bottleneck is happening.
Try another network card. Maybe even a cheap USB-dongle and see if that makes a difference. This could indicate something with the igb-device.
I noticed you use the same igb0 for WAN (via VLAN) as well for LAN access. Is it possible to add a second network card to split WAN and LAN traffic physically?
Regards,
Ronald.
Van: "Ivan K." <[email protected]>
Datum: woensdag, 25 februari 2026 10:08
Aan: Mark Saad <[email protected]>
CC: [email protected]
Onderwerp: Re: Asymmetric forwarded TCP TX collapse over one GRE path on FreeBSD 14.3-RELEASE-p8 (PF independent, offloads disabled)
>
> The problem is that I've tried everything I know/found on the internet/was recommended by Codex and GPT, and I (and LLVMs) have run out of ideas :(
> >>
>> 25 . 2026., 10:55, Ivan K. <[email protected]> ():
>>
>> I don't quite understand your question.
>> iperf3 shows 200-300 Mbit through VPS0 and 100 kbit through VPS2 at the same time when I run them from my laptop.
>> If I run them from the router, the speeds are good on both tunnels.
>> >>>
>>> 25 . 2026., 01:12, Mark Saad <[email protected]> ():
>>>
>>>
>>>
>>> On Tue, Feb 24, 2026 at 3:35PM Ivan K. <[email protected]> wrote:
>>>>
>>>> Yes, I wrote about this in my original letter. It is an Intel Gigabit server network card. I disabled everything possible on it with the command:
>>>>
>>>> ifconfig igb0 -mextpg -vlanhwtag -vlanhwfilter -vlanhwtso -vlanhwcsum -rxcsum -txcsum -rxcsum6 -txcsum6 -tso -lro
>>>>
>>>> The options that remained are: VLAN_MTU, JUMBO_MTU, WOL_MAGIC, HWSTATS
>>>> >>>>>
>>>>> 24 . 2026., 23:23, Mark Saad <[email protected]> ():
>>>>>
>>>>>
>>>>>
>>>>> On Tue, Feb 24, 2026 at 3:02PM Ivan K. <[email protected]> wrote:
>>>>>>
>>>>>> Each real life - 1500, gre 1472.
>>>>>> I checked this. I tried setting mtu 1200 on the client device. I have mss clamping at 1400 in pf.
>>>>>> If it's mtu, it doesn't explain why the identical tunnel on the same machine functions normally.
>>>>>>
>>>>>> >>>>>>>
>>>>>>> 24 . 2026., 21:10, Mark Saad <[email protected]> ():
>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> On Tue, Feb 24, 2026 at 12:19PM Ivan K. <[email protected]> wrote:
>>>>>>>>
>>>>>>>> Hello, everyone. The text below was written by me and “polished” by a neural network, which I have been using to research this problem for three days now.
>>>>>>>>
>>>>>>>> I am investigating a reproducible GRE issue and would appreciate advice on further debugging.
>>>>>>>>
>>>>>>>> Versions / hosts
>>>>>>>> - Router: FreeBSD 14.3-RELEASE-p8
>>>>>>>> - VPS0: FreeBSD 14.3-RELEASE-p8
>>>>>>>> - VPS2: FreeBSD 14.3-RELEASE-p8
>>>>>>>>
>>>>>>>> Topology
>>>>>>>> - Router has WAN on VLAN interface `onp` (parent `igb0`) and LAN on `igb0`.
>>>>>>>> - Two GRE tunnels from the same router:
>>>>>>>> - GRE to VPS0: healthy
>>>>>>>> - GRE to VPS2: problematic
>>>>>>>> - There is also an ng tunnel to VPS2 (works better than bad GRE, but not full speed).
>>>>>>>> - Clients behind router tested via both Wi-Fi and wired Ethernet (same problem).
>>>>>>>> - Also tested traffic from a VM on the router: VM path does not reproduce the severe collapse.
>>>>>>>>
>>>>>>>> Core symptom (asymmetric)
>>>>>>>> - Only TX direction is affected (client uploads through GRE->VPS2).
>>>>>>>> - RX/download direction is not affected.
>>>>>>>> - Through GRE->VPS2, TCP upload collapses to about ~100 Kbit/s per flow.
>>>>>>>> - 5 parallel TCP flows produce roughly 5 x ~100 Kbit/s.
>>>>>>>> - iperf3 TCP shows many retransmissions and very small cwnd (~2.7 KB).
>>>>>>>> - Same client via GRE->VPS0 is fast (normal high throughput).
>>>>>>>> - Traffic generated locally on router through GRE->VPS2 is also fast.
>>>>>>>>
>>>>>>>> Firewall / policy checks
>>>>>>>> - PF disabled completely (`pfctl -d`) for tests: problem remains.
>>>>>>>> - Therefore this is not PF rule/policy dependent in current testing.
>>>>>>>> - Also checked for ipfw/dummynet:
>>>>>>>> - `kldstat | egrep 'ipfw|dummynet|pf'` shows only `pf.ko`.
>>>>>>>> - `ipfw show` returns “Protocol not available”.
>>>>>>>> - `net.inet.ip.fw.enable` OID not present.
>>>>>>>>
>>>>>>>> Offload / NIC checks performed
>>>>>>>> - Explicitly disabled offload features (multiple attempts), including:
>>>>>>>> - `-lro`
>>>>>>>> - `-tso` / tso-related toggles
>>>>>>>> - `-rxcsum` / `-txcsum` / v6 checksum toggles
>>>>>>>> - VLAN hardware offload related toggles where supported
>>>>>>>> - No improvement.
>>>>>>>> - Current interface snapshot during testing:
>>>>>>>> - `ifconfig igb0` shows options: `VLAN_MTU,JUMBO_MTU,WOL_MAGIC,HWSTATS` (no LRO/TSO flags active)
>>>>>>>> - `ifconfig onp` shows `options=0`
>>>>>>>> - Reboots of router and VPS2 done after changes: no improvement.
>>>>>>>> - Also tested with minimal manual GRE setup (without tunnel-management scripts, only one GRE up): no improvement.
>>>>>>>>
>>>>>>>> Queue/counters/sysctl checks
>>>>>>>> - `netstat -Q` shows no queue drops (`QDrops=0`).
>>>>>>>> - `netstat -iW -I ...` during tests shows no obvious interface error growth.
>>>>>>>> - `sysctl net.inet.ip.fastforwarding` does not exist on this release.
>>>>>>>>
>>>>>>>> UDP vs TCP behavior
>>>>>>>> - UDP over problematic GRE path can run at high throughput (similar order from client and from router), i.e., no analogous dramatic collapse.
>>>>>>>> - TCP over same path collapses hard in TX direction from forwarded clients.
>>>>>>>>
>>>>>>>> Packet capture findings
>>>>>>>> - Collected captures at multiple points (router and VPS side).
>>>>>>>> - In bad runs, client-side TCP packets are seen, but many segments are retransmitted repeatedly.
>>>>>>>> - Comparing capture points suggests segments often appear on VPS GRE side only much later / after retransmission, consistent with loss on the GRE path for this forwarded TCP case.
>>>>>>>> - In control runs (local router-generated traffic), packet timing/count alignment is normal.
>>>>>>>>
>>>>>>>> What makes this puzzling
>>>>>>>> - Same router, same OS version, same general config:
>>>>>>>> - GRE to VPS0 works well.
>>>>>>>> - GRE to VPS2 fails for forwarded TCP TX only.
>>>>>>>> - PF disabled does not change behavior.
>>>>>>>> - Offloads disabled does not change behavior.
>>>>>>>> - No obvious drops in standard queue/interface counters.
>>>>>>>>
>>>>>>>> Request
>>>>>>>> Any suggestions on what to inspect next for FreeBSD 14.3-p8 in this scenario:
>>>>>>>> - netisr / forwarding path diagnostics
>>>>>>>> - GRE-specific kernel instrumentation
>>>>>>>> - known igb/VLAN/GRE forwarding edge cases
>>>>>>>> - any tunables or debug knobs that could explain “forwarded TCP TX collapse only on one GRE peer/path”
>>>>>>>>
>>>>>>>> Thanks.
>>>>>>>>
>>>>>>>
>>>>>>>
>>>>>>> Ivan
>>>>>>> Just a guess but it sounds like you may have a MTU issue . What is the MTU settings on each of the real interfaces and the gre tunnels ?
>>>>>>> --
>>>>>>> mark saad | [email protected]
>>>>>>
>>>>>
>>>>> Ivan
>>>>> What sort of network cards are you using? This is another long shot but do you have anything that does GRE offload? Some of the intel
>>>>> and mellanox cards support some sort of checksum offloads.
>>>>> --
>>>>> mark saad | [email protected]
>>>>
>>>
>>> Ivan
>>> So back to how you are testing. Do you have iperf / iperf3 blasting traffic from router to target VPS0 and VPS2 at the same time ?
>>>
>>> --
>>> mark saad | [email protected]
>>
>