RE: Why is MTU an issue? (RE: Please read - proposed W G termination)
Dror Goldenberg <[email protected]> Fri, 2 Sep 2005 22:31:12 +0300
| Newsgroups | gmane.ietf.ipoib |
|---|---|
| Message-ID | <[email protected]> |
> From: Mark Townsley [mailto:[email protected]] > Sent: Friday, September 02, 2005 4:21 PM > > Harald Tveit Alvestrand wrote: > > > --On 1. september 2005 11:19 -0700 "H.K. Jerry Chu" > > <[email protected]> wrote: > > > >> What's more hairy to the OS stack is the per-destination MTU and > >> different > >> MTU for multicast than for unicast inherited in IPoIB CM. > > > > > > I don't quite understand why this should be an issue. > > At the moment, TCP/IPv4 stacks have 3 options: > > > > - Set MTU to 576 bytes (not b* likely) > > - Set MTU to physical interface MTU and depend on > fragmentation in the > > network to save their bacon in the case of smaller MTUs > (*curse* PPPoE) > > - Set initial MTU to physical interface MTU, set DF bit, > and support > > Path MTU discovery > > > > I *think* all "modern" IP stacks take the last choice. Which means > > that they have built-in support for variable MTU already. > > Many do, but then ICMP is filtered with reckless abandon on > the Internet > (big I), breaking the feedback loop for a perfectly good PMTU enabled > stack. MTU must still hand-managed *way* more than it should be. > > - Mark > I think that there is an even more fundamental issue than ICMPs being filtered. ICMPs for "Fragmentation Needed and DF Set" are generated by routers. So, for local subnet traffic, RFC 1191 (PMTU Discovery) won't work to my understanding. PMTU discovery is a layer 3 protocol, not layer 2. For the specific case of ipoib-cm with local subnet traffic. A local "huge" frame will be handed over to the layer 2 IB interface for transmission. If it doesn't manage to send that large MTU because there isn't such a large MTU between the local node and its peer, then the frame is silenty discarded by the TRANSMITTER. One way to go now is to generate a local ICMP message saying that this was a too large frame. But I don't think it's the right way to go. I think that the right approach is to have a per-neighbor MTU in the L2 interface. And furthermore, MTU can be different between multicast and unicast. We're actually trying to see whether a per neighbor MTU can be added to Linux. -Dror _______________________________________________ IPoverIB mailing list [email protected] https://www1.ietf.org/mailman/listinfo/ipoverib