RE: comments on draft-kashyap-ipoib-connected-mode-02.txt

Michael Krause <[email protected]> Mon, 13 Dec 2004 10:07:02 -0800
Newsgroups gmane.ietf.ipoib
Message-ID <[email protected]>
At 06:53 AM 12/11/2004, Bill Strahm wrote:
>OK - a couple of misconceptions
>
>IB Layer 3 is NOT IPv6.  It has this thing that looks ALMOST like an
>IPv6 header - but it does routing completely differently.  Infact it
>isn't until you do IB Routing that you would even put a GRH (Global
>Route Header) that is what ends up looking like a IPv6 header.  If you
>are staying on the local IB subnet - you are only required to put a LRH
>(Local Route Header) that looks nothing like IPv6.
>
>The reliability mechanism is not TCP therefore.  I will agree with
>Michael on this one.  That said - the same problems that you discuss can
>happen between the layer 4 IB and the Layer 4 TCP/IP.
>
>I do not believe there will be a problem however because the timers are
>on such a different scale.  Let me try and give an example... There are
>retries in Ethernet (I am talking about collision detection and back
>off) - What if we decided to worry about TCP retransmitting because the
>Layer 2 packet hasn't been able to get on the wire and the TCP timers
>went off and started retransmitting ?  People don't worry because the
>timers are SO different - this will be even more pronounced in the IB
>world.

True.  One can configure the IB timers to be rather small or even 
infinite.  In practice, the timers will be no more than a couple of seconds 
worst case but had been envisioned as measured in milli-seconds in practice 
as the bandwidths combined with the switch latencies are such that large 
timers do not make a lot of sense.  When we were developing the IB 
transports, a lot of discussion went into how things change when moving 
from a Gbps world to N GBps world.  Things like congestion management 
become more complicated as attempting to detect and adjust injection rates 
for something that may be momentary burst can cause such oscillations in 
the performance that one needs to consider two approaches.  Detect and 
monitor for N events in a given period of time.  If sustained, then examine 
injection rates or adjust path selection to reduce the number of 
events.   This is done on more of a global basis as one may not want to 
limit this on a per endnode basis depending upon what type and priority a 
given endnode operates.  Or, detect and monitor and adjust on a per endnode 
basis assuming that all endnodes are of equal priority.  One can also 
adjust the VL arbitrations, etc. as well to effect change without impacting 
injection rate simply because of the service rates for the VL and the 
implicit flow control that occurs which will cause appropriate 
back-pressure.  The IBTA completed the congestion spec a couple of months 
ago and it is worth a read if people have interest.


>I don't think we will be able to move the MTU above 64K - just because I
>believe the IP packet size is (was ??? Has the size parameter gone up ?)
>64K so don't worry about someone trying to stick a 2G IP packet onto the
>wire.  At the same time if it is possible to specify a single IP packet
>that is 2G, maybe we need some text staying the MTU should be limited to
>something sane

TCP Jumbo I thought could go to 256K.

Mike


>Bill
>
>-----Original Message-----
>From: [email protected] [mailto:[email protected]] On
>Behalf Of Margaret Wasserman
>Sent: Saturday, December 11, 2004 4:47 AM
>To: [email protected]
>Subject: Re: [Ipoverib] comments on
>draft-kashyap-ipoib-connected-mode-02.txt
>
>Michael Krause wrote:
> >No.  It uses InfiniBand RC or UC to communicate IP datagrams (v4 /
> >v6) between connected endnodes.
>
>And what is InfiniBand RC, under the covers?
>
> >Rest is based on a misconception as there is only IB below IP and no
> >tunneling of TCP over TCP occurs.
>
>Well, I know that there is another IP(v6) in there, as IB is very
>closely based on IPv6.  Is RC based on TCP? SCTP?  Or did IB define a
>different reliable, connection-oriented IP protocol?  If the latter,
>what type of retransmission and congestion control algorithms are
>used by IB RC?  And, has there been any study regarding how they
>would interact with the retransmission and congestion control
>algorithms of TCP or other reliable upper-layer transports?
>
>Even if this isn't TCP, per se, there may still be issues, but it
>will take more careful research to determine what they are.
>
>Margaret
>
>
>
>_______________________________________________
>IPoverIB mailing list
>[email protected]
>https://www1.ietf.org/mailman/listinfo/ipoverib
>
>
>_______________________________________________
>IPoverIB mailing list
>[email protected]
>https://www1.ietf.org/mailman/listinfo/ipoverib

_______________________________________________
IPoverIB mailing list
[email protected]
https://www1.ietf.org/mailman/listinfo/ipoverib