Re: small change in the connected mode draft

Michael Krause <[email protected]> Thu, 23 Mar 2006 12:57:59 -0800
Newsgroups gmane.ietf.ipoib
Message-ID <[email protected]>
--===============1622897801==
Content-Type: multipart/alternative;
	boundary="=====================_572946904==.ALT"

--=====================_572946904==.ALT
Content-Type: text/plain; charset="us-ascii"; format=flowed

At 04:04 PM 3/21/2006, H.K. Jerry Chu wrote:
>Hi folks,
>
>Allison Mankin, the transport area AD who brought up some
>concern during IESG review of the connected draft regarding
>simultaneous retransmissions at different layers, has suggested
>and Vivek agreed to the following change to section 7.1
>"A Cautionary Note on IPoIB-RC".
>
>The revised section reads like this:
>
>         The RC mode of InfiniBand guarantees in-order delivery of
>         packets. Every message transmitted over the RC connection is
>         broken into physical MTU sized packets by the RC connection. If
>         any packet is lost, it is retransmitted until the complete
>         message is exchanged. Therefore, there is a possibility of an
>         upper transport layer experiencing a timeout, while the RC layer
>         is still in the process of transferring the complete message.
>         TCP will view the timeout as an indicator of congestion and
>         enter slow-start thereby affecting throughput drastically
>         [RFC2581]. Other upper layer protocols might insert
>         retransmissions into the fabric adding to the already existing
>         congestion.
>
>         The applicability of Infiniband reliability is on a fabric
>         with short latencies (not wide area).  Therefore, the RC timer
>         values should be short compared with the starting minimum
>         time values used by the upper end-to-end transports.  In
>         addition, because the RC mode does not have measurement
>         based reliable transmission, its use over fabrics with long
>         latency or very dynamic latency may be a concern for congestion-
>         aware traffic traversing those fabrics.
>
>If you have any comment/issue on the proposed change please post them
>to the list before COB this Friday (3/24/06). I'd like to move the
>draft past IESG review so we can wrap up the WG asap.
>
>BTW, please be informed that my email address will change after this
>Friday. My new address is [email protected].

What about RNR which can go into an infinite timer state?  It should also 
be noted that IB does not mandate timer ranges that necessarily correspond 
to TCP timers.  In fact, in the face of real IB congestion or a port / VL 
arbitration policy that places such traffic in a best effort QoS slot, then 
it is quite possible to see delays that would be interpreted by any ULP 
such as TCP as congestion.  Is this really an issue as the applications 
continue to operate albeit at a slower rate?

Mike 
--=====================_572946904==.ALT
Content-Type: text/html; charset="us-ascii"

<html>
<body>
<font size=3>At 04:04 PM 3/21/2006, H.K. Jerry Chu wrote:<br>
<blockquote type=cite class=cite cite="">Hi folks,<br><br>
Allison Mankin, the transport area AD who brought up some<br>
concern during IESG review of the connected draft regarding<br>
simultaneous retransmissions at different layers, has suggested<br>
and Vivek agreed to the following change to section 7.1<br>
&quot;A Cautionary Note on IPoIB-RC&quot;.<br><br>
The revised section reads like this:<br><br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>The RC
mode of InfiniBand guarantees in-order delivery of<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>packets.
Every message transmitted over the RC connection is<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>broken
into physical MTU sized packets by the RC connection. If<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>any packet
is lost, it is retransmitted until the complete<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>message is
exchanged. Therefore, there is a possibility of an<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>upper
transport layer experiencing a timeout, while the RC layer<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>is still
in the process of transferring the complete message.<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>TCP will
view the timeout as an indicator of congestion and<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>enter
slow-start thereby affecting throughput drastically<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>[RFC2581].
Other upper layer protocols might insert<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>
retransmissions into the fabric adding to the already existing<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>
congestion.<br><br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>The
applicability of Infiniband reliability is on a fabric<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>with short
latencies (not wide area).&nbsp; Therefore, the RC timer<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>values
should be short compared with the starting minimum<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>time
values used by the upper end-to-end transports.&nbsp; In<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>addition,
because the RC mode does not have measurement<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>based
reliable transmission, its use over fabrics with long<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>latency or
very dynamic latency may be a concern for congestion-<br>
<x-tab>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</x-tab>aware
traffic traversing those fabrics.<br><br>
If you have any comment/issue on the proposed change please post
them<br>
to the list before COB this Friday (3/24/06). I'd like to move the<br>
draft past IESG review so we can wrap up the WG asap.<br><br>
BTW, please be informed that my email address will change after this<br>
Friday. My new address is [email protected].</blockquote><br>
What about RNR which can go into an infinite timer state?&nbsp; It should
also be noted that IB does not mandate timer ranges that necessarily
correspond to TCP timers.&nbsp; In fact, in the face of real IB
congestion or a port / VL arbitration policy that places such traffic in
a best effort QoS slot, then it is quite possible to see delays that
would be interpreted by any ULP such as TCP as congestion.&nbsp; Is this
really an issue as the applications continue to operate albeit at a
slower rate?<br><br>
Mike</font></body>
</html>

--=====================_572946904==.ALT--




--===============1622897801==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
IPoverIB mailing list
[email protected]
https://www1.ietf.org/mailman/listinfo/ipoverib

--===============1622897801==--