RE: Please read - proposed WG termination
Bernard King-Smith <[email protected]> Wed, 31 Aug 2005 20:10:51 -0400
| Newsgroups | gmane.ietf.ipoib |
|---|---|
| Message-ID | <OFEAD798BC.5BD12489-ON8525706E.00447E83-8525706F.0000FE88@us.ibm.com> |
Having IPoIB-CM is a very important feature to make IB a viable interconnect in clustered systems. Without IPoIB-CM for HPC clusters, and commercial clusters using IB to SAN, you need two networks for good total cluster performance, one for IB ( non-IP traffic ) and an IP performance network like GigE. This means that IB is not cost effective as the GigE ( or 10 GigE) network which handles both types of traffic reasonably. In the HPC world most clusters use the cluster fabric ( and IB is the future direction ) for both MPI and IP traffic. The IP traffic is usually for parallel file systems and system management and control. This high bandwidth IP network is required in most production HPC clusters. With the current IPoIB only using UD, the performance is dismal. Our simulations using the small packet MTU of IB says that the parallel file systems ( GPFS, PVFS, Lustre etc ) can only get 25% of a 4X IB link today and at 12X it will be about 10%. The problem is that the IP drivers are single threaded per adapter. Also the CPU utilization of TCP/IP at a MTU of the IB link very high because of the per packet stack processing. Going to IPoIB-CM means we can cut down the number of TCP/IP stack traversals from 32 to 1 for a 60K IP packet. This means that you have 30 times as much data transmitted per device driver call. This will enable IP to show similar bandwidth with multiple sockets as other protocols that can use RC or fragment within the device driver. For commercial clusters, if IB is used for storage, then you save a network by having fast IP performance and can use the IB network for both. Why use IB and another network for the commercial cluster, when the other network supports similar bandwidth for storage and IP. Implementing IPoIB-CM makes IB viable in the HPC cluster and some commercial clusters. Otherwise I don't think it competes economically with other network technologies. Regards. Bernie King-Smith IBM Corporation Server Group Cluster System Performance [email protected] (845)433-8483 Tie. 293-8483 or wombat2 on NOTES "We are not responsible for the world we are born into, only for the world we leave when we die. So we have to accept what has gone before us and work to change the only thing we can, -- The Future." William Shatner Dror Goldenberg <[email protected] o.il> To Sent by: [email protected], ipoverib-bounces@ "H.K. Jerry Chu" ietf.org <[email protected]> cc [email protected], 08/30/2005 09:32 [email protected], AM [email protected] Subject RE: [Ipoverib] Please read - proposed WG termination > From: Vivek Kashyap [mailto:[email protected]] > Sent: Tuesday, August 30, 2005 8:39 AM > > On Mon, 29 Aug 2005, H.K. Jerry Chu wrote: > <snip> > > 1. IPoIB connected mode draft-ietf-ipoib-connected-mode-00.txt > > updated recently > > Well, in recent days there has been a discussion going on > based on Dror's input. I also made some updates after some > discussion on OpenIB (not on > IETF though). This draft itself became a working group draft > this february > after some lively discussion just before that. It appears to > me that we > should be possible to finalise this draft soon enough. > > 20th sept. might be long enough to know one way or the other... > > vivek > We would like to see IPoIB-CM being finalized in IETF. We see great value in having a standard for connected mode which effectively increases the MTU. We are willing to contribute to the standardization effort. We're also looking at the implementation of IPoIB-CM in Linux. -Dror _______________________________________________ IPoverIB mailing list [email protected] https://www1.ietf.org/mailman/listinfo/ipoverib