RE: Please read - proposed WG termination
Bernard King-Smith <[email protected]> Thu, 1 Sep 2005 15:43:26 -0400
| Newsgroups | gmane.ietf.ipoib |
|---|---|
| Message-ID | <OF5AB40320.F29BE5A6-ON8525706F.006B5976-8525706F.006C58C9@us.ibm.com> |
> At 05:10 PM 8/31/2005, Bernard King-Smith wrote:
>
>
>
>
> Having IPoIB-CM is a very important feature to make IB a viable
> interconnect in clustered systems. Without IPoIB-CM for HPC
clusters, and
> commercial clusters using IB to SAN, you need two networks for
good total
> cluster performance, one for IB ( non-IP traffic ) and an IP
performance
> network like GigE. This means that IB is not cost effective as
the GigE (
> or 10 GigE) network which handles both types of traffic
reasonably.
>
> In the HPC world most clusters use the cluster fabric ( and IB is
the
> future direction ) for both MPI and IP traffic. The IP traffic is
usually
> for parallel file systems and system management and control. This
high
> bandwidth IP network is required in most production HPC clusters.
With the
> current IPoIB only using UD, the performance is dismal. Our
simulations
> using the small packet MTU of IB says that the parallel file
systems (
> GPFS, PVFS, Lustre etc ) can only get 25% of a 4X IB link today
and at 12X
> it will be about 10%. The problem is that the IP drivers are
single
> threaded per adapter. Also the CPU utilization of TCP/IP at a MTU
of the IB
> link very high because of the per packet stack processing. Going
to
> IPoIB-CM means we can cut down the number of TCP/IP stack
traversals from
> 32 to 1 for a 60K IP packet. This means that you have 30 times as
much data
> transmitted per device driver call. This will enable IP to show
similar
> bandwidth with multiple sockets as other protocols that can use RC
or
> fragment within the device driver.
>
> These performance problems are primarily implementation-specific and
have little to do with IB technology itself. In > > addition, nearly
all IB solutions use a 2KB not the smallest MTU to transfer data - no
different than Ethernet. As I and > others have raised over the years,
the enablement of IP over IB to perform well is a local HCA issue not a
standards > > issue. Addition of checksum off-load support to the HCA is
rather trivial and does not require standardization (this is> > what is
done for Ethernet today and is non-standard). Addition of large send
off-load support is a local HCA issue not > a standards issue and
effectively provides the same benefit as connected mode. The use of
multiple QP to spread work > across CPU for both send / receive ala the
multi-queue support I've worked with various Ethernet IHV to get in place
is > again a local HCA issue (does not have to be visible as part of the
layer 2 address resolution). One can construct a very > nice performing
IP over IB solution but there hasn't been much public progress to implement
these de facto capabilities > > found in Ethernet solutions on IB.
Getting these into a HCA implementation is a heck of a lot easier and
faster to do > > than to develop a standard and getting all of the OS
changes made (the HCA implementation issues can all be done > > >
underneath the IP stack just like with Ethernet so no real OS impacts).
>
I understand the point of putting expansions into the HCA or stack that
don't already exist, just like jumboframes and large_send were added by
vendors for GigE. However, where I lost your logic was when you recommend
IP not to use the RC function, already implemented in all adapters, but
instead ask vendors to add an additional (duplicate) function to IB
adapters to provide RC equivalence specifically for IP.
>
> For commercial clusters, if IB is used for storage, then you save
a network
> by having fast IP performance and can use the IB network for both.
Why use
> IB and another network for the commercial cluster, when the other
network
> supports similar bandwidth for storage and IP.
>
> There will always be Ethernet in any cluster so the fabric is there.
The question is whether it is just for low-bandwidth /
> management services or for applications. For storage, need to separate
the discussion into whether it is block or file. For > block, IB
gateways to Fibre Channel, etc. can and are being used today quite nicely.
Performance is reasonable and the > > ecosystem costs, target
availability, customer "pain", etc. are much lower than attempting to move
to native IB storage. The > same applies to file based where IB gateways
to Ethernet which then attaches to file servers works quite nicely. In
fact, the > original vision of IB was that of an I/O fabric to create
modular server solutions. The addition of IPC came later in the > > > >
process when it was found to be relatively low cost to define. So, IB is
successful in the HPC world and slowly entering > > > some commercial
solutions. To state that its future relies on getting an IP over IB RC
solution is perhaps blowing it a bit out > of proportion. The easier
path for all is to simply use the techniques I and others have advocated
for years now and solve > the problems within the HCA implementation.
Much lower costs and will result in delivering a good performance solution.
>
> BTW, RNIC / Ethernet solutions implement these techniques today. With
the arrival of 10 GbE and the lower prices of RNIC > and 10 GbE switch
ports, lower latency switches (competitive enough with IB for commercial
and many HPC clusters), etc. > > the success of IB must lie elsewhere and
not on an IETF spec. This was noted at the recent IEEE Hot Interconnects
> > > > conference as well so isn't just my opinion.
>
Mike
Regards.
Bernie King-Smith
IBM Corporation
Server Group
Cluster System Performance
[email protected] (845)433-8483
Tie. 293-8483 or wombat2 on NOTES
"We are not responsible for the world we are born into, only for the world
we leave when we die.
So we have to accept what has gone before us and work to change the only
thing we can,
-- The Future." William Shatner