[ippm] Soliciting comments on draft-ye-ippm-switching-efficien cy
Weiqiang Sun <[email protected]> Sun, 26 Jul 2026 10:54:24 +0800
| Newsgroups | gmane.ietf.ippm,gmane.ietf.bmwg |
|---|---|
| Message-ID | <[email protected]> |
--===============4081426002303708968== Content-Type: multipart/alternative; boundary="Apple-Mail=_4F1B16A5-842F-4BE8-BF3B-5792FCF7D503" --Apple-Mail=_4F1B16A5-842F-4BE8-BF3B-5792FCF7D503 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=utf-8 Hello IPPM/BMWG! This is a much belated solicitation for comments on = https://datatracker.ietf.org/doc/draft-ye-ippm-switching-efficiency/. = The draft was uploaded in April and presented on IETF 126 in Viena. The questions we ask is: when ASICs with different capacities (radix = times port capacity) are deployed in AI fabrics following different = design philosophies, i.e., fabric architectures (spine/leaf, dragonfly, = or 3d-torus), are they efficiently utilized? Are there spaces to make = the AI fabric leaner and greener? Switching efficiency, which we shall = now call Fabric Switching Efficiency =E2=80=93 FSE, measures = computationally effective data throughput (analogous to algorithmic = bandwidth) per unit of deployed switching capacity. FSE characterizes = how the AI fabric design (topology, routing and use of = in-network-computing) fit the tasks run on the fabric. One interesting = analogy, which I should borrow from one expert I talked to these days, = is PUE (Power Usage Effectiveness). PUE is a high level metric. It does = not take into account the complexity of the operations inside the DCs, = but it effectively captures whether energy/power is efficiently used for = compute in DCs. With FSE, we aim to provide a measure that can be easily = understood and communicated just like PUE. One interesting feature of = FSE is that it can be very naturally decomposed into 3 finer granular = metrics: data efficiency, port utilization and routing efficiency, each = capturing one specific AI fabric design target. FSE can be used at least = by two groups of people: 1) AIDC network architects who would like to = improve the fabric design. 2) anyone who likes to operate a leaner / = greener AIDC network.=20 We are aware that the set of AI fabric benchmarking documents are = already pretty complete with well defined scope/terminology, and = methodologies (for inference and training). The documents provides = performance metrics (end to end) for AI fabrics from the eyes of model = users who runs inference or training jobs. FSE operates at higher = levels. We don=E2=80=99t intend to provide performance metrics to be = used by LLM model developer/deployment people.=20 The natural next step is to turn this concept into methodologies that = IPPMers/BMWGers/BPMers are interested in and would be willing to = contribute to. We will be happy to receive any feedbacks and work with = the community to move this work forward. For folks who are interested, please also read our paper on arXiv - = arXiv:2604.14690 (https://arxiv.org/abs/2604.14690), in which the = rationale and numerical results are discussed in much more detail. We = are still improving this paper by adding more packet level simulations = to reveal how the metrics behave when different compute / communication = kernels are executed and different routing schemes are applied. Much thanks, Weiqiang -- Weiqiang Sun Professor, Shanghai Jiao Tong University Mobile: +86-13801847900 Email: [email protected] Addr: Room 5-513, SEIEE Building, Shanghai Jiao Tong University, = Minhang, Shanghai, 200240, China --Apple-Mail=_4F1B16A5-842F-4BE8-BF3B-5792FCF7D503 Content-Transfer-Encoding: quoted-printable Content-Type: text/html; charset=utf-8 <html aria-label=3D"message body"><head><meta http-equiv=3D"content-type" = content=3D"text/html; charset=3Dutf-8"></head><body = style=3D"overflow-wrap: break-word; -webkit-nbsp-mode: space; = line-break: after-white-space;"><p class=3D"MsoNormal" style=3D"margin: = 0in 0in 8pt; line-height: 18.4px; text-align: justify;"><font = face=3D"Calibri" style=3D"font-size: 14px;">Hello = IPPM/BMWG!<o:p></o:p></font></p><p class=3D"MsoNormal" style=3D"margin: = 0in 0in 8pt; line-height: 18.4px; text-align: justify;"><font = face=3D"Calibri" style=3D"font-size: 14px;">This is a much belated = solicitation for comments on <a = href=3D"https://datatracker.ietf.org/doc/draft-ye-ippm-switching-efficienc= y/" style=3D"color: rgb(149, 79, = 114);">https://datatracker.ietf.org/doc/draft-ye-ippm-switching-efficiency= /</a>. The draft was uploaded in April and presented on IETF 126 in = Viena.<o:p></o:p></font></p><p class=3D"MsoNormal" style=3D"margin: 0in = 0in 8pt; line-height: 18.4px; text-align: justify;"><font = face=3D"Calibri"><span style=3D"font-size: 14px;">The questions we ask = is: when ASICs with different capacities (radix times port capacity) are = deployed in AI fabrics following different design philosophies, i.e., = fabric architectures (spine/leaf, dragonfly, or 3d-torus), are they = efficiently utilized? Are there spaces to make the AI fabric leaner and = greener? Switching efficiency, </span></font><span = style=3D"font-family: Calibri; font-size: 14px;">which we shall now call = Fabric Switching Efficiency =E2=80=93 FSE, </span><span = style=3D"font-size: 14px; font-family: Calibri;">measures = computationally effective data throughput (analogous to algorithmic = bandwidth) per unit of deployed switching capacity.</span><span = style=3D"font-size: 14px; font-family: Calibri;"> FSE = characterizes how the AI fabric design (topology, routing and use of = in-network-computing) fit the tasks run on the fabric. One interesting = analogy, which I should borrow from one expert I talked to these days, = is PUE (Power Usage Effectiveness). PUE is a high level metric. It does = not take into account the complexity of the operations inside the DCs, = but it effectively captures whether energy/power is efficiently used for = compute in DCs. With FSE, we aim to provide a measure that can be easily = understood and communicated just like PUE. One interesting feature of = FSE is that it can be very naturally decomposed into 3 finer granular = metrics: data efficiency, port utilization and routing efficiency, each = capturing one specific AI fabric design target. FSE can be used at least = by two groups of people: 1) AIDC network architects who would like to = improve the fabric design. 2) anyone who likes to operate a leaner = / greener AIDC network.</span><span style=3D"font-size: 14px; = font-family: Calibri;"> </span></p><p class=3D"MsoNormal" = style=3D"margin: 0in 0in 8pt; line-height: 18.4px; text-align: = justify;"><font face=3D"Calibri" style=3D"font-size: 14px;">We are aware = that the set of AI fabric benchmarking documents are already pretty = complete with well defined scope/terminology, and methodologies (for = inference and training). The documents provides performance metrics (end = to end) for AI fabrics from the eyes of model users who runs inference = or training jobs. FSE operates at higher levels. We don=E2=80=99= t intend to provide performance metrics to be used by LLM model = developer/deployment people. <o:p></o:p></font></p><p = class=3D"MsoNormal" style=3D"margin: 0in 0in 8pt; line-height: 18.4px; = text-align: justify;"><font face=3D"Calibri" style=3D"font-size: = 14px;">The natural next step is to turn this concept into methodologies = that IPPMers/BMWGers/BPMers are interested in and would be willing to = contribute to. We will be happy to receive any feedbacks and = work with the community to move this work = forward.<o:p></o:p></font></p><p class=3D"MsoNormal" style=3D"margin: = 0in 0in 8pt; line-height: 18.4px; text-align: justify;"><font = face=3D"Calibri" style=3D"font-size: 14px;">For folks who are = interested, please also read our paper on arXiv - arXiv:2604.14690 (<a = href=3D"https://arxiv.org/abs/2604.14690" style=3D"color: rgb(149, 79, = 114);">https://arxiv.org/abs/2604.14690</a>), in which the rationale and = numerical results are discussed in much more detail. We are still = improving this paper by adding more packet level simulations to reveal = how the metrics behave when different compute / communication kernels = are executed and different routing schemes are = applied.<o:p></o:p></font></p><p class=3D"MsoNormal" style=3D"margin: = 0in 0in 8pt; line-height: 18.4px; text-align: justify;"><span = style=3D"font-size: 14px; font-family: Calibri;"><br></span></p><p = class=3D"MsoNormal" style=3D"margin: 0in 0in 8pt; line-height: 18.4px; = text-align: justify;"><span style=3D"font-size: 14px; font-family: = Calibri;">Much thanks,</span></p><p class=3D"MsoNormal" style=3D"margin: = 0in 0in 8pt; line-height: 18.4px; text-align: justify;"><span = style=3D"font-size: 14px; font-family: Calibri;">Weiqiang</span></p><p = class=3D"MsoNormal" style=3D"margin: 0in 0in 8pt; line-height: 18.4px; = text-align: justify;"><br></p><div>--</div><div>Weiqiang = Sun</div><div>Professor, Shanghai Jiao Tong University</div><div>Mobile: = +86-13801847900</div><div>Email: [email protected]</div><div>Addr: Room = 5-513, SEIEE Building, Shanghai Jiao Tong University, Minhang, Shanghai, = 200240, China</div> <br></body></html>= --Apple-Mail=_4F1B16A5-842F-4BE8-BF3B-5792FCF7D503-- --===============4081426002303708968== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KaXBwbSBtYWls aW5nIGxpc3QgLS0gaXBwbUBpZXRmLm9yZwpUbyB1bnN1YnNjcmliZSBzZW5kIGFuIGVtYWlsIHRv IGlwcG0tbGVhdmVAaWV0Zi5vcmcK --===============4081426002303708968==--