Re: [PATCH net-next v4 3/3] net/smc: transition to RDMA core CQ pooling
Mahanta Jambigi <[email protected]> Thu, 30 Jul 2026 12:55:00 +0530
| Newsgroups | org.kernel.vger.linux-rdma,org.kernel.vger.linux-s390,org.kernel.vger.netdev |
|---|---|
| Message-ID | <[email protected]> |
On 30/07/26 11:13 am, D. Wythe wrote: > On Wed, Jul 22, 2026 at 05:17:15PM +0530, Mahanta Jambigi wrote: >> >> >> On 16/07/26 5:07 pm, D. Wythe wrote: >>> Performance Test: redis-benchmark with max 32 connections per QP >>> Data format: Requests Per Second (RPS), Percentage in brackets >>> represents the gain/loss compared to TCP. >>> >>> | Clients | TCP | SMC (original) | SMC (cq_pool) | >>> |---------|----------|---------------------|---------------------| >>> | c = 1 | 24449 | 31172 (+27%) | 34039 (+39%) | >>> | c = 2 | 46420 | 53216 (+14%) | 64391 (+38%) | >>> | c = 16 | 159673 | 83668 (-48%) <-- | 216947 (+36%) | >>> | c = 32 | 164956 | 97631 (-41%) <-- | 249376 (+51%) | >>> | c = 64 | 166322 | 118192 (-29%) <-- | 249488 (+50%) | >>> | c = 128 | 167700 | 121497 (-27%) <-- | 249480 (+48%) | >>> | c = 256 | 175021 | 146109 (-16%) <-- | 240384 (+37%) | >>> | c = 512 | 168987 | 101479 (-40%) <-- | 226634 (+34%) | >>> >>> The results demonstrate that this optimization effectively resolves the >>> scalability bottleneck, with RPS increasing by over 110% at c=64 >>> compared to the original implementation. >> >> Thanks for the performance numbers — the scalability improvement at high >> client counts is striking. A few questions to help reproduce these results: >> >> 1) Was net.smc.smcr_max_conns_per_lgr=32 the value used? The "max 32 >> connections per QP" in the description suggests so, but the sysctl >> wasn't explicitly listed. >> >> 2) Were net.smc.smcr_max_send_wr and net.smc.smcr_max_recv_wr left at >> their defaults, or tuned? With per-link CQs the optimal WR counts may >> differ from the global-CQ case, so knowing whether these were touched >> would help interpret the numbers. >> >> 3) Was this SMC-R v2? The v2 TX path (smc_wr_tx_get_v2_slot, >> wr_tx_v2_ib) has different concurrency characteristics, so confirming >> the version matters for the contention analysis in patch 2/3 as well. > > In our tests, we set net.smc.smcr_max_conns_per_lgr=32 via sysctl. We > also tuned net.smc.smcr_max_send_wr and net.smc.smcr_max_recv_wr from > their default to 64 to provide sufficient work request depth. > > All tests were run with SMC-R v2. > > And the performance numbers do need a refresh: with recent improvements—including > thundering herd mitigation in earlier patches, the original SMC implementation’s > regression in the c=16–512 range is expected to be largely resolved. I believe that Patch 2/3, "net/smc: reduce TX slot contention with exclusive wait", largely mitigates the thundering herd effect. > >> >> 4) Was uperf tested for v4 as well, as we did for v1[1] in >> February-2026? uperf exercises a wider range of message sizes and thread >> counts, so its numbers would strengthen the case — particularly for the >> intermediate client counts (c=16 to c=64) where the original >> implementation regresses below TCP. > > Of course uperf was tested. We’ll include a comprehensive uperf > evaluation in the next version. > > Given that this CQ-pooling patch may require additional review time, I > plan to send the first two patches of the series separately. > They’ve already been reviewed and are ready to merge. This patch > will follow as a standalone submission once the uperf results are > ready. Sounds good. I'll be looking forward to the standalone patch that introduces CQ pooling and the accompanying uperf results. > > D. Wythe > >> >> [1] >> https://lore.kernel.org/netdev/[email protected]/