Re: [PATCH net-next v4 3/3] net/smc: transition to RDMA core CQ pooling

Mahanta Jambigi <[email protected]>
Newsgroups org.kernel.vger.linux-s390,org.kernel.vger.linux-rdma,org.kernel.vger.netdev
Message-ID <[email protected]>

On 30/07/26 11:13 am, D. Wythe wrote:
> On Wed, Jul 22, 2026 at 05:17:15PM +0530, Mahanta Jambigi wrote:
>>
>>
>> On 16/07/26 5:07 pm, D. Wythe wrote:
>>> Performance Test: redis-benchmark with max 32 connections per QP
>>> Data format: Requests Per Second (RPS), Percentage in brackets
>>> represents the gain/loss compared to TCP.
>>>
>>> | Clients | TCP      | SMC (original)      | SMC (cq_pool)       |
>>> |---------|----------|---------------------|---------------------|
>>> | c = 1   | 24449    | 31172  (+27%)       | 34039  (+39%)       |
>>> | c = 2   | 46420    | 53216  (+14%)       | 64391  (+38%)       |
>>> | c = 16  | 159673   | 83668  (-48%)  <--  | 216947 (+36%)       |
>>> | c = 32  | 164956   | 97631  (-41%)  <--  | 249376 (+51%)       |
>>> | c = 64  | 166322   | 118192 (-29%)  <--  | 249488 (+50%)       |
>>> | c = 128 | 167700   | 121497 (-27%)  <--  | 249480 (+48%)       |
>>> | c = 256 | 175021   | 146109 (-16%)  <--  | 240384 (+37%)       |
>>> | c = 512 | 168987   | 101479 (-40%)  <--  | 226634 (+34%)       |
>>>
>>> The results demonstrate that this optimization effectively resolves the
>>> scalability bottleneck, with RPS increasing by over 110% at c=64
>>> compared to the original implementation.
>>
>> Thanks for the performance numbers — the scalability improvement at high
>> client counts is striking. A few questions to help reproduce these results:
>>
>> 1) Was net.smc.smcr_max_conns_per_lgr=32 the value used? The "max 32
>> connections per QP" in the description suggests so, but the sysctl
>> wasn't explicitly listed.
>>
>> 2) Were net.smc.smcr_max_send_wr and net.smc.smcr_max_recv_wr left at
>> their defaults, or tuned? With per-link CQs the optimal WR counts may
>> differ from the global-CQ case, so knowing whether these were touched
>> would help interpret the numbers.
>>
>> 3) Was this SMC-R v2? The v2 TX path (smc_wr_tx_get_v2_slot,
>> wr_tx_v2_ib) has different concurrency characteristics, so confirming
>> the version matters for the contention analysis in patch 2/3 as well.
> 
> In our tests, we set net.smc.smcr_max_conns_per_lgr=32 via sysctl. We
> also tuned net.smc.smcr_max_send_wr and net.smc.smcr_max_recv_wr from
> their default to 64 to provide sufficient work request depth. 
> 
> All tests were run with SMC-R v2.
> 
> And the performance numbers do need a refresh: with recent improvements—including
> thundering herd mitigation in earlier patches, the original SMC implementation’s
> regression in the c=16–512 range is expected to be largely resolved.

I believe that Patch 2/3, "net/smc: reduce TX slot contention with
exclusive wait", largely mitigates the thundering herd effect.

> 
>>
>> 4) Was uperf tested for v4 as well, as we did for v1[1] in
>> February-2026? uperf exercises a wider range of message sizes and thread
>> counts, so its numbers would strengthen the case — particularly for the
>> intermediate client counts (c=16 to c=64) where the original
>> implementation regresses below TCP.
> 
> Of course uperf was tested. We’ll include a comprehensive uperf
> evaluation in the next version. 
> 
> Given that this CQ-pooling patch may require additional review time, I
> plan to send the first two patches of the series separately.
> They’ve already been reviewed and are ready to merge. This patch
> will follow as a standalone submission once the uperf results are
> ready.

Sounds good. I'll be looking forward to the standalone patch that
introduces CQ pooling and the accompanying uperf results.

> 
> D. Wythe
> 
>>
>> [1]
>> https://lore.kernel.org/netdev/[email protected]/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.