Re: [PATCH v2] nvme-tcp: pin io_cpu to submitter cpu

Saravanan D <[email protected]>
Newsgroups org.infradead.lists.linux-nvme,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
On Sun, 23 Aug 2026 00:20:21 +0300 Sagi Grimberg <[email protected]> wrote:
> Can you share a concrete example?
> 1. lscpu
> 2. grep -r ''  /sys/block/nvmeXnY/mq/*/cpu_list
> Also, what would be the VMs cpusets on such a server?

Very grateful for your review and suggestions.

Concrete example from one of these hosts, where there are more cpus than
the io queues NVMe-oF targets offer.

lscpu (trimmed): AMD EPYC 9654, 2 sockets, 2 threads/core, 384 CPUs
  NUMA node0 CPUs: 0-95,192-287
  NUMA node1 CPUs: 96-191,288-383

Each nvme-tcp controller exposes 128 io queues (queue_count 129) against
384 CPUs, so blk_mq_map_queues() folds three CPUs into every hctx. cpu_list for the namespace's path device, around one tenant's boundary:

  hctx67:  6, 7, 198
  hctx68:  8, 9, 200
  hctx69:  10, 11, 202
  hctx75:  22, 23, 214
  hctx76:  24, 25, 216
  hctx113: 199, 201, 203
  hctx114: 205, 207, 209
  hctx115: 211, 213, 215

Two VMs are live on this same shared controller:
  VM A cpuset: 8-23,200-215
  VM B cpuset: 104-191,296-383

VM A owns 200-215, so hctx113's group {199,201,203} has 201 and 203 inside
VM A and 199 outside it. The connect time pick can select 199 as that
queue's io_cpu, so VM A's I/O through hctx113 runs its socket work on 199,
outside its cpuset for its entire lifetime. Adoption will pull it back 
to 201 or 203 when VM A submits I/O.

> In an ideal setting, the nvme controller would support $NR_CPUS worth of
> io queues [...] map cpu-queue with perfect alignment.

Right, and that is the case we cannot reach here. The controller tops out
at 128 io queues while the host has 384 CPUs, so the three to one folding
is unavoidable and some groups end up straddling a tenant boundary.

> I think that for controllers with low queue count this approach would not
> work very well. Plus the cpu<->queue mapping by the block layer is not
> guaranteed to align with how workloads/VMs are scheduled on the host...

Agreed on both. A queue whose cpu group straddles two busy tenants is only
partially addressed, since adoption then selects which tenant hosts that
shared queue rather than removing the sharing. It fixes the common case
where cpu group sits within one VM's cpuset and the case where only one
tenant on the queue is active, and it is strictly better than the connect
time pick, which chooses that queue's io_cpu before any I/O and can land on
any member. Solving the common case also lets the control plane, which owns
the VM cpusets, affinitize each VM's virtio-blk submitter threads away from
CPUs whose mapped queue would land socket work on a neighbor.

> To me the problem statement sounds to me like you are looking to get a
> "private" controller (set of IO queues) for each VM. [...]
> nvme connect -t tcp -a <traddr> --cpu-groups=0-63,64-127,128-191,192-255

A private controller or a cpu-groups partition fits workloads where the CPU
partition is known and stable at connect time. Our production constraint 
does not afford us that luxury. The nvme-tcp controllers are host wide and
shared, the tenant VMs are namespaces under one multipath subsystem, and the
control plane creates, destroys and bin packs VMs continuously on these core
dense compute hosts. A cpu-groups mask fixed at connect would be stale as
soon as the tenant mix changes, and re-partitioning would mean reconnecting
a controller that many live tenants depend on. We also cannot give each
tenant its own controller, because the overhead on the host adds up and we
run the risk of exceeding the target's connection limits. Adopting the
submitting CPU needs no per tenant configuration and follows the tenant
churn on its own, which is why it fits a long lived shared connection whose
tenant set is not known at connect time.

Next the fleet operators who care about VM steal time already steer NIC IRQs,
NAPI to a dedicated housekeeping CPU pool outside the tenant cpusets, so the
common network processing overhead Nilay raised does not run on tenant CPUs.
With that in place the remaining contributor that still lands on a
tenant's CPU is nvme_tcp_io_work, which this patch pins to the submitting
cpu and keeps within the tenant's cpuset.

Thanks,
Saravanan D.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.