Re: [PATCH v2] nvme-tcp: pin io_cpu to submitter cpu
Saravanan D <[email protected]>
| Newsgroups | org.infradead.lists.linux-nvme,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Sun, 23 Aug 2026 00:20:21 +0300 Sagi Grimberg <[email protected]> wrote: > Can you share a concrete example? > 1. lscpu > 2. grep -r '' /sys/block/nvmeXnY/mq/*/cpu_list > Also, what would be the VMs cpusets on such a server? Very grateful for your review and suggestions. Concrete example from one of these hosts, where there are more cpus than the io queues NVMe-oF targets offer. lscpu (trimmed): AMD EPYC 9654, 2 sockets, 2 threads/core, 384 CPUs NUMA node0 CPUs: 0-95,192-287 NUMA node1 CPUs: 96-191,288-383 Each nvme-tcp controller exposes 128 io queues (queue_count 129) against 384 CPUs, so blk_mq_map_queues() folds three CPUs into every hctx. cpu_list for the namespace's path device, around one tenant's boundary: hctx67: 6, 7, 198 hctx68: 8, 9, 200 hctx69: 10, 11, 202 hctx75: 22, 23, 214 hctx76: 24, 25, 216 hctx113: 199, 201, 203 hctx114: 205, 207, 209 hctx115: 211, 213, 215 Two VMs are live on this same shared controller: VM A cpuset: 8-23,200-215 VM B cpuset: 104-191,296-383 VM A owns 200-215, so hctx113's group {199,201,203} has 201 and 203 inside VM A and 199 outside it. The connect time pick can select 199 as that queue's io_cpu, so VM A's I/O through hctx113 runs its socket work on 199, outside its cpuset for its entire lifetime. Adoption will pull it back to 201 or 203 when VM A submits I/O. > In an ideal setting, the nvme controller would support $NR_CPUS worth of > io queues [...] map cpu-queue with perfect alignment. Right, and that is the case we cannot reach here. The controller tops out at 128 io queues while the host has 384 CPUs, so the three to one folding is unavoidable and some groups end up straddling a tenant boundary. > I think that for controllers with low queue count this approach would not > work very well. Plus the cpu<->queue mapping by the block layer is not > guaranteed to align with how workloads/VMs are scheduled on the host... Agreed on both. A queue whose cpu group straddles two busy tenants is only partially addressed, since adoption then selects which tenant hosts that shared queue rather than removing the sharing. It fixes the common case where cpu group sits within one VM's cpuset and the case where only one tenant on the queue is active, and it is strictly better than the connect time pick, which chooses that queue's io_cpu before any I/O and can land on any member. Solving the common case also lets the control plane, which owns the VM cpusets, affinitize each VM's virtio-blk submitter threads away from CPUs whose mapped queue would land socket work on a neighbor. > To me the problem statement sounds to me like you are looking to get a > "private" controller (set of IO queues) for each VM. [...] > nvme connect -t tcp -a <traddr> --cpu-groups=0-63,64-127,128-191,192-255 A private controller or a cpu-groups partition fits workloads where the CPU partition is known and stable at connect time. Our production constraint does not afford us that luxury. The nvme-tcp controllers are host wide and shared, the tenant VMs are namespaces under one multipath subsystem, and the control plane creates, destroys and bin packs VMs continuously on these core dense compute hosts. A cpu-groups mask fixed at connect would be stale as soon as the tenant mix changes, and re-partitioning would mean reconnecting a controller that many live tenants depend on. We also cannot give each tenant its own controller, because the overhead on the host adds up and we run the risk of exceeding the target's connection limits. Adopting the submitting CPU needs no per tenant configuration and follows the tenant churn on its own, which is why it fits a long lived shared connection whose tenant set is not known at connect time. Next the fleet operators who care about VM steal time already steer NIC IRQs, NAPI to a dedicated housekeeping CPU pool outside the tenant cpusets, so the common network processing overhead Nilay raised does not run on tenant CPUs. With that in place the remaining contributor that still lands on a tenant's CPU is nvme_tcp_io_work, which this patch pins to the submitting cpu and keeps within the tenant's cpuset. Thanks, Saravanan D.