[PATCH] nvme-tcp: pin io_cpu to submitter cpu
Saravanan D <[email protected]> Wed, 5 Aug 2026 19:39:44 -0700
| Newsgroups | org.infradead.lists.linux-nvme,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
nvme_tcp_set_queue_io_cpu() picks each queue's io_cpu at connect time, before any I/O exists, as the least loaded CPU in the queue's blk-mq map group, and all socket work then runs there for the connection's lifetime. On hosts that partition CPUs between pinned workloads a map group can straddle a partition boundary, so the pick can land one workload's socket processing on CPUs owned by another. On a 384 cpu multi tenant host with one VM driving ~1.1 GB/s of writes, 9% of nvme_tcp_io_work executions ran outside the submitting VM's cpuset, all on io_cpus of boundary straddling map groups, observed by the neighbor as steal time it did not cause. Adopt the submitting CPU of non passthrough requests as io_cpu. The submitter is a member of the map group by construction, and the nvme_tcp_cpu_queues accounting moves with each adoption. Passthrough commands do not participate since the io queue Connect arrives from an arbitrary group CPU via blk_mq_alloc_request_hctx(), so the first real read or write decides. Queues outlive the workloads that submit through them, so adoption re-arms after 30 seconds of queue quiet. An idle queue is reclaimed by its next submitter, while a busy queue keeps a stable io_cpu and cannot ping pong between two live submitters. Concurrent writers on different CPUs serialize on a cmpxchg on io_cpu. The behavior is opt in via the new wq_adopt module parameter, default off and runtime writable. wq_unbound takes precedence when both are set. Signed-off-by: Saravanan D <[email protected]> --- drivers/nvme/host/tcp.c | 71 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 71 insertions(+) diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c index 87d8067f3283..6509273f478c 100644 --- a/drivers/nvme/host/tcp.c +++ b/drivers/nvme/host/tcp.c @@ -44,6 +44,10 @@ static bool wq_unbound; module_param(wq_unbound, bool, 0644); MODULE_PARM_DESC(wq_unbound, "Use unbound workqueue for nvme-tcp IO context (default false)"); +static bool wq_adopt; +module_param(wq_adopt, bool, 0644); +MODULE_PARM_DESC(wq_adopt, "Adopt the submitting cpu as queue io_cpu (default false)"); + /* * TLS handshake timeout */ @@ -92,6 +96,7 @@ enum nvme_tcp_queue_flags { NVME_TCP_Q_LIVE = 1, NVME_TCP_Q_POLLING = 2, NVME_TCP_Q_IO_CPU_SET = 3, + NVME_TCP_Q_IO_CPU_ADOPTED = 4, }; enum nvme_tcp_recv_state { @@ -105,6 +110,7 @@ struct nvme_tcp_queue { struct socket *sock; struct work_struct io_work; int io_cpu; + unsigned long last_data; struct mutex queue_lock; struct mutex send_mutex; @@ -2783,6 +2789,69 @@ static void nvme_tcp_commit_rqs(struct blk_mq_hw_ctx *hctx) queue_work_on(queue->io_cpu, nvme_tcp_wq, &queue->io_work); } +/* Re-adopt io_cpu on the first data request after this much queue idle time */ +#define NVME_TCP_READOPT_IDLE (30 * HZ) + +/* + * Adopt the CPU of the current data submission as the queue's io_cpu. + * + * The connect time choice in nvme_tcp_set_queue_io_cpu() picks the least + * loaded CPU in the queue's mq_map group before any I/O exists, so it + * cannot know which side of the group the actual submitters live on. On + * hosts that partition CPUs between pinned workloads a group that + * straddles a partition boundary can get an io_cpu on CPUs the submitting + * workload does not own, and its network processing then preempts an + * unrelated workload. The submitting CPU is in the queue's mq_map group + * by construction, so adopting it preserves the spreading property while + * landing the work on the side that generates it. + * + * Queues belong to the controller connection and outlive the workloads + * that submit through them, so adoption re-arms after NVME_TCP_READOPT_IDLE + * of queue quiet. A successor workload reclaims an idle queue with its + * first data request, while a continuously busy queue keeps a stable + * io_cpu and cannot ping pong between two live submitters. + * + * Passthrough commands (the io queue Connect in particular) are submitted + * from an arbitrary group CPU by blk_mq_alloc_request_hctx() and do not + * represent the data path, so they are skipped and the first real read + * or write decides. + * + * Adoption is opt in via the wq_adopt module parameter and is bypassed + * when wq_unbound is set. + */ +static void nvme_tcp_adopt_io_cpu(struct nvme_tcp_queue *queue, + struct request *rq) +{ + int old, new; + + if (!wq_adopt || wq_unbound || blk_rq_is_passthrough(rq)) + return; + if (!nvme_tcp_queue_id(queue)) + return; + + if (test_bit(NVME_TCP_Q_IO_CPU_ADOPTED, &queue->flags) && + time_before(jiffies, READ_ONCE(queue->last_data) + + NVME_TCP_READOPT_IDLE)) { + WRITE_ONCE(queue->last_data, jiffies); + return; + } + + WRITE_ONCE(queue->last_data, jiffies); + set_bit(NVME_TCP_Q_IO_CPU_ADOPTED, &queue->flags); + + old = READ_ONCE(queue->io_cpu); + new = raw_smp_processor_id(); + if (old == new || !try_cmpxchg(&queue->io_cpu, &old, new)) + return; + + if (test_bit(NVME_TCP_Q_IO_CPU_SET, &queue->flags)) { + atomic_dec(&nvme_tcp_cpu_queues[old]); + atomic_inc(&nvme_tcp_cpu_queues[new]); + } + dev_dbg(queue->ctrl->ctrl.device, "queue %d: adopted io_cpu %d\n", + nvme_tcp_queue_id(queue), new); +} + static blk_status_t nvme_tcp_queue_rq(struct blk_mq_hw_ctx *hctx, const struct blk_mq_queue_data *bd) { @@ -2802,6 +2871,8 @@ static blk_status_t nvme_tcp_queue_rq(struct blk_mq_hw_ctx *hctx, nvme_start_request(rq); + nvme_tcp_adopt_io_cpu(queue, rq); + nvme_tcp_queue_request(req, bd->last); return BLK_STS_OK; base-commit: bf881dd20062db5e951a0d0703cb476df8c9fdee -- 2.53.0