Re: [PATCH v8 00/10] nvme-multipath: introduce latency I/O policy

Sagi Grimberg <[email protected]>
Newsgroups org.infradead.lists.linux-nvme
Message-ID <[email protected]>

On 25/08/2026 8:05, Nilay Shroff wrote:
> On 8/23/26 4:17 AM, Sagi Grimberg wrote:
>>
>>
>> On 15/08/2026 20:34, Nilay Shroff wrote:
>>> Hi,
>>>
>>> This series introduces a new latency I/O policy for NVMe native
>>> multipath. Existing policies such as numa, round-robin, and queue-depth
>>> are static and do not adapt to real-time transport performance. The 
>>> numa
>>> selects the path closest to the NUMA node of the current CPU, 
>>> optimizing
>>> memory and path locality, but ignores actual path performance. The
>>> round-robin distributes I/O evenly across all paths, providing fairness
>>> but not performance awareness. The queue-depth reacts to instantaneous
>>> queue occupancy, avoiding heavily loaded paths, but does not account 
>>> for
>>> actual latency, throughput, or link speed.
>>>
>>> The new latency policy addresses these gaps selecting paths dynamically
>>> based on measured I/O latency for both PCIe and fabrics. Latency is
>>> derived by passively sampling I/O completions. Each path is assigned a
>>> weight proportional to its latency score, and I/Os are then forwarded
>>> accordingly. As condition changes (e.g. latency spikes, bandwidth
>>> differences), path weights are updated, automatically steering traffic
>>> toward better-performing paths.
>>>
>>> Early results show reduced tail latency under mixed workloads and
>>> improved throughput by exploiting higher-speed links more effectively.
>>> For example, with NVMf/TCP using two paths (one throttled with ~30 ms
>>> delay), fio results with random read/write/rw workloads (direct I/O)
>>> showed:
>>
>> TBH, I do not know if this measurement represent any real-life
>> scenario. I do think that occasional packet drops are a real-life 
>> scenario, and
>> it would be a worthy use-case to optimize for. Can you perhaps measure
>> how the path selectors compare in this case?
>>
>
> Okay so I have now measured another workload where I simulated packet 
> loss/drops
> which are occasional and compared different path selectors. In this 
> measurement
> I have a shared NVMe namespace configured which is reachable over two 
> tcp paths.
> Now to simulate the occasional packet loss/drop I have configured one 
> of the paths
> to experience occasional packet loss/drop as shown below over the 
> period of
> 480 seconds:
>
> 0            120           240           360           480
> |-------------|-------------|-------------|-------------|
> |<--5% drop-->|<--no drop-->|<--3% drop-->|<--no drop-->|
>
> As shown above for the first 120 seconds the path experiences 5% 
> packet drop,
> for the next 120 seconds path sees no packet drop and again for 
> subsequent 120
> seconds path experiences 3% packet drop and for the rest of the 
> duration (during
> last 120 seconds) there's no packet drop observed by the path. With 
> this simulation,
> I ran fio workload for 480 seconds leveraging direct I/O, bs=4k, 
> iodepth=64,
> numjobs=32 and ioengine=io_uring. Shown below is bw observed running 
> fio test
> using different I/O policies:
>
>              numa     round-robin queue-depth latency
>              (MiB/s)  (MiB/s)     (MiB/s)     (MiB/s)
>              -------  ----------- ----------- ---------
> randread:    1288     1202        1464        1642
> randwrite:   1456     1493        1779        1944
> randrw:      R:623    R:594       R:750       R:822
>              W:623    W:594       W:750       W:822
>
>>>
>>>          numa         round-robin   queue-depth  adaptive
>>>          -----------  -----------   -----------  ---------
>>> READ:   50.0 MiB/s   105 MiB/s     230 MiB/s    350 MiB/s
>>> WRITE:  65.9 MiB/s   125 MiB/s     385 MiB/s    446 MiB/s
>>> RW:     R:30.6 MiB/s R:56.5 MiB/s  R:122 MiB/s  R:175 MiB/s
>>>          W:30.7 MiB/s W:56.5 MiB/s  W:122 MiB/s  W:175 MiB/s
>>
>> And I'm assuming there are zero downsides for the normal
>> case?
>>
> For the normal case where all paths are symmetric I saw
> queue-depth, latency and round-robin policies yielding
> nearly same bandwidth. However for numa policy, it depends
> on CPU/numa locality.
>
>>>
>>> This pathcset includes totla 8 patches:
>>> [PATCH 1/10] block: expose blk_stat_{enable,disable}_accounting()
>>>    - Make blk_stat APIs available to block drivers.
>>>    - Needed for per-path latency measurement.
>>>
>>> [PATCH 2/10] block: record I/O request start time for passthru request
>>>    - Record I/O start time for I/O passthru requests.
>>>    - This is prep patch which allows measuring I/O completion latency
>>>      for passthru requests.
>>>
>>> [PATCH 3/10] block: support nesting for blk-mq flag 
>>> QUEUE_FLAG_SAME_FORCE
>>>    - Support nesting for QUEUE_FLAG_SAME_FORCE as multiple users
>>>      could toggle QUEUE_FLAG_SAME_FORCE.
>>>
>>> [PATCH 4/10] nvme-multipath: pass I/O type to nvme_find_path()
>>>    - This is the prep patch which updates nvme_find_path() signature
>>> [PATCH 5/10] nvme-multipath: add latency I/O policy
>>>    - Implement path scoring based on latency (EWMA).
>>>    - Distribute I/O proportionally to per-path weights.
>>>
>>> [PATCH 6/10] nvme: add generic debugfs support
>>>    - Introduce generic debugfs support for NVMe module
>>>
>>> [PATCH 7/10] nvme-multipath: add debugfs attribute latency_ewma_shift
>>>    - Adds a debugfs attribute to control ewma shift
>>>
>>> [PATCH 8/10] nvme-multipath: add debugfs attribute 
>>> latency_batch_timeout
>>>    - Adds a debugfs attribute to control latency batch window interval
>>>
>>> [PATCH 9/10] nvme-multipath: add debugfs attribute latency_stat
>>>    - Add “latency_stat” under per-path and head debugfs directories to
>>>      expose latency policy state and statistics.
>>>
>>> [PATCH 10/10] nvme-multipath: add documentation for latency I/O policy
>>>    - Includes documentation for latency I/O multipath policy.
>>>
>>> LSFMM discussion:
>>> =================
>>> During lsfmm 2026, it was decided to rename this I/O policy from
>>> "adaptive" to "latency". This series reflects that rename.
>>>
>>> The discussion at lsfmm also focused extensively on the latency
>>> measurement model, including whether latency should be tracked
>>> per-CPU or per-NUMA, and whether separate I/O-size buckets should
>>> be maintained for different request sizes.
>>>
>>> After detailed discussion and evaluation of throughput results, the
>>> consensus was to initially measure I/O completion latency on a
>>> per-CPU basis. The available performance data showed that the
>>> per-CPU implementation already provides sufficient averaging across
>>> CPUs while keeping the design relatively simple.
>>>
>>> The use of additional I/O-size buckets did not demonstrate meaningful
>>> throughput improvement in the general case and would introduce extra
>>> complexity into the fast path and accounting logic. As a result, the
>>> consensus was to avoid I/O-size bucketing for now and keep the policy
>>> focused on per-CPU latency measurement.
>>>
>>> If future real-world workloads demonstrate a clear benefit from
>>> I/O-size-aware latency accounting, the policy can be extended later
>>> to support it.
>>>
>>> As ususal, feedback and suggestions are most welcome!
>>
>> Nilay, do we have evidence that round-robin/queue-depth are better
>> for any workload? As a user, I would be very confused with the amount
>> of path selectors I have available and which should I choose.
>
> From my experiments, when the paths are symmetric, both round-robin and
> queue-depth (and for that matter latency) exhibit similar behavior, with
> the workload being distributed roughly equally across the active paths.
>
> When the paths are asymmetric, I found queue-depth to perform better than
> round-robin. Queue-depth tries to steer I/O toward the less-loaded path
> based on the number of in-flight I/Os on each path, whereas round-robin
> continues to distribute I/O evenly across all active paths.
>
> However, queue-depth still has a limitation in this scenario. It uses
> the number of in-flight I/Os as an indirect indication of path
> performance, it does not have a direct signal of the actual I/O
> completion latency. For example, if one path starts experiencing packet
> loss, I/O completion on that path can become significantly slower.
> Queue-depth can react to this as the path accumulates more outstanding
> I/Os, but it can still continue sending I/O to the degraded path as long
> as its queue depth remains comparable to the healthy path. In other
> words, it can reduce the amount of I/O sent to the degraded path, but it
> cannot directly account for how much slower that path has become.
>
> The latency policy uses I/O completion latency as the signal instead.
> When one path becomes degraded, its observed latency increases and its
> path score/weight decreases. Consequently, the policy shifts more I/O 
> towards
> the healthy path. This allows the healthy path to sustain a higher queue
> depth while the degraded path receives substantially less I/O, rather
> than trying to maintain a similar queue depth across both paths.
>
> This is also reflected in the packet-loss experiment above. Round-robin
> continues to distribute I/O across both paths, while queue-depth does a
> better job by reacting to the increased queue occupancy of the degraded
> path. The latency policy goes one step further by directly using the
> increased completion latency as a signal and therefore steers more I/O
> toward the healthy path, resulting in higher throughput.
>
> So based on the results I have so far, I would characterize the existing
> policies as follows: round-robin is useful when paths are symmetric and
> equal distribution is desired. The queue-depth is preferable when 
> paths are
> asymmetric and queue occupancy provides a useful indication of path
> load and the latency policy is intended for cases where path performance
> can vary dynamically and we want the policy to adapt based on actual
> observed latency.

If you have policies A, B, and C and you say:
- In certain conditions policy C > A, B
- In some conditions C = B > A
- In all other cases C = B = A

This means that C should always be used, and A, B should never be used.

Hence I ask, should we really have this as an option? or should we deprecate
round-robin/queue-depth and have only numa|latency?

I would like to avoid introducing this as a config knob if it is always 
behaves
better.

What do others think?
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.