Re: [PATCH V2] nvme: Add module reference counting for multipath devices

Nilay Shroff <[email protected]>
Newsgroups org.infradead.lists.linux-nvme
Message-ID <[email protected]>
On 8/26/26 10:06 PM, Keith Busch wrote:
> On Wed, Aug 26, 2026 at 11:18:16AM -0500, Wen Xiong wrote:
>> On 2026-08-26 07:40, Nilay Shroff wrote:
>>
>>> Also, since a multipath head can have paths through different transports
>>> , we should not take a reference only to the transport of one
>>> namespace/path
>>> found through nvme_find_path(), as was done in v1. Instead, when opening
>>> the
>>> head, take a module reference iterating through each controller
>>> reachable
>>> from the corresponding NVMe subsystem. This ensures that every transport
>>> that can service I/O for the active multipath namespace remains loaded
>>> for
>>> the duration of its use.
>>>
>> I will look into iterating though each controller/each namespace from nvme
>> subsystem.
> 
> This is not viable. You can add and remove paths to a namespace at any
> time such that the transports counted on open are not the namespace's
> transports on close.

Yes correct, and I think we need some additional change in the code to
handle this gracefully. I though about it have some initial idea to
address this:

1. Add nr_openers to struct nvme_ns_head.

2. When the ns head is opened, iterate through each namespace associated
    with the head and increment the reference count of its underlying
    transport module. Then increment nr_openers.

3. If a new ns/path is added while the head is open, check nr_openers and,
    if it is non-zero, increment the reference count of the corresponding
    transport module nr_openers times.

4. If an existing ns/path is removed while the head is open, check nr_openers
    and decrement the reference count of the corresponding transport module
    nr_openers times.

5. When the ns head is closed, iterate through the namespaces associated
    with the head and decrement the reference count of each underlying
    transport module. Then decrement nr_openers.

The above operations are serialized by subsys->lock, so nr_openers serves
as the number of users that have actually opened the head node.

For example, suppose we have a shared namespace reachable through
TCP and RDMA paths:

1. User opens the head node:
    head->nr_openers = 1; tcp_module_ref_count = 1; rdma_module_ref_count = 1
2. The RDMA path is removed:
    head->nr_openers = 1; tcp_module_ref_count = 1; rdma_module_ref_count = 0
3. User closes the head node:
    head->nr_openers = 0; tcp_module_ref_count = 0; rdma_module_ref_count = 0

Another example with multiple openers:

1. User A opens the head node:
    head->nr_openers = 1; tcp_module_ref_count = 1; rdma_module_ref_count = 1
2. A new TCP path is added and linked to the head:
    head->nr_openers = 1; tcp_module_ref_count = 2; rdma_module_ref_count = 1
3. User B opens the head node:
    head->nr_openers = 2; tcp_module_ref_count = 4; rdma_module_ref_count = 2
4. User A closes the head node:
    head->nr_openers = 1; tcp_module_ref_count = 2; rdma_module_ref_count = 1
5. The RDMA path is removed:
    head->nr_openers = 1; tcp_module_ref_count = 2; rdma_module_ref_count = 0
6. User B closes the head node:
    head->nr_openers = 0; tcp_module_ref_count = 0; rdma_module_ref_count = 0

This way, the transport module references track the actual number of openers
and the set of paths associated with the head, even when paths are dynamically
added or removed while the head remains open.

Thanks,
--Nilay
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.