Re: [PATCH] virtio_blk: add use_irq_affinity module parameter

Stefan Hajnoczi <[email protected]>
Newsgroups dev.linux.lists.virtualization
Message-ID <CAJSP0QU0Ys6Kb+hbftW1CaWT9pyZ31VzUJfTYmVZGqHGubgsXw@mail.gmail.com>
On Tue, Aug 11, 2026 at 10:25 PM Liu, Changcheng <[email protected]> wrote:
>
> On Tue, Aug 11, 2026 at 10:37:28AM -0400, Stefan Hajnoczi wrote:
> > External email: Use caution opening links or attachments
> >
> >
> > On Wed, Aug 5, 2026 at 10:40 PM Liu, Changcheng <[email protected]> wrote:
> > >
> > > On Wed, Aug 05, 2026 at 01:45:58PM -0400, Stefan Hajnoczi wrote:
> > > >
> > > > On Sat, Aug 1, 2026 at 2:21 AM Liu, Changcheng <[email protected]> wrote:
> > > > >
> > > > > When creating many virtio-blk devices, probe starts failing with
> > > > > -ENOSPC (-28) because the system runs out of interrupt vectors:
> > > > >
> > > > >   virtio_blk virtioNNN: probe with driver virtio_blk failed with error -28
> > > > >
> > > > > By default virtio-blk uses managed IRQ affinity, which reserves an
> > > > > interrupt vector on every CPU for each device (about nr_cpus vectors per
> > > > > device). On a host with many CPUs and many devices this exhausts the
> > > > > vectors long before all devices are probed.
> > > > >
> > > > > Add use_irq_affinity (default true, no behaviour change). Set it to 0 to
> > > > > use unmanaged interrupts, so each device only uses a couple of vectors
> > > > > instead of one per CPU, allowing far more devices to probe.
> > > > >
> > > > > Signed-off-by: Liu, Changcheng <[email protected]>
> > > >
> > > > There is already a num_request_queues module parameter for cases where
> > > > the user wishes to reduce the number of virtqueues. Did you benchmark
> > > > that and decide the performance of many queues sharing a single irq
> > > > makes it worth adding another module parameter?
> > > >
> > > > Stefan
> > >
> > > num_request_queues does not address this case because each virtio-blk device
> > > already has only one request virtqueue when the failure occurs, so the queue
> > > count cannot be reduced further.
> >
> > Okay.
> >
> > > The issue is reproduced after probing approximately 800 virtio-blk devices on
> > > a 64-core bare-metal host. With managed IRQ affinity, the IRQ-vector reservations
> > > are eventually exhausted and probing additional devices fails with -ENOSPC.
> > > Disabling managed IRQ affinity avoids these per-CPU vector reservations and
> > > allows more devices to be probed successfully.
> >
> > I don't follow. My understanding was that non-managed IRQ vectors are
> > allocated by request_irq(), which is called during virtio_find_vqs().
> > Probing 800 devices would still require 800 * (1 virtqueue irq + 1
> > config change irq) = 1,600 vectors. How come non-managed IRQs do not
> > hit the limit here?
> >
> > Thanks,
> > Stefan
> >
>
> With one request queue, virtio-blk allocates two MSI-X interrupts: an
> unmanaged configuration interrupt and a managed request-queue interrupt.
>
> The request-queue interrupt's managed affinity mask covers all possible
> CPUs. On x86, irq_matrix_reserve_managed() reserves one vector slot on
> every CPU in that mask so the interrupt can migrate safely during CPU
> hotplug. Consequently, every virtio-blk device consumes one managed
> vector slot on every CPU for its lifetime.
>
> Once any CPU exhausts its roughly 200 usable vector slots, managed vector
> allocation fails even if other CPUs still have capacity. virtio-pci then
> falls back to its shared-vector MSI-X allocation, for which the affinity
> descriptor is discarded and the interrupts are unmanaged. Allocation can
> continue using capacity on the other CPUs, explaining the observed limit
> of approximately 800 devices.
>
> When managed affinity is disabled, both interrupts are activated on
> individual CPUs and distributed by the vector allocator. APIC vector
> numbers are per-CPU, so the same vector number can be reused on different
> CPUs. On a 64-CPU system, 800 devices therefore consume approximately
> 800 * 2 / 64 = 25 vector slots per CPU.

Hi Changcheng,
Thanks for the explanation. Can you update the commit message to
mention that non-managed irqs stay below the limit (800 * 2 / 64 = 25
vector slots per CPU) because they are allocated to the CPU with the
lowest number by the vector allocator?

Other than that:
Reviewed-by: Stefan Hajnoczi <[email protected]>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.