Re: [PATCH v3] iommu/arm-smmu-v3: Shrink command/event/PRI queues in kdump kernel
Will Deacon <[email protected]> Tue, 4 Aug 2026 15:02:46 +0100
| Newsgroups | dev.linux.lists.iommu,org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <anHxBh-xGW9JXsa4@willie-the-truck> |
On Tue, Jul 28, 2026 at 02:53:16PM +0100, Kiryl Shutsemau wrote: > On Tue, Jul 28, 2026 at 11:16:10AM +0100, Will Deacon wrote: > > On Mon, Jul 06, 2026 at 09:47:08AM +0100, Kiryl Shutsemau (Meta) wrote: > > > All SMMU queues are sized from the maxima the hardware advertises in IDR1, > > > which can be several megabytes each, and are allocated at probe. The kdump > > > kernel already disables the event and PRI queues (arm_smmu_device_reset() > > > drops CR0_EVTQEN/CR0_PRIQEN) but still allocates them at full size. On > > > systems with many SMMUv3 instances that cost is paid per instance and adds > > > up to tens of megabytes of coherent DMA in the capture kernel. > > > > > > A kdump capture kernel runs from a small crashkernel reservation and only > > > has to drive the few devices used to save the dump, so deep queues serve > > > no purpose. The queues are not on the DMA data path, so dump throughput is > > > unaffected; a shallower command queue only bounds how many commands may be > > > in flight before a sync, which does not matter for the capture kernel's > > > small device count and modest I/O. > > > > > > Clamp every queue to a single page when is_kdump_kernel() is true. Doing > > > it in arm_smmu_init_one_queue() covers the command, event and PRI queues > > > in one place. The command queue still holds at least one batch plus a sync > > > (256 entries on a 4K-page kernel, well above CMDQ_BATCH_ENTRIES), so > > > command batching keeps working. > > > > Wouldn't we be better of not allocating unused queues in the first place? > > That's what this patch does: > > > > https://lore.kernel.org/r/0b035c53cb401acde8244b805d4b6a0312b83708.1782799827.git.nicolinc@nvidia.com > > Maybe I'm reading it wrong, but I couldn't find the allocation change in > there. That patch seems to only touch arm_smmu_device_reset() -- skipping > the EVTQ/PRIQ base/prod/cons writes and the CR0_EVTQEN/CR0_PRIQEN enables > instead of the current enable-then-mask-out. > > I went through the rest of the series too and didn't see > arm_smmu_init_queues() or arm_smmu_init_one_queue() touched anywhere, so > as far as I can tell the kdump kernel still allocates full-size > evtq/priq, and the cmdq is not affected either. > > Is the allocation side handled somewhere else that I've missed? Sorry, yes, it would need extending to avoid allocating the queues now that they're not used at all. I was hoping you and Nicolin could work together on that. > > If you want to reduce the cmdq size, I'd prefer a cmdline option rather > > than special-casing kdump (as I've had other folks ask about configuring > > the cmdq size for other reasons). > > Fair enough for the cmdq. I'm less sure what to do about evtq/priq: they > get disabled in the kdump kernel anyway, so I don't know what a size knob > would mean for them, other than every kdump config having to pass it to > get back the memory. > > Would skipping their allocation when is_kdump_kernel() be acceptable on > its own, or would you rather see it done differently? With Nicolin's patch, we can just avoid allocating the priq and the evtq entirely in the kdump case. We can then add a cmdline option to control the maximum size of the cmdq, which is useful regardless of kdump. Will