Re: dma_opt_mapping_size returns way too low sizes when using IOMMU
Christoph Hellwig <[email protected]>
| Newsgroups | org.infradead.lists.linux-nvme,dev.linux.lists.iommu |
|---|---|
| Message-ID | <[email protected]> |
On Mon, Aug 17, 2026 at 10:12:25AM +0100, John Garry wrote: > On 17/08/2026 09:36, Christoph Hellwig wrote: >> Hi all, >> >> I got reports that NVMe devices were arbitrarily limited to 128kiB >> transfers in recent kernel. > > How recent a kernel? This NVMe and DMA mapping code has not changed in > years as far as I know. This was hardware QA moving from an old distro kernel to a "recent" (aka still old) one. But I've actually reproduced it locally on an upstream kernel. I guess most kernel developers or power users simply do not run with IOMMU enabled. >> Both the NVMe performance numbers and common sense suggest that this >> is NOT the optimal DMA mapping granularity. Can we pick a saner value >> for iommu_dma_opt_mapping_size that does not restrict common I/O sizes? > > Note that SCSI does not use iommu_dma_opt_mapping_size() for clamping max > HW sectors, but instead sets opt size / max sectors from this value (so it > is not a hard limit there). Could we consider similar for NVMe? We could consider that, but it would still reduce performane.. > The reason for which we have iommu_dma_opt_mapping_size() is that > performance can go through the floor we can't use the rcache for getting > the IOVA, i.e. we need to always alloc and dealloc an IOVA from the RB tree > for each mapping, and this can greatly reduce performance when the IOVA > space fills. > > If we increase IOVA_RANGE_CACHE_MAX_SIZE, then we just get caching of > larger IOVAs and I am not sure that is a great idea. At least on the four different nvme devices I tested, the larger I/O sizes made up for this. But maybe the details depend on other factors as well.