Re: dma_opt_mapping_size returns way too low sizes when using IOMMU

John Garry <[email protected]>
Newsgroups org.infradead.lists.linux-nvme,dev.linux.lists.iommu
Organization Oracle Corporation
Message-ID <[email protected]>
On 17/08/2026 10:18, Christoph Hellwig wrote:
> On Mon, Aug 17, 2026 at 10:12:25AM +0100, John Garry wrote:
>> On 17/08/2026 09:36, Christoph Hellwig wrote:
>>> Hi all,
>>>
>>> I got reports that NVMe devices were arbitrarily limited to 128kiB
>>> transfers in recent kernel. 
>>
>> How recent a kernel? This NVMe and DMA mapping code has not changed in 
>> years as far as I know.
> 
> This was hardware QA moving from an old distro kernel to a "recent" (aka
> still old) one.   But I've actually reproduced it locally on an
> upstream kernel.  I guess most kernel developers or power users simply
> do not run with IOMMU enabled.

Yeah, they don't like any performance hit.

> 
>>> Both the NVMe performance numbers and common sense suggest that this
>>> is NOT the optimal DMA mapping granularity.  Can we pick a saner value
>>> for iommu_dma_opt_mapping_size that does not restrict common I/O sizes?
>>
>> Note that SCSI does not use iommu_dma_opt_mapping_size() for clamping max 
>> HW sectors, but instead sets opt size / max sectors from this value (so it 
>> is not a hard limit there). Could we consider similar for NVMe?
> 
> We could consider that, but it would still reduce performane..
 > >> The reason for which we have iommu_dma_opt_mapping_size() is that
>> performance can go through the floor we can't use the rcache for getting 
>> the IOVA, i.e. we need to always alloc and dealloc an IOVA from the RB tree 
>> for each mapping, and this can greatly reduce performance when the IOVA 
>> space fills.
>>
>> If we increase IOVA_RANGE_CACHE_MAX_SIZE, then we just get caching of 
>> larger IOVAs and I am not sure that is a great idea.
> 
> At least on the four different nvme devices I tested, the larger I/O
> sizes made up for this.  But maybe the details depend on other
> factors as well.
Are you saying that you tried increasing IOVA_RANGE_CACHE_MAX_SIZE and 
got better performance?

As I remember, I was told that the value of 6 for 
IOVA_RANGE_CACHE_MAX_SIZE was originally chosen from the value then in 
max page order, i.e. the idea was that we should not be getting 
streaming IOs larger than that value. But in looking at lore, 8 was very 
originally proposed, but I can't see any discussion on why that changed 
or any relation to page max order.

Some time ago I did try some work to allow IOVA_RANGE_CACHE_MAX_SIZE be 
set per IOMMU domain, but it was not merged. We went with 
dma_opt_mapping_size() solution instead.

https://lore.kernel.org/linux-scsi/[email protected]/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.