Re: [PATCH v4 07/10] dma-buf: heaps: Add support for Tegra VPR

Vincent Donnefort <[email protected]>
Newsgroups org.kernel.vger.linux-tegra,dev.linux.lists.iommu,org.freedesktop.lists.dri-devel,org.kernel.vger.linux-devicetree,org.kernel.vger.linux-kernel,org.kernel.vger.linux-media,org.kernel.vger.linux-s390,org.kernel.vger.linux-trace-kernel,org.kvack.linux-mm
Message-ID <[email protected]>
On Fri, Aug 14, 2026 at 02:56:05PM +0200, Thierry Reding wrote:
> On Thu, Aug 13, 2026 at 07:25:20PM +0100, Vincent Donnefort wrote:
> > On Fri, Aug 07, 2026 at 05:54:27PM +0200, Thierry Reding wrote:
> > > From: Thierry Reding <[email protected]>
> > > 
> > > NVIDIA Tegra SoCs commonly define a Video-Protection-Region, which is a
> > > region of memory dedicated to content-protected video decode and
> > > playback. This memory cannot be accessed by the CPU and only certain
> > > hardware devices have access to it.
> > > 
> > > Expose the VPR as a DMA heap so that applications and drivers can
> > > allocate buffers from this region for use-cases that require this kind
> > > of protected memory.
> > > 
> > > VPR has a few very critical peculiarities. First, it must be a single
> > > contiguous region of memory (there is a single pair of registers that
> > > set the base address and size of the region), which is configured by
> > > calling back into the secure monitor. The memory region also needs to
> > > quite large for some use-cases because it needs to fit multiple video
> > > frames (8K video should be supported), so VPR sizes of ~2 GiB are
> > > expected. However, some devices cannot afford to reserve this amount
> > > of memory for a particular use-case, and therefore the VPR must be
> > > resizable.
> > > 
> > > Unfortunately, resizing the VPR is slightly tricky because the GPU found
> > > on Tegra SoCs must be in reset during the VPR resize operation. This is
> > > currently implemented by freezing all userspace processes and calling
> > > invoking the GPU's freeze() implementation, resizing and the thawing the
> > > GPU and userspace processes. This is quite heavy-handed, so eventually
> > > it might be better to implement thawing/freezing in the GPU driver in
> > > such a way that they block accesses to the GPU so that the VPR resize
> > > operation can happen without suspending all userspace.
> > > 
> > > In order to balance the memory usage versus the amount of resizing that
> > > needs to happen, the VPR is divided into multiple chunks. Each chunk is
> > > implemented as a CMA area that is completely allocated on first use to
> > > guarantee the contiguity of the VPR. Once all buffers from a chunk have
> > > been freed, the CMA area is deallocated and the memory returned to the
> > > system.
> > 
> > Hi,
> > 
> > I believe we (the Android team) are trying to solve similar issue to yours: Arm
> > CPUs can still speculatively read memory after it has been transitioned to the
> > Secure state, as long as they retain a cacheable mapping to it.
> > 
> > As modifying the direct mapping is difficult, the workaround ended up in the
> > hypervisor which unmaps the pages from the host stage-2 on intercepted FF-A Lend
> > invocations.
> > 
> > If convenient, this is nonetheless the wrong place to do it. Scattering the host
> > stage-2 is really terrible for performance and we would like to move it where it
> > should be, directly into the kernel...
> > 
> > This is what I thought was the attempt in v3, but I am now a bit confused
> > because I see you are using set_direct_map_invalid_noflush()
> > set_direct_map_default_noflush(), but I am not sure that works if rodata=full is
> > not set? So does the memory for the NVIDIA IP still need to be unmapped?
> 
> Yeah, this currently relies on the circumstances being such that
> can_set_direct_map() returns true, and in the case where we want to use
> the resizable VPR functionality, we're going to have page-granular
> mappings anyway.
> 
> > If so, how about having an option where the CMA allocation is backed by a direct
> > map with the same page granularity, or at least a smaller and aligned granule?
> > 
> > With that, we know that for whatever CMA allocation we do, we can safely unmap
> > from the direct map without risking splitting blocks. On CMA free, we can safely
> > remap into the direct map as I do not believe we coalesce yet.
> 
> This sounds intriguing. For VPR we could possibly make the size a
> multiple of the memblock size (or whatever might be appropriate). If we
> can create a page-granular linear mapping specifically for that region,
> that'd be ideal. I don't know if the linear mapping can be subdivided in
> this fashion, though.

Actually I don't think modifying the memblock is necessary at all!

Here's what I have so far:

diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
index d4de88770ecf..a9f81413b70d 100644
--- a/arch/arm64/mm/mmu.c
+++ b/arch/arm64/mm/mmu.c
@@ -22,6 +22,7 @@
 #include <linux/fs.h>
 #include <linux/io.h>
 #include <linux/mm.h>
+#include <linux/of_fdt.h>
 #include <linux/vmalloc.h>
 #include <linux/set_memory.h>
 #include <linux/kfence.h>
@@ -1138,6 +1139,27 @@ static inline void arm64_kfence_map_pool(void) { }
 
 #endif /* CONFIG_KFENCE */
 
+#define MAX_FORCE_PTE_REGIONS 64 /* Greater or equal to MAX_RESERVED_REGIONS */
+
+static struct {
+       phys_addr_t start;
+       phys_addr_t end;
+} force_pte_regions[MAX_FORCE_PTE_REGIONS] __initdata;
+
+void __init early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end)
+{
+       static int count;
+
+       if (force_pte_mapping())
+               return;
+
+       if (count >= MAX_FORCE_PTE_REGIONS)
+               return;
+
+       force_pte_regions[count].start = start;
+       force_pte_regions[count++].end = end;
+}
+
 static void __init map_mem(void)
 {
        static const u64 direct_map_end = _PAGE_END(VA_BITS_MIN);
@@ -1186,6 +1208,15 @@ static void __init map_mem(void)
        __map_memblock(init_end, kernel_end, pgprot_tagged(PAGE_KERNEL),
                       flags);
 
+       for (i = 0; i < ARRAY_SIZE(force_pte_regions); i++) {
+               if (!force_pte_regions[i].end)
+                       break;
+
+               __map_memblock(force_pte_regions[i].start, force_pte_regions[i].end,
+                              pgprot_tagged(PAGE_KERNEL),
+                              flags | NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS);
+       }
+
        /* map all the memory banks */
        for_each_mem_range(i, &start, &end) {
                /*
diff --git a/drivers/of/of_reserved_mem.c b/drivers/of/of_reserved_mem.c
index 42e3e2d8a2b8..445f204f01e4 100644
--- a/drivers/of/of_reserved_mem.c
+++ b/drivers/of/of_reserved_mem.c
@@ -136,6 +136,8 @@ static int __init early_init_dt_reserve_memory(phys_addr_t base,
        return memblock_reserve(base, size);
 }
 
+void __weak early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end) { }
+
 /*
  * __reserved_mem_reserve_reg() - reserve memory described in the
  * first entry in 'reg' property
@@ -168,6 +170,9 @@ static int __init __reserved_mem_reserve_reg(unsigned long node,
        size = s;
 
        if (size && early_init_dt_reserve_memory(base, size, nomap) == 0) {
+               if (of_get_flat_dt_prop(node, "force-pte", NULL))
+                       early_init_dt_force_pte_arch(base, base + size);
+
                fdt_fixup_reserved_mem_node(node, base, size);
                pr_debug("Reserved memory: reserved region for node '%s': base %pa, size %lu MiB\n",
                         uname, &base, (unsigned long)(size / SZ_1M));
@@ -517,6 +522,9 @@ static int __init __reserved_mem_alloc_size(unsigned long node, const char *unam
                return -ENOMEM;
        }
 
+       if (of_get_flat_dt_prop(node, "force-pte", NULL))
+               early_init_dt_force_pte_arch(base, base + size);
+
        fdt_fixup_reserved_mem_node(node, base, size);
        fdt_init_reserved_mem_node(node, uname, base, size);
diff --git a/include/linux/of_fdt.h b/include/linux/of_fdt.h
index 51dadbaa3d63..6c73c76dee1b 100644
--- a/include/linux/of_fdt.h
+++ b/include/linux/of_fdt.h
@@ -74,6 +74,7 @@ extern void early_init_dt_check_for_usable_mem_range(void);
 extern int early_init_dt_scan_chosen_stdout(void);
 extern void early_init_fdt_scan_reserved_mem(void);
 extern void early_init_fdt_reserve_self(void);
+extern void early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end);
 extern void early_init_dt_add_memory_arch(u64 base, u64 size);
 extern u64 dt_mem_next_cell(int s, const __be32 **cellp);
 
> 
> > (This alignment is not necessary for upcoming systems with BBML3)
> > 
> > However, to implement this, we'd need to split memblocks and I believe the only
> > way at the moment is temporarily mark it as "nomap" which didn't seem very
> > popular in the comments on v3.
> 
> Could this be simplified if this type of allocation is always memblock
> aligned?
> 
> Looking at map_mem(), it seems like we could add some sort of special-
> casing in the for_each_mem_range() block to check if the memory is VPR
> (or generic, page-granular carveout, or whatever we want to call it) and
> set NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS in that case.
> 
> If that works, set_direct_map_*() could be enhanced to detect such cases
> and always work. Or perhaps a more specific API could be introduced.
> 
> Thierry

for set_direct_map() I think we could extend it to check if it is mapped at the
PTE-level and if it is we can proceed?

I am currently looking at extending CMA with an option "unmap-on-alloc;" that
would only be available if CONFIG_ARCH_HAS_SET_DIRECT_MAP, or
cma_set_unmap_on_alloc, or if "force-pte;" is set. 

Hopefully I can share something this week.

-- 
Vincent
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.