Re: [PATCH v4 07/10] dma-buf: heaps: Add support for Tegra VPR
Thierry Reding <[email protected]>
| Newsgroups | org.kernel.vger.linux-trace-kernel,dev.linux.lists.iommu,org.freedesktop.lists.dri-devel,org.kernel.vger.linux-devicetree,org.kernel.vger.linux-kernel,org.kernel.vger.linux-media,org.kernel.vger.linux-s390,org.kernel.vger.linux-tegra,org.kvack.linux-mm |
|---|---|
| Message-ID | <an8K2gb1pIqbDj6-@orome> |
On Thu, Aug 13, 2026 at 07:25:20PM +0100, Vincent Donnefort wrote: > On Fri, Aug 07, 2026 at 05:54:27PM +0200, Thierry Reding wrote: > > From: Thierry Reding <[email protected]> > > > > NVIDIA Tegra SoCs commonly define a Video-Protection-Region, which is a > > region of memory dedicated to content-protected video decode and > > playback. This memory cannot be accessed by the CPU and only certain > > hardware devices have access to it. > > > > Expose the VPR as a DMA heap so that applications and drivers can > > allocate buffers from this region for use-cases that require this kind > > of protected memory. > > > > VPR has a few very critical peculiarities. First, it must be a single > > contiguous region of memory (there is a single pair of registers that > > set the base address and size of the region), which is configured by > > calling back into the secure monitor. The memory region also needs to > > quite large for some use-cases because it needs to fit multiple video > > frames (8K video should be supported), so VPR sizes of ~2 GiB are > > expected. However, some devices cannot afford to reserve this amount > > of memory for a particular use-case, and therefore the VPR must be > > resizable. > > > > Unfortunately, resizing the VPR is slightly tricky because the GPU found > > on Tegra SoCs must be in reset during the VPR resize operation. This is > > currently implemented by freezing all userspace processes and calling > > invoking the GPU's freeze() implementation, resizing and the thawing the > > GPU and userspace processes. This is quite heavy-handed, so eventually > > it might be better to implement thawing/freezing in the GPU driver in > > such a way that they block accesses to the GPU so that the VPR resize > > operation can happen without suspending all userspace. > > > > In order to balance the memory usage versus the amount of resizing that > > needs to happen, the VPR is divided into multiple chunks. Each chunk is > > implemented as a CMA area that is completely allocated on first use to > > guarantee the contiguity of the VPR. Once all buffers from a chunk have > > been freed, the CMA area is deallocated and the memory returned to the > > system. > > Hi, > > I believe we (the Android team) are trying to solve similar issue to yours: Arm > CPUs can still speculatively read memory after it has been transitioned to the > Secure state, as long as they retain a cacheable mapping to it. > > As modifying the direct mapping is difficult, the workaround ended up in the > hypervisor which unmaps the pages from the host stage-2 on intercepted FF-A Lend > invocations. > > If convenient, this is nonetheless the wrong place to do it. Scattering the host > stage-2 is really terrible for performance and we would like to move it where it > should be, directly into the kernel... > > This is what I thought was the attempt in v3, but I am now a bit confused > because I see you are using set_direct_map_invalid_noflush() > set_direct_map_default_noflush(), but I am not sure that works if rodata=full is > not set? So does the memory for the NVIDIA IP still need to be unmapped? Yeah, this currently relies on the circumstances being such that can_set_direct_map() returns true, and in the case where we want to use the resizable VPR functionality, we're going to have page-granular mappings anyway. > If so, how about having an option where the CMA allocation is backed by a direct > map with the same page granularity, or at least a smaller and aligned granule? > > With that, we know that for whatever CMA allocation we do, we can safely unmap > from the direct map without risking splitting blocks. On CMA free, we can safely > remap into the direct map as I do not believe we coalesce yet. This sounds intriguing. For VPR we could possibly make the size a multiple of the memblock size (or whatever might be appropriate). If we can create a page-granular linear mapping specifically for that region, that'd be ideal. I don't know if the linear mapping can be subdivided in this fashion, though. > (This alignment is not necessary for upcoming systems with BBML3) > > However, to implement this, we'd need to split memblocks and I believe the only > way at the moment is temporarily mark it as "nomap" which didn't seem very > popular in the comments on v3. Could this be simplified if this type of allocation is always memblock aligned? Looking at map_mem(), it seems like we could add some sort of special- casing in the for_each_mem_range() block to check if the memory is VPR (or generic, page-granular carveout, or whatever we want to call it) and set NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS in that case. If that works, set_direct_map_*() could be enhanced to detect such cases and always work. Or perhaps a more specific API could be introduced. Thierry
signature.asc
(application/pgp-signature, 833 B)
-----BEGIN PGP SIGNATURE----- iQIzBAABCgAdFiEEiOrDCAFJzPfAjcif3SOs138+s6EFAmp/EGIACgkQ3SOs138+ s6FIfA/7BDMg2roMkVSyRLt06etjxYFBWHpNzxj/SvOzYROjoeOaUOJR15Xodk18 HI650LP1luxhtmIKPHRjGxZK30ROvHB6QTsl0lYM0cQZ4tVrv5WTMcL9uCui0WzA LkF9SVH4YiJdbE4tCnWc9eJSxxedpCsS8d5cHLfDyI2cKvjAxgTxll8INqCMafAU wehMyhd65emp5OYmvSuftsNzYf/uM0Oke8Rf64zZjOyZzpBi7AZG2CdROlrN/v1V xODXrjdPaZvdq6HkzFVW0GYsouBaWb2rqzw/IsgO5RbT6vg/oOwo+4z50ZDefimA x2mLXvg+gzqjY20hfJ82Bar0sCQ8nNlO6rLPpxwXMn9ICDpFV5aaiGvbaZ9ibDxR yScvx4k0i/k6NPFoyuKJFSqNn9s4WGqU6p0j+un2K97HMEPtJ3onKs3Pjj63E38i EkVAJZkrbt+PNYGiILRWWnU7dTqd5/CreBR7F0zFvVz0+JYjgfmlWcTdLKPAazag PpRU7SZKUfHRd+LnsS54706H2dOP1rq5NXx+AGSdRJcwgw30XbIhZ3CmqYBVjbdN Nh/5Y19dBDV5UhRx4o3Cqcb1BV/cuu2w0TpGyxoyIUUTFxgDp59uBi2Qi6nY/8kY KdeMys4RVWucUUwpbtpKJamulwDRGtlLjyLEIqVeYZY/FOBiaKU= =ERUC -----END PGP SIGNATURE-----