Re: [PATCH 2/2] net: wwan: qcom_bam_dmux: Alloc RX buffers as a single coherent block
Vishnu Santhosh <[email protected]>
| Newsgroups | org.kernel.vger.linux-arm-msm,org.kernel.vger.linux-devicetree,org.kernel.vger.linux-kernel,org.kernel.vger.netdev |
|---|---|
| Message-ID | <[email protected]> |
On 24-07-2026 03:04 pm, Stephan Gerhold wrote: > On Fri, Jul 24, 2026 at 10:16:31AM +0530, Vishnu Santhosh wrote: >> On 14-07-2026 01:05 pm, Stephan Gerhold wrote: >>> On Tue, Jul 14, 2026 at 11:02:32AM +0530, Vishnu Santhosh wrote: >>>> On Qualcomm SoCs where the modem (e.g. the mDSP on Shikra, VMID 43 / >>>> NAV) is the AXI master for BAM-DMUX RX transfers and the XPU enforces >>>> per-region access control, each individually DMA-mapped RX buffer >>>> requires its own XPU resource group (RG). With ~16 RGs available, the >>>> 32 per-buffer dma_map_single() calls exhaust the table and the first >>>> inbound transfer faults with an XPU violation. >>>> >>>> BAM-DMUX is a singleton (exactly one instance per SoC), so the >>>> destination VMID does not need to be a DT property; it is looked up >>>> from the compatible string's match data instead. Add struct >>>> bam_dmux_data with a single vmid field, and a shikra_data instance >>>> hardcoding QCOM_SCM_VMID_NAV for qcom,shikra-bam-dmux. >>>> >>>> When match data is present, allocate all BAM_DMUX_NUM_SKB RX buffers as >>>> a single contiguous dma_alloc_coherent() block and SCM-assign that >>>> block to HLOS plus the VMID once at probe. This reduces RG consumption >>>> from 32 to 1. The block is never reclaimed across a modem power cycle >>>> (bam_dmux_power_off() does not touch it), so the probe-time assignment >>>> covers every subsequent restart without re-assigning or reclaiming. It >>>> is reclaimed to HLOS only once, at remove or on a probe error, and if >>>> that reclaim fails it is leaked rather than returned to the page >>>> allocator. >>>> >>>> Each rx_skbs[] slot is pre-assigned its virtual and DMA address from >>>> the block, so no per-buffer mapping is needed at power-on. Because the >>>> coherent block is not page-backed, received payload is copied into a >>>> regular netdev skb before handoff to the network stack; this is an >>>> unavoidable extra copy on the XPU-enforced RX path. >>>> >>>> Platforms without match data are unaffected: rx_virt stays NULL, no >>>> coherent memory is allocated, and the per-buffer dma_map_single() path >>>> is unchanged. >>>> >>>> Co-developed-by: Deepak Kumar Singh <[email protected]> >>>> Signed-off-by: Deepak Kumar Singh <[email protected]> >>>> Signed-off-by: Vishnu Santhosh <[email protected]> >>> So how do you handle TX buffers? Right now, they are just passed on from >>> the net subsystem. There can be up to 32 TX buffers in progress as well. >>> >>> Overall, I have mixed feelings about this patch. It looks reasonably >>> simple, but fundamentally I don't understand why we need to go back to >>> the old days of implementing protection using a highly limited MPU (in >>> your case: the xPU). >>> >>> Why does the setup of BAM-DMUX differ e.g. from the setup for the crypto >>> engine? Crypto is also using bam-dma, but it avoids this inflexibility >>> by making use of the &apps_smmu. Is BAM-DMUX not covered by the SMMU? Or >>> did you just decide to bypass the SMMU in this case? (If so: Why?) >> I checked with secure systems team on this. Crypto BAM is >> behind apps_smmu, so protection is enforced through the SMMU's Stage-2 >> page tables. >> >> A2 BAM (used by BAM-DMUX) is present in secure domain and does not >> support Stage-2 translation on this SoC, and there is no IOMMU domain >> that can be attached to it. The only protection mechanism available is >> the xPU. >> > Thanks for investigating this! > > So is this a hardware limitation or something you could change with a > firmware update? Could you move the A2 BAM out of the secure domain and > protect it via the IOMMU instead of the xPU mechanism? The other modern > platforms with IPA do not have this limitation, they can use the IOMMU > for this. > > We can try to support the xPU protection mechanism in the BAM-DMUX > driver, but it's pretty bad from a performance and memory usage point of > view if you need to copy buffers around multiple times. So if you have > some way to change this in the firmware (and there is still time to do > so before production boards ship), I would strongly recommend to > investigate that. > > Thanks, > Stephan Based on what we confirmed with the Secure Systems team, this is a limitation of the current Shikra platform rather than something that can be addressed through a firmware-only update. The A2 BAM used by BAM-DMUX is not connected to an SMMU/IOMMU domain on Shikra, which means Stage-2 translation is not available for this path. As a result, it is not possible to move this path behind an IOMMU. Modern IPA-based platforms differ because their data paths are physically routed through SMMU interfaces and therefore do not rely on VMID/xPU ownership assignment for this type of access control. On Shikra, the A2 BAM path is protected using the xPU3 VM-based access-control model, where DDR memory access is restricted and granted through the request-based Hypervisor VM assignment framework. In contrast, older targets relied on the earlier xPU2 resource-sharing model, in which modem access did not require this type of explicit VM ownership configuration. This architectural difference explains why the issue does not occur on those older platforms. So, for this path on Shikra, SCM-driven VMID assignment remains the only practical solution. Thanks, Vishnu