Re: [PATCH v8] net: airoha: npu: use cacheline-sized buffers for mailbox DMA

Daniel Pawlik <[email protected]>
Newsgroups org.infradead.lists.linux-mediatek,org.infradead.lists.linux-arm-kernel,org.kernel.vger.netdev
Message-ID <CALC2ApiCSY8DkVpd7puTssbJZMWBHGVtFCnG66fiK0PX7-u3Hw@mail.gmail.com>
Hi,

TL;DR: The kernel forces swiotlb bounce for sub-cacheline-sized DMA on
non-coherent ARM64. The EN7581 NPU cannot write to the swiotlb bounce
buffer region. Cache-line-aligned buffers bypass swiotlb entirely and
streaming DMA works — NPU version reads 0.1111, WiFi fully functional.

Test progression (EN7581 / Gemtek W1700K, kernel 6.18, cache_align=64):

Boot 1 - stock (dma_set_coherent_mask only):
  91/91 swiotlb bounce, NPU hangs after ~41 calls, version 0.0

Boot 2 - + dma_set_mask_and_coherent(DMA_BIT_MASK(32)):
  Still 91/91 swiotlb. Hang gone (all 116 succeed), version still 0.0.
  Bounce-direct read via phys_to_virt matches pre-map — NPU never wrote.

Boot 3 - + dcache_inval_poc before bounce read:
  Bounce buffer still matches pre-map after cache invalidation.
  Rules out CPU cache — NPU genuinely never writes to bounce region.

Boot 4 - + phys/mask logging:
  dma_mask=0xffffffff, phys fits. size=24 fails IS_ALIGNED(size, 64)
  in dma_direct_map_page() — this is why swiotlb is forced.

Boot 5 - kmalloc(ALIGN(size, 64), GFP_KERNEL):
  is_swiotlb=0 on all 169 mappings. Version 0.1111. 0 timeouts.
  mt7996e firmware loaded, WiFi works.

  Version query (size=12):
    pre-map:    30 00 00 00 0a 00 00 00 00 00 00 00
    post-unmap: 30 00 00 00 0a 00 00 00 57 04 00 00
    Bytes 8-11 now contain response data (always zeros in boots 1-4).

Root cause:

dma_direct_map_page() forces swiotlb when:

  !dev_is_dma_coherent(dev) &&
  !IS_ALIGNED(phys | size, dma_get_cache_alignment())

size=24 < cache_line=64 fails this regardless of address alignment.
The swiotlb sync code itself is fine, but the NPU cannot DMA-write to
the bounce buffer address range (it reads commands fine, never writes
responses back). This is probably an EN7581 platform limitation.

Testing a fix:

Align the mailbox DMA buffer to cache line size:

  aligned_size = ALIGN(size, dma_get_cache_alignment());
  dma_buf = kmalloc(aligned_size, GFP_KERNEL);
  dma_map_single(dev, dma_buf, aligned_size, DMA_BIDIRECTIONAL);

This keeps streaming DMA, eliminates swiotlb, and works. The separate
dma_set_mask_and_coherent() fix is still needed (fixes hang after ~41
calls) but alone does not prevent bounce — the sub-cacheline size
triggers it regardless of mask.

Debug commit:
https://github.com/openwrt/openwrt/commit/c9363b2597528eb3badef3bb04608c0d21c81b6f

Boot 5 dmesg available if needed.

Regards,
Dan

śr., 19 sie 2026 o 14:27 Daniel Pawlik <[email protected]> napisał(a):
>
> Hi,
>
> I instrumented airoha_npu_send_msg() on EN7581 (Gemtek W1700K) with
> DMA_BIDIRECTIONAL and the original unaligned 24-byte payload. Here is
> what I found.
>
> 1) Every dma_map_single() call lands in swiotlb bounce buffers.
>    The driver only calls dma_set_coherent_mask(DMA_BIT_MASK(32)) at
>    probe but never calls dma_set_mask(). On this non-coherent ARM64
>    platform that appears to force all streaming DMA through swiotlb.
>
> 2) The NPU permanently hangs after ~41 mailbox calls.
>    The first 41 PPE SRAM-init calls succeed (mbox_status=0x7, ~10ms
>    each). Starting at the 42nd, the NPU never sets the DONE bit —
>    every subsequent call times out at 100ms (ret=-110,
>    mbox_status=0x1), including the WLAN version query that triggers
>    the "failed getting NPU fw version" probe failure.
>
> 3) Response data is never synced back through the bounce buffer.
>    post-sync and post-unmap hex dumps are byte-identical to pre-map
>    for both successful and timed-out calls.
>
> 4) Buffer pointer is 64-byte aligned (p_aligned=1), size is 24
>    (sub-cacheline, cache_align=64). The swiotlb slot size is 2048,
>    so the bounce buffer itself is always page-aligned.
>
> 5) dma_alloc_coherent mappings (shown by DMA-API debug dump) map
>    P==D at addresses around 0x91xxxxxx, bypassing swiotlb entirely.
>    The mailbox works reliably through that path.
>
> Representative trace (last success / first failure):
>
>   [19.541] func_id=0 p=...dfc0 size=24 cache_align=64 p_aligned=1
>   [19.545] mapped: dma=0xff765000 dma_aligned=1 is_swiotlb=1
>   [19.541] mbox done: ret=0 mbox_status=0x7      <-- last success
>
>   [19.588] func_id=0 p=...dfc0 size=24 cache_align=64 p_aligned=1
>   [19.618] mapped: dma=0xff765800 dma_aligned=1 is_swiotlb=1
>   [19.728] mbox done: ret=-110 mbox_status=0x1    <-- first timeout
>   (all subsequent calls also timeout, NPU never recovers)
>
> My working theory is that the missing dma_set_mask() is the root cause.
> Without it the kernel bounces every streaming mapping through swiotlb,
> and either the bounce-buffer sync path on this non-coherent platform
> has a bug with sub-cacheline transfers, or the NPU firmware cannot
> handle the latency/address-range change that bouncing introduces.
>
> I plan to test next with dma_set_mask_and_coherent(DMA_BIT_MASK(32))
> to see whether that eliminates the swiotlb bounce and restores the
> mailbox. If it does, that narrows the bug to either the swiotlb sync
> implementation or the platform DMA ops for non-coherent devices.
>
> Debug instrumentation patch (applied on top of OpenWrt's 6.18.44):
> https://github.com/openwrt/openwrt/commit/13ea79d9a411dc497882e820e4c881443457bd97
>
> Full dmesg from the instrumented boot:
> https://gist.github.com/danpawlik/b351e8a06218a3aaf68a2cb31a32ea8f
>
> Thanks for help,
> Dan
>
>
> wt., 18 sie 2026 o 17:29 Jakub Kicinski <[email protected]> napisał(a):
> >
> > On Tue, 18 Aug 2026 08:23:33 +0200 Daniel Pawlik wrote:
> > > v8 does not switch to coherent DMA; it only rounds up mailbox payload
> > > allocations (dma_get_cache_alignment() in v9) and maps the same rounded
> > > length with DMA_BIDIRECTIONAL, while the mailbox length register still
> > > uses the original payload size. On this board, mapping only the payload
> > > size regresses probe back to 0.0 even with the larger allocation, so the
> > > rounded map length seems required here.
> >
> > Please trace into the dma API implementation on this platform where the
> > alignment makes a difference. The "unaligned length" path must be buggy.
> > We should fix it there, not in all the drivers (my concern being that
> > this is not the only driver that hits the issue).
>
>
>
> --
> Z poważaniem,
> Daniel Pawlik



-- 
Z poważaniem,
Daniel Pawlik
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.