Re: [PATCH v8 02/12] accel/rocket: wait for a running IRQ handler before resetting a core
Igor Paunovic <[email protected]>
| Newsgroups | org.infradead.lists.linux-rockchip,dev.linux.lists.iommu,org.freedesktop.lists.dri-devel,org.infradead.lists.linux-arm-kernel,org.kernel.vger.linux-kernel,org.kernel.vger.linux-pm |
|---|---|
| Message-ID | <[email protected]> |
Hi Jiaxing, Thank you for the round history and for running my archive check - the zero stall/paging count telling us the MMU is not responding at all is a better characterization than anything I had. put_noidle vs put_autosuspend with a separately forced suspend/resume sounds like the right de-confounding split; I will watch for the result. Here is the induced reset test I promised, run this morning. Setup: RK3588 (Orange Pi 5 Plus), all three cores bound. My 7.2-rc6 tree with exactly two rocket changes from your series - 1/12 and 2/12 - plus one local test-only patch lowering JOB_TIMEOUT_MS to 2 ms so that healthy jobs (~5 ms at this clock) cross the timeout deterministically. No other rocket changes; in particular my lifecycle series is not applied. PROVE_LOCKING=y and DEBUG_ATOMIC_SLEEP=y. Serial console captured on a second machine for the whole session. Protocol, built around the trap you described - the RK3576 symptom emits from rk_iommu_enable() on the next attach, not from the reset itself: 20 scheduler-driven runs over the model set with all three cores active, a follow-up inference after every induced reset, then a forced autosuspend cycle and one more inference. Two full passes, at console_loglevel 8 and 4, because synchronous serial printing on this path can perturb the timing. Results: - Pass 1 (loglevel 8): 12 induced resets. Pass 2 (loglevel 4): 8. - Every reset recovered. Zero MMU_DTE_ADDR, zero "Error during raw reset", zero lockdep or atomic-sleep hits across both passes. - Outputs matched the oracle in 48/48 checks per pass, including the inference after the forced suspend/resume. - All three cores returned to runtime-suspended between rounds; the domain did drop and come back cleanly after every reset. So on RK3588 with 1/12+2/12 the block comes back every time, and your non-recovery does not reproduce. Combined with your archive result this is consistent with the failure being RK3576-specific on the platform/IOMMU side rather than rocket-wide. Two honest limits on what this run shows: 1. All resets ran with three cores bound. Isolating a single core requires unbinding the other two, and without my pending lifecycle fixes that path is not safe on this tree (the list corruption I reported on Aug 12), so I skipped it deliberately rather than test through a known-broken path. 2. This run alone says nothing about the race 1/12+2/12 close. The same protocol on the base without those two patches is queued as a separate build; only that differential earns a Tested-by, and when it lands the tag will carry its conditions: # RK3588, three cores, induced reset, JOB_TIMEOUT_MS=2 Raw logs (dmesg, per-run outputs, serial capture) are kept; happy to share any of it on request. Regards, Igor _______________________________________________ Linux-rockchip mailing list [email protected] http://lists.infradead.org/mailman/listinfo/linux-rockchip