Re: [PATCH] drm/msm/dp: skip PUSH_IDLE when the link was never enabled

Konrad Dybcio <[email protected]>
Newsgroups org.kernel.vger.linux-arm-msm,org.freedesktop.lists.dri-devel,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
On 8/8/26 7:13 PM, Jesse Casco wrote:
> msm_dp_display_atomic_enable() returns early when link training fails,
> leaving ->power_on false and the main link down.
> msm_dp_display_atomic_disable() nevertheless writes DP_STATE_CTRL_PUSH_IDLE
> and waits for an idle-pattern completion that cannot arrive, so every failed
> enable is followed by "PUSH_IDLE pattern timedout".
> 
> Every other step of the teardown is already gated on that flag:
> msm_dp_display_disable(), called from .atomic_post_disable(), returns early
> on !power_on. The PUSH_IDLE write is the only one that is not, so the
> controller's runtime-PM reference is then dropped without the link having
> been taken down.
> 
> On glymur (Snapdragon X2 Elite) the consequence is not a warning. The SoC
> does not survive it: TrustZone force-stops the SOCCP and ADSP remote
> processors and the machine resets silently about 50 ms later, with no oops
> and no panic. On an ASUS Zenbook A16 (UX3607OA), whose eDP panel does not
> currently train, this reproduces without any compositor or GPU involvement:
> 
>   # eDP enable has already failed with "Failed link training (rc=-104)"
>   echo 1 > /sys/class/graphics/fb0/blank
> 
>   [535.645455] === marker ===
>   [535.694833] qcom_q6v5_pas d00000.remoteproc: fatal error received: \
>                  sys_m_smsm.c:512:TZ force stop
>   [535.694875] remoteproc remoteproc0: crash detected in soccp: type fatal error
>   [535.728857] qcom_q6v5_pas 6800000.remoteproc: fatal error received: \
>                  sys_m_smsm.c:783:err fatal notification received from TZ
>   <SoC reset>

FWIW the rproc crashes are only a symptom - if the system is going down due
to (among other causes) unclocked access, all remote processors will be
forced to go down and this is what Linux sees

[...]

> +	/*
> +	 * If .atomic_enable() bailed out - link training failure is the common
> +	 * case - the mainlink was never brought up and ->power_on stayed false.
> +	 * Driving the PUSH_IDLE pattern into a controller that was never
> +	 * enabled times out, and .atomic_post_disable() then drops the
> +	 * controller's runtime-PM reference without tearing the PHY back down,
> +	 * because msm_dp_display_disable() returns early on !power_on.  On
> +	 * glymur (Snapdragon X2 Elite) that combination is answered by a
> +	 * TrustZone-level SOCCP/ADSP force-stop and a silent SoC reset.
> +	 * There is nothing to push idle, so leave it alone.
> +	 */
> +	if (!dp->power_on)
> +		return;

Dmitry, should we even be reaching atomic_disable if lt fails before?

Konrad
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.