Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] Mutex deadlock during fwnode_pcs_del_provider() due to notifier callback re-acquiring fwnode_pcs_mutex.
- [High] Concurrent list modification in phylink_create() error path due to missing mutex and incorrect cleanup ordering.
--
--- Patch [7]: [PATCH net-next v11 07/12] net: Document PCS subsystem ---
commit 0d23601c1b08a3129af4d3f510928e8e13e5be2d
Author: Christian Marangi <[email protected]>
net: Document PCS subsystem
Add extensive documentation of the new PCS subsystem and the fwnode
implementation with producer/consumer API.
Also update the sfp-phylink migration guide.
> diff --git a/drivers/net/pcs/pcs.c b/drivers/net/pcs/pcs.c
> new file mode 100644
> index 0000000000000..79fa596c2900d
> --- /dev/null
> +++ b/drivers/net/pcs/pcs.c
> @@ -0,0 +1,294 @@
> +// SPDX-License-Identifier: GPL-2.0-or-later
[ ... ]
> +void fwnode_pcs_del_provider(struct fwnode_pcs_provider *pp)
> +{
> + if (IS_ERR_OR_NULL(pp))
> + return;
> +
> + mutex_lock(&fwnode_pcs_mutex);
> +
> + /* Signal phylink to release any PCS from this provider */
> + blocking_notifier_call_chain(&fwnode_pcs_notify_list,
> + FWNODE_PCS_PROVIDER_DEL,
> + pp->fwnode);
[Severity: High]
Does this code cause a mutex deadlock?
When fwnode_pcs_del_provider() is called, it acquires fwnode_pcs_mutex to
protect the providers list before triggering the blocking notifier.
The registered notifier pcs_provider_notify() runs in the same thread context.
It handles the event by calling fwnode_phylink_pcs_get_from_fwnode(), which
cascades down to fwnode_pcs_get_from_pcsspec():
pcs_provider_notify()
fwnode_phylink_pcs_get_from_fwnode()
fwnode_pcs_get()
__fwnode_pcs_get()
fwnode_pcs_get_from_pcsspec()
mutex_lock(&fwnode_pcs_mutex);
Since fwnode_pcs_mutex is a standard (non-recursive) mutex, attempting to lock
it twice would result in a deterministic deadlock during normal teardown paths.
> +
> + list_del(&pp->link);
> + fwnode_dev_initialized(pp->fwnode, false);
> + fwnode_handle_put(pp->fwnode);
> + kfree(pp);
> +
> + mutex_unlock(&fwnode_pcs_mutex);
> +}
> +EXPORT_SYMBOL_GPL(fwnode_pcs_del_provider);
[ ... ]
> diff --git a/drivers/net/phy/phylink.c b/drivers/net/phy/phylink.c
> index b241768edbcba..69d4f6675db0f 100644
> --- a/drivers/net/phy/phylink.c
> +++ b/drivers/net/phy/phylink.c
[ ... ]
> @@ -1928,6 +2171,13 @@ struct phylink *phylink_create(struct phylink_config *config,
> release_link_gpio:
> if (pl->link_gpio)
> gpiod_put(pl->link_gpio);
> +unlink_pcs_list:
> + list_for_each_entry_safe(pcs, tmp, &pl->pcs_list, list) {
> + list_del(&pcs->list);
> + pcs->phylink = NULL;
> + }
> +unregister_pcs_notify:
> + unregister_fwnode_pcs_notifier(&pl->fwnode_pcs_nb);
[Severity: High]
Can this error path lead to concurrent list modification?
In this error path, pl->pcs_list is manipulated without holding
pl->state_mutex, and this occurs before the PCS notifier (pl->fwnode_pcs_nb)
is unregistered.
If a provider addition or deletion event fires concurrently,
pcs_provider_notify() will be invoked. The notifier safely acquires
pl->state_mutex to modify pl->pcs_list, but because this teardown path does
not hold the mutex while executing list_del(&pcs->list), it creates an unsafe
data race.
To safely fix this, could the teardown unregister the notifier first and then
hold pl->state_mutex while iterating the list, similar to how it is handled in
phylink_destroy()?
> free_pl:
> kfree(pl);
> return ERR_PTR(ret);
--
Sashiko AI review · https://sashiko.dev/#/patchset/[email protected]?part=7
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.