Re: [PATCH] mm/vmalloc: fix vmap_purge_lock livelock under memory pressure

Uladzislau Rezki <[email protected]>
Newsgroups org.kvack.linux-mm,org.kernel.vger.linux-kernel
Message-ID <aoxE7rw6gru13DB5@milan>
On Mon, Aug 24, 2026 at 05:50:19PM +0800, Ye Liu wrote:
> From: Ye Liu <[email protected]>
> 
> The vmap_purge_lock mutex can be held for an extended period by
> __purge_vmap_area_lazy() which calls flush_work() to wait for
> purge_vmap_node workers while holding the lock.  Under memory
> pressure, those workers may themselves be blocked in direct
> reclaim trying to acquire the same lock via the
> vmap_node_shrink_scan() shrinker callback, creating a circular
> dependency that deadlocks the entire system.
> 
> Two changes:
> 
> 1. vmap_node_shrink_scan(): replace blocking guard(mutex) with
>    mutex_trylock().  This is a shrinker that only decays the vmap
>    pool and returns SHRINK_STOP without freeing memory; skipping a
>    decay cycle when the lock is contended is harmless and prevents
>    tasks from piling up on the mutex in the direct reclaim path.
> 
> 2. __purge_vmap_area_lazy(): purge all vmap nodes inline instead of
>    scheduling purge_vmap_node via workqueue and calling flush_work()
>    while holding vmap_purge_lock.  The workqueue dispatch + flush
>    pattern under a mutex is the deadlock trigger: if no free worker
>    threads are available (all blocked on the same lock), flush_work()
>    never returns and the lock is held indefinitely.
> 
> Fixes: 7679ba6b36db ("mm: vmalloc: add a shrinker to drain vmap pools")
> Signed-off-by: Ye Liu <[email protected]>
> ---
>  mm/vmalloc.c | 51 +++++++++++++++++++++------------------------------
>  1 file changed, 21 insertions(+), 30 deletions(-)
> 
> diff --git a/mm/vmalloc.c b/mm/vmalloc.c
> index bea9f76ed7e7..41443708d22d 100644
> --- a/mm/vmalloc.c
> +++ b/mm/vmalloc.c
> @@ -2358,7 +2358,6 @@ static bool __purge_vmap_area_lazy(unsigned long start, unsigned long end,
>  		bool full_pool_decay)
>  {
>  	unsigned long nr_purged_areas = 0;
> -	unsigned int nr_purge_helpers;
>  	static cpumask_t purge_nodes;
>  	unsigned int nr_purge_nodes;
>  	struct vmap_node *vn;
> @@ -2397,36 +2396,17 @@ static bool __purge_vmap_area_lazy(unsigned long start, unsigned long end,
>  	if (nr_purge_nodes > 0) {
>  		flush_tlb_kernel_range(start, end);
>  
> -		/* One extra worker is per a lazy_max_pages() full set minus one. */
> -		nr_purge_helpers = atomic_long_read(&vmap_lazy_nr) / lazy_max_pages();
> -		nr_purge_helpers = clamp(nr_purge_helpers, 1U, nr_purge_nodes) - 1;
> -
> -		for_each_cpu(i, &purge_nodes) {
> -			vn = &vmap_nodes[i];
> -
> -			if (nr_purge_helpers > 0) {
> -				INIT_WORK(&vn->purge_work, purge_vmap_node);
> -
> -				if (cpumask_test_cpu(i, cpu_online_mask))
> -					schedule_work_on(i, &vn->purge_work);
> -				else
> -					schedule_work(&vn->purge_work);
> -
> -				nr_purge_helpers--;
> -			} else {
> -				vn->purge_work.func = NULL;
> -				purge_vmap_node(&vn->purge_work);
> -				nr_purged_areas += vn->nr_purged;
> -			}
> -		}
> -
> +		/*
> +		 * Purge all nodes inline.  Do not schedule_work() and
> +		 * flush_work() here: flush_work() while holding
> +		 * vmap_purge_lock can deadlock if the worker pool is
> +		 * starved (e.g. all workers blocked on this same lock
> +		 * in the direct reclaim path via vmap_node_shrink_scan).
> +		 */
>  		for_each_cpu(i, &purge_nodes) {
>  			vn = &vmap_nodes[i];
> -
> -			if (vn->purge_work.func) {
> -				flush_work(&vn->purge_work);
> -				nr_purged_areas += vn->nr_purged;
> -			}
> +			purge_vmap_node(&vn->purge_work);
> +			nr_purged_areas += vn->nr_purged;
>  		}
>  	}
>  
> @@ -5519,10 +5499,21 @@ vmap_node_shrink_scan(struct shrinker *shrink, struct shrink_control *sc)
>  {
>  	struct vmap_node *vn;
>  
> -	guard(mutex)(&vmap_purge_lock);
> +	/*
> +	 * This shrinker is invoked from direct reclaim path where memory
> +	 * pressure is already high.  Blocking on vmap_purge_lock here can
> +	 * cause a pile-up of tasks all waiting for the same mutex while
> +	 * the lock holder may itself be blocked in flush_work() waiting for
> +	 * a worker that is stuck in the same reclaim path.  Use trylock to
> +	 * avoid this; skipping a pool decay cycle is harmless.
> +	 */
> +	if (!mutex_trylock(&vmap_purge_lock))
> +		return SHRINK_STOP;
> +
>  	for_each_vmap_node(vn)
>  		decay_va_pool_node(vn, true);
>  
> +	mutex_unlock(&vmap_purge_lock);
>  	return SHRINK_STOP;
>  }
>  
> -- 
> 2.25.1
> 
Have you seen any report about the problem you described?

--
Uladzislau Rezki
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.