Re: [PATCH] md: fix soft lockup during resync when sync is repeatedly skipped

"yu kuai" <[email protected]> Sun, 19 Jul 2026 18:45:25 +0800
Newsgroups org.kernel.vger.linux-raid,org.kernel.vger.linux-kernel
Message-ID <[email protected]>
Hi,

在 2026/7/17 14:27, Yunye Zhao 写道:
> md_do_sync()'s main loop advances io_sectors only when I/O is actually
> issued (skipped == 0).  When sync_request() keeps returning skipped == 1,
> io_sectors never increases, the "last_check + window > io_sectors" test
> stays true, and every iteration takes the continue branch:

That's not expected, io_sectors should always increase in the skip case.

>
> 	sectors = mddev->pers->sync_request(mddev, j, max_sectors, &skipped);
> 	...
> 	if (!skipped)
> 		io_sectors += sectors;
> 	j += sectors;
> 	...
> 	if (last_check + window > io_sectors || j == max_sectors)
> 		continue;
>
> During recovery or resync of a large array with a sparse bitmap, many
> regions that need no syncing are skipped:
>
> 	raid10_sync_request()
> 	  md_bitmap_start_sync()	-> must_sync = false (no bitmap page)
> 	  /* every mirror skipped */
> 	  biolist == NULL -> *skipped = 1; return max_sync;
>
> j then traverses the whole skipped range while io_sectors stays
> unchanged.  On a non-preemptive kernel the resync thread (mdX_resync)
> hogs the CPU for a long time and eventually triggers a soft lockup:
>
>    watchdog: BUG: soft lockup - CPU#149 stuck for 313s! [mdX_resync]
>     md_bitmap_start_sync+0x6f/0xe0
>     raid10_sync_request+0x2c9/0x1530 [raid10]
>     md_do_sync+0x810/0x1030
>     md_thread+0xa7/0x150

What kernel version you're testing? If this is latest kernel, bitmap_start_sync()
need to be fixed. It can't return skip while setting skipping sectors to 0. And
since this is dead loop, a cond_resched() will not fix anything.

>
> Add cond_resched() to this continue path.
>
> Signed-off-by: Yunye Zhao <[email protected]>
> ---
>   drivers/md/md.c | 5 +++--
>   1 file changed, 3 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/md/md.c b/drivers/md/md.c
> index d1465bcd86c8..e7411b033490 100644
> --- a/drivers/md/md.c
> +++ b/drivers/md/md.c
> @@ -9881,9 +9881,10 @@ void md_do_sync(struct md_thread *thread)
>   			 */
>   			md_new_event();
>   
> -		if (last_check + window > io_sectors || j == max_sectors)
> +		if (last_check + window > io_sectors || j == max_sectors) {
> +			cond_resched();
>   			continue;
> -
> +		}
>   		last_check = io_sectors;
>   	repeat:
>   		if (time_after_eq(jiffies, mark[last_mark] + SYNC_MARK_STEP )) {

-- 
Thanks,
Kuai