[PATCH] md/raid5: Fix bio retry on interrupted reshape

Nigel Croxon <[email protected]>
Newsgroups gmane.linux.raid
Message-ID <[email protected]>
When a bio encounters LOC_INSIDE_RESHAPE during a reshape that is
interrupted (stopped or unable to progress), the code sets
bi->bi_status = BLK_STS_RESOURCE to signal the block layer for retry.
However, bio_endio() is never called, so the block layer never
receives the completion notification and the retry never happens.

This causes I/O to hang when a filesystem is layered over RAID5 and
reshape gets stuck.

Fix this by calling bio_endio(bi) before md_free_cloned_bio(bi) so
the block layer is properly notified of the BLK_STS_RESOURCE status
and can retry the request.

Tested stripes and stripe size conversions under load comparing
files multiple times during each conversion (i.e. MD reshape) on
ext4 after dropping caches degrading the RaidLV each time and
no data corruption.

Fixes: https://lwn.net/Articles/757123/

Signed-off-by: Nigel Croxon <[email protected]>
---
  drivers/md/raid5.c | 1 +
  1 file changed, 1 insertion(+)

diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c
index 6e79829c5acb..9a3475429ef4 100644
--- a/drivers/md/raid5.c
+++ b/drivers/md/raid5.c
@@ -6217,6 +6217,7 @@ static bool raid5_make_request(struct mddev 
*mddev, struct bio * bi)

      mempool_free(ctx, conf->ctx_pool);
      if (res == STRIPE_WAIT_RESHAPE) {
+        bio_endio(bi);
          md_free_cloned_bio(bi);
          return false;
      }
-- 
2.47.3
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.