Re: Subject: RFC: Read repair for md RAID1 after mirror read failures

"G.W. Kant - Hunenet B.V." <[email protected]>
Newsgroups gmane.linux.raid
Message-ID <[email protected]>
On 7/15/26 7:27 AM, Abd-Alrhman Masalkhi wrote:
>> In other words, the first successfully recovered read request could
>> automatically become a repair opportunity. The repair could even be
>> scheduled asynchronously, so the successful read is returned immediately
>> while the rewrite is performed in the background. Unlike a periodic
>> resync, this repair would be driven by an actual read failure, making it
>> targeted rather than rewriting the entire mirror.
>>
> Yes, md has had this for a long time. Look at fix_read_error() in
> raid1.c. It is called from handle_read_error() on any failed read. It
> reads from a healthy mirror and rewrites the bad region on the failing
> device, giving the drive a chance to rewrite or remap the sector. If the
> rewrite fails, it records a bad block. md does this synchronously under
> a frozen array, so it is not a missing feature.
> 
> The likely reason you didn't see it is that your array was already
> degraded, so there was no healthy in-array copy for fix_read_error() to
> recover from. In your case, you were likely able to retrieve the data
> due to btrfs level redundancy, and md can't repair across arrays.
> 
>> With today's 18–24 TB HDDs and backup/archive workloads, where data may
>> remain unchanged for years, latent media degradation seems increasingly
>> relevant. A successful read from the alternate mirror may be one of the
>> last opportunities to refresh such a sector before it becomes
>> permanently unreadable.
>>
> And Check/Repair is the right defense for cold archival data on large
> drives.

Thank you. fix_read_error() was exactly the piece I was looking for. I 
had missed that md already performs read repair when another healthy 
mirror is available.

My situation was indeed different because the array had already lost one 
member, so the successful read originated from Btrfs redundancy across 
another md RAID1 rather than from the surviving md mirror.

One question remains, though. check/repair periodically reads all 
sectors, but as far as I understand it, successfully read sectors are 
not rewritten. From a media aging perspective this is a different 
problem than recovering from a read error. Have there ever been 
discussions about a true media refresh pass that rewrites successfully 
read sectors to refresh long-lived magnetic recordings?

Best regards,
Dion Kant
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.