Re: SMART statistics for 40, 000 disk drives, from Backblaze

Bruce Allen <[email protected]>
Newsgroups gmane.linux.utilities.smartmontools
Message-ID <[email protected]>
Hi John,

<SNIP>

> I think it'd be useful to automate the recovery of initial drive
> failures in a replicated environment.  For example, if you have RAID,
> online backups, or dual copies of items, then when a sector or group
> of sectors fails, they can be overwritten with the correct data, and
> thus remapped by the drive.  This *should* be an automatic process,
> particularly at scale when you have many thousands of drives, but at
> the moment it isn't.

<SNIP>

> Does Linux software RAID, or ZFS, already do this "write over a bad
> sector" recovery?

I have had correspondence with Neil Brown over the years concerning this feature of the mds (software raid) driver.  Nine years ago, this was near the top of Neil's "to-do list":

http://neil.brown.name/blog/20050727141521

(see item "Don't kick drives on read errors"). My understanding is that was in fact implemented and is now part of the standard mds driver.  But I'm not sure, perhaps someone else here can confirm or deny that.

Cheers,
	Bruce



------------------------------------------------------------------------------
Download BIRT iHub F-Type - The Free Enterprise-Grade BIRT Server
from Actuate! Instantly Supercharge Your Business Reports and Dashboards
with Interactivity, Sharing, Native Excel Exports, App Integration & more
Get technology previously reserved for billion-dollar corporations, FREE
http://pubads.g.doubleclick.net/gampad/clk?id=157005751&iu=/4140/ostg.clktrk
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.