Re: SAS preventive disk replacement

"L.A. Walsh" <[email protected]> Sat, 26 Nov 2016 13:44:14 -0800
Newsgroups gmane.linux.utilities.smartmontools
Message-ID <[email protected]>
Gandalf Corvotempesta wrote:
>
> I'm still here asking the following question because month ago nobody 
> replied
>
> In a SAS disks,  which values should i look for when prevenyively 
> replacing a disk?
>
----
    Do SAS disks have the same or similar fields for 'smartmon'
as ATA/SATA disks?  FWIW -- nobody may have responded because,
like me, no one really knew an authoritative answer.  That said,
I'll just spout off "unauthoritatively"... (YMMV)...

> Should i look for "elements in grown defect list"?
> Should i look for the uncorrected errors in the below table reporting 
> writes/reads/verifies?
>
> Should i look for something else in the "-x" output?
>
----
    Dang... that's one thing about smartmon, is that for better
or worse, it makes the "call" based on its recorded data.  If SAS
doesn't have similar, someone would have to know how the various
parameters collected affect failure rate.

    I think I read a report by google that said the single biggest
correlating factor in failed disks was temperature -- though I don't
know if it was 'max temperature' or 'daily-max-averaged' or what...

> Can someone explain this to me?
> Docs on smartmon page is not detailed about sas
>
---
    smartmon was originally invented for consumer ATA disks which
eventually became SATA disks.  I don't know that the regulating
committees for SAS disks did the same or adopted the same numbers
and failure guidelines as for SATA disks.

    That's all from 10-15+ YO-memory as I don't do alot of smartmon
monitoring as my disks are behind a RAID controller that will kick
the disk out as "bad" if it gets out-of-sync with the rest of the disks
due to sector remappping.  I.e. pretty much the 1st time a sector
gets remapped and causes a slowdown -- if it exceeds some threshold
in the LSI controller, it will just mark it as bad. 

    Before that -- any time I noticed, or "heard" a disk (before SSD's)
doing a "retry" -- I scheduled it for replacement.

    Had an interesting data point when I accidently got a
load of consumer-grade disks instead of enterprise and decided to
try them anyway.  Out of 24 disks, only 3 were marked good.  It wasn't
because of remapping, but the disks' RPMs: they varied by 10-15% from
slowest to fastest.

    When I got the replacements, they were all, right on 7200RPM.

    Good luck in finding your answers.  Did you google?  ;-)

-l




------------------------------------------------------------------------------