Re: HD dying?

"Robin H. Johnson" <[email protected]> Fri, 7 Apr 2017 20:55:01 +0000
Newsgroups gmane.linux.utilities.smartmontools
Message-ID <[email protected]>
On Fri, Apr 07, 2017 at 04:28:43PM -0400, David Niklas wrote:
...
> > If you can repeat it, consider some of the following to get a better
> > insight as to what's going on.
> > - set up serial kernel console or network kernel console logging.
> > - set up kdump or similar.
> No, It's random so far.
Ok, get yourself network console logging, since networking was still
working, and you can just let the kernel send a copy of all klog entries
over the network.

See in the kernel sources, see Documentation/networking/netconsole.txt
or examples in the Ubuntu & Arch wikis.

> > That's not to say that the drive isn't the source of the problem, just
> > that it's not likely based on the output you've shown.
> Why not?
> What else causes all writes to the drive to stop except a problem with
> the drive or MB (my laptop has not cabling)?
Most failure modes of a spinning drive would cause various error
counters to be incremented. The few that I could think of that wouldn't
involve specific component failures on the drive PCB.

Drive PCB-originating failures should NOT cause your video to lock up,
but may stop the logging to disk of any errors.

I can start up a linux system, running off a sata drive, open a
terminal, suddenly disconnect the drive, and still be able to run dmesg
and/or see live kernel log entries (Provided that dmesg itself is at
least already cached and running doesn't need anything to be read off
disk).

So what we're looking for as root cause is some manner of error that
causes both video & drive to become unresponsive, but the kernel to
still respond to ICMP ping (ergo network stack is operational).

That root cause COULD have other effects (like a power spike that then
damages the drive PCB), but it's the root cause we care about.

Overheating causing a component fault (like causing a capacitor to go
out of tolerance or fail) on one of the PCI/PCIe busses, and therein
affecting the drive & graphics. The networking might be on a different
bus, and continues to function.

> > You say this is a laptop, and the drive by power hours has racked up
> > ~1.5 years of usage, so it possibly hasn't been opened in at least that
> > long. How much dust has built up inside it? Overheating of the graphics
> > CAN cause the symptoms you've described.
> The laptop is my primary way to get online, it's not be left off for more
> than 2 days unless it's HW failed (the original drive died).
> 
> So, I'm not misreading the S.M.A.R.T. data? No values that aught to be
> interpreted in HEX, OCTAL or something?
No, the drive data seems good, and representative of a health &
well-used drive. No reallocated sectors, no other issues, not that many
power cycles even for a laptop drive w/ aggressive power saving.

-- 
Robin Hugh Johnson
Gentoo Linux: Dev, Infra Lead, Foundation Trustee & Treasurer
E-Mail   : [email protected]
GnuPG FP : 11ACBA4F 4778E3F6 E4EDF38E B27B944E 34884E85
GnuPG FP : 7D0B3CEB E9B85B1F 825BCECF EE05E6F6 A48F6136

------------------------------------------------------------------------------
Check out the vibrant tech community on one of the world's most
engaging tech sites, Slashdot.org! http://sdm.link/slashdot