System Panic Makes My Life Easier

Joseph Brenner <[email protected]>
Newsgroups gmane.org.user-groups.linux.svlug
Message-ID <CAFfgvXUNSNKxKPJ5YbtUfYb57xjyjU8KwYpg7Lit-+J39AVRfA@mail.gmail.com>
I've been puzzling over a sick linux box for a little while lately.

It's a dual-Opteron box (over ten years old now? wow...) that I've been
upgrading off-and-on (bigger disks, a new video card...),
but after a recent round of software upgrades it had been
acting incredibly flaky, with uptimes of only a few days.  It would
totally lock-up and require a hard reboot... couldn't even ssh into it.

I was trying to get an idea of what software change could've caused
this problem-- the list of possibilities was long-- but of late the
problem has gotten far worse, and it throws system panics and won't
boot at all, so it's almost certainly a hardware problem. The cpu fan
has been in bad shape for some time... it's got some cooling now, but
I can easily imagine it's lifespan was shortened by overheating in the
past.

But, I've been wondering about how one would narrow down one's
suspicions about flaky software, and I thought I would ask how one
would go about it, even though it's just academic now.  Look through
system logs?  Play with dtrace?

I've been running ubuntu for years, but I've been wondering about some
odd decisions coming from them of late, so rather than go with the
latest ubuntu I've been switching back to Debian.  Support for
reiserfs seems like a second-class citizen these days, so I've been
trying btrfs.  So the suspect list included:

  btrfs
  systemd
  Debian amd64 binary packages (the same for intel and amd? Um...)

And we might throw in firefox, which I've got running pretty often
and loves to torture it's users with automatic upgrades.

In the old days, my first guess actually would've been that linux
itself is rock solid, but I'm afraid linux has seemed increasingly
flaky of late.  I've seen this hard lock-up symptom on a number of
thinkpads, particularly when running a media player like vlc or totem.


(By the way: there's a Fred Moyer talk at SF Perl
that's more in the sysadmin/devops direction:
"Better Service Monitoring Through Histograms"
https://www.meetup.com/San-Francisco-Perl-Mongers/events/232569513/
RSVP before 3pm if you're interested. This one is located
in the Union Square area of SF.)
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.