System Panic Makes My Life Easier
Joseph Brenner <[email protected]>
| Newsgroups | gmane.org.user-groups.linux.svlug |
|---|---|
| Message-ID | <CAFfgvXUNSNKxKPJ5YbtUfYb57xjyjU8KwYpg7Lit-+J39AVRfA@mail.gmail.com> |
I've been puzzling over a sick linux box for a little while lately. It's a dual-Opteron box (over ten years old now? wow...) that I've been upgrading off-and-on (bigger disks, a new video card...), but after a recent round of software upgrades it had been acting incredibly flaky, with uptimes of only a few days. It would totally lock-up and require a hard reboot... couldn't even ssh into it. I was trying to get an idea of what software change could've caused this problem-- the list of possibilities was long-- but of late the problem has gotten far worse, and it throws system panics and won't boot at all, so it's almost certainly a hardware problem. The cpu fan has been in bad shape for some time... it's got some cooling now, but I can easily imagine it's lifespan was shortened by overheating in the past. But, I've been wondering about how one would narrow down one's suspicions about flaky software, and I thought I would ask how one would go about it, even though it's just academic now. Look through system logs? Play with dtrace? I've been running ubuntu for years, but I've been wondering about some odd decisions coming from them of late, so rather than go with the latest ubuntu I've been switching back to Debian. Support for reiserfs seems like a second-class citizen these days, so I've been trying btrfs. So the suspect list included: btrfs systemd Debian amd64 binary packages (the same for intel and amd? Um...) And we might throw in firefox, which I've got running pretty often and loves to torture it's users with automatic upgrades. In the old days, my first guess actually would've been that linux itself is rock solid, but I'm afraid linux has seemed increasingly flaky of late. I've seen this hard lock-up symptom on a number of thinkpads, particularly when running a media player like vlc or totem. (By the way: there's a Fred Moyer talk at SF Perl that's more in the sysadmin/devops direction: "Better Service Monitoring Through Histograms" https://www.meetup.com/San-Francisco-Perl-Mongers/events/232569513/ RSVP before 3pm if you're interested. This one is located in the Union Square area of SF.)