Re: Random system crashes
Robert Freiberger <[email protected]>
| Newsgroups | gmane.org.user-groups.linux.svlug |
|---|---|
| Message-ID | <CAK5y5r74CBB5F1cutihiuspQ5YeXBjSemzFNzh98VKYUfJ3myg@mail.gmail.com> |
I had this similar issue before and it was related to an SSD controller locking up the drive (those horrible OCZ drives). The weird part is the system would sometimes work but the background processes would hang, other times it would simply lock up. Since the drive was locked, no logs were written and upon reboot, the system appeared fine. If you are running into issues that fail to mention anything on the system logs, I would look into the core components and run them doing a classic A/B swap. Say boot the system from USB live distro and see if it crashes in the same manner, then swap out all of the memory except for one stick, etc. On Thu, Apr 26, 2018 at 10:46 AM Michael Eager <[email protected]> wrote: > On 04/26/2018 10:39 AM, Dan Ritter wrote: > > On Thu, Apr 26, 2018 at 10:29:41AM -0700, Michael Eager wrote: > >> On 04/26/2018 10:03 AM, Dan Ritter wrote: > >>> On Thu, Apr 26, 2018 at 09:42:42AM -0700, Michael Eager wrote: > >>>> I have a server which periodically freezes and becomes unresponsive. > >>>> The system is AMD FX-8320, 32Gb, CentOS 7, kernel-4.15. > >>>> > >>>> There is nothing in the system log which indicates any problem. > >>>> The freezes appear to be random and happen when the system is > >>>> heavily loaded and when it is idle. Even when heavily loaded, > >>>> system temperature seems reasonable. > >>> > >>> Does it come back without intervention after the freeze, or does > >>> it die and need a reboot or power-cycle? > >> > >> Never recovers. I have to reset or power-cycle. > > > > If it only happened under heavy load, I would guess a power > > supply problem. > > I thought the same. I may have replaced the PSU, not sure I recall. > > > Sadly, I think you have a serious hardware fault in the motherboard or > > CPU. Sorry. > > I may put the system on the workbench and throw a heat gun > at it to see if something starts to fail consitently. > > > Diagnosis: let it run for N hours with a CPU exerciser like > > stress-ng, set to thrash the CPU but not memory or disk. > > I'll give that a try. > > > > > Semi-good news: you can replace both of them for around $250 at > > NewEgg. > > No problem that throwing more money at won't fix. :-( > > -- > Michael Eager [email protected] > 1960 Park Blvd., Palo Alto, CA 94306 > > _______________________________________________ > svlug mailing list > [email protected] > http://lists.svlug.org/lists/listinfo/svlug > -- Robert Freiberger 510-936-1210 _______________________________________________ svlug mailing list [email protected] http://lists.svlug.org/lists/listinfo/svlug