RE: Problems with server stability
"Winston Gutkowski" <[email protected]>
| Newsgroups | gmane.linux.redhat.release.enigma |
|---|---|
| Message-ID | <[email protected]> |
Hi Chris, Sorry for delay in replying. Can up2date be run securely? Currently these servers are isolated from the outside world and only allow ssh access. If it is possible, I would prefer to run it on a separate machine and then download whatever patches it produces myself. Is this possible? Also, I wasn't quite sure what you mean by "don't forget the kernel". Is this something you have to specify separately? Thanks for your suggestion. Winston -----Original Message----- From: [email protected] [mailto:[email protected]]On Behalf Of Reynolds, Chris Sent: Friday, March 28, 2003 6:48 To: [email protected] Subject: RE: Problems with server stability Have you updated to the latest set of patches. Try running "up2date" and don't forget the kernel. -Chris -----Original Message----- From: Winston Gutkowski [mailto:[email protected]] Sent: Thursday, March 27, 2003 7:19 PM To: [email protected] Subject: Problems with server stability Hi, I'm new to this list, and I'm hoping you may be able to help me out. We have 3 generic Intel servers running RH7.2, and 2 of them have major stability problems. They are as follows: 1. firewall - no problems. This is a small (1Gb memory) machine with 2 software RAID 1'd IDE drives, performing firewall and DNS services. 2. Application routing server - constant problems Medium-sized (2Gb memory) machine with 2 software RAID 1'd IDE drives, running Jetty HTTP server serving Java servlets. 3. Database server - intermittent problems Big server (3Gb memory) + Disk Array: 2 onboard and 6 external SCSI disks, all software RAID 1'd, running Jetty HTTP server and Oracle 8.1.7. In addition to the above mentioned software, all machines also run tripwire intrusion detection software as well. On both machines that have problems, the system either freezes or crashes just after 4AM, making me think that it may have something to do with the cron.daily jobs. I found some pages on the Web from people who had problems with machines crashing when makewhatis runs, but supposedly the problem was fixed in 7.1. On the database server I would sometimes get messages saying that the system did not have enough memory to execute a fork() shortly before the system froze (it didn't halt; the console merely stopped responding). On the Application server, the system would simply power itself down (badly; as though someone had simply turned it off at the wall) shortly after 4AM every 3rd or 4th day. I am completely mystified by the problem as there are no diagnostics or log messages to indicate that any problem is occurring; the system simply freezes or shuts down. It does not happen *every* day, but consistently once or twice a week and most often on Sunday mornings, when the weekly jobs also run (but *before* they do). After I set the database server to reboot itself once a week the problem occurs far less frequently, but it still happens from time to time. Has anyone else run into this problem? I understand that we are probably unusual in running software RAID, but it is essential for us to have servers which can be quickly recovered in the event of a disk crash and hardware RAID cards are expensive (and a single point of failure). Alternatively, does anyone know of any monitoring tools which might be of help in uncovering this maddeningly un-reproducable mystery? Hope you can help me because I'm rapidly losing what hair I have left! :-) Thanks in advance Winston Gutkowski _______________________________________________ enigma-list mailing list [email protected] https://listman.redhat.com/mailman/listinfo/enigma-list _______________________________________________ enigma-list mailing list [email protected] https://listman.redhat.com/mailman/listinfo/enigma-list