RE: Problems with server stability
"Stewart M. Ives" <[email protected]>
| Newsgroups | gmane.linux.redhat.release.enigma |
|---|---|
| Message-ID | <[email protected]> |
See questions in-line (a top poster I'm learning not to be): [email protected] wrote: > Hi, > > I'm new to this list, and I'm hoping you may be able to help me out. > > We have 3 generic Intel servers running RH7.2, and 2 of them have > major stability problems. They are as follows: > 1. firewall - no problems. > This is a small (1Gb memory) machine with 2 software RAID 1'd IDE > drives, performing firewall and DNS services. Just out of curiosity, how do you have these IDE drives hooked up?? Do you have a single drive on a single IDE channel or do you have 2 drives on a single IDE channel. I don't think this has any bearing on your problem but I have always read that in a software RAID configuration w IDE drives, we only should put a single IDE drive on a single IDE channel. > 2. Application routing server - constant problems > Medium-sized (2Gb memory) machine with 2 software RAID 1'd IDE > drives, running Jetty HTTP server serving Java servlets. > 3. Database server - intermittent problems > Big server (3Gb memory) + Disk Array: 2 onboard and 6 external SCSI > disks, all software RAID 1'd, running Jetty HTTP server and Oracle > 8.1.7. On #2 & #3 are you running the BIG KERNEL or the basic kernel?? > > In addition to the above mentioned software, all machines also run > tripwire intrusion detection software as well. > > On both machines that have problems, the system either freezes or > crashes just after 4AM, making me think that it may have something to > do with the cron.daily jobs. I found some pages on the Web from > people who had problems with machines crashing when makewhatis runs, > but supposedly the problem was fixed in 7.1. On the database server I > would sometimes get messages saying that the system did not have > enough memory to execute a fork() shortly before the system froze (it > didn't halt; the console merely stopped responding). On the > Application server, the system would simply power itself down (badly; > as though someone had simply turned it off at the wall) shortly after > 4AM every 3rd or 4th day. > > I am completely mystified by the problem as there are no diagnostics > or log messages to indicate that any problem is occurring; the system > simply freezes or shuts down. It does not happen *every* day, but > consistently once or twice a week and most often on Sunday mornings, > when the weekly jobs also run (but *before* they do). After I set the > database server to reboot itself once a week the problem occurs far > less frequently, but it still happens from time to time. > > Has anyone else run into this problem? I understand that we are > probably unusual in running software RAID, but it is essential for us > to have servers which can be quickly recovered in the event of a disk > crash and hardware RAID cards are expensive (and a single point of > failure). > > Alternatively, does anyone know of any monitoring tools which might > be of help in uncovering this maddeningly un-reproducable mystery? > > Hope you can help me because I'm rapidly losing what hair I have > left! :-) > > Thanks in advance > > Winston Gutkowski > > > > _______________________________________________ > enigma-list mailing list > [email protected] > https://listman.redhat.com/mailman/listinfo/enigma-list stew --- Outgoing SofTEC USA mail is certified Virus Free. Checked by AVG anti-virus system (http://www.grisoft.com). Version: 6.0.465 / Virus Database: 263 - Release Date: 3/25/2003