RE: Problems with server stability

"Stewart M. Ives" <[email protected]>
Newsgroups gmane.linux.redhat.release.enigma
Message-ID <[email protected]>
See questions in-line (a top poster I'm learning not to be):

[email protected] wrote:
> Hi,
>
> I'm new to this list, and I'm hoping you may be able to help me
out.
>
> We have 3 generic Intel servers running RH7.2, and 2 of them
have
> major stability problems. They are as follows:
> 1. firewall - no problems.
> This is a small (1Gb memory) machine with 2 software RAID 1'd
IDE
> drives, performing firewall and DNS services.

Just out of curiosity, how do you have these IDE drives hooked
up??  Do you have a single drive on a single IDE channel or do
you have 2 drives on a single IDE channel.  I don't think this
has any bearing on your problem but I have always read that in a
software RAID configuration w IDE drives, we only should put a
single IDE drive on a single IDE channel.

> 2. Application routing server - constant problems
> Medium-sized (2Gb memory) machine with 2 software RAID 1'd IDE
> drives, running Jetty HTTP server serving Java servlets.
> 3. Database server - intermittent problems
> Big server (3Gb memory) + Disk Array: 2 onboard and 6 external
SCSI
> disks, all software RAID 1'd, running Jetty HTTP server and
Oracle
> 8.1.7.

On #2 & #3 are you running the BIG KERNEL or the basic kernel??

>
> In addition to the above mentioned software, all machines also
run
> tripwire intrusion detection software as well.
>
> On both machines that have problems, the system either freezes
or
> crashes just after 4AM, making me think that it may have
something to
> do with the cron.daily jobs. I found some pages on the Web from
> people who had problems with machines crashing when makewhatis
runs,
> but supposedly the problem was fixed in 7.1. On the database
server I
> would sometimes get messages saying that the system did not
have
> enough memory to execute a fork() shortly before the system
froze (it
> didn't halt; the console merely stopped responding). On the
> Application server, the system would simply power itself down
(badly;
> as though someone had simply turned it off at the wall) shortly
after
> 4AM every 3rd or 4th day.
>
> I am completely mystified by the problem as there are no
diagnostics
> or log messages to indicate that any problem is occurring; the
system
> simply freezes or shuts down. It does not happen *every* day,
but
> consistently once or twice a week and most often on Sunday
mornings,
> when the weekly jobs also run (but *before* they do). After I
set the
> database server to reboot itself once a week the problem occurs
far
> less frequently, but it still happens from time to time.
>
> Has anyone else run into this problem? I understand that we are
> probably unusual in running software RAID, but it is essential
for us
> to have servers which can be quickly recovered in the event of
a disk
> crash and hardware RAID cards are expensive (and a single point
of
> failure).
>
> Alternatively, does anyone know of any monitoring tools which
might
> be of help in uncovering this maddeningly un-reproducable
mystery?
>
> Hope you can help me because I'm rapidly losing what hair I
have
> left! :-)
>
> Thanks in advance
>
> Winston Gutkowski
>
>
>
> _______________________________________________
> enigma-list mailing list
> [email protected]
> https://listman.redhat.com/mailman/listinfo/enigma-list

stew

---
Outgoing SofTEC USA mail is certified Virus Free.
Checked by AVG anti-virus system (http://www.grisoft.com).
Version: 6.0.465 / Virus Database: 263 - Release Date: 3/25/2003
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.