RE: PERC3/Di reset on 2650 server running Red Hat 7.3

sean.upton-lttx/[email protected]
Newsgroups gmane.linux.drivers.aacraid.devel
Message-ID <AA7A72A46469D411B8B300508BE329500AE4D6FD@desi2>
For what it is worth, all 4 reports of these types of problems that I've
seen in the last month (myself included) on the aacraid list are on HT-aware
Xeon boxes and recent 2.4 kernels, which may just be a coincidence, but it's
worth noting that I have had these problems with HyperThreading enabled as
well; I have not yet had the chance to test this with logical processors
disabled in the setup.  

My problem can't be reproduced under artificial stress (bonnie++), but
running large amounts of MySQL INSERT queries that have data, logging, and
temp files stored on another RAID volume (same controller) seem to trigger
the problem (hundreds-of-thousands of INSERT queries in a 15 minute period
every hour - box dies after about 8-12 hours).  It should be mentioned that
I have swap on the volume reporting problems, and I haven't yet tried moving
swap to another volume (though I can to remove any/all write pressure from
the problem volume and try to see if I can duplicate, I think).

Later this week, I will try testing with HT off and swap migrated to another
volume.

Sean

-----Original Message-----
From: [email protected] [mailto:[email protected]]
Sent: Monday, April 28, 2003 5:22 PM
To: [email protected]
Cc: [email protected]
Subject: RE: PERC3/Di reset on 2650 server running Red Hat 7.3


Forwarding to aacraid list.
Steve

-----Original Message-----
From: Carl Litt [mailto:[email protected]]
Sent: Monday, April 28, 2003 4:10 PM
To: Linux PowerEdge
Subject: Re: PERC3/Di reset on 2650 server running Red Hat 7.3


[ Sorry for the quote style, I'm replying from the archives ]

I am getting something similar on reiserfs except I actually get a
kernel panic and the system freezes.  I can reproduce this overnight,
all I have to do is load them down with 4 setiathome processes and come
back the next day.  They do not exhibit any problems if I let them sit
idle, nor if I run bonnie++ for a few hours.  The OS is running from a
fresh install with nothing extra installed except the latest kernel and
glibc.

I have opened a Bugzilla report here:
https://bugzilla.redhat.com/bugzilla/show_bug.cgi?id=89844.  Please
refer there for snapshots of the kernel panic taken from the RAC.  If
anyone feels their problem is related to this, please add a comment.

ObSpecs: Three PE2650's (sequential Service Tags), dual 2.4 Xeon (BIOS
A10) with HyperThreading enabled, PERC 3/Di (firmware 2.7-1), 3x36G 15K
Seagate, Red Hat 7.3 base updated to kernel-2.4.18-27.7xsmp and
glibc-2.2.5-43, all filesystems reiserfs 3.6.25.  Each server was
updated and OS installed from the same flash & boot disks and kickstart
file.

Carl Litt
Network Administrator
Execulink Internet

--- eric.rietjens-//[email protected] wrote:

We have three Dell PowerEdge 2650 systems running RedHat 7.3 that are
used
exclusively as MySQL server.
During startup of our system, several gigabytes of data are retrieved
from
the SQL server.

Two of our servers sometimes fail with the following message:
      "aacraid: Host adapter reset request. SCSI hang ?"
After this message, we get loads of error messages from the EXT3
filesystem (it is not possible to access the filesystem anymore).

Before we get this message, the disc activity already seems to have
halted
for more than 10 seconds (the activity LEDs are not flashing any more).

The configuration of the servers that show the problem as follows:
- 2 x Pentium 2.4
- 2GB internal memory
- 5x70GB harddrive (RAID5)
- Additional Intel PRO/100S network adapter

The 2650 that does not show the problem does not have the additional
network adapter. The configuration of the RAID controller is the
same for all three 2650 servers. The problem occurs for both the
2.4.18-19.7.xsmp kernel and the 2.4.18-27.7.xsmp kernel.

The 2650 system were installed with the standard RedHat CDs (obtained
via
DELL), after which I installed the RedHat
updates (using the rpm -Fvh *.rpm) command.

I experimented with a RedHat installation (from the standard CDs) on
which
I only installed an updated kernel (I did this with kernel 2.4.18-4.smp
and 2.4.18-27.7.xsmp).
I did not manager to reproduce the probem on these configurations. As
soon
as I installed the remaining updates, the problem occurred again.
However, it takes some time (10 minutes .. 25 hours) and 'luck' to
reproduce the problem, so I am not sure whether the absence of the
probem
has anything to do with not installing the RedHat updates.

Any explanation or solution would be most welcome!



_______________________________________________
Linux-PowerEdge mailing list
[email protected]
http://lists.us.dell.com/mailman/listinfo/linux-poweredge
Please read the FAQ at http://lists.us.dell.com/faq or search the list
archives at http://lists.us.dell.com/htdig/

_______________________________________________
Linux-aacraid-devel mailing list
[email protected]
http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel
Please read the FAQ at http://lists.us.dell.com/faq or search the list
archives at http://lists.us.dell.com/htdig/

_______________________________________________
Linux-aacraid-devel mailing list
[email protected]
http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel
Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.