RE: PERC3/Di reset on 2650 server running Red Hat 7.3
| Newsgroups | gmane.linux.drivers.aacraid.devel |
|---|---|
| Message-ID | <F777C1664B68704A87FA70E524FD0A2D741602@ausx2kmpc108.aus.amer.dell.com> |
Forwarding to aacraid list. Steve -----Original Message----- From: Carl Litt [mailto:[email protected]] Sent: Monday, April 28, 2003 4:10 PM To: Linux PowerEdge Subject: Re: PERC3/Di reset on 2650 server running Red Hat 7.3 [ Sorry for the quote style, I'm replying from the archives ] I am getting something similar on reiserfs except I actually get a kernel panic and the system freezes. I can reproduce this overnight, all I have to do is load them down with 4 setiathome processes and come back the next day. They do not exhibit any problems if I let them sit idle, nor if I run bonnie++ for a few hours. The OS is running from a fresh install with nothing extra installed except the latest kernel and glibc. I have opened a Bugzilla report here: https://bugzilla.redhat.com/bugzilla/show_bug.cgi?id=89844. Please refer there for snapshots of the kernel panic taken from the RAC. If anyone feels their problem is related to this, please add a comment. ObSpecs: Three PE2650's (sequential Service Tags), dual 2.4 Xeon (BIOS A10) with HyperThreading enabled, PERC 3/Di (firmware 2.7-1), 3x36G 15K Seagate, Red Hat 7.3 base updated to kernel-2.4.18-27.7xsmp and glibc-2.2.5-43, all filesystems reiserfs 3.6.25. Each server was updated and OS installed from the same flash & boot disks and kickstart file. Carl Litt Network Administrator Execulink Internet --- eric.rietjens-//[email protected] wrote: We have three Dell PowerEdge 2650 systems running RedHat 7.3 that are used exclusively as MySQL server. During startup of our system, several gigabytes of data are retrieved from the SQL server. Two of our servers sometimes fail with the following message: "aacraid: Host adapter reset request. SCSI hang ?" After this message, we get loads of error messages from the EXT3 filesystem (it is not possible to access the filesystem anymore). Before we get this message, the disc activity already seems to have halted for more than 10 seconds (the activity LEDs are not flashing any more). The configuration of the servers that show the problem as follows: - 2 x Pentium 2.4 - 2GB internal memory - 5x70GB harddrive (RAID5) - Additional Intel PRO/100S network adapter The 2650 that does not show the problem does not have the additional network adapter. The configuration of the RAID controller is the same for all three 2650 servers. The problem occurs for both the 2.4.18-19.7.xsmp kernel and the 2.4.18-27.7.xsmp kernel. The 2650 system were installed with the standard RedHat CDs (obtained via DELL), after which I installed the RedHat updates (using the rpm -Fvh *.rpm) command. I experimented with a RedHat installation (from the standard CDs) on which I only installed an updated kernel (I did this with kernel 2.4.18-4.smp and 2.4.18-27.7.xsmp). I did not manager to reproduce the probem on these configurations. As soon as I installed the remaining updates, the problem occurred again. However, it takes some time (10 minutes .. 25 hours) and 'luck' to reproduce the problem, so I am not sure whether the absence of the probem has anything to do with not installing the RedHat updates. Any explanation or solution would be most welcome! _______________________________________________ Linux-PowerEdge mailing list [email protected] http://lists.us.dell.com/mailman/listinfo/linux-poweredge Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/ _______________________________________________ Linux-aacraid-devel mailing list [email protected] http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/