SCSI problems with a 2650
Russell Stuart <[email protected]> 22 Jul 2003 16:21:43 +1000
| Newsgroups | gmane.linux.drivers.aacraid.devel,gmane.linux.hardware.dell.poweredge |
|---|---|
| Organization | Lube Mobile |
| Message-ID | <1058854902.23853.243.camel@ras> |
Hi all,
This is a long post because it represents 6 months worth of effort in
trying to solve the problem myself. I have now decided its a hardware
problem. What I am after is confirmation, or failing that suggestions
on what to do next. Preferably this should come from someone with
"@dell.com" or "@redhat.com" in their email address for reasons that
will be obvious later.
The basics:
Dell 2650, Dual 1.8Ghz Xeon, 2Gb RAM
OS: RedHat, currently kernel-2.4.20-18.7smp
Drives:
/dev/sda: Segate ST336752LC x 2, Raid 1
/dev/sdb: Segate ST336706LC
Firmware:
Embedded Server Management Firmware: 2.10
Backplane Firmware: 0.25
Perc 3/Di Raid controller: 2.7-1 [Build 3170]
Embedded Remote Access Controller: 2.10
Supplier: Dell for all bar the ST336706LC.
This box is a hot spare. There is an identical box (except for possibly
older firmware) which is the main server. The main server has no
problems. I have an standard install CD (which I call an SOE) I have
created which installs and configures itself. It is just a tar image of
a master machine, plus some software I wrote copies the tar image onto
the drive and configures it. The implication of this is that I am
certain the software is identical on all machines.
It all started when I was upgrading the software on the machine. I
bumped the IDE CDROM+Floppy eject button with my knee, and the machine
when down in a heap. I don't know whether that is relevant or not.
After fixing that the upgrade went smoothly. To this day I have not
upgraded the main machine. I can't until I have something that runs
smoothly on the hot spare.
The next weekend the machine crashed. The kernel oops said something
about to nested interrupts in ext3. I am now certain it crashed during
an automated backup that happens on Sunday night. The backup:
1. mkfs's /dev/sdb1 with ext3.
2. Copies (using pax) / (on /dev/sda1) to /dev/sdb1.
3. Runs grub.
After this I am left with a bootable /dev/sdb, which is identical to
what is on /dev/sda.
The next weekend it crashed again. Similar oops message on console. I
changed the backup schedule to every 1/2 hour. To my surprise it did
not crash again until the next weekend. So I put the ethernet under
substantial load. It crashed the next weekend. I put it under a large
CPU load with memory pressure so ensure all of RAM was being exercised.
It crashed the next weekend. To this day I have no idea what was so
special about the weekend.
However by this time I had seen other messages in the kernel oops. The
SCSI driver often appeared in there. Since I had an identical machine
using the same kernel and doing the same backups at the same time on
Sunday, I rang Dell and reported a hardware problem. The Dell support
person asked me to send him the kernel oops. I explained it only
appeared on the monitor in the early morning on Sunday - there was no
disk log. But he insisted. So I waited until it crashed the next
weekend and wrote it all down and emailed it. Then he wanted the logs.
Then he told me he was an NT guy and had no idea what he was looking at,
but nonetheless he thought is was a software problem.
A couple of weeks of haggling passed during which I mirrored the working
server onto non-working one, and verified the problem still happened.
Eventually he relented and agreed to send a Dell engineer out to replace
bits. But he did so only on the condition that if the problem
re-occurred I would would not call Dell again. Instead I would report
it to the operating system vendor, ie Red Hat, and get them to fix the
problem. I guess it was a reasonable thing to insist on if you believed
it was a software problem. Note: this is why a response from someone at
Dell or RedHat would be nice.
The Dell engineer came replaced most the boards in the machine, upgraded
all the firmware and left. A couple of days later it the machine
crashed. Same problem, but now it no longer happens only weekends.
This trend - the machine staying up for shorter and shorter periods of
time, has continued and so that now it rarely survives one day. This
time there were messages about /dev/sdb in the oops. The drives were
one of the few things not replaced by the Dell engineer. This was one
non-Dell thing in the machine, so I sent it in for warranty repair. A
new drive came back 3 months later. During those 3 months the machine
did not crash once. The backups were not being done, of course, as the
backup drive was not present.
After installing the new drive the machine crashed a day or so later.
However by the time I was beginning to see numerous messages on this
list about problems with the aacraid and bcm5700 drivers. So for the
first time I considered the possibility that the Dell support person was
right after all. Perhaps it was a software problem. The next month or
so (until now) has been spent running different kernels in different
configurations, trying to isolate the problem. Here are the tests and
results. Unless states the machine is unloaded, except for running my
backup script every 1/2 an hour from cron.
Test: Booted off RedHat 8.0 Install CD, kernel-2.4.18-3BOOT, running:
badblocks -n -c $((128*1024*1024)) /dev/sd?
on both /dev/sda and /dev/sdb simultaneously and continuously.
Result: Ran perfectly for 2 weeks. No disk errors reported. Nothing
in /var/log/messages. Several hundred passes completed.
Test: Booted off RedHat 8.0 Install CD, kernel-2.4.18-3BOOT, chrooted
to normal root on /dev/sda1, running my backup script
repeatedly.
Result: Failed after a long time (over a week) - machine frooze without
an oops. SCSI timeout errors would appear in /var/log/messages
every 48 hours or so. The timeouts happen on both sda and sdb.
Test: kernel-2.4.18-3 (mirror of working machine). Tested with
SMP+SMT on, SMP+SMT off, no SMP.
Result: Failed after 3-4 days with ext3 error messages about I/O
errors. SCSI timeouts appear in /var/log/messages every
3-4 hours. The timeouts happen on both sda and sdb.
Test: 2.4.20-18.7. Tqested with SMP+SMT on, SMP+SMT off, no SMP.
Result: Machine usually dead by next morning. Nothing unusual in
/var/log/messages, except for one day which had this:
Jul 8 06:30:30 mephisto kernel: kjournald starting. \
Commit interval 5 seconds
Jul 8 06:30:30 mephisto kernel: EXT3 FS 2.4-0.9.19, \
19 August 2002 on sd(8,17),
Jul 8 06:30:30 mephisto kernel: EXT3-fs: mounted filesystem with \
ordered data mode.
Jul 8 06:34:09 mephisto kernel: aacraid: Host adapter reset request. \
SCSI hang ?
Jul 8 06:34:19 mephisto kernel: aacraid: Host adapter reset request. \
SCSI hang ?
Jul 8 06:35:39 mephisto last message repeated 8 times
Jul 8 06:35:49 mephisto kernel: scsi: device set offline - command \
error recover failed: host 0 channel 0 id 1 lun 0
Jul 8 06:35:49 mephisto kernel: SCSI disk error : host 0 channel 0 \
id 1 lun 0 return code = 6000000
Jul 8 06:35:49 mephisto kernel: I/O error: dev 08:11, sector 22903784
Jul 8 06:35:49 mephisto kernel: I/O error: dev 08:11, sector 22903792
(similar lines to the last one repeated many, many times)
Here is an example of the SCSI timeout messages that appear in
/var/log/messages under old kernel versions:
Jul 22 15:30:21 mephisto kernel: kjournald starting. \
Commit interval 5 seconds
Jul 22 15:30:21 mephisto kernel: EXT3 FS 2.4-0.9.17, 10 Jan 2002 \
on sd(8,17), internal journal
Jul 22 15:30:21 mephisto kernel: EXT3-fs: mounted filesystem \
with ordered data mode.
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31256, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d1 37 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31257, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d1 b7 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31258, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d2 37 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31259, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d2 b7 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31260, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d3 37 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31261, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d3 b7 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31262, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d4 37 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31263, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d4 b7 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31264, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d5 37 00 00 80 00
Jul 22 15:31:48 mephisto kernel: scsi : aborting command due to timeout
: pid 31265, scsi0, channel 0, id 1, lun 0 Write (10) 00
00 00 d5 b7 00 00 80 00
Jul 22 15:31:49 mephisto kernel: scsi : aborting command due to timeout
: pid 31266, scsi0, channel 0, id 0, lun 0 Write (10) 00
00 6c 36 df 00 00 80 00
Jul 22 15:31:49 mephisto kernel: scsi : aborting command due to timeout
: pid 31267, scsi0, channel 0, id 0, lun 0 Write (10) 00
00 6c 37 5f 00 00 80 00
Jul 22 15:31:49 mephisto kernel: scsi : aborting command due to timeout
: pid 31268, scsi0, channel 0, id 0, lun 0 Write (10) 00
00 6c 37 df 00 00 38 00
Jul 22 15:31:55 mephisto kernel: scsi : aborting command due to timeout
: pid 31269, scsi0, channel 0, id 0, lun 0 Write (10) 00
00 e0 40 77 00 00 80 00
Jul 22 15:31:56 mephisto kernel: scsi : aborting command due to timeout
: pid 31270, scsi0, channel 0, id 0, lun 0 Write (10) 00
00 e0 40 f7 00 00 38 00
(no more error messages - machine continues to work normally)
Currently I think there could only be two possibilities:
1. There is a problem with the aacraid driver until all versions of
Linux I have tried.
2. There is a hardware problem somewhere in the SCSI hardware, and
its unlikely to be the non-Dell drive because it has been
replaced.
I have read numerous assurances from people on these lists that the
aaraid driver is OK, so I am discounting that. That leaves a hardware
problem. Is this a reasonable conclusion?
--
Regards,
Russell Stuart
IT, Lube Mobile
http://www.lubemobile.com.au
Ph: +61 (7) 3344 1311
_______________________________________________
Linux-aacraid-devel mailing list
[email protected]
http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel
Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/