aacraid Abort timeout - Benign messages?

sean.upton-lttx/[email protected]
Newsgroups gmane.linux.drivers.aacraid.devel
Message-ID <AA7A72A46469D411B8B300508BE329500AE4D655@desi2>
I've done some searching, but seem to be unable to answer this (perhaps
dumb) question: under stress (i.e. doing bonnie++ benchmarks using
unbuffered i/o) are occasional 'Abort Time-out' messages expected and/or
benign using aacraid-supported adapters?  Should I be worried?

I'm using aacraid from stock 2.4.20 kernel (as packaged by Debian
maintainers in the kernel-image-2.4.20-686-smp package) on a dual Xeon
2.0GHz box (HT on) using an Adaptec 2200S and an Adaptec (nStor made) 312R
Ultra160 disk enclosure that contains 8 73GB Seagate ST373405LC 73GB disks.
The 2200S is connected to the enclosure using a single channel and a single
cable.  Under most circumstances, this seems to be able to push up to 50-70
MB/s combined i/o from this volume using bonnie++ with the -b (unbuffered)
setting.

I get a few 'Abort timeout' messages in /var/log/messages during bonnie++
random creation and deletion tests (up to a few million small files on a
ReiserFS filesystem on an LVM volume on a PV that sits on a
controller-defined RAID 10 of 8 disks in above-mentioned enclosure).  I've
also seen this with bonnie++ rewriting tests another fs (ext3) on the same
PV as above.

Apr  1 09:09:45 homer kernel: aacraid:ID(1:00:0) Abort Time-out. Resetting
bus.
Apr  1 09:09:45 homer kernel: aacraid:SCSI bus reset issued on channel 1
Apr  1 09:24:34 homer -- MARK --
Apr  1 09:44:34 homer -- MARK --
Apr  1 10:04:34 homer -- MARK --
Apr  1 10:24:34 homer -- MARK --
Apr  1 10:40:21 homer kernel: aacraid:ID(1:00:0) Abort Time-out. Resetting
bus.
Apr  1 10:40:22 homer kernel: aacraid:SCSI bus reset issued on channel 1

The only other think worth mentioning is that there is currently a failed
disk in the enclosure that is set to offline (this volume is now running off
a hot-spare, rebuild to the spare finished this weekend, when the failure
occurred; it is at a remote location, so we have not yet had a chance to
remove the failed disk); we plan to remove this disk later this week.  Could
this be causing problems that degrade the SCSI bus in the enclosure?

Any thoughts would be appreciated.

Sean

_______________________________________________
Linux-aacraid-devel mailing list
[email protected]
http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel
Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.