aacraid Abort timeout - Benign messages?
Leon Toh <[email protected]>
| Newsgroups | gmane.linux.drivers.aacraid.devel |
|---|---|
| Message-ID | <[email protected]> |
Sean, Not knowing the configuration of your array setup and SCSI devices it's difficult to ascertain why you receive this *Abort* command. >aacraid:ID(1:00:0) Abort Time-out. Resetting bus. The above message tells me that you happen to have a SCSI device on bus 1 with target ID of 0 and LUN 0 attach to a controller which happen to use *aacraid* as it's driver. This SCSI device receive a command request to abort whatever it's currently processing. As per SCSI specification, after an *Abort* command received by the device concern it has to be followed up by a complete *Reset* bus command. As it happens *Abort* command happens to be a commonly used command. Whenever you stop a process prematurely it trigger's *Abort* command to complete current I/O process. The other culprit that also trigger *Abort* command is benchmark programs as it happens to be the best method to clean out any outstanding I/O commands before servicing the next test. If you are testing the performance of your RAID array I don't believe bonnie is the right benchmark to used. I understand that it benchmark physical drive I/O's rather than the overall array I/O's instead. Benchmark which I use is Intel IOMeter and this benchmark the actual array I/O's instead. By the way it's a waste of time to perform benchmark test on an array which is in degraded mode or even on a array which is currently being initialised. The results which you going to obtain won't be the best results. Furthermore if the array happens to be running in degraded mode and has vital data and you than perform a benchmark test at the same time where at the worst case this could lead to lost of another drive within the degraded array and this would than most likely cause a complete lost of data which I'm sure you won't be happy with. I have seen far too many of these happen and often caused by physical hardware's / configuration issue especially when someone cutting corner's or after a cheap RAID solution. By the way Adaptec has a message board, http://mbserver.adaptec.com which you can utilised to post your questions. Best Regards, Leon >Date: Tue, 01 Apr 2003 11:48:38 -0800 >From: sean.upton-lttx/[email protected] >Subject: aacraid Abort timeout - Benign messages? >To: [email protected] > >I've done some searching, but seem to be unable to answer this (perhaps >dumb) question: under stress (i.e. doing bonnie++ benchmarks using >unbuffered i/o) are occasional 'Abort Time-out' messages expected and/or >benign using aacraid-supported adapters? Should I be worried? > >I'm using aacraid from stock 2.4.20 kernel (as packaged by Debian >maintainers in the kernel-image-2.4.20-686-smp package) on a dual Xeon >2.0GHz box (HT on) using an Adaptec 2200S and an Adaptec (nStor made) 312R >Ultra160 disk enclosure that contains 8 73GB Seagate ST373405LC 73GB disks. >The 2200S is connected to the enclosure using a single channel and a single >cable. Under most circumstances, this seems to be able to push up to 50-70 >MB/s combined i/o from this volume using bonnie++ with the -b (unbuffered) >setting. > >I get a few 'Abort timeout' messages in /var/log/messages during bonnie++ >random creation and deletion tests (up to a few million small files on a >ReiserFS filesystem on an LVM volume on a PV that sits on a >controller-defined RAID 10 of 8 disks in above-mentioned enclosure). I've >also seen this with bonnie++ rewriting tests another fs (ext3) on the same >PV as above. > >Apr 1 09:09:45 homer kernel: aacraid:ID(1:00:0) Abort Time-out. Resetting >bus. >Apr 1 09:09:45 homer kernel: aacraid:SCSI bus reset issued on channel 1 >Apr 1 09:24:34 homer -- MARK -- >Apr 1 09:44:34 homer -- MARK -- >Apr 1 10:04:34 homer -- MARK -- >Apr 1 10:24:34 homer -- MARK -- >Apr 1 10:40:21 homer kernel: aacraid:ID(1:00:0) Abort Time-out. Resetting >bus. >Apr 1 10:40:22 homer kernel: aacraid:SCSI bus reset issued on channel 1 > >The only other think worth mentioning is that there is currently a failed >disk in the enclosure that is set to offline (this volume is now running off >a hot-spare, rebuild to the spare finished this weekend, when the failure >occurred; it is at a remote location, so we have not yet had a chance to >remove the failed disk); we plan to remove this disk later this week. Could >this be causing problems that degrade the SCSI bus in the enclosure? > >Any thoughts would be appreciated. > >Sean _______________________________________________ Linux-aacraid-devel mailing list [email protected] http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/