SCSI Timeout: errors resulting in system crash
"j.b. zimmerman" <[email protected]>
| Newsgroups | gmane.linux.drivers.aacraid.devel |
|---|---|
| Organization | Ximian |
| Message-ID | <[email protected]> |
greetings all! We have a Dell 2650 which we've been testing prior to putting it into production as a mail server. We're running Debian 3.0, with a 2.4.18 kernel, and have a 2-drive mirror for the system drive and a 3-drive RAID5 set for the mail directories. After only a few days in live mode, we lost the system; it crashed, requiring a powercycle. Upon rebooting (which went fine) we found entries in the logs looking like this: ---cut--- Feb 25 01:19:00 skeptopotamus kernel: aacraid:ID(1:03:0) Timeout detected on cmd[0x28] Feb 25 01:19:00 skeptopotamus kernel: aacraid:SCSI Channel[1]: Timeout Detected On 2 Command(s) Feb 25 01:19:13 skeptopotamus kernel: scsi : aborting command due to timeout : pid 50212683, scsi0, channel 0, id 1, lun 0 Read (10) 00 00 48 01 8f 00 00 80 00 Feb 25 01:19:13 skeptopotamus kernel: scsi : aborting command due to timeout : pid 50212684, scsi0, channel 0, id 1, lun 0 Read (10) 00 00 48 02 0f 00 00 80 00 Feb 25 01:19:15 skeptopotamus kernel: aacraid:SCSI Channel[1]: Timeout Detected On 4 Command(s) Feb 25 01:19:25 skeptopotamus kernel: aacraid:ID(1:03:0) Timeout detected on cmd[0x28] Feb 25 01:19:25 skeptopotamus kernel: aacraid:SCSI Channel[1]: Timeout Detected On 7 Command(s) ---cut--- ...after getting the machine back up, I ran afacli and got the following: ---cut--- Executing: diagnostic show history /old=TRUE *** HISTORY BUFFER FROM LAST RUN *** [00]: <...repeats 1 more times> [01]: SCSI Channel[1]: Timeout Detected On 4 Command(s) [02]: ID(1:03:0) Timeout detected on cmd[0x28] [03]: ID(1:03:0): Timeout detected on cmd[0x28] [04]: ID(1:03:0): Timeout detected on cmd[0x2a] [05]: <...repeats 1 more times> [06]: ID(1:03:0): Timeout detected on cmd[0x28] [07]: <...repeats 1 more times> [08]: ID(1:03:0): Timeout detected on cmd[0x2a] [09]: SCSI Channel[1]: Timeout Detected On 7 Command(s) [10]: ID(1:03:0): Timeout detected on cmd[0x28] [11]: ID(1:03:0): Timeout detected on cmd[0x2a] [12]: ID(1:03:0): Timeout detected on cmd[0x28] [13]: <...repeats 1 more times> [14]: ID(1:03:0): Timeout detected on cmd[0x2a] [15]: ID(1:03:0): Timeout detected on cmd[0x28] [16]: ID(1:03:0): Timeout detected on cmd[0x2a] [17]: ID(1:03:0): Timeout detected on cmd[0x28] [18]: <...repeats 5 more times> [19]: SCSI Channel[1]: Timeout Detected On 13 Command(s) [20]: ID(1:03:0): Timeout detected on cmd[0x2a] [21]: ID(1:03:0): Timeout detected on cmd[0x28] [22]: <...repeats 2 more times> [23]: ID(1:03:0): Timeout detected on cmd[0x2a] [24]: ID(1:03:0): Timeout detected on cmd[0x28] [25]: <...repeats 1 more times> [26]: ID(1:03:0): Timeout detected on cmd[0x2a] [27]: ID(1:03:0): Timeout detected on cmd[0x28] [28]: <...repeats 4 more times> [29]: SCSI Channel[1]: Timeout Detected On 13 Command(s) [30]: ID(1:03:0): Timeout detected on cmd[0x28] [31]: <...repeats 3 more times> [32]: ID(1:03:0): Timeout detected on cmd[0x2a] [33]: ID(1:03:0): Timeout detected on cmd[0x28] [34]: <...repeats 2 more times> [35]: ID(1:03:0): Timeout detected on cmd[0x2a] [36]: ID(1:03:0): Timeout detected on cmd[0x28] [37]: <...repeats 3 more times> [38]: ID(1:03:0): Timeout detected on cmd[0x2a] [39]: SCSI Channel[1]: Timeout Detected On 14 Command(s) [40]: ID(1:03:0): Timeout detected on cmd[0x28] [41]: ID(1:03:0): Timeout detected on cmd[0x2a] [42]: ID(1:03:0): Timeout detected on cmd[0x28] [43]: <...repeats 5 more times> [44]: ID(1:03:0): Timeout detected on cmd[0x2a] [45]: ID(1:03:0): Timeout detected on cmd[0x28] [46]: ID(1:03:0): Timeout detected on cmd[0x2a] [47]: ID(1:03:0): Timeout detected on cmd[0x28] [48]: <...repeats 2 more times> [49]: SCSI Channel[1]: Timeout Detected On 14 Command(s) [50]: ID(1:03:0) Timeout detected on cmd[0x28] [51]: ID(1:03:0): Timeout detected on cmd[0x28] [52]: <...repeats 5 more times> [53]: ID(1:03:0): Timeout detected on cmd[0x2a] [54]: ID(1:03:0): Timeout detected on cmd[0x28] [55]: ID(1:03:0): Timeout detected on cmd[0x2a] [56]: ID(1:03:0): Timeout detected on cmd[0x28] [57]: ID(1:03:0): Timeout detected on cmd[0x2a] [58]: <...repeats 1 more times> [59]: ID(1:03:0): Timeout detected on cmd[0x28] [60]: <...repeats 1 more times> [61]: SCSI Channel[1]: Timeout Detected On 15 Command(s) [62]: ID(1:03:0): Timeout detected on cmd[0x2a] [63]: ID(1:03:0): Timeout detected on cmd[0x28] [64]: <...repeats 1 more times> [65]: ID(1:03:0): Timeout detected on cmd[0x2a] [66]: ID(1:03:0): Timeout detected on cmd[0x28] [67]: <...repeats 2 more times> [68]: ID(1:03:0): Timeout detected on cmd[0x2a] [69]: ID(1:03:0): Timeout detected on cmd[0x28] [70]: <...repeats 2 more times> [71]: ID(1:03:0): Timeout detected on cmd[0x2a] [72]: ID(1:03:0): Timeout detected on cmd[0x28] [73]: <...repeats 2 more times> [74]: SCSI Channel[1]: Timeout Detected On 15 Command(s) [75]: ID(1:03:0): Timeout detected on cmd[0x2a] [76]: ID(1:03:0): Timeout detected on cmd[0x28] [77]: <...repeats 1 more times> [78]: ID(1:03:0): Timeout detected on cmd[0x2a] [79]: ID(1:03:0): Timeout detected on cmd[0x28] [80]: ID(1:03:0): Timeout detected on cmd[0x2a] [81]: ID(1:03:0): Timeout detected on cmd[0x28] [82]: <...repeats 2 more times> [83]: ID(1:03:0): Timeout detected on cmd[0x2a] [84]: ID(1:03:0): Timeout detected on cmd[0x28] [85]: <...repeats 4 more times> [86]: SCSI Channel[1]: Timeout Detected On 15 Command(s) [87]: ID(1:03:0): Timeout detected on cmd[0x28] [88]: <...repeats 1 more times> [89]: ID(1:03:0): Timeout detected on cmd[0x2a] [90]: ID(1:03:0): Timeout detected on cmd[0x28] [91]: <...repeats 4 more times> [92]: ID(1:03:0): Timeout detected on cmd[0x2a] [93]: ID(1:03:0): Timeout detected on cmd[0x28] [94]: ID(1:03:0): Timeout detected on cmd[0x2a] [95]: <...repeats 2 more times> [96]: ID(1:03:0): Timeout detected on cmd[0x28] [97]: <...repeats 1 more times> [98]: SCSI Channel[1]: Timeout Detected On 15 Command(s) [99]: ======================== History Output Complete. ---cut--- My question: Is this a drive error? If so, why would it bring down the system, given that both RAID sets are redundant? If not, what are my options? Thanks very much to all. -J.B. Zimmerman -- ------------------------------------- J.B. Zimmerman [email protected] Network Administrator Ximian, Inc. - http://www.ximian.com
signature.asc
(application/pgp-signature, 232 B)
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.0.6 (GNU/Linux) Comment: For info see http://www.gnupg.org iD8DBQA+W5HL+Tg/H/6mPr0RAndfAJ9EbNyoHoH5YY5Lkpc7MdKbAtaHHgCffZip XpuMr1EJCoZ+meUXeLyyoYs= =jJ4r -----END PGP SIGNATURE-----