RE: PE 1650 with Perc 3/Di

sean.upton-lttx/[email protected] Fri, 05 Mar 2004 12:18:02 -0800
Newsgroups gmane.linux.drivers.aacraid.devel
Message-ID <AA7A72A46469D411B8B300508BE329500AE4E3CB@desi2>
I believe that this is an as-of-yet unresolved issue that has been known to
be this way for close a year now.  I don't know if your initial problem and
your latest problem are related.  It appears that, from what I've seen on my
own boxes (using Adaptec 2200S PCI controllers), and from what I have read
on this list, that there is some edge condition that when the CPU load is
high enough along with high write pressure, things start to go nuts between
the driver and the controller: we are able to cause this with commits of big
database transactions, but only intermittently.

My (now skeptical) take on this subject is that you should only use internal
hardware raid controllers for your bootdisks, if at all, and do software
RAID or DAS for everything else.

Sean

> -----Original Message-----
> From: Matt Soccio [mailto:[email protected]] 
> Sent: Friday, March 05, 2004 11:12 AM
> To: [email protected]
> Subject: PE 1650 with Perc 3/Di
> 
> I have a Poweredge 1650 with the Perc 3/Di raid controller, 
> and have not been able to keep it up an running for more than 
> a week or so at a time.  The mainboard was recently replaced 
> on a pre-emptive basis for an unrelated problem, but the 
> problem existed before the replacement as well.  After 
> replacement, I updated all bios and esm firmware files that 
> were available on dell's website.
> 
> I was running a 2.4.22 kernel with the stock aacraid driver 
> and it would crash about every 8-12 days with errors like this:
> 
> Feb 29 18:43:55 phoenix kernel: aacraid:        NMI ISR:
> NMI_SECONDARY_ATU_ERROR
> Feb 29 18:47:13 phoenix kernel: scsi: device set offline - 
> command error recover failed: host 0 channel 0 id 0 lun 0 Feb 
> 29 18:47:13 phoenix kernel: SCSI disk error : host 0 channel 
> 0 id 0 lun 0 return code = 6000000 Feb 29 18:47:13 phoenix 
> kernel:  I/O error: dev 08:04, sector 15466552 Feb 29 
> 18:47:13 phoenix kernel:  I/O error: dev 08:02, sector 3211272
> 
> Based on info on this list, I moved to a 2.4.25 kernel but kept the
> 2.4.22 aacraid driver.  It crashed the same way.
> 
> I have since got the latest (1.1.5) version of the driver 
> from Adaptec's site and it crashed in less than 48 hours with 
> different errors, but I could not grab them as it locked up 
> harder than usual.
> 
> The errors flooding the console were about scsi hang? and 
> scsi bus not responding.
> 
> The really strange part of this is that I can't reproduce the 
> problem, it just happens randomly.  If I run bonnie++ for a 
> couple hours with a few gigs of I/O it chugs right along, but 
> then it'll be sitting there doing nothing the next day and 
> just stop responding.  I just can't find the trigger for the failure.
> 
> Is there any way to run linux with this hardware?  Is there 
> some way to bypass this hardware and run a software raid, or 
> would I have to buy a new scsi controller?  With the disk 
> backplane the way it is in the server, would I even be able 
> to put a new controller in there?
> 
> Any suggestions or pointers are appreciated.
> 
> Matt
> 
> _______________________________________________
> Linux-aacraid-devel mailing list
> [email protected]
> http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel
> Please read the FAQ at http://lists.us.dell.com/faq or search 
> the list archives at http://lists.us.dell.com/htdig/
> 

_______________________________________________
Linux-aacraid-devel mailing list
[email protected]
http://lists.us.dell.com/mailman/listinfo/linux-aacraid-devel
Please read the FAQ at http://lists.us.dell.com/faq or search the list archives at http://lists.us.dell.com/htdig/