Wondering why Veritas didn't hot swap

"Robert M. Martel - CSU" <[email protected]>
Newsgroups gmane.os.solaris.managers
Organization Maxine Goodman Levin College of Urban Affairs
Message-ID <[email protected]>
Greetings,

This is not a critical situation - the machine is question is 
functional, the filesystem affected is almost unused - and I have a 
known-good backup of it anyway...But I want to understand what is happening.

Sun Sparc running Solaris 9 with Veritas volume manager 3.5 and Veritas 
File System 3.5 connected to an A5200 array.  (Golden Oldies)

The machine panicked on Dec 24th reporting "WARNING: [AFT1] E$ Tag scrub 
event on CPU3 at TL=0, errID 0x00307ef4.203c63a6"  I do not know if the 
panic-restart has anything to do with what followed, but there were no 
unusual messages during reboot or after the system recovered.

A couple hours after rebooting  I see unrecoverable read errors on one 
of the disks in the A5200 array that is part of a RAID 5 volume - with 
hot-swap spare available.  After a few sets of SCSI warning messages I see:

Dec 24 21:01:40 wolf vxio: [ID 686135 kern.warning] WARNING: vxvm:vxio: 
object urbanmk11-01 detached from RAID-5 Staff at column 2 offset 80420928
Dec 24 21:01:40 wolf vxio: [ID 354480 kern.warning] WARNING: vxvm:vxio: 
RAID-5 Staff entering degraded mode operation
Dec 24 21:01:40 wolf vxio: [ID 680909 kern.warning] WARNING: vxvm:vxio: 
RAID-5 Staff was degraded with stale parity.  Volume is now unusable and 
will be disabled.
Dec 24 21:01:47 wolf vxfs: [ID 702911 kern.warning] WARNING: msgcnt 1 
vxfs: mesg 025: vx_wsuper - /dev/vx/dsk/urbanmk2/Staff file system 
super-block update failed
Dec 24 21:01:47 wolf vxfs: [ID 702911 kern.warning] WARNING: msgcnt 2 
vxfs: mesg 031: vx_disable - /dev/vx/dsk/urbanmk2/Staff file system disabled

In every instance of failed disks in the A5200 array I've dealt with in 
the last ten years Veritas has *always* swapped in the spare disk and 
the system kept chugging along.

What attempts I've made to coax the volume back to life have failed - I 
can gain access to the filesystem for a limited time by forcing the 
volume to start, but in the end it ends-up detached again.

So I am wondering why Veritas didn't hot-swap the disk as it always had 
in the past - opting instead to detach/disable the volume.  If it is 
taking issue with the "stale parity", how did it get that way in the 
first place?

As I started this messages with, the data is not critical, no one is 
hounding me to regain access, I have good backups, I just want to 
understand what happened and if recovery is possible short of destroying 
the volume and re-building it.

Thanks,
Bob


-- 
***********************************************************************
Robert M. Martel                    Pushing myself and this old machine
System Administrator                Burning fumes
Levin College of Urban Affairs      and what's left of my dreams
Cleveland State University
(216) 687-2214
[email protected]
***********************************************************************
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.