>>> WAITED TOO LONG FOR A ROW CACHE ENQUEUE LOCK!

"Hahn, Klaus" <[email protected]>
Newsgroups gmane.linux.suse.oracle.general
Message-ID <9CA76282E1E3B44A96EE19AC3B199E960111356E@mlhw163a.ww002.siemens.net>
Hi list,

I'm testing RAC + ASM + iSCSI-shared-storage on SUSE SLES10 SP1.

Configuration of RAC:
1. 2 nodes (a + b) with 2 instances TEST1 and TEST2

Configuration of ASM:
1. normal redundancy
2. Diskgroup data  with 2 failure groups st01_data and st02_data.
    Each failure group has 1 disk (/dev/raw/raw11 and  /dev/raw/raw12)
    raw11 is an iSCSI-Target from node st01.
    raw12 is an iSCSI-Target from node st02.
3. Diskgroup reco (recovery area)  with 2 failure groups st01_reco and
st02_reco.
    Each failure group has 1 disk (/dev/raw/raw6 and  /dev/raw/raw7)
    raw6 is an iSCSI-Target from node st01.
    raw7 is an iSCSI-Target from node st02.

Taking offline of the iSCSI-Target node st01 produce a strange error
within the database:
- no commit possible
- alter system switch logfile not possible
- database seems to be readonly

srvctl says: All RAC-Services and Instances are running.


 
alert.log of RAC-Instance TEST1:

>>>>>..
ORA-27091: I/O kann nicht in Queue gestellt werden
ORA-27072: Datei-I/O-Fehler
Linux-x86_64 Error: 5: Input/output error
Additional information: 4
Additional information: 192544
Additional information: -1
Tue Oct 30 11:57:10 2007
Errors in file /opt/oracle/admin/TEST/udump/test1_ora_30342.trc:
ORA-27091: I/O kann nicht in Queue gestellt werden
ORA-27072: Datei-I/O-Fehler
Linux-x86_64 Error: 5: Input/output error
Additional information: 4
Additional information: 192544
Additional information: -1
WARNING: offlining disk 2.2526723344 (RECO_0002) with mask 0x3
Tue Oct 30 11:57:19 2007
Errors in file /opt/oracle/admin/TEST/bdump/test1_ckpt_15122.trc:
ORA-27091: I/O kann nicht in Queue gestellt werden
ORA-27072: Datei-I/O-Fehler
Linux-x86_64 Error: 5: Input/output error
Additional information: 4
Additional information: 192608
Additional information: -1
WARNING: offlining disk 2.2526723344 (RECO_0002) with mask 0x3
Tue Oct 30 12:34:16 2007
>>> WAITED TOO LONG FOR A ROW CACHE ENQUEUE LOCK! pid=16
System State dumped to trace file
/opt/oracle/admin/TEST/bdump/test1_mmon_15153.trc
Tue Oct 30 12:47:03 2007
>>> WAITED TOO LONG FOR A ROW CACHE ENQUEUE LOCK! pid=37
System State dumped to trace file
/opt/oracle/admin/TEST/udump/test1_ora_23960.trc
Tue Oct 30 12:50:53 2007
>>> WAITED TOO LONG FOR A ROW CACHE ENQUEUE LOCK! pid=54
System State dumped to trace file
/opt/oracle/admin/TEST/udump/test1_ora_14206.trc

<<<<<<

alert.log of ASM-Instance: +ASM1

<<<<<<<<<<<<
ORA-27091: unable to queue I/O
ORA-27072: File I/O error
Linux-x86_64 Error: 5: Input/output error
Additional information: 4
Additional information: 2056
Additional information: -1
Tue Oct 30 11:57:19 2007
Errors in file /opt/oracle/admin/+ASM/bdump/+asm1_gmon_28227.trc:
ORA-27091: unable to queue I/O
ORA-27072: File I/O error
Linux-x86_64 Error: 5: Input/output error
Additional information: 4
Additional information: 2048
Additional information: -1
Tue Oct 30 11:57:19 2007
Errors in file /opt/oracle/admin/+ASM/bdump/+asm1_gmon_28227.trc:
ORA-27091: unable to queue I/O
ORA-27072: File I/O error
Linux-x86_64 Error: 5: Input/output error
Additional information: 4
Additional information: 2056
Additional information: -1
NOTE: cache closing disk 2 of grp 2: RECO_0002
Tue Oct 30 11:57:19 2007
NOTE: group RECO: relocated PST to: disk 0000 (PST copy 0)
Tue Oct 30 11:57:22 2007
NOTE: PST update: grp = 2, dsk = 2, mode = 0x4
Tue Oct 30 11:57:22 2007
NOTE: group RECO: relocated PST to: disk 0000 (PST copy 0)
NOTE: cache closing disk 2 of grp 2: RECO_0002
Tue Oct 30 11:57:52 2007
NOTE: PST refresh pending for group 1/0x175a4df9 (DATA)
SUCCESS: refreshed PST for 1/0x175a4df9 (DATA)
NOTE: PST refresh pending for group 2/0x506a4dfe (RECO)
SUCCESS: refreshed PST for 2/0x506a4dfe (RECO)
Tue Oct 30 11:59:24 2007
SUCCESS: refreshed PST for 1/0x175a4df9 (DATA)
SUCCESS: refreshed PST for 2/0x506a4dfe (RECO)
<<<<<<<<<<<

Reconfiguration of ASM  by " alter diskgroup add failgroup ..."  (after
deleting the asm-metadata on the raw device)
takes no effect.

After restart of the RAC-instances all work fine !

Any hint ?

Regards
Klaus



-- 
To unsubscribe, email: [email protected]
For additional commands, email: [email protected]
Please see http://www.suse.com/oracle/ before posting
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.