SLES9 SP4 (x86_64), ora10.2 and aio == > disaster. ATTENTION!

"Alexei_Roudnev" <[email protected]>
Newsgroups gmane.linux.suse.oracle.general,gmane.linux.suse.sles-e
Message-ID <[email protected]>
ATTENTION! Don't upgrade to SLES9 SP4 if you have an Oracle 10.2.0.3 (may 
be, others as well), use async io, without a very (VERY!) careful testing 
(including FAL, if you use it, and rman, if you use it).

We made upgrade on one of our staging systems.
Results:

- After the upgrade, async_io doesn't work anymore in some applications. It 
works in core oracle, but it doesn't work with archive logs, for example:

   * rman, --> for example, backup validate archivelog all -> can't read 
archive log(s).
   * log writer - can't ship archive logs;
   * FAL can't ship archive logs.

Experiments shows that:
   - If filesystemio_options = SETALL, both components don't work.
   - If filesystemio_options = ASYNCH, only RMAN doesn't work (Log writer 
can ship logs).
   - if I boot old kernel (286 instead of 308) everything works well.
   - oracle relink(ing) doesn't help. Kernel 286 works well with the new 
libaio shared library.

I did not made careful testing on VMWare guest (which allows me to rollback 
system in a few minutes and try different scenarios), BUT it is really a 
disaster for many, many oracle shops (and it can be a very serious bug).

It is serious because you can upgrade and Oracle works well. rman works well 
if it have a very few files to transfer (so it works while you have a light 
system load). StandBy can work if you have not gaps.

But if system fail and you need more archive logs OR standby needs to take 
some file by FAL or DataGuard tries to reinstate an instance - you end up 
with frozen dataguard,  database, backups and standby's. We did oracle patch 
few hours before this change (patch is rolled back now), so it took us a few 
days to understand what really went wrong.

I can open it as a bug with Novell, but I hit it on the staging system so we 
are not in the urgent need to fix it. Anyway, it is a VERY SERIOUS - even if 
something got wrong on our side, I have now Oracle @ Linux which works well 
on kernel 286 and refuse to work properly (rman + log shipping) on the 
kernel 308.

PS. System - SLES9 SP4, x86_64 (DELL 2850), Oracle 10.2.0.3, we tried 
interim patch 6035495 but rolled it back now (it don't make much sense - 
system should not behave differently on the kernel(s) 286 and 308)

Any idea, what is broken?? Any idea, how to report this bug? (System is 
staging, not on direct support, but we have a partnernet agreement so we 
have support access).

Alexei Roudnev



-- 
To unsubscribe, email: [email protected]
For additional commands, email: [email protected]
Please see http://www.suse.com/oracle/ before posting
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.