RE: RHEL3 hangs during shutdown (apparently unmounting /dev/md0)

"Coe, Colin C. (Unix Engineer)" <Colin.Coe-YYMz0cdUNTk0n/[email protected]> Tue, 26 Feb 2008 07:51:01 +0900
Newsgroups gmane.linux.redhat.release.taroon.general
Message-ID <BAC1D28A5AB852439A6914CA7AB4E63F05192ABD@permls05.wde.woodside.com.au>
Nope, you're not the only one :)

Have you done an 'lsof' to see if anything is using or running from this
mount point?

CC

> -----Original Message-----
> From: [email protected]=20
> [mailto:[email protected]] On Behalf Of Michael O'Donnell
> Sent: Tuesday, 26 February 2008 4:11 AM
> To: Michael O'Donnell; [email protected]
> Subject: RHEL3 hangs during shutdown (apparently unmounting /dev/md0)
>=20
> Hello,
>=20
> I hope I'm not the only one on the planet still using RHEL3...
>=20
> If anybody's listening, I wonder if you might have any clues
> about the following:
>=20
>     Our RHEL3 update9 systems hang during shutdown.  Since very
>     few processes are still running at the time time I've
>     instrumented the /etc/init.d halt script (to spawn an
>     instance of /bin/bash on /dev/tty2) and poked around while
>     the shutdown script was hung.  ps shows the umount command
>     as basically the only one running and attempts to use strace
>     to monitor its activities just hang, so I assume the process
>     is in kernel space and getting nowhere.  During this time
>     the system goes unresponsive for a minute or more, briefly
>     shows signs of life and then going unresponsive again.
>     If left alone it seems to always recover.
>=20
>     Our RAID1 on /dev/md0 contains an ext3 filesystem mounted
>     on /mnt/MD0 and the unmount operation itself seems to have
>     succeeded at the time we notice the hang.  At least, ls
>     shows only the naked mountpoint directory there, not the
>     contents of the mounted filesystem.  It's as if after the
>     unmount has completed the kernel code gets stuck trying to
>     do some sort of cleanup operation.
>=20
>     using strace to run ps shows that every open of /proc/nnn
>     takes approx 1 second.
>=20
>     Attempts to use mdadm to --examine either of the RAID halves
>     or to --query /dev/md0 always hang when they get to the end
>     where they'd normally name the /dev/ entries that the RAID
>     halves live on.
>=20
> ...so this smells strongly of an MD bug...
>=20
> --=20
>=20
> Michael O'Donnell   WSI Corp.   978-983-6613
>=20
> --
> Taroon-list mailing list
> [email protected]
> https://www.redhat.com/mailman/listinfo/taroon-list
>=20

NOTICE: This email and any attachments are confidential.=20
They may contain legally privileged information or=20
copyright material. You must not read, copy, use or=20
disclose them without authorisation. If you are not an=20
intended recipient, please contact us at once by return=20
email and then delete both messages and all attachments.


--
Taroon-list mailing list
[email protected]
https://www.redhat.com/mailman/listinfo/taroon-list