RE: RE: [Opendlm-devel] ODLM/OGFS Recovery
Stanley Wang <[email protected]> Sat, 01 May 2004 10:14:38 +0000
| Newsgroups | gmane.comp.file-systems.opengfs.devel |
|---|---|
| Message-ID | <1083406478.3444.37.camel@stanw> |
--=-wCb3xf36FK3ZMcCXvKgX
Content-Type: text/plain
Content-Transfer-Encoding: quoted-printable
On Fri, 2004-04-30 at 14:32, Zickus II, Don wrote:
[snip]
> From what I understand of the code, when the holder of the lock (via proc=
ess or group) disappears unexpectedly (or expectedly and forgots to release=
the locks), will _NOT_ cause openDLM to immediately release the locks. In=
fact the locks will stick around for up to 3 seconds (or 1 second with our=
forthcoming patch). The reason for this is that as soon as the holder die=
s, its pid is put on a queue. Later on an asynchronous thread (clm_master_=
loop() inside clm_main.c) will have its timer expire and check for work. I=
f it finds a pid then it will perform the dlm_purge(). =20
> To prove this you can write a quick little program that creates 100,000 r=
esources but doesn't unlock them. After the program is finished (it will p=
robably take over a minute), for the next 3 seconds the system will be fair=
ly responsive (ie ls is quick). Then all of a sudden the machine will beco=
me extremely sluggish as it purges all the resources. =20
> Of course this doesn't cover node death as it is very difficult to make t=
hose locks persistent without replicating its info on all the nodes (which =
kind of defeats the point of distributed). =20
Thanks for your comments!
You are totally right. For client(process) failure cases, a internal
purge request with PURGE_DEAD will be issues. And only orphan/persistent
locks can survive this purge. For node fauilure cases no locks
(including orphan/persistent locks) can survive the recovery processs.
But we can work around this issue by combining LVB and DLM_VALNOTVALID
etc.=20
I also prefer not to change current codes much if it can fullfill our
requirement. And it seems OpenDLM can now :) Thanks very much for these
days' good discussion on this topic!
BTW, did you notice there is a little pitfall in the mechanism of using
DLM_VALNOTVALID? That is:
After node failure event, if there is a block lock request (on a=20
persistent lock resource and blocked by a client on the died node) that
requests PW or EX mode lock, and when it is granted "DLM_VALNOTVALID"
will NOT be returned. (Is that true? I got my conclusion from
"valueblock()", if I miss sth, please correct me.)
Is it a problem for you? Thanks!
Best Regards,
Stan
--=20
Opinions expressed are those of the author and do not represent Intel
Corporation
=20
"gpg --recv-keys --keyserver wwwkeys.pgp.net E1390A7F"
{E1390A7F:3AD1 1B0C 2019 E183 0CFF 55E8 369A 8B75 E139 0A7F}
--=-wCb3xf36FK3ZMcCXvKgX
Content-Type: application/pgp-signature; name=signature.asc
Content-Description: This is a digitally signed message part
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.2.3 (GNU/Linux)
iD4DBQBAk3iONpqLdeE5Cn8RAjmrAKCYcauyob0vUvmKrWq1lah2FuSB2ACWMQVF
xHp6xOXE1Fmbx7Rv+Paz0A==
=M39J
-----END PGP SIGNATURE-----
--=-wCb3xf36FK3ZMcCXvKgX--
-------------------------------------------------------
This SF.Net email is sponsored by: Oracle 10g
Get certified on the hottest thing ever to hit the market... Oracle 10g.
Take an Oracle 10g class now, and we'll give you the exam FREE.
http://ads.osdn.com/?ad_id=3149&alloc_id=8166&op=click