RE: RE: [Opendlm-devel] ODLM/OGFS Recovery

Daniel McNeil <[email protected]> 27 Apr 2004 17:08:38 -0700
Newsgroups gmane.comp.file-systems.opengfs.devel
Message-ID <[email protected]>
On Tue, 2004-04-27 at 08:51, Stanley Wang wrote:
> On 26 Apr 2004, Daniel McNeil wrote:
> 
> > This DLM/recovery discussion has been a very interesting.  I just
> > thought I would add my 2 cents worth from previous experience with
> > a cluster file system from several years ago.  This is from my
> > memory, so this might not be exact.
> 
> Thanks for your help :) You do help us a lot (deadman lock are introduced to 
> OpenGFS by your suggestion :)
>  
> > We had to work out these same issues.  We also used a DLM that provided
> > persistent locks (obviously not open DLM).  The way the DLM worked was
> > that after a crash (node crash), persistent locks for a particular
> > lock domain would continue to return DUBIOUS even after DLM recovery
> > until the application made a call to re-validate all persistent locks
> > for the domain.
> 
> Yes, we met the same problem with one difference: orphan lock in OpenDLM is NOT 
> persistent in node failure case. That is what made us headache. We are now 
> trying to using LVB or some other way to work around this pitfall. Any 
> suggestion?
> 
> Best Regards,
> Stan


I did a little googling on DLM persistent locks and found this man
page.  It looks like it is for trucluster.  It describes persistent
locks and how recovery was done.  This one is similar to what
I remember (which makes since since persistent locks were implemented
for a large database software company and many vendors implemented
what they asked for :) )

http://h30097.www3.hp.com/docs/cluster_doc/cluster_16/MAN/MAN3/0024____.HTM

I am a bit surprised that openDLM's persistent locks do not provide
the same kind of functionality.  The whole point was to give an
application a chance to do a recovery before granting more locks
without having to keep the locks open.

Are there any openDLM folks who can explain how these ORPHAN can be used
in a distributed application if it cannot handle node death?  I thought
node death was one of the big reasons for clusters and DLMs...

Depending on the order of lock recovery seems to be asking for trouble.

Daniel



-------------------------------------------------------
This SF.Net email is sponsored by: Oracle 10g
Get certified on the hottest thing ever to hit the market... Oracle 10g. 
Take an Oracle 10g class now, and we'll give you the exam FREE.
http://ads.osdn.com/?ad_id=3149&alloc_id=8166&op=click