[ ssic-linux-Bugs-992652 ] Nodes don't always know that the cluster has kicked them out

"SourceForge.net" <[email protected]>
Newsgroups gmane.linux.cluster.ssic.devel
Message-ID <[email protected]>
Bugs item #992652, was opened at 2004-07-16 20:30
Message generated for change (Settings changed) made by rogertsang
You can respond by visiting: 
https://sourceforge.net/tracker/?func=detail&atid=405834&aid=992652&group_id=32541

Please note that this message will contain a full copy of the comment thread,
including the initial issue submission, for this request,
not just the latest update.
Category: Miscellaneous
Group: None
>Status: Closed
Resolution: Works For Me
Priority: 5
Private: No
Submitted By: David Zafman (dzafman)
Assigned to: John Byrne (jlbyrne)
Summary: Nodes don't always know that the cluster has kicked them out

Initial Comment:
I have a 4 node failover cluster.  I dropped node 1 into the 
debugger to perform a failover to node 2.  Node 4 hit a breakpoint 
which I didn't get to in time.  I continued node 4 and it output 
"nm_nodedown_daemon: Node 1 went down" and just sat there.  
The cluster ended up consisting of nodes 2 and 3, but node 4 never 
realized that it had been declared down.

In a production environment if a node is declared down by the 
cluster it is in, if it is still alive it should reboot itself.  This way it 
might be able to rejoin the cluster.

According to John this is a problem in the node monitoring 
algorithm.

----------------------------------------------------------------------

Comment By: Roger Tsang (rogertsang)
Date: 2007-10-11 22:45

Message:
Logged In: YES 
user_id=1246761
Originator: NO

In SSI-1.9.3 nodes panic (or can reboot if you like) when they cannot talk
to or cannot find any CLMS master after the configurable CLMS timeout.

----------------------------------------------------------------------

Comment By: Roger Tsang (rogertsang)
Date: 2005-08-26 18:07

Message:
Logged In: YES 
user_id=1246761

Can someone verify this on SSI-1.2.2?  Thanks.

----------------------------------------------------------------------

Comment By: David Zafman (dzafman)
Date: 2004-07-23 19:41

Message:
Logged In: YES 
user_id=297844

I hit the following scenario on my 4 node failover cluster.  I took node 1

down, and node 2 crashed trying to take over.  In my case the 
cfs_setroot didn't get done for some reason on node 2.  The remaining 
cluster just waited doing nothing.  I don't know what would happen when 
node 1 reboots.  It will start looking for a CLMS master since node 3 & 4

are clueless.

----------------------------------------------------------------------

You can respond by visiting: 
https://sourceforge.net/tracker/?func=detail&atid=405834&aid=992652&group_id=32541

-------------------------------------------------------------------------
This SF.net email is sponsored by: Microsoft
Defy all challenges. Microsoft(R) Visual Studio 2005.
http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.