[ ssic-linux-Bugs-992652 ] Nodes don't always know that the cluster has kicked them out
"SourceForge.net" <[email protected]>
| Newsgroups | gmane.linux.cluster.ssic.devel |
|---|---|
| Message-ID | <[email protected]> |
Bugs item #992652, was opened at 2004-07-16 20:30 Message generated for change (Settings changed) made by rogertsang You can respond by visiting: https://sourceforge.net/tracker/?func=detail&atid=405834&aid=992652&group_id=32541 Please note that this message will contain a full copy of the comment thread, including the initial issue submission, for this request, not just the latest update. Category: Miscellaneous Group: None >Status: Closed Resolution: Works For Me Priority: 5 Private: No Submitted By: David Zafman (dzafman) Assigned to: John Byrne (jlbyrne) Summary: Nodes don't always know that the cluster has kicked them out Initial Comment: I have a 4 node failover cluster. I dropped node 1 into the debugger to perform a failover to node 2. Node 4 hit a breakpoint which I didn't get to in time. I continued node 4 and it output "nm_nodedown_daemon: Node 1 went down" and just sat there. The cluster ended up consisting of nodes 2 and 3, but node 4 never realized that it had been declared down. In a production environment if a node is declared down by the cluster it is in, if it is still alive it should reboot itself. This way it might be able to rejoin the cluster. According to John this is a problem in the node monitoring algorithm. ---------------------------------------------------------------------- Comment By: Roger Tsang (rogertsang) Date: 2007-10-11 22:45 Message: Logged In: YES user_id=1246761 Originator: NO In SSI-1.9.3 nodes panic (or can reboot if you like) when they cannot talk to or cannot find any CLMS master after the configurable CLMS timeout. ---------------------------------------------------------------------- Comment By: Roger Tsang (rogertsang) Date: 2005-08-26 18:07 Message: Logged In: YES user_id=1246761 Can someone verify this on SSI-1.2.2? Thanks. ---------------------------------------------------------------------- Comment By: David Zafman (dzafman) Date: 2004-07-23 19:41 Message: Logged In: YES user_id=297844 I hit the following scenario on my 4 node failover cluster. I took node 1 down, and node 2 crashed trying to take over. In my case the cfs_setroot didn't get done for some reason on node 2. The remaining cluster just waited doing nothing. I don't know what would happen when node 1 reboots. It will start looking for a CLMS master since node 3 & 4 are clueless. ---------------------------------------------------------------------- You can respond by visiting: https://sourceforge.net/tracker/?func=detail&atid=405834&aid=992652&group_id=32541 ------------------------------------------------------------------------- This SF.net email is sponsored by: Microsoft Defy all challenges. Microsoft(R) Visual Studio 2005. http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/