Re: Heartbeat not reclaiming IP properly

Matthew Newton <[email protected]>
Newsgroups gmane.linux.highavailability.ultramonkey
Message-ID <[email protected]>
Hi,

Thanks for the reply.

On Mon, Jun 26, 2006 at 05:02:04PM +0900, Horms wrote:
> In article <[email protected]> you wrote:
> > We now have, seemingly, lb2 handing control back to lb1, after
> > which lb2 stops. However, lb1 never brings the vIP back up again,
> > even though it seems to think that it is now the master. If you
> > then do:
> 
> What does the heartbeat log looklike on lb1 at this time.
> Does it claim that it is taking the IP back?

Here you get the following on lb1:

heartbeat: 2006/06/26_10:01:41 info: Received shutdown notice from 'lb2.test'.
heartbeat: 2006/06/26_10:01:41 info: Resources being acquired from lb2.test.
heartbeat: 2006/06/26_10:01:41 info: Running /etc/ha.d/rc.d/status status
heartbeat: 2006/06/26_10:01:42 info: /usr/lib/heartbeat/mach_down: nice_failback: foreign resources acquired
heartbeat: 2006/06/26_10:01:42 info: mach_down takeover complete for node lb2.test.
heartbeat: 2006/06/26_10:01:42 info: Exiting status process 5435 returned rc 0.
heartbeat: 2006/06/26_10:01:42 info: 1 local resources from [/usr/lib/heartbeat/ResourceManager listkeys lb1.test]
heartbeat: 2006/06/26_10:01:42 info: Local Resource acquisition completed.
heartbeat: 2006/06/26_10:01:42 info: Exiting req_our_resources(ask) process 5436 returned rc 0.
heartbeat: 2006/06/26_10:01:56 WARN: node network: is dead
heartbeat: 2006/06/26_10:01:56 info: Local status now set to: 'active'
heartbeat: 2006/06/26_10:01:56 info: Starting child client "/usr/lib/heartbeat/ipfail" (1002,104)
heartbeat: 2006/06/26_10:01:56 info: Starting "/usr/lib/heartbeat/ipfail" as uid 1002  gid 104 (pid 5487)
heartbeat: 2006/06/26_10:01:56 info: Running /etc/ha.d/rc.d/status status
heartbeat: 2006/06/26_10:01:56 info: Exiting status process 5486 returned rc 0.
heartbeat: 2006/06/26_10:02:14 WARN: node lb2.test: is dead
heartbeat: 2006/06/26_10:02:14 info: Dead node lb2.test gave up resources.
heartbeat: 2006/06/26_10:02:14 info: Link lb2.test:eth0 dead.

I've put all the log files and configs at the following URL (with
a couple of extra blank lines inserted in some logs, hopefully for
clarity):

  http://www.le.ac.uk/cc/mcn4/ultramonkey/

This particular extract is from:

  http://www.le.ac.uk/cc/mcn4/ultramonkey/with-ipfail_with-fallbackoff/lb2-to-lb1/ha-log.lb1

I've run two different configs, both with fallback off. One is
with ipfail, and the other without. lb1-to-lb2 means that:

  lb1 was started, and left to bring up the vIP
  lb2 was then started, and
  lb1 was taken down before lb2 came up properly (i.e. after about 5s)
  finally, lb2 was stopped

The log files were wiped before this, so the log is exactly this
sequence and nothing else. lb2-to-lb1 is the opposite, of course.

lb1-to-lb2 always works fine. lb2-to-lb1 has the problem, when
using ipfail.

> This sounds like a split-brain kind of problem. Especially
> as you only have one communication channel between lb1 and lb2 (right?).

[Yes right, although the network is definitely reliable at this
stage.]

> Are you using ipfail?  That would probably help.

When ipfail is taken out of the configuration, it now seems to
work fine. The log line

heartbeat: 2006/06/26_10:01:56 WARN: node network: is dead

seems relevant, but I don't know why the "network" node should be
dead. The ping group config line is:

  ping_group network 143.210.5.1 143.210.5.2 143.210.5.9 143.210.5.10

which are switches on the near network; they are all up and
(AFAICT) reliably pingable.

> Also, I would be interested to know if changing the value of
> auto_failback affects this behaviour.

Haven't tried this yet, as ipfail seems to be causing the problem;
will give it a go if required!

> > I can post configs and logs if needed.
> 
> Yes, please do.

I have also tried the heartbeat-2 package (sorry, not sure if I
mentioned; this is Debian sarge, hb2 from etch) using the same
configuration. Heartbeat-2 worked properly in this situation
(using ipfail), but would very easily bring _both_ machines up as
master. This is slightly less of a problem as at least service
continues that way, but it is still very worrying ;-).

Leaving heartbeat-2 out for now, have I got something wrong with
the ping_group command?

I can rig up two links between the boxes for testing if necessary.
If this goes live I definitely will aim to get a second
link in, but as they are in separate buildings it involves
ethernet-and-fibre, rather than serial.

Thanks!

Matthew


-- 
Matthew Newton <[email protected]>

UNIX and e-mail Systems Administrator, Network Support Section,
Computer Centre, University of Leicester,
Leicester LE1 7RH, United Kingdom


-- 
Ultra Monkey - http://www.ultramonkey.org/
To UNSUBSCRIBE, email to [email protected], with a body:
unsubscribe ultramonkey-users [email protected]
where "[email protected]" is YOUR email address.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.