Spurious failovers

John Horne <[email protected]>
Newsgroups gmane.linux.keepalived.devel
Organization Plymouth University
Message-ID <[email protected]>
Hello,

We have been experiencing a problem with two servers running keepalived,
so I sent a message to the CentOS users list to see if anyone could
help. (I wasn't sure if it was a general CentOS problem or not.) The
message, with more details, can be found at
http://lists.centos.org/pipermail/centos/2014-November/147933.html

We had another failover last night, at a slightly different time
(03:42). What we saw (by monitoring the VRRP traffic with tcpdump) was a
'hang' on the master server for around 20 seconds. During that time, the
secondary server took the VIP but when the interface became active again
the original master server took the VIP back again. So, in effect, we
had two failovers (although I gather that the second failover may have
been due to keepalived comparing the IP addresses and selecting the
higher one as master given that the priorities are equal.)

I don't think this has anything to do with Apache or Tomcat - there is
absolutely nothing logged to indicate a problem with them. Likewise I
don't think it is any nighttime process such as log rotation, updates or
backups (none of those were running at the hang time).

Some output of the VRRP traffic:

===================
03:42:11.430827 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement,
vrid 51, prio 100, authtype simple, intvl 3s, length 20
03:42:14.431230 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement,
vrid 51, prio 100, authtype simple, intvl 3s, length 20
03:42:17.431975 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement,
vrid 51, prio 100, authtype simple, intvl 3s, length 20
03:42:37.315743 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement,
vrid 51, prio 100, authtype simple, intvl 3s, length 20
03:42:37.316554 IP 141.163.66.143 > 224.0.0.18: VRRPv2, Advertisement,
vrid 51, prio 100, authtype simple, intvl 3s, length 20
03:42:40.316780 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement,
vrid 51, prio 100, authtype simple, intvl 3s, length 20
===================

The advertisement interval is 3 seconds, but note the 20 second gap at
03:42:17. Initially 66.144 has the VIP, but during the 20 second gap the
other server (66.143) takes it. When the interface comes back up both
servers are advertising, so keepalived selects the master based on the
IP address (since the prio is equal for both servers).

What I cannot work out is whether the interface has hung because of
keepalived or something else on the system. There is nothing logged in
any log file (that I can find) to indicate a problem with the network
interface. Has anyone else seen anything like this?



Thanks,

John.

-- 
John Horne                   Tel: +44 (0)1752 587287
Plymouth University, UK


------------------------------------------------------------------------------
Comprehensive Server Monitoring with Site24x7.
Monitor 10 servers for $9/Month.
Get alerted through email, SMS, voice calls or mobile push notifications.
Take corrective actions from your mobile device.
http://pubads.g.doubleclick.net/gampad/clk?id=154624111&iu=/4140/ostg.clktrk
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.