Spurious failovers
John Horne <[email protected]>
| Newsgroups | gmane.linux.keepalived.devel |
|---|---|
| Organization | Plymouth University |
| Message-ID | <[email protected]> |
Hello, We have been experiencing a problem with two servers running keepalived, so I sent a message to the CentOS users list to see if anyone could help. (I wasn't sure if it was a general CentOS problem or not.) The message, with more details, can be found at http://lists.centos.org/pipermail/centos/2014-November/147933.html We had another failover last night, at a slightly different time (03:42). What we saw (by monitoring the VRRP traffic with tcpdump) was a 'hang' on the master server for around 20 seconds. During that time, the secondary server took the VIP but when the interface became active again the original master server took the VIP back again. So, in effect, we had two failovers (although I gather that the second failover may have been due to keepalived comparing the IP addresses and selecting the higher one as master given that the priorities are equal.) I don't think this has anything to do with Apache or Tomcat - there is absolutely nothing logged to indicate a problem with them. Likewise I don't think it is any nighttime process such as log rotation, updates or backups (none of those were running at the hang time). Some output of the VRRP traffic: =================== 03:42:11.430827 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement, vrid 51, prio 100, authtype simple, intvl 3s, length 20 03:42:14.431230 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement, vrid 51, prio 100, authtype simple, intvl 3s, length 20 03:42:17.431975 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement, vrid 51, prio 100, authtype simple, intvl 3s, length 20 03:42:37.315743 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement, vrid 51, prio 100, authtype simple, intvl 3s, length 20 03:42:37.316554 IP 141.163.66.143 > 224.0.0.18: VRRPv2, Advertisement, vrid 51, prio 100, authtype simple, intvl 3s, length 20 03:42:40.316780 IP 141.163.66.144 > 224.0.0.18: VRRPv2, Advertisement, vrid 51, prio 100, authtype simple, intvl 3s, length 20 =================== The advertisement interval is 3 seconds, but note the 20 second gap at 03:42:17. Initially 66.144 has the VIP, but during the 20 second gap the other server (66.143) takes it. When the interface comes back up both servers are advertising, so keepalived selects the master based on the IP address (since the prio is equal for both servers). What I cannot work out is whether the interface has hung because of keepalived or something else on the system. There is nothing logged in any log file (that I can find) to indicate a problem with the network interface. Has anyone else seen anything like this? Thanks, John. -- John Horne Tel: +44 (0)1752 587287 Plymouth University, UK ------------------------------------------------------------------------------ Comprehensive Server Monitoring with Site24x7. Monitor 10 servers for $9/Month. Get alerted through email, SMS, voice calls or mobile push notifications. Take corrective actions from your mobile device. http://pubads.g.doubleclick.net/gampad/clk?id=154624111&iu=/4140/ostg.clktrk