Re: Spurious failovers

Graeme Fowler <[email protected]>
Newsgroups gmane.linux.keepalived.devel
Message-ID <[email protected]>
On Thu, 2014-11-20 at 18:03 +0000, John Horne wrote:
> Well in the end we decided to run an strace on the keepalive processes
> overnight. The problem occurred again last night, and the culprit seems
> to have been within a 'select' statement:
<snip>

Are the affected machines VMs? (goes to read thread on CentOS list...)
Ah, yes they are.

Check your virtualisation platform. This smacks of a VM snapshot being
taken by a backup system, or the VMs being moved from host to host for
maintenance or system load reasons.

We have a big VMware platform here and I've had to turn down our HA
sensitivity (both in keepalived and heartbeat, where used) because
snapshots can cause the VMs to "pause" for a few seconds when they are
very near completion; the VM state is quiesced before the task
completes. On most quiet systems this is not noticeable but on systems
with a high rate of data change or time-sensitive processes it can be a
problem, causing as you've seen spurious failovers.

Graeme


------------------------------------------------------------------------------
Download BIRT iHub F-Type - The Free Enterprise-Grade BIRT Server
from Actuate! Instantly Supercharge Your Business Reports and Dashboards
with Interactivity, Sharing, Native Excel Exports, App Integration & more
Get technology previously reserved for billion-dollar corporations, FREE
http://pubads.g.doubleclick.net/gampad/clk?id=157005751&iu=/4140/ostg.clktrk
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.