Re: Spurious failovers
Graeme Fowler <[email protected]>
| Newsgroups | gmane.linux.keepalived.devel |
|---|---|
| Message-ID | <[email protected]> |
On Thu, 2014-11-20 at 18:03 +0000, John Horne wrote: > Well in the end we decided to run an strace on the keepalive processes > overnight. The problem occurred again last night, and the culprit seems > to have been within a 'select' statement: <snip> Are the affected machines VMs? (goes to read thread on CentOS list...) Ah, yes they are. Check your virtualisation platform. This smacks of a VM snapshot being taken by a backup system, or the VMs being moved from host to host for maintenance or system load reasons. We have a big VMware platform here and I've had to turn down our HA sensitivity (both in keepalived and heartbeat, where used) because snapshots can cause the VMs to "pause" for a few seconds when they are very near completion; the VM state is quiesced before the task completes. On most quiet systems this is not noticeable but on systems with a high rate of data change or time-sensitive processes it can be a problem, causing as you've seen spurious failovers. Graeme ------------------------------------------------------------------------------ Download BIRT iHub F-Type - The Free Enterprise-Grade BIRT Server from Actuate! Instantly Supercharge Your Business Reports and Dashboards with Interactivity, Sharing, Native Excel Exports, App Integration & more Get technology previously reserved for billion-dollar corporations, FREE http://pubads.g.doubleclick.net/gampad/clk?id=157005751&iu=/4140/ostg.clktrk