Re: Spurious failovers
John Horne <[email protected]>
| Newsgroups | gmane.linux.keepalived.devel |
|---|---|
| Organization | Plymouth University |
| Message-ID | <[email protected]> |
On Fri, 2014-11-21 at 11:30 +0000, Graeme Fowler wrote: > On Thu, 2014-11-20 at 18:03 +0000, John Horne wrote: > > Well in the end we decided to run an strace on the keepalive processes > > overnight. The problem occurred again last night, and the culprit seems > > to have been within a 'select' statement: > <snip> > > > Check your virtualisation platform. This smacks of a VM snapshot being > taken by a backup system, or the VMs being moved from host to host for > maintenance or system load reasons. > Hello, Thanks for that. I think you are right about snapshots being taken. My understanding was that this wasn't happening, but having checked with our Infrastructure guys apparently it is. I have sent them some dates/times we have had a problem, and they are going to check to see if the snapshots occurred at those times. I have to admit that finding 'select' have a problem with a one second timeout seemed really bizarre. Given how much select is used I would have thought that someone would have noticed if there was some fundamental problem with it. John. -- John Horne Tel: +44 (0)1752 587287 Plymouth University, UK ------------------------------------------------------------------------------ Download BIRT iHub F-Type - The Free Enterprise-Grade BIRT Server from Actuate! Instantly Supercharge Your Business Reports and Dashboards with Interactivity, Sharing, Native Excel Exports, App Integration & more Get technology previously reserved for billion-dollar corporations, FREE http://pubads.g.doubleclick.net/gampad/clk?id=157005751&iu=/4140/ostg.clktrk