Re: Spurious failovers

John Horne <[email protected]>
Newsgroups gmane.linux.keepalived.devel
Organization Plymouth University
Message-ID <[email protected]>
On Fri, 2014-11-21 at 11:30 +0000, Graeme Fowler wrote:
> On Thu, 2014-11-20 at 18:03 +0000, John Horne wrote:
> > Well in the end we decided to run an strace on the keepalive processes
> > overnight. The problem occurred again last night, and the culprit seems
> > to have been within a 'select' statement:
> <snip>
> 

> 
> Check your virtualisation platform. This smacks of a VM snapshot being
> taken by a backup system, or the VMs being moved from host to host for
> maintenance or system load reasons.
> 
Hello,

Thanks for that. I think you are right about snapshots being taken. My
understanding was that this wasn't happening, but having checked with
our Infrastructure guys apparently it is. I have sent them some
dates/times we have had a problem, and they are going to check to see if
the snapshots occurred at those times.

I have to admit that finding 'select' have a problem with a one second
timeout seemed really bizarre. Given how much select is used I would
have thought that someone would have noticed if there was some
fundamental problem with it.




John.

-- 
John Horne                   Tel: +44 (0)1752 587287
Plymouth University, UK


------------------------------------------------------------------------------
Download BIRT iHub F-Type - The Free Enterprise-Grade BIRT Server
from Actuate! Instantly Supercharge Your Business Reports and Dashboards
with Interactivity, Sharing, Native Excel Exports, App Integration & more
Get technology previously reserved for billion-dollar corporations, FREE
http://pubads.g.doubleclick.net/gampad/clk?id=157005751&iu=/4140/ostg.clktrk
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.