Re: obscure problem, any clues?

Simon Luff <[email protected]>
Newsgroups gmane.linux.highavailability.ultramonkey
Message-ID <[email protected]>
hi all,

Sorry for long quote, just in case you didn't see the original.

FYI, eventually tracked this down to a problem with CEF (Cisco Express 
Forwarding) on the upstream routers. Seems some of the cache data 
remains stale after a heartbeat IPAddr2 type IP change. I tried various 
different explicit cache flushing commands but to no avail, could 
confirm the cause by demonstrating problem/resolution with CEF 
enabled/disabled on the routers.

So on a vaguely related note, is anyone using an IP failover scheme on 
linux that uses a floating MAC address (like Cisco HSRP I believe)?

best,
Simon.

  Simon Luff wrote:
> hi all,
> 
> We are running a pair of masquerading ultramonkey load balancers, with 
> about 20 internet IPs spread in different arrangements across about 20 
> real servers on private IPs. The balancers also run quagga BGP daemons 
> in order to route external traffic via one of several IP providers.
> 
> We're having a problem where we seem to be dropping some traffic on one 
> of our sites from one particular balancer and we can't figure out why. I 
> haven't yet come up with any specific test that fails, but when we 
> failover to the other balancer we see a 50% increase in traffic for this 
> particular site.
> 
> In normal operation 1 cluster operates from one balancer and all the 
> rest run from the other. Nominal traffic is about 10Mb/s with a few 
> peaks around 80Mb/s. Although the problem is demonstrable when the 
> problem site is doing 2-3 Mb/s. The balancers can flood ping >90Mb/s out 
> front and back interfaces without loss.
> 
> The internet VIPs are configured on the loopback interfaces of the 
> balancers, and there is a floating public and private address for each 
> balancer, with static routes on the routers and real servers to assign 
> some of the VIPs and some of the real servers to each balancer.
> 
> I believe both balancers are configured the same, running on modern 
> major-vendor servers, a mature FC3 updates kernel, idle normally > 95%, 
> ip_conntrack_count shows way below ip_conntrack_max, and I see no 
> indication of network errors in the  interface counters or in the system 
> logs.
> 
> Anyone have any clues or ideas as to how I can further diagnose this? 
> The only evidence so far seems to be that for this one site I can show 
> it does higher traffic levels on one balancer than the other.
> 
> best regards,
> Simon.
> 


-- 
Ultra Monkey - http://www.ultramonkey.org/
To UNSUBSCRIBE, email to [email protected], with a body:
unsubscribe ultramonkey-users [email protected]
where "[email protected]" is YOUR email address.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.