Odd weight and healthcheck issue?

Daniel Selans <[email protected]> Wed, 5 Sep 2007 13:31:03 -0400
Newsgroups gmane.linux.keepalived.announce
Message-ID <[email protected]>
Hello list,

This morning I woke up to a call from one of one of our techs at the  
company saying that our DNS cluster was down. After quickly checking  
the ipvs table on the director, I saw something really odd:

   -> RemoteAddress:Port           Forward Weight ActiveConn InActConn
UDP  dns1.dimenoc.com:domain wlc persistent 30
   -> 10.85.83.49:domain           Masq    1      0          450573
   -> 10.85.83.48:domain           Masq    2      0          10504
   -> 10.85.83.47:domain           Masq    2      0          2574
   -> 10.85.83.46:domain           Masq    2      0          1842

Let me give you some insight on the setup - .46, .47, .48 are dual  
xeons, thus the weight of 2 and .49 is a P4. All of the dual xeon  
machines were up and responding, although the .49 was clearly  
overloaded and not even responding to ssh. Only after clearing the  
ipvs tables, and restarting keepalived was I able to get those  
numbers to start balancing out.

How is this possible? Keepalived should've removed the .49 machine,  
and should've started distributing load across the machines with  
barely any connections. Yet .49 was left around and continued taking  
on connections while the dual xeons didn't get any work at all.

If anyone has any input on this issue, I'd greatly appreciate it.

The LVS director is on a dual xeon as well, on a 2.6.9-smp kernel, if  
this matters, and same goes for the rest of the machines.

Daniel Selans
Sr. Systems Administrator/Network Engineer
Hostdime.com, Inc. | http://www.hostdime.com/



-------------------------------------------------------------------------
This SF.net email is sponsored by: Splunk Inc.
Still grepping through log files to find problems?  Stop.
Now Search log events and configuration files using AJAX and a browser.
Download your FREE copy of Splunk now >>  http://get.splunk.com/