Re: UDP packet loss when server health check fails

Ed Ravin <[email protected]> Thu, 6 Jul 2017 08:40:16 -0400
Newsgroups gmane.linux.keepalived.devel
Message-ID <[email protected]>
Quentin, thank you for the thoughtful response!  Here is some more
information:

* You are exactly right, keepalived isn't the problem. I can reproduce the
same delay when manually removing a server out of the farm with ipvsadm.

* It turns out I wasn't using dnsperf correctly, and it was hiding the
true length of the outage until I told it to allow more "outstanding
requests" (the -q option).  The period of packet loss is actully an
entire second when testing at 10,000 queries per second.

* When I dialed the query rate down, the packet loss period got smaller.
It disappeared entirely at 400 queries per second.

* As I'm using source hashing for the distribution algorithm, I don't
think changing the weights is meaningful unless I switch to round-robin.
When I tried inhibit_on_failure with source hashing, it looked like IPVS
was blackholing traffic for the failed server so the clients got
no responses at all.

I think the next line of inquiry is to look at what IPVS does when there
is a farm reconfiguration event, and how sensitive that process is to
the rate of incoming traffic.  I also want to compare against using
round-robin and see whether that makes a difference.

Thanks again,

	-- Ed


On Wed, Jul 05, 2017 at 02:04:06PM +0100, Quentin Armitage wrote:
> On Tue, 2017-07-04 at 14:11 -0400, Ed Ravin wrote:
> > I'm using keepalived to distribute DNS requests (UDP port 53) to a
> > group of DNS servers.  The farm is using source hashing.  Environment
> > is RHEL7.2, with the stock keepalived and IPVS.
> > 
> > I'm testing what happens when a health check fails and one host is
> > taken out of the farm.  My test bed has two farm servers and one
> > keepalived server.  The keepalived server is using two ethernet
> > adapters in a bond interface as its primary adapter.  For testing,
> > I'm using a fake health check of "sh -c '! test -f FLAGFILE'", which
> > returns success as long as the file doesn't exist, and I create the
> > fileto provoke a health check failure.
> > 
> > Using dnsperf to generate a stream of 10,000 queries/second, I create the
> > flag file and keepalived reports taking the real server out of the farm.
> > dnsperf then reports losing around 100 queries (around 10 milliseconds)
> > during the transition.
> > 
> > I ran tcpdump to capture the traffic, and I can see on successful queries
> > the packet is received on bond0 and then re-transmitted out on bond0 with
> > the destination server's MAC address.  The configuration is using direct
> > server response so the source and destination IP addresses in the packet
> > are unchanged when it is transmitted to the farm server.
> > 
> > I checked several of the query-ids that dnsperf reported as missing. tcpdump
> > saw all of their query packets arriving, but did not show any of them
> > getting re-sent out the interface.
> > 
> > My questions are:
> > 
> > * Is it realistic to expect that no packets will be dropped during a farm
> > reconfigure transition?
> > 
> > * If it's not realistic, what can I do to minimize the drops?  10 ms is
> > not a lot by some standards, but in my environment it could be 100-200
> > requests that I'd rather see answered.
> > 
> > * My theory is either keepalived dropped the requests, IPVS dropped them,
> > or something further down the network stack dropped them.  I looked into
> > the IPVS counters in /proc but didn't see anything that keeps track of errors.
> > Can anyone suggest other avenues of visibility into finding where the
> > requests or responses are being lost?
> > 
> > Thanks,
> > 
> > 	-- Ed
> > 
> For the reasons given below, I think the first two questions are really
> for the IPVS people.
> 
> In relation to your third question, first of all, keepalived doesn't
> see the DNS packets; keepalived is simply managing the configuration of
> the IPVS service, in other words adding, removing and configuring the
> real and virtual servers. So the glib, but unhelpful, answer is that
> keepalived cannot be droping the packets. On the other hand, it is
> possible that keepalived is the cause of dropping the packets when it
> reconfigures the real server(s).
> 
> There are two different ways that keepalived can manage a real server
> in the event of a health check failure, depending on the setting of
> inhibit_on_failure. If inhibit_on_failure is set in a real server, then
> if a health check fails, the priority of the real server is set to 0;
> if inhibit_on_failure is not set, then the real server is removed from
> the farm. It always strikes me that setting the priority to 0 must be
> less disruptive than removing a real server, so if you are not already
> doing so it might be worth setting inhibit_on_failure.
> 
> In order to determine whether or not keepalived is the cause of the
> problem, I suggest you set up the virtual/real server config without
> keepalived, using ipvsadm. Then you could try and either remove a real
> server or setting its priority to 0 using ipvsadm and see if you still
> get the packet loss. That would confirm whether the problem lies on the
> keepalived or IPVS side of the fence. If you find that the problem
> doesn't occur without keepalived then we will need to have a look at
> keepalived, but we will need rather more details, such as keepalived
> configuration, detailed network interface configuration etc.
> 
> If you reduce the rate that dnsperf is sending queries to say 5000/sec,
> does the number of lost queries halve also? This would give an
> indication of whether the problem is simply that packets are lost for
> around 10ms, or whether it is some sort of queue overload issue.
> 
> It might be worth considering using the iptables TRACE target to see if
> that helps in identifying where the DNS packets are being lost. Given
> the rates at which you are sending packets, it might be necessary to
> send the output to nf_log and write a program to read to nf_log output
> rather then sending everything to the system log. If the packet loss
> does appear to be due to a 10ms gap, then reducing the packet rate to
> 1000/sec might make it easier to see what is going on.
> 
> I hope that helps.
> 
> Quentin Armitage

------------------------------------------------------------------------------
Check out the vibrant tech community on one of the world's most
engaging tech sites, Slashdot.org! http://sdm.link/slashdot