Re: UDP packet loss when server health check fails
Ed Ravin <[email protected]> Thu, 6 Jul 2017 08:40:16 -0400
| Newsgroups | gmane.linux.keepalived.devel |
|---|---|
| Message-ID | <[email protected]> |
Quentin, thank you for the thoughtful response! Here is some more information: * You are exactly right, keepalived isn't the problem. I can reproduce the same delay when manually removing a server out of the farm with ipvsadm. * It turns out I wasn't using dnsperf correctly, and it was hiding the true length of the outage until I told it to allow more "outstanding requests" (the -q option). The period of packet loss is actully an entire second when testing at 10,000 queries per second. * When I dialed the query rate down, the packet loss period got smaller. It disappeared entirely at 400 queries per second. * As I'm using source hashing for the distribution algorithm, I don't think changing the weights is meaningful unless I switch to round-robin. When I tried inhibit_on_failure with source hashing, it looked like IPVS was blackholing traffic for the failed server so the clients got no responses at all. I think the next line of inquiry is to look at what IPVS does when there is a farm reconfiguration event, and how sensitive that process is to the rate of incoming traffic. I also want to compare against using round-robin and see whether that makes a difference. Thanks again, -- Ed On Wed, Jul 05, 2017 at 02:04:06PM +0100, Quentin Armitage wrote: > On Tue, 2017-07-04 at 14:11 -0400, Ed Ravin wrote: > > I'm using keepalived to distribute DNS requests (UDP port 53) to a > > group of DNS servers. The farm is using source hashing. Environment > > is RHEL7.2, with the stock keepalived and IPVS. > > > > I'm testing what happens when a health check fails and one host is > > taken out of the farm. My test bed has two farm servers and one > > keepalived server. The keepalived server is using two ethernet > > adapters in a bond interface as its primary adapter. For testing, > > I'm using a fake health check of "sh -c '! test -f FLAGFILE'", which > > returns success as long as the file doesn't exist, and I create the > > fileto provoke a health check failure. > > > > Using dnsperf to generate a stream of 10,000 queries/second, I create the > > flag file and keepalived reports taking the real server out of the farm. > > dnsperf then reports losing around 100 queries (around 10 milliseconds) > > during the transition. > > > > I ran tcpdump to capture the traffic, and I can see on successful queries > > the packet is received on bond0 and then re-transmitted out on bond0 with > > the destination server's MAC address. The configuration is using direct > > server response so the source and destination IP addresses in the packet > > are unchanged when it is transmitted to the farm server. > > > > I checked several of the query-ids that dnsperf reported as missing. tcpdump > > saw all of their query packets arriving, but did not show any of them > > getting re-sent out the interface. > > > > My questions are: > > > > * Is it realistic to expect that no packets will be dropped during a farm > > reconfigure transition? > > > > * If it's not realistic, what can I do to minimize the drops? 10 ms is > > not a lot by some standards, but in my environment it could be 100-200 > > requests that I'd rather see answered. > > > > * My theory is either keepalived dropped the requests, IPVS dropped them, > > or something further down the network stack dropped them. I looked into > > the IPVS counters in /proc but didn't see anything that keeps track of errors. > > Can anyone suggest other avenues of visibility into finding where the > > requests or responses are being lost? > > > > Thanks, > > > > -- Ed > > > For the reasons given below, I think the first two questions are really > for the IPVS people. > > In relation to your third question, first of all, keepalived doesn't > see the DNS packets; keepalived is simply managing the configuration of > the IPVS service, in other words adding, removing and configuring the > real and virtual servers. So the glib, but unhelpful, answer is that > keepalived cannot be droping the packets. On the other hand, it is > possible that keepalived is the cause of dropping the packets when it > reconfigures the real server(s). > > There are two different ways that keepalived can manage a real server > in the event of a health check failure, depending on the setting of > inhibit_on_failure. If inhibit_on_failure is set in a real server, then > if a health check fails, the priority of the real server is set to 0; > if inhibit_on_failure is not set, then the real server is removed from > the farm. It always strikes me that setting the priority to 0 must be > less disruptive than removing a real server, so if you are not already > doing so it might be worth setting inhibit_on_failure. > > In order to determine whether or not keepalived is the cause of the > problem, I suggest you set up the virtual/real server config without > keepalived, using ipvsadm. Then you could try and either remove a real > server or setting its priority to 0 using ipvsadm and see if you still > get the packet loss. That would confirm whether the problem lies on the > keepalived or IPVS side of the fence. If you find that the problem > doesn't occur without keepalived then we will need to have a look at > keepalived, but we will need rather more details, such as keepalived > configuration, detailed network interface configuration etc. > > If you reduce the rate that dnsperf is sending queries to say 5000/sec, > does the number of lost queries halve also? This would give an > indication of whether the problem is simply that packets are lost for > around 10ms, or whether it is some sort of queue overload issue. > > It might be worth considering using the iptables TRACE target to see if > that helps in identifying where the DNS packets are being lost. Given > the rates at which you are sending packets, it might be necessary to > send the output to nf_log and write a program to read to nf_log output > rather then sending everything to the system log. If the packet loss > does appear to be due to a 10ms gap, then reducing the packet rate to > 1000/sec might make it easier to see what is going on. > > I hope that helps. > > Quentin Armitage ------------------------------------------------------------------------------ Check out the vibrant tech community on one of the world's most engaging tech sites, Slashdot.org! http://sdm.link/slashdot