Failure of secondaries
Hal Burgiss <[email protected]> Sat, 11 Sep 2010 15:44:16 -0400
| Newsgroups | gmane.network.djbdns |
|---|---|
| Message-ID | <[email protected]> |
--001485f6c762adf89b0490011658 Content-Type: text/plain; charset=ISO-8859-1 Hello, I am trying understand a predicament I found myself in today. As background, my environment is that I work for a small web hosting company. We handle the authoritative DNS for most of our clients, using djbdns/tinydns. So we have ns1, ns2, and ns3 type setup. data.cdb is shared among the 3 when the Makefile is executed so that everything stays in sync. This is a non-clustered set up, with one ip address per server. This has seemed to work flawlessly for years now. Last night though someone inadvertantly disconnected the wrong server, and unplugged the ns1 system. The eventual impact of that one mistake was that the dns for the hosted domains all went down totally. The ns2 and n3 systems were never queried. Direct querying during testing showed they were responding normally (eg dig blah.com @ns2). Yet, for all practical purposes they might as well been unplugged too since they were totally quiet. I had been under the false assumption that should ns1 go down, that the others would automatically come into play. What am I missing? Secondly, when I realized what happened and that the two secondary systems were totally useless, I moved the ip address from the ns1 to ns3, and changed the tinydns configs, restarted the service, verified that tinydns was listening on the correct ip and port, and direct test queries worked fine. I am doing all this remotely, and did not have the ability to reconnect the original system. I was assuming the ip move would be a reasonable hotfix. But this did not work. Some 2 hours later the original system was reconnected, and within mintues all started working normally again. Help me understand this so I can avoid this kind of headache in the future! Thank you. -- Hal --001485f6c762adf89b0490011658 Content-Type: text/html; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable Hello,<div><br></div><div>I am trying understand a predicament I found myse= lf in today. As background, my environment is that I work for a small web h= osting company. We handle the authoritative DNS for most of our clients, us= ing djbdns/tinydns. So we have ns1, ns2, and ns3 type setup. data.cdb is sh= ared among the 3 when the Makefile is executed so that everything stays in = sync. This is a non-clustered set up, with one ip address per server.</div> <div><br></div><div>This has seemed to work flawlessly for years now. Last = night though someone inadvertantly disconnected the wrong server, and unplu= gged the ns1 system. The eventual impact of that one mistake was that the d= ns for the hosted domains all went down totally. The ns2 and n3 systems wer= e never queried. Direct querying during testing showed they were responding= normally (eg dig <a href=3D"http://blah.com">blah.com</a> @ns2). =A0Yet, f= or all practical purposes they might as well been unplugged too since they = were totally quiet. I had been under the false assumption that should ns1 g= o down, that the others would automatically come into play. What am I missi= ng?</div> <div><br></div><div>Secondly, when I realized what happened and that the tw= o secondary systems were totally useless, I moved the ip address from the n= s1 to ns3, and changed the tinydns configs, restarted the service, verified= that tinydns was listening on the correct ip and port, and direct test que= ries worked fine. I am doing all this remotely, and did not have the abilit= y to reconnect the original system. I was assuming the ip move would be a r= easonable hotfix. But this did not work. Some 2 hours later the original sy= stem was reconnected, and within mintues all started working normally again= . Help me understand this so I can avoid this kind of headache in the futur= e!</div> <div><br></div><div>Thank you.</div><div><br></div><div>-- <br>Hal<br> </div> --001485f6c762adf89b0490011658--