Failure of secondaries

Hal Burgiss <[email protected]> Sat, 11 Sep 2010 15:44:16 -0400
Newsgroups gmane.network.djbdns
Message-ID <[email protected]>
--001485f6c762adf89b0490011658
Content-Type: text/plain; charset=ISO-8859-1

Hello,

I am trying understand a predicament I found myself in today. As background,
my environment is that I work for a small web hosting company. We handle the
authoritative DNS for most of our clients, using djbdns/tinydns. So we have
ns1, ns2, and ns3 type setup. data.cdb is shared among the 3 when the
Makefile is executed so that everything stays in sync. This is a
non-clustered set up, with one ip address per server.

This has seemed to work flawlessly for years now. Last night though someone
inadvertantly disconnected the wrong server, and unplugged the ns1 system.
The eventual impact of that one mistake was that the dns for the hosted
domains all went down totally. The ns2 and n3 systems were never queried.
Direct querying during testing showed they were responding normally (eg dig
blah.com @ns2).  Yet, for all practical purposes they might as well been
unplugged too since they were totally quiet. I had been under the false
assumption that should ns1 go down, that the others would automatically come
into play. What am I missing?

Secondly, when I realized what happened and that the two secondary systems
were totally useless, I moved the ip address from the ns1 to ns3, and
changed the tinydns configs, restarted the service, verified that tinydns
was listening on the correct ip and port, and direct test queries worked
fine. I am doing all this remotely, and did not have the ability to
reconnect the original system. I was assuming the ip move would be a
reasonable hotfix. But this did not work. Some 2 hours later the original
system was reconnected, and within mintues all started working normally
again. Help me understand this so I can avoid this kind of headache in the
future!

Thank you.

-- 
Hal

--001485f6c762adf89b0490011658
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

Hello,<div><br></div><div>I am trying understand a predicament I found myse=
lf in today. As background, my environment is that I work for a small web h=
osting company. We handle the authoritative DNS for most of our clients, us=
ing djbdns/tinydns. So we have ns1, ns2, and ns3 type setup. data.cdb is sh=
ared among the 3 when the Makefile is executed so that everything stays in =
sync. This is a non-clustered set up, with one ip address per server.</div>
<div><br></div><div>This has seemed to work flawlessly for years now. Last =
night though someone inadvertantly disconnected the wrong server, and unplu=
gged the ns1 system. The eventual impact of that one mistake was that the d=
ns for the hosted domains all went down totally. The ns2 and n3 systems wer=
e never queried. Direct querying during testing showed they were responding=
 normally (eg dig <a href=3D"http://blah.com">blah.com</a> @ns2). =A0Yet, f=
or all practical purposes they might as well been unplugged too since they =
were totally quiet. I had been under the false assumption that should ns1 g=
o down, that the others would automatically come into play. What am I missi=
ng?</div>
<div><br></div><div>Secondly, when I realized what happened and that the tw=
o secondary systems were totally useless, I moved the ip address from the n=
s1 to ns3, and changed the tinydns configs, restarted the service, verified=
 that tinydns was listening on the correct ip and port, and direct test que=
ries worked fine. I am doing all this remotely, and did not have the abilit=
y to reconnect the original system. I was assuming the ip move would be a r=
easonable hotfix. But this did not work. Some 2 hours later the original sy=
stem was reconnected, and within mintues all started working normally again=
. Help me understand this so I can avoid this kind of headache in the futur=
e!</div>
<div><br></div><div>Thank you.</div><div><br></div><div>-- <br>Hal<br>
</div>

--001485f6c762adf89b0490011658--