archiving suspension over

Jeff Breidenbach <[email protected]> Sun, 24 Nov 2013 21:34:35 -0800
Newsgroups gmane.mail.archives.mail-archive
Message-ID <CAHjiUboJJpTW91omOVdTT7kfcH1MQ4JGG5znJ2NDBx6dS3PtiA@mail.gmail.com>
--===============0772788545073892500==
Content-Type: multipart/alternative; boundary=047d7b6d95ee21c55004ebf9b804

--047d7b6d95ee21c55004ebf9b804
Content-Type: text/plain; charset=ISO-8859-1

You may have noticed that archiving was suspended at
The Mail Archive recently. Things are fine now, read on
if you want gory details.

We are hosted at a professional datacenter, complete
with building wide uninterruptible power supply (UPS) and
backup generators. About 5 years ago, the datacenter
botched some maintenance work and accidentally cut
power. In response, we deployed an individual UPS on
our primary server to make things even safer.

That turned out to be a mistake. On Thursday that UPS
failed in the worst possible way, abruptly dropping power
while a lot of data was in flight. This is normally just an
annoyance, but in this case it caused enough damage to
the filesystem that we had suspend archiving and switch
to read-only mode.

To fix things right, we express ordered and installed some
additional storage. Because everything is now on large solid
state drives (SSD) we can afford switching to a filesystem
that is tuned more towards robustness, at the cost of some
performance. For the filesystem junkies that means migrating
from the fairly exotic XFS setup below to a fairly stock EXT4
setup. The new filesystem is about 20% less space efficient,
but that's okay. We now have enough room for years to come.

It took almost a day to fully diagnose, an overnight parts
delivery,  a few hours to get everything set up correctly, then
10 more hours to move all the data. We did not have to
resort to restoring from backups, but they are certainly there if
we need it.

I'd like to emphasize that the data is safe. We were able to
reconstruct everything that was in flight at the time of the outage.
And while we had archiving paused, inbound mail was queuing
up patiently. The system is crazy fast and we burned off the
archiving backlog in just a few hours.

Thanks for your patience and I hope you enjoyed this peek
into what goes on behind the scenes. I think the biggest benefit
of using a service like The Mail Archive is we get the fun of
dealing with problems like this so you don't have to.

Cheers,
Jeff

===

mkfs.xfs -n size=16k -i attr=2 -l lazy-count=1,version=2,size=32m \
-b size=512 noatime,logbufs=8,logbsize=256k

mount -onoatime,logbufs=8,logbsize=256k

--047d7b6d95ee21c55004ebf9b804
Content-Type: text/html; charset=ISO-8859-1
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">You may have noticed that archiving was suspended at<div>T=
he Mail Archive recently. Things are fine now, read on=A0</div><div>if you =
want gory details.<div><div><br></div><div>We are hosted at a professional =
datacenter, complete</div>
<div>with building wide uninterruptible power supply (UPS) and=A0</div><div=
>backup generators. About 5 years ago, the datacenter</div><div>botched som=
e maintenance work and accidentally cut</div><div>power. In response, we de=
ployed an individual UPS on</div>
<div>our primary server to make things even safer.</div><div><br></div><div=
>That turned out to be a mistake. On Thursday that UPS=A0</div><div>failed =
in the worst possible way, abruptly dropping power=A0</div><div>while a lot=
 of data was in flight. This is normally just an</div>
<div>annoyance, but in this case it caused enough damage to</div><div>the f=
ilesystem that we had suspend archiving and switch=A0</div><div>to read-onl=
y mode.=A0</div><div><br></div><div>To fix things right, we express ordered=
 and installed some=A0</div>
<div>additional storage. Because everything is now on large solid</div><div=
>state drives (SSD) we can afford switching to a filesystem=A0</div><div>th=
at is tuned more towards robustness, at the cost of some</div><div>performa=
nce. For the filesystem junkies that means migrating=A0</div>
<div>from the fairly exotic XFS setup below to a fairly stock EXT4</div><di=
v>setup. The new filesystem is about 20% less space efficient,</div><div>bu=
t that&#39;s okay. We now have enough room for years to come.</div><div>
<div><br></div><div>It took almost a day to fully diagnose, an overnight pa=
rts=A0</div><div>delivery, =A0a few hours to get everything set up correctl=
y, then</div><div>10 more hours to move all the data. We did not have to=A0=
</div>
<div>resort to restoring from backups, but they are certainly there if=A0</=
div><div>we need it.</div><div><br></div><div>I&#39;d like to emphasize tha=
t the data is safe. We were able to</div><div>reconstruct everything that w=
as in flight at the time of the outage.</div>
<div>And while we had archiving paused, inbound mail was queuing</div><div>=
up patiently. The system is crazy fast and we burned off the=A0</div><div>a=
rchiving backlog in just a few hours.</div><div><br></div><div>Thanks for y=
our patience and I hope you enjoyed this peek</div>
<div>into what goes on behind the scenes. I think the biggest benefit</div>=
<div>of using a service like The Mail Archive is we get the fun of=A0</div>=
<div>dealing with problems like this so you don&#39;t have to.</div><div>
<br></div><div>Cheers,</div><div>Jeff</div><div><br></div><div>=3D=3D=3D</d=
iv><div><br></div><div>mkfs.xfs -n size=3D16k -i attr=3D2 -l lazy-count=3D1=
,version=3D2,size=3D32m \</div><div>-b size=3D512 noatime,logbufs=3D8,logbs=
ize=3D256k</div></div>
<div><br></div><div>mount -onoatime,logbufs=3D8,logbsize=3D256k</div><div><=
br></div></div></div></div>

--047d7b6d95ee21c55004ebf9b804--


--===============0772788545073892500==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
Gossip mailing list
https://www.mail-archive.com/gossip-piJqB+B6DfPJghKJT/[email protected]
http://mail-archive.com/cgi-bin/mailman/options/gossip
--===============0772788545073892500==--