mass storage recommendations

"Richard Pickett" <[email protected]> Tue, 30 May 2006 13:53:57 -0500
Newsgroups gmane.linux.enbd.general
Message-ID <0f1d01c6841a$69cebb50$2001a8c0@foxtrot1>
We are looking into the best solution for doing mass storage in-house. I
really like the prospects of enbd and figure you guys have handled
things like what we are wanting to do and would have some great advice.
 
Our solution is currently all linux and we want it to remain that way.
 
Currently we have a web interface that serves up files from a number of
raid boxes behind it (`mount`ed on the web server). Currently the raids
are 4T (3.5T usable). It’s a simple enough solution, if a raid box
starts acting flaky we pull it out and can drop a replacement in, copy
the data over and it’s back up and running. The down side is the
management. Each box can host a limited number of clients, and if one
client starts growing beyond the size of that box we’re in trouble (or
at least in for a management headache). It would be easier to have one
huge raid, say 100T, and then clients that grow wouldn’t really be a
management problem.
 
With that in mind we were exploring the idea of taking a layered
approach to the raids and their devices. We could easily have a “front”
box that enbd-pulls in a large number of devices that are running on a
number of “back” boxes, then raid across those devices.
 
Another step after this would be to make an LVM that runs on top of
several of these raided systems, so if we need to grow the usable space
we just add more hardware to the LVM.
 
It is easier to manage because all of the hosted files would exist on
one huge mount point. The problem is the fault-tolerance and recovery.
With our current solution we can replace an existing box, files and all,
within a few hours. With the LVM-on-RAID-on-ENBD if two units (whether
that’s two HD if the “front” box enbd-pulls each device in, or two
systems if the “back” boxes enbd-extend all it’s drives as one drive)
fail the entire LVM is down, meaning all the data is gone.
 
So we see these two factors opposing each other: (1) large
easy-to-manage partition and (2) failover / fault tolerance.
 
I’m sure you guys have come across similar scenarios and look forward to
your advice. Maybe it’s not even feasible because network latency would
prohibit such a large system from running efficiently.
 
Thanks for your time.

Richard W. Pickett, Jr.
President, CSR Technologies .com, Inc.
[email protected]
Office - (270) 746-0324
Cell    – (270) 303-9154