Re: Uses for a cluster

"Evan Turner" <[email protected]> Thu, 29 Jun 2006 09:58:03 -0500
Newsgroups gmane.linux.cluster.openmosix.general
Message-ID <[email protected]>
Lars,

Failover
--------

Hardware failures will always occur, no matter how much backup you have.
Also if you have swappable hardware, if a node dies you still lost your
job.. regardless if now your cluster is back in operation after a quick
fix.

My advice would be to look into software fault tolerance.  Does your
code support checkpoint-restart?  Can you recover your application after
a critical failure? This is a very, very, hot topic right now in the
academic circles.  This is necessary to deploy large problems over
huge-normous HPC grids.  (hint the powers at be in the OM community
should devise how to start & stop processes to make a snapshot)

Manforce
--------

You hit on a good topic there.  I think from about the money you spend
on support crew you could have bought an IBM.  However, these machines
do not have a 1-time cost.  You also have to "buy" warrantee support
from IBM/SGI/SUN and its very very very expensive.  We're talking 100k a
year expensive.  You could have hired someone with that (or two
someone's).

My old lab I had the philosophy of throw money at people not hardware.
It didn't matter that we spent our whole grant on computers, if we
didn't have money to hire people to run them, program on them, promote
the office.  (plus I needed a job, I didn't want to work for free)

This is basically up to your 'management' decisions.  

7 machines
----------

Firstly, why seven specifically.  Most of what we do here goes by powers
of 2... everything tends to work out better that way.  And the power of
2 applies to the # of processors as well as nodes.

Honestly, with such a small setup, those machines could very well be on
for days-weeks at a time, especially if you're just running a
single-sane application on them.  Your points of failure are much lower.
Commodity hardware is continually becoming more reliable over time.
(I'd still buy 'real' servers from a company and not build them from
spare parts, no telling the quality there).

My 12 node dell 2650 OM cluster ran without reboots or failures for
several weeks at a time.  Usually we rebooted whenever we decided to
monkey with software changes and the like.  In fact, after a 2 years and
several upgrades the most we ever lost was one of the hot-swap
harddrives.

-evan T.

-----Original Message-----
From: openmosix-general-bounces-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org
[mailto:openmosix-general-bounces-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org] On Behalf Of
Lars O. Grobe
Sent: Thursday, June 29, 2006 9:43 AM
To: openmosix-general-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org
Subject: Re: [openMosix-general] Uses for a cluster

Hi Evan!

> Our largest production cluster is a ~1000 node dual-xeon blade system.

I think you play in a different league... still if you need some 
processing load for your cluster feel free to ask ;-)

> We never purchase failover hardware because we always want to get the
> maximum system out there.  At any one time we might have 1 or 2 nodes
> out for maintenance, but that does not impact our running.

Actually, openmosix allows clusters without failover hardware (at least 
the damage is limited). I was thinking of the network infrastructure, or

connectivity, be it Ethernet or whatever. These switches and Co can get 
really expensive, and if one fails, all your cluster is gone.

> Even though we have a large system, we have a mandatory 24hr runtime
> limit on our cluster systems because of stability.

Such a limit can be very useful. Maybe I should implement it into our 
environment, too. At the moment I have to restart after failures.

> What you describe as buying a single machine to save on cooling &
power
> the industry buzz around here calls the 'price performance
comparison'.
> Basically it's a marketing spreadsheet that says FLOPS per dollar of
the
> entire solution of hardware/networking/storage/cooling/power (you
> provide your own people usually.

In fact, I just wanted to mention the availability of parallel computing

platforms like the cell architecture as an alternative for certain 
applications.

> accomplish, obviously.  We have a few sun quad socket, DC machines and
> they start at 24k a piece!  The IBM, SGI, and SUN supermachine options
> are nice, but you can easily bust a million dollars on one rack.

I think it also depends on the manforce available to fix problems and 
the tolerance of the clients who need results. I can get a cheap cluster

running, but if I had to give a warranty for results in time... Those 
supermachines are both expensive and high-quality (=stable and fixeable,

service, documented, ...) imho. But I am sure Evan has much more 
experience with the different platforms.

> If you're going to stick to 'little' man you can get away with
probably
> about 12-16 nodes before you have to really worry about cooling.

Here it depends on how little you are. If you even have no dedicated 
room, these things can get crucial. Looking at non-x86-platforms might 
make sense, again, for power consumption etc.

My response was directly connected to th "7 PCs" scenario. In general I 
doubt that seven standard PCs are a very useful setup for a 24/7 
cluster, so that explains my doubts. ;-) For experiments it is fine
however.

CU Lars.


Using Tomcat but need to do more? Need to support web services,
security?
Get stuff done quickly with pre-integrated technology to make your job
easier
Download IBM WebSphere Application Server v.1.0.1 based on Apache
Geronimo
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=120709&bid=263057&dat=121642
_______________________________________________
openMosix-general mailing list
openMosix-general-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org
https://lists.sourceforge.net/lists/listinfo/openmosix-general

Using Tomcat but need to do more? Need to support web services, security?
Get stuff done quickly with pre-integrated technology to make your job easier
Download IBM WebSphere Application Server v.1.0.1 based on Apache Geronimo
http://sel.as-us.falkag.net/sel?cmd=lnk&kid=120709&bid=263057&dat=121642