Re: Uses for a cluster
"Evan Turner" <[email protected]> Thu, 29 Jun 2006 09:58:03 -0500
| Newsgroups | gmane.linux.cluster.openmosix.general |
|---|---|
| Message-ID | <[email protected]> |
Lars, Failover -------- Hardware failures will always occur, no matter how much backup you have. Also if you have swappable hardware, if a node dies you still lost your job.. regardless if now your cluster is back in operation after a quick fix. My advice would be to look into software fault tolerance. Does your code support checkpoint-restart? Can you recover your application after a critical failure? This is a very, very, hot topic right now in the academic circles. This is necessary to deploy large problems over huge-normous HPC grids. (hint the powers at be in the OM community should devise how to start & stop processes to make a snapshot) Manforce -------- You hit on a good topic there. I think from about the money you spend on support crew you could have bought an IBM. However, these machines do not have a 1-time cost. You also have to "buy" warrantee support from IBM/SGI/SUN and its very very very expensive. We're talking 100k a year expensive. You could have hired someone with that (or two someone's). My old lab I had the philosophy of throw money at people not hardware. It didn't matter that we spent our whole grant on computers, if we didn't have money to hire people to run them, program on them, promote the office. (plus I needed a job, I didn't want to work for free) This is basically up to your 'management' decisions. 7 machines ---------- Firstly, why seven specifically. Most of what we do here goes by powers of 2... everything tends to work out better that way. And the power of 2 applies to the # of processors as well as nodes. Honestly, with such a small setup, those machines could very well be on for days-weeks at a time, especially if you're just running a single-sane application on them. Your points of failure are much lower. Commodity hardware is continually becoming more reliable over time. (I'd still buy 'real' servers from a company and not build them from spare parts, no telling the quality there). My 12 node dell 2650 OM cluster ran without reboots or failures for several weeks at a time. Usually we rebooted whenever we decided to monkey with software changes and the like. In fact, after a 2 years and several upgrades the most we ever lost was one of the hot-swap harddrives. -evan T. -----Original Message----- From: openmosix-general-bounces-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org [mailto:openmosix-general-bounces-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org] On Behalf Of Lars O. Grobe Sent: Thursday, June 29, 2006 9:43 AM To: openmosix-general-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org Subject: Re: [openMosix-general] Uses for a cluster Hi Evan! > Our largest production cluster is a ~1000 node dual-xeon blade system. I think you play in a different league... still if you need some processing load for your cluster feel free to ask ;-) > We never purchase failover hardware because we always want to get the > maximum system out there. At any one time we might have 1 or 2 nodes > out for maintenance, but that does not impact our running. Actually, openmosix allows clusters without failover hardware (at least the damage is limited). I was thinking of the network infrastructure, or connectivity, be it Ethernet or whatever. These switches and Co can get really expensive, and if one fails, all your cluster is gone. > Even though we have a large system, we have a mandatory 24hr runtime > limit on our cluster systems because of stability. Such a limit can be very useful. Maybe I should implement it into our environment, too. At the moment I have to restart after failures. > What you describe as buying a single machine to save on cooling & power > the industry buzz around here calls the 'price performance comparison'. > Basically it's a marketing spreadsheet that says FLOPS per dollar of the > entire solution of hardware/networking/storage/cooling/power (you > provide your own people usually. In fact, I just wanted to mention the availability of parallel computing platforms like the cell architecture as an alternative for certain applications. > accomplish, obviously. We have a few sun quad socket, DC machines and > they start at 24k a piece! The IBM, SGI, and SUN supermachine options > are nice, but you can easily bust a million dollars on one rack. I think it also depends on the manforce available to fix problems and the tolerance of the clients who need results. I can get a cheap cluster running, but if I had to give a warranty for results in time... Those supermachines are both expensive and high-quality (=stable and fixeable, service, documented, ...) imho. But I am sure Evan has much more experience with the different platforms. > If you're going to stick to 'little' man you can get away with probably > about 12-16 nodes before you have to really worry about cooling. Here it depends on how little you are. If you even have no dedicated room, these things can get crucial. Looking at non-x86-platforms might make sense, again, for power consumption etc. My response was directly connected to th "7 PCs" scenario. In general I doubt that seven standard PCs are a very useful setup for a 24/7 cluster, so that explains my doubts. ;-) For experiments it is fine however. CU Lars. Using Tomcat but need to do more? Need to support web services, security? Get stuff done quickly with pre-integrated technology to make your job easier Download IBM WebSphere Application Server v.1.0.1 based on Apache Geronimo http://sel.as-us.falkag.net/sel?cmd=lnk&kid=120709&bid=263057&dat=121642 _______________________________________________ openMosix-general mailing list openMosix-general-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org https://lists.sourceforge.net/lists/listinfo/openmosix-general Using Tomcat but need to do more? Need to support web services, security? Get stuff done quickly with pre-integrated technology to make your job easier Download IBM WebSphere Application Server v.1.0.1 based on Apache Geronimo http://sel.as-us.falkag.net/sel?cmd=lnk&kid=120709&bid=263057&dat=121642