Re: Failure on OM + 2.4.31
Peter Cordes <[email protected]>
| Newsgroups | gmane.linux.cluster.openmosix.general |
|---|---|
| Message-ID | <[email protected]> |
On Wed, Nov 16, 2005 at 09:36:07AM -0500, Terry Gliedt wrote: > Just to followup on this. After a huge amount of messing around with > hardware, I was able to replace the SATA drives (which required the > 2.4.31 kernel) and use IDE drives. This allowed me to use a 2.4.27 > kernel with patch-2.4.27-om-20041102.bz2 (224713 bytes Nov 6 2004). This > is a kernel we've been using on another set of hardware. Interesting. On my dual Opterons with 2GB of RAM (Tyan S2881) and IDE disks (the cluster compute nodes), Tab's 2.4.27-om20041102 is stable. I have had to turn the master node into an NFS server and have people log in to what was one of the compute nodes. I could not get the master node to be stable with that kernel. It's a Tyan S2882 with dual Opterons, 4GB of RAM, and two SATA disks (using md for three RAID1 partitions and a RAID0 partition (/, /usr/local, /home, and /data respectively)). The 2882 is a very similar mobo, but it has a different BIOS, and I'm using SATA on it. It also features a 10/100 e100 NIC built in besides the two tg3 gigE ports they all have, but it was unstable even when not using that. I've tried: memtest86+ (it passes many runs, and has ECC RAM). booting with acpi=off (and various combinations of ACPI options in the BIOS) booting with mem=2048M (I haven't tried physically taking RAM out of it). upgrading the BIOS from 2.04 to 3.05. various BIOS setting changes (including very conservative memory and other bus speeds with no effect even on the frequency of lockups, so I'm confident it's not one of those can't-run-the-ram-that-fast problems that was getting me on my desktop at home...) I compiled the kernel with highmem4G, not 64GB, since Moshe says that's not supported :( I used gcc 3.3.5 from Debian, and I'm running Debian sarge on the machines. Is that really the case? And is there a problem with having exactly 4GB? I was suspicious of my BIOS maybe letting Linux stomp on some address ranges that weren't just memory, but I don't know. I haven't tried a non-SMP kernel. I haven't tried repartitioning to not use md, because that would suck. I have another cluster with a master node that has a Tyan S2881 mobo (and 6GB of RAM), so I might be able to do some messing around and see how it does for stability with various configurations. The compute nodes are blades, so there's not much scope for putting two disks into one of them. I could try md on the network block device, though... > This older kernel has proven to be completely stable. I can run the full > range of openMosix Stress-Test tests for hours on end. So my expereience > is that the later kernel is not stable with any serious amount of work > thrown at it. That's reassuring. This trial and error with stability has been really frustrating. I think that's what I found booting 2.4.31 2005 05 something from voxus on two of my compute nodes. I just booted CHAOS on them, and ran a few invocations of the infinite-loop testapp, and the machines came to a grinding halt. I think there was an oops message on the remote node. It would be nice if there was a stable openMosix for 2.4.31, because of the security fixes and other bugfixes in the rest of the kernel between 2.4.27 and .31. -- #define X(x,y) x##y Peter Cordes ; e-mail: X(peter@cor , des.ca) "The gods confound the man who first found out how to distinguish the hours! Confound him, too, who in this place set up a sundial, to cut and hack my day so wretchedly into small pieces!" -- Plautus, 200 BC ------------------------------------------------------- This SF.net email is sponsored by: Splunk Inc. Do you grep through log files for problems? Stop! Download the new AJAX search engine that makes searching your log files as easy as surfing the web. DOWNLOAD SPLUNK! http://ads.osdn.com/?ad_id=7637&alloc_id=16865&op=click