Re: kernel BUG with 2.4.27-om20041102-tab on SMP when doing heavy I/O
Peter Cordes <[email protected]>
| Newsgroups | gmane.linux.cluster.openmosix.devel |
|---|---|
| Message-ID | <[email protected]> |
On Tue, Sep 27, 2005 at 03:13:32PM -0300, Peter Cordes wrote: > I have been running 2.4.27-om20041102-tab (peter-OPrYhD/TZC/[email protected]) > (gcc version 3.3.5 (Debian 1:3.3.5-13)) #1 SMP Wed Sep 7 on my dual Opteron > cluster. The master node has 4GB of RAM, and the other 7 nodes have 2GB. > All with gigabit ethernet. The master has an Intel e100 built in to the > mobo too, which is used to connect to the outside world. > > My kernel config is > CONFIG_MOSIX=y > # CONFIG_MOSIX_TOPOLOGY is not set > CONFIG_MOSIX_SECUREPORTS=y > CONFIG_MOSIX_DISCLOSURE=3 > # CONFIG_MOSIX_FS is not set > CONFIG_MOSIX_PIPE_EXCEPTIONS=y > # CONFIG_MOSIX_NO_OOM is not set > CONFIG_MOSIX_EXT_LOCALTIME=y > > My filesystems are JFS on RAID0 and RAID1 partitions of two SATA drives, > connected to the SATA_SIL ports on the Tyan S2882 motherboard. Upgrading the BIOS from 2.04 to 3.04 on my Tyan S2882 motherboard seems to have resolved this completely. I hadn't run the machine for long or under heavy loads with non-OM kernels, so I assumed the problem was oM's fault... I think I'm still running that machine with acpi=off. Other machines in the cluster are using Tyan S2881 motherboards with 2.05 BIOS, and didn't have problems. OTOH, they weren't running much on their own, just one ATA disk and not having processes started on them. Anyway, situation is under control. (But I'm still figuring out how to do BIOS upgrades on a whole cluster of machines. Apparently someone got Tyan S2880 mobos to upgrade their BIOS booting a disk image with PXELinux and memdisk, using a flash utility other than Tyan's AMIFLASH.COM which doesn't run from memdisk. aminf335 doesn't work (too old), and I haven't tried aminf341 yet. I'm messing around with bochs and scripts to make disk images, just so I don't have to do some much messing around every time I want something different on a disk image. Tyan mobos boot off USB memory sticks ok, though, which is nice when you actually need to do something even if you have to go one machine at a time.) I did get an oops when I was memtest'ing (and therefore had mlock()ed) more than 3.5 out of 4GB of RAM, and then did a big rsync and formatdb. I don't know what to make of that. There was an actual oops on the screen, not just locked up thrashing. I think I saved it the kernel log, but I'm not going jump to the conclusion that oM is to blame for it. -- #define X(x,y) x##y Peter Cordes ; e-mail: X(peter@cor , des.ca) "The gods confound the man who first found out how to distinguish the hours! Confound him, too, who in this place set up a sundial, to cut and hack my day so wretchedly into small pieces!" -- Plautus, 200 BC
signature.asc
(application/pgp-signature, 351 B)
-----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.1 (GNU/Linux) iQC1AwUBQ2OkoQWkmhLkWuRTAQJIZwUAu+sglt+q+nd5jYI1T5+8zAADNy1JHUP4 oakBWpqqKsCWoz1TU7qfyrVGqnuK3CLXuVvLySV2wNDRRcAkzhX5wGFOpPkZwII3 MCz8/0tTvWtsyJZ7nt8uaV3DL3kMWg+xrK2n9CyMslQ6pt6EKYWeok0636DlNf/f +dTPoVLiN1qlwl8RmeGlO/UTXwO6J+O7hpQCAuFmbVx5NOMnypP86g== =+v9k -----END PGP SIGNATURE-----