RE: [myrinet] Re: GM-1.5, Linux-2.4.14, gm_board_info hangs
"Ritch, David" <[email protected]>
| Newsgroups | gmane.network.myrinet.general |
|---|---|
| Message-ID | <E2F90491BAC4D511AD8C00105AC5C9A007061E@HPTI_MAIN.hpti.com> |
I did a bit of work to try to clean up and implement all the suggestions
that were made. I discovered that I had compiled the kernel with 4GB memory
support, which I thought I had turned off. In the current configuration, it
is turned off. Of course, I also recompiled and reinstalled gm.
The system seemed to be OK for a while. Then, I ran a user code from node
1, and hit Ctrl-C. The process attached to the terminal apparently exited -
I got control of the terminal back. However, I noticed later that there
were several MPI jobs still hanging around. I tried to kill them, and they
ignored it. I used kill -9, and they promptly exited. I was a bit worried,
so I ran gm_board_info. It hung in disk wait.
I was still able ping node 2 from node 1 via myrinet. I logged into node 2,
and found that I could not ping node 1 via myrinet. In addition, the two
processes started for my job by MPI on node 2 were in disk wait. I tried
killing them without and with "-9", and they did not die.
Here is the ksymoops output on node 1 (called t0001);
[dbritch@t0001 ruc40]$ dmesg | ksymoops
ksymoops 2.4.0 on i686 2.4.14. Options used
-V (default)
-k /proc/ksyms (default)
-l /proc/modules (default)
-o /lib/modules/2.4.14/ (default)
-m /boot/System.map-2.4.14 (default)
Warning: You did not tell me where to find symbol information. I will
assume that the log matches the kernel and modules that are running
right now and I'll use the default options above for symbol resolution.
If the current kernel and/or modules do not match the log, you can get
more accurate output by telling me the kernel version and where to find
map, modules, ksyms etc. ksymoops -h explains the options.
Warning (compare_maps): mismatch on symbol GM_PAGE_LEN , gm says f891dc28,
/usr/local/gm-1.5/sbin/gm says f891bfa8. Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol _gm_lanai_globals_saved , gm
says f891dc2c, /usr/local/gm-1.5/sbin/gm says f891bfac. Ignoring
/usr/local/gm-1.5/sbin/gm entry
Warning (compare_maps): mismatch on symbol gm_in_intr , gm says f891dc20,
/usr/local/gm-1.5/sbin/gm says f891bfa0. Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol gm_linux_intr_lock , gm says
f891dc24, /usr/local/gm-1.5/sbin/gm says f891bfa4. Ignoring
/usr/local/gm-1.5/sbin/gm entry
cpu: 0, clocks: 991260, slice: 330420
cpu: 1, clocks: 991260, slice: 330420
Receiver lock-up bug exists -- enabling work-around.
invalid operand: 0000
CPU: 1
EIP: 0010:[<c012de1c>] Not tainted
Using defaults from ksymoops -t elf32-i386 -a i386
EFLAGS: 00010202
eax: 00000090 ebx: 00000000 ecx: c1000000 edx: 00000000
esi: c1b83e40 edi: f9add104 ebp: 00000000 esp: d93c3d38
ds: 0018 es: 0018 ss: 0018
Process hybcst (pid: 7350, stackpage=d93c3000)
Stack: f9a77088 f9a7708c f9a77090 f9a77090 00000000 c1b83e40 f9add104
d93c3d78
f88b7076 f6c1a000 2e0f9000 00000000 f6c1a0f4 f6c1a000 f5bfc010
00000000
d93c3dc0 f88aebb9 f9add0f8 00000002 40bbd000 f6c1a000 00000000
2e0f9000
Call Trace: [<f88b7076>] [<f88aebb9>] [<f88aefb0>] [<f88b5c19>] [<f88b55e6>]
[<f88b8598>] [<f88b8535>] [<c0134cde>] [<c0133b3d>] [<c01195dd>]
[<c0119ddf>]
[<c011f173>] [<c011f22d>] [<c0106dc4>] [<c01058fd>] [<c0114169>]
[<c0111ccc>]
[<c0106f54>]
Code: 0f 0b 83 66 18 eb b8 00 e0 ff ff 21 e0 f6 40 05 20 0f 85 31
>>EIP; c012de1c <__free_pages_ok+5c/200> <=====
Trace; f88b7076 <[gm]gm_arch_unlock_user_buffer_page+196/1a4>
Trace; f88aebb9 <[gm]gm_dereference_mapping+1ed/430>
Trace; f88aefb0 <[gm]gm_disable_port_DMAs+1b4/280>
Trace; f88b5c19 <[gm]gm_port_state_close+91/158>
Trace; f88b55e6 <[gm]gm_minor_free+5a/88>
Trace; f88b8598 <[gm]gm_linux_port_close+30/44>
Trace; f88b8535 <[gm]gm_linux_close+7d/b0>
Trace; c0134cde <fput+4e/f0>
Trace; c0133b3d <filp_close+8d/a0>
Trace; c01195dd <put_files_struct+4d/d0>
Trace; c0119ddf <do_exit+10f/240>
Trace; c011f173 <collect_signal+93/e0>
Trace; c011f22d <dequeue_signal+6d/b0>
Trace; c0106dc4 <do_signal+244/2b0>
Trace; c01058fd <__switch_to+3d/100>
Trace; c0114169 <schedule+3a9/5c0>
Trace; c0111ccc <smp_apic_timer_interrupt+ec/110>
Trace; c0106f54 <signal_return+14/18>
Code; c012de1c <__free_pages_ok+5c/200>
00000000 <_EIP>:
Code; c012de1c <__free_pages_ok+5c/200> <=====
0: 0f 0b ud2a <=====
Code; c012de1e <__free_pages_ok+5e/200>
2: 83 66 18 eb andl $0xffffffeb,0x18(%esi)
Code; c012de22 <__free_pages_ok+62/200>
6: b8 00 e0 ff ff mov $0xffffe000,%eax
Code; c012de27 <__free_pages_ok+67/200>
b: 21 e0 and %esp,%eax
Code; c012de29 <__free_pages_ok+69/200>
d: f6 40 05 20 testb $0x20,0x5(%eax)
Code; c012de2d <__free_pages_ok+6d/200>
11: 0f 85 31 00 00 00 jne 48 <_EIP+0x48> c012de64
<__free_pages_ok+a4/200>
5 warnings issued. Results may not be reliable.
At the same time that this problem occured on node 1, node 2 syslogged the
following:
Nov 16 06:52:17 t0002 kernel: invalid operand: 0000
Nov 16 06:52:17 t0002 kernel: CPU: 0
Nov 16 06:52:17 t0002 kernel: EIP: 0010:[__free_pages_ok+92/512] Not
tainted
Nov 16 06:52:17 t0002 kernel: EIP: 0010:[<c012de1c>] Not tainted
Nov 16 06:52:17 t0002 kernel: EFLAGS: 00010202
Nov 16 06:52:17 t0002 kernel: eax: 00000098 ebx: 00000000 ecx: c1000000
edx: 00000000
Nov 16 06:52:17 t0002 kernel: esi: c192c400 edi: f9af9104 ebp: 00000000
esp: e6631d38
Nov 16 06:52:17 t0002 kernel: ds: 0018 es: 0018 ss: 0018
Nov 16 06:52:17 t0002 kernel: Process hybcst (pid: 7590, stackpage=e6631000)
Nov 16 06:52:17 t0002 kernel: Stack: f9af6088 f9af608c f9af6090 f9af6090
00000000 c192c400 f9af9104 e6631d78
Nov 16 06:52:17 t0002 kernel: f88b3076 ea4f2000 24b10000 00000000
ea4f20f4 ea4f2000 e93a7010 00000000
Nov 16 06:52:17 t0002 kernel: e6631dc0 f88aabb9 f9af90f8 00000002
40bbd000 ea4f2000 00000000 24b10000
Nov 16 06:52:17 t0002 rsh(pam_unix)[7549]: session closed for user dbritch
Nov 16 06:52:17 t0002 rsh(pam_unix)[7550]: session closed for user dbritch
Nov 16 06:52:17 t0002 kernel: Call Trace: [<f88b3076>] [<f88aabb9>]
[<f88aafb0>] [<f88b1c19>] [<f88b15e6>]
Nov 16 06:52:17 t0002 kernel: [<f88b4598>] [<f88b4535>] [fput+78/240]
[filp_close+141/160] [put_files_struct+77/208] [do_exit+271/576]
Nov 16 06:52:17 t0002 kernel: [<f88b4598>] [<f88b4535>] [<c0134cde>]
[<c0133b3d>] [<c01195dd>] [<c0119ddf>]
Nov 16 06:52:17 t0002 kernel: [collect_signal+147/224]
[dequeue_signal+109/176] [do_signal+580/688] [do_softirq+123/224]
[do_IRQ+221/240] [signal_return+20/2
4]
Nov 16 06:52:17 t0002 kernel: [<c011f173>] [<c011f22d>] [<c0106dc4>]
[<c011b0ab>] [<c01089bd>] [<c0106f54>]
Nov 16 06:52:17 t0002 kernel:
Nov 16 06:52:17 t0002 kernel: Code: 0f 0b 83 66 18 eb b8 00 e0 ff ff 21 e0
f6 40 05 20 0f 85 31
After that, it began syslogging the following approximately every 10
seconds.
Nov 16 06:54:11 t0002 kernel: NETDEV WATCHDOG: myri0: transmit timed out
Nov 16 06:54:11 t0002 kernel: GM: WARNING:
Nov 16 06:54:11 t0002 kernel: GM: myri0: gmip_timeout called!
Node 1 is not logging this.
Here is the output from dmesg | ksymoops on node 2:
[dbritch@t0002 ~]$ dmesg | ksymoops
ksymoops 2.4.0 on i686 2.4.14. Options used
-V (default)
-k /proc/ksyms (default)
-l /proc/modules (default)
-o /lib/modules/2.4.14/ (default)
-m /boot/System.map-2.4.14 (default)
Warning: You did not tell me where to find symbol information. I will
assume that the log matches the kernel and modules that are running
right now and I'll use the default options above for symbol resolution.
If the current kernel and/or modules do not match the log, you can get
more accurate output by telling me the kernel version and where to find
map, modules, ksyms etc. ksymoops -h explains the options.
Warning (compare_maps): mismatch on symbol GM_PAGE_LEN , gm says f8919c28,
/usr/local/gm-1.5/sbin/gm says f8917fa8. Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol _gm_lanai_globals_saved , gm
says f8919c2c, /usr/local/gm-1.5/sbin/gm says f8917fac. Ignoring
/usr/local/gm-1.5/sbin/gm entry
Warning (compare_maps): mismatch on symbol gm_in_intr , gm says f8919c20,
/usr/local/gm-1.5/sbin/gm says f8917fa0. Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol gm_linux_intr_lock , gm says
f8919c24, /usr/local/gm-1.5/sbin/gm says f8917fa4. Ignoring
/usr/local/gm-1.5/sbin/gm entry
5 warnings issued. Results may not be reliable.
[dbritch@t0002 ~]$
Thank you for your help on this!
David
> -----Original Message-----
> From: Bob Felderman [mailto:[email protected]]
> Sent: Thursday, November 15, 2001 4:41 PM
> To: Susan Blackford
> Cc: Ritch, David; [email protected]; [email protected]
> Subject: Re: [myrinet] Re: GM-1.5, Linux-2.4.14, gm_board_info hangs
>
>
> You might also want to try
>
> dmesg | ksymoops
>
> to see if the backtrace it gives you provides any useful information.
>
>
>
> On Thu, 15 Nov 2001, Susan Blackford wrote:
> >Hi,
> >
> >This is a kernel "oops"/seg fault.
> >
> >Did you see any strange messages in the kernel log (dmesg) when you
> >GM_INSTALL?
> >
> >What about when you ran the mapper, did you see any strange messages?
> >
> >Did the gm_board_info hang happen on all of the nodes in
> your cluster,
> >or only a specific node?
> >
> >Would it be possible for us to have remote access to this machine and
> >be able to load/unload the GM module (root) and experiment?
> >
> >This is very strange.
> >
> >Susan
> >
> >
> >On Thursday 15 November 2001 02:48 pm, Ritch, David wrote:
> >> I'm a relative newcomer to myrinet, and I have a question.
> >>
> >> I'm attempting to run GM-1.5 on a cluster of
> dual-processor Pentium-4 nodes
> >> under RedHat-7.1. I compiled the 2.4.14 kernel with kgcc,
> and compiled gm
> >> with CC="kgcc -I/usr/src/linux/include", so as to get the
> include files
> >> associated with the correct kernel. I'm also running mpi-1.2.1..7,
> >> compiled with gcc.
> >>
> >> GM appears to work OK at first, but usually when I run a
> real application
> >> on it, a gm_board_info process is started by root, and
> goes into a disk
> >> hang (ps reports status of D). At about the same time,
> the following
> >> appears in /var/log/messages:
> >>
> >> Nov 13 17:42:35 t0001 kernel: invalid operand: 0000
> >> Nov 13 17:42:35 t0001 kernel: CPU: 0
> >> Nov 13 17:42:35 t0001 kernel: EIP:
> 0010:[__free_pages_ok+92/512] Not
> >> tainted
> >> Nov 13 17:42:35 t0001 kernel: EIP: 0010:[<c012ea1c>]
> Not tainted
> >> Nov 13 17:42:35 t0001 kernel: EFLAGS: 00010202
> >> Nov 13 17:42:35 t0001 kernel: eax: 00000890 ebx:
> 00000000 ecx: c1000000
> >> edx: 00000000
> >> Nov 13 17:42:35 t0001 kernel: esi: c1e3e540 edi:
> f8920c24 ebp: 00000000
> >> esp: d0c91d38
> >> Nov 13 17:42:35 t0001 kernel: ds: 0018 es: 0018 ss: 0018
> >> Nov 13 17:42:35 t0001 kernel: Process systest (pid: 21412,
> >> stackpage=d0c91000)
> >> Nov 13 17:42:35 t0001 kernel: Stack: f891fe08 f891fe0c
> f891fe10 f891fe10
> >> 00000000 c1e3e540 f8920c24 d0c91d78
> >> Nov 13 17:42:35 t0001 kernel: f88b4076 f7a82000
> 38f95000 00000000
> >> f7a820f4 f7a82000 d0ba0000 00000000
> >> Nov 13 17:42:35 t0001 kernel: d0c91dc0 f88abbb9
> f8920c18 00000004
> >> 080d1000 f7a82000 00000000 38f95000
> >>
> >> I'm not sure what is going on, but my myrinet does not appear to be
> >> working. Could someone help to enlighten me on this?
> >>
> >> Thanks,
> >>
> >> dbr
> >> --
> >> David B. Ritch
> >> High Performance Technologies, Inc.
> >> [email protected]
> >
> >----------------------------------------
> >Content-Type: text/html; charset="iso-8859-1"; name="Attachment: 1"
> >Content-Transfer-Encoding: quoted-printable
> >Content-Description:
> >----------------------------------------
> >
> >
>