RE: [myrinet] Re: GM-1.5, Linux-2.4.14, gm_board_info hangs

"Ritch, David" <[email protected]>
Newsgroups gmane.network.myrinet.general
Message-ID <E2F90491BAC4D511AD8C00105AC5C9A007061E@HPTI_MAIN.hpti.com>
I did a bit of work to try to clean up and implement all the suggestions
that were made.  I discovered that I had compiled the kernel with 4GB memory
support, which I thought I had turned off.  In the current configuration, it
is turned off.  Of course, I also recompiled and reinstalled gm.

The system seemed to be OK for a while.  Then, I ran a user code from node
1, and hit Ctrl-C.  The process attached to the terminal apparently exited -
I got control of the terminal back.  However, I noticed later that there
were several MPI jobs still hanging around.  I tried to kill them, and they
ignored it.  I used kill -9, and they promptly exited.  I was a bit worried,
so I ran gm_board_info.  It hung in disk wait.

I was still able ping node 2 from node 1 via myrinet.  I logged into node 2,
and found that I could not ping node 1 via myrinet.  In addition, the two
processes started for my job by MPI on node 2 were in disk wait.  I tried
killing them without and with "-9", and they did not die.

Here is the ksymoops output on node 1 (called t0001);

[dbritch@t0001 ruc40]$ dmesg | ksymoops
ksymoops 2.4.0 on i686 2.4.14.  Options used
     -V (default)
     -k /proc/ksyms (default)
     -l /proc/modules (default)
     -o /lib/modules/2.4.14/ (default)
     -m /boot/System.map-2.4.14 (default)

Warning: You did not tell me where to find symbol information.  I will
assume that the log matches the kernel and modules that are running
right now and I'll use the default options above for symbol resolution.
If the current kernel and/or modules do not match the log, you can get
more accurate output by telling me the kernel version and where to find
map, modules, ksyms etc.  ksymoops -h explains the options.

Warning (compare_maps): mismatch on symbol GM_PAGE_LEN  , gm says f891dc28,
/usr/local/gm-1.5/sbin/gm says f891bfa8.  Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol _gm_lanai_globals_saved  , gm
says f891dc2c, /usr/local/gm-1.5/sbin/gm says f891bfac.  Ignoring
/usr/local/gm-1.5/sbin/gm entry
Warning (compare_maps): mismatch on symbol gm_in_intr  , gm says f891dc20,
/usr/local/gm-1.5/sbin/gm says f891bfa0.  Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol gm_linux_intr_lock  , gm says
f891dc24, /usr/local/gm-1.5/sbin/gm says f891bfa4.  Ignoring
/usr/local/gm-1.5/sbin/gm entry
cpu: 0, clocks: 991260, slice: 330420
cpu: 1, clocks: 991260, slice: 330420
  Receiver lock-up bug exists -- enabling work-around.
invalid operand: 0000
CPU:    1
EIP:    0010:[<c012de1c>]    Not tainted
Using defaults from ksymoops -t elf32-i386 -a i386
EFLAGS: 00010202
eax: 00000090   ebx: 00000000   ecx: c1000000   edx: 00000000
esi: c1b83e40   edi: f9add104   ebp: 00000000   esp: d93c3d38
ds: 0018   es: 0018   ss: 0018
Process hybcst (pid: 7350, stackpage=d93c3000)
Stack: f9a77088 f9a7708c f9a77090 f9a77090 00000000 c1b83e40 f9add104
d93c3d78 
       f88b7076 f6c1a000 2e0f9000 00000000 f6c1a0f4 f6c1a000 f5bfc010
00000000 
       d93c3dc0 f88aebb9 f9add0f8 00000002 40bbd000 f6c1a000 00000000
2e0f9000 
Call Trace: [<f88b7076>] [<f88aebb9>] [<f88aefb0>] [<f88b5c19>] [<f88b55e6>]

   [<f88b8598>] [<f88b8535>] [<c0134cde>] [<c0133b3d>] [<c01195dd>]
[<c0119ddf>] 
   [<c011f173>] [<c011f22d>] [<c0106dc4>] [<c01058fd>] [<c0114169>]
[<c0111ccc>] 
   [<c0106f54>] 
Code: 0f 0b 83 66 18 eb b8 00 e0 ff ff 21 e0 f6 40 05 20 0f 85 31 

>>EIP; c012de1c <__free_pages_ok+5c/200>   <=====
Trace; f88b7076 <[gm]gm_arch_unlock_user_buffer_page+196/1a4>
Trace; f88aebb9 <[gm]gm_dereference_mapping+1ed/430>
Trace; f88aefb0 <[gm]gm_disable_port_DMAs+1b4/280>
Trace; f88b5c19 <[gm]gm_port_state_close+91/158>
Trace; f88b55e6 <[gm]gm_minor_free+5a/88>
Trace; f88b8598 <[gm]gm_linux_port_close+30/44>
Trace; f88b8535 <[gm]gm_linux_close+7d/b0>
Trace; c0134cde <fput+4e/f0>
Trace; c0133b3d <filp_close+8d/a0>
Trace; c01195dd <put_files_struct+4d/d0>
Trace; c0119ddf <do_exit+10f/240>
Trace; c011f173 <collect_signal+93/e0>
Trace; c011f22d <dequeue_signal+6d/b0>
Trace; c0106dc4 <do_signal+244/2b0>
Trace; c01058fd <__switch_to+3d/100>
Trace; c0114169 <schedule+3a9/5c0>
Trace; c0111ccc <smp_apic_timer_interrupt+ec/110>
Trace; c0106f54 <signal_return+14/18>
Code;  c012de1c <__free_pages_ok+5c/200>
00000000 <_EIP>:
Code;  c012de1c <__free_pages_ok+5c/200>   <=====
   0:   0f 0b                     ud2a      <=====
Code;  c012de1e <__free_pages_ok+5e/200>
   2:   83 66 18 eb               andl   $0xffffffeb,0x18(%esi)
Code;  c012de22 <__free_pages_ok+62/200>
   6:   b8 00 e0 ff ff            mov    $0xffffe000,%eax
Code;  c012de27 <__free_pages_ok+67/200>
   b:   21 e0                     and    %esp,%eax
Code;  c012de29 <__free_pages_ok+69/200>
   d:   f6 40 05 20               testb  $0x20,0x5(%eax)
Code;  c012de2d <__free_pages_ok+6d/200>
  11:   0f 85 31 00 00 00         jne    48 <_EIP+0x48> c012de64
<__free_pages_ok+a4/200>


5 warnings issued.  Results may not be reliable.


At the same time that this problem occured on node 1, node 2 syslogged the
following:

Nov 16 06:52:17 t0002 kernel: invalid operand: 0000
Nov 16 06:52:17 t0002 kernel: CPU:    0
Nov 16 06:52:17 t0002 kernel: EIP:    0010:[__free_pages_ok+92/512]    Not
tainted
Nov 16 06:52:17 t0002 kernel: EIP:    0010:[<c012de1c>]    Not tainted
Nov 16 06:52:17 t0002 kernel: EFLAGS: 00010202
Nov 16 06:52:17 t0002 kernel: eax: 00000098   ebx: 00000000   ecx: c1000000
edx: 00000000
Nov 16 06:52:17 t0002 kernel: esi: c192c400   edi: f9af9104   ebp: 00000000
esp: e6631d38
Nov 16 06:52:17 t0002 kernel: ds: 0018   es: 0018   ss: 0018
Nov 16 06:52:17 t0002 kernel: Process hybcst (pid: 7590, stackpage=e6631000)
Nov 16 06:52:17 t0002 kernel: Stack: f9af6088 f9af608c f9af6090 f9af6090
00000000 c192c400 f9af9104 e6631d78 
Nov 16 06:52:17 t0002 kernel:        f88b3076 ea4f2000 24b10000 00000000
ea4f20f4 ea4f2000 e93a7010 00000000 
Nov 16 06:52:17 t0002 kernel:        e6631dc0 f88aabb9 f9af90f8 00000002
40bbd000 ea4f2000 00000000 24b10000 
Nov 16 06:52:17 t0002 rsh(pam_unix)[7549]: session closed for user dbritch
Nov 16 06:52:17 t0002 rsh(pam_unix)[7550]: session closed for user dbritch
Nov 16 06:52:17 t0002 kernel: Call Trace: [<f88b3076>] [<f88aabb9>]
[<f88aafb0>] [<f88b1c19>] [<f88b15e6>] 
Nov 16 06:52:17 t0002 kernel:    [<f88b4598>] [<f88b4535>] [fput+78/240]
[filp_close+141/160] [put_files_struct+77/208] [do_exit+271/576] 
Nov 16 06:52:17 t0002 kernel:    [<f88b4598>] [<f88b4535>] [<c0134cde>]
[<c0133b3d>] [<c01195dd>] [<c0119ddf>] 
Nov 16 06:52:17 t0002 kernel:    [collect_signal+147/224]
[dequeue_signal+109/176] [do_signal+580/688] [do_softirq+123/224]
[do_IRQ+221/240] [signal_return+20/2
4] 
Nov 16 06:52:17 t0002 kernel:    [<c011f173>] [<c011f22d>] [<c0106dc4>]
[<c011b0ab>] [<c01089bd>] [<c0106f54>] 
Nov 16 06:52:17 t0002 kernel: 
Nov 16 06:52:17 t0002 kernel: Code: 0f 0b 83 66 18 eb b8 00 e0 ff ff 21 e0
f6 40 05 20 0f 85 31 

After that, it began syslogging the following approximately every 10
seconds.

Nov 16 06:54:11 t0002 kernel: NETDEV WATCHDOG: myri0: transmit timed out
Nov 16 06:54:11 t0002 kernel: GM: WARNING: 
Nov 16 06:54:11 t0002 kernel: GM: myri0: gmip_timeout called!

Node 1 is not logging this.

Here is the output from dmesg | ksymoops on node 2:

[dbritch@t0002 ~]$ dmesg | ksymoops
ksymoops 2.4.0 on i686 2.4.14.  Options used
     -V (default)
     -k /proc/ksyms (default)
     -l /proc/modules (default)
     -o /lib/modules/2.4.14/ (default)
     -m /boot/System.map-2.4.14 (default)

Warning: You did not tell me where to find symbol information.  I will
assume that the log matches the kernel and modules that are running
right now and I'll use the default options above for symbol resolution.
If the current kernel and/or modules do not match the log, you can get
more accurate output by telling me the kernel version and where to find
map, modules, ksyms etc.  ksymoops -h explains the options.

Warning (compare_maps): mismatch on symbol GM_PAGE_LEN  , gm says f8919c28,
/usr/local/gm-1.5/sbin/gm says f8917fa8.  Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol _gm_lanai_globals_saved  , gm
says f8919c2c, /usr/local/gm-1.5/sbin/gm says f8917fac.  Ignoring
/usr/local/gm-1.5/sbin/gm entry
Warning (compare_maps): mismatch on symbol gm_in_intr  , gm says f8919c20,
/usr/local/gm-1.5/sbin/gm says f8917fa0.  Ignoring /usr/local/gm-1.5/sbin/gm
entry
Warning (compare_maps): mismatch on symbol gm_linux_intr_lock  , gm says
f8919c24, /usr/local/gm-1.5/sbin/gm says f8917fa4.  Ignoring
/usr/local/gm-1.5/sbin/gm entry

5 warnings issued.  Results may not be reliable.
[dbritch@t0002 ~]$ 


Thank you for your help on this!

David

> -----Original Message-----
> From: Bob Felderman [mailto:[email protected]]
> Sent: Thursday, November 15, 2001 4:41 PM
> To: Susan Blackford
> Cc: Ritch, David; [email protected]; [email protected]
> Subject: Re: [myrinet] Re: GM-1.5, Linux-2.4.14, gm_board_info hangs
> 
> 
> You might also want to try
> 
> 	dmesg | ksymoops
> 
> to see if the backtrace it gives you provides any useful information.
> 
> 
> 
> On Thu, 15 Nov 2001, Susan Blackford wrote:
> >Hi,
> >
> >This is a kernel "oops"/seg fault.
> >
> >Did you see any strange messages in the kernel log (dmesg) when you
> >GM_INSTALL?
> >
> >What about when you ran the mapper, did you see any strange messages?
> >
> >Did the gm_board_info hang happen on all of the nodes in 
> your cluster,
> >or only a specific node?
> >
> >Would it be possible for us to have remote access to this machine and
> >be able to load/unload the GM module (root) and experiment?
> >
> >This is very strange.
> >
> >Susan
> >
> >
> >On Thursday 15 November 2001 02:48 pm, Ritch, David wrote:
> >> I'm a relative newcomer to myrinet, and I have a question.
> >>
> >> I'm attempting to run GM-1.5 on a cluster of 
> dual-processor Pentium-4 nodes
> >> under RedHat-7.1.  I compiled the 2.4.14 kernel with kgcc, 
> and compiled gm
> >> with CC="kgcc -I/usr/src/linux/include", so as to get the 
> include files
> >> associated with the correct kernel.  I'm also running mpi-1.2.1..7,
> >> compiled with gcc.
> >>
> >> GM appears to work OK at first, but usually when I run a 
> real application
> >> on it, a gm_board_info process is started by root, and 
> goes into a disk
> >> hang (ps reports status of D).  At about the same time, 
> the following
> >> appears in /var/log/messages:
> >>
> >> Nov 13 17:42:35 t0001 kernel: invalid operand: 0000
> >> Nov 13 17:42:35 t0001 kernel: CPU:    0
> >> Nov 13 17:42:35 t0001 kernel: EIP:    
> 0010:[__free_pages_ok+92/512]    Not
> >> tainted
> >> Nov 13 17:42:35 t0001 kernel: EIP:    0010:[<c012ea1c>]    
> Not tainted
> >> Nov 13 17:42:35 t0001 kernel: EFLAGS: 00010202
> >> Nov 13 17:42:35 t0001 kernel: eax: 00000890   ebx: 
> 00000000   ecx: c1000000
> >> edx: 00000000
> >> Nov 13 17:42:35 t0001 kernel: esi: c1e3e540   edi: 
> f8920c24   ebp: 00000000
> >> esp: d0c91d38
> >> Nov 13 17:42:35 t0001 kernel: ds: 0018   es: 0018   ss: 0018
> >> Nov 13 17:42:35 t0001 kernel: Process systest (pid: 21412,
> >> stackpage=d0c91000)
> >> Nov 13 17:42:35 t0001 kernel: Stack: f891fe08 f891fe0c 
> f891fe10 f891fe10
> >> 00000000 c1e3e540 f8920c24 d0c91d78
> >> Nov 13 17:42:35 t0001 kernel:        f88b4076 f7a82000 
> 38f95000 00000000
> >> f7a820f4 f7a82000 d0ba0000 00000000
> >> Nov 13 17:42:35 t0001 kernel:        d0c91dc0 f88abbb9 
> f8920c18 00000004
> >> 080d1000 f7a82000 00000000 38f95000
> >>
> >> I'm not sure what is going on, but my myrinet does not appear to be
> >> working. Could someone help to enlighten me on this?
> >>
> >> Thanks,
> >>
> >> dbr
> >> --
> >> David B. Ritch
> >> High Performance Technologies, Inc.
> >> [email protected]
> >
> >----------------------------------------
> >Content-Type: text/html; charset="iso-8859-1"; name="Attachment: 1"
> >Content-Transfer-Encoding: quoted-printable
> >Content-Description:
> >----------------------------------------
> >
> >
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.