Re: gm_board_info only shows the node itself

[email protected] Tue, 13 Jan 2004 09:37:57 -0500
Newsgroups gmane.network.myrinet.general
Message-ID <[email protected]>

> But I couldn't make mpirun work because of the message below.

> >[root@hgt001 basic]# mpirun -np 4 ./cpi
> >Process 2 of 4 on hgt001.linux.coc
> >[3]: alloc failed, not enough memory (Fatal Error)
> >Context: <(gmpi_init) gmpi_dma_alloc: dma send buffers>
> >[1]: alloc failed, not enough memory (Fatal Error)
> >Context: <(gmpi_init) gmpi_dma_alloc: dma send buffers>

> I confirmed the memory is not full...

This resembles what I saw when I first started using mpich v2.5.10
and didn't realize that rather than using ~/.gmpi/conf as a machines
file by default,  mpich now uses

whatever_your_mpi_path/share/machines.ch_gm.LINUX

as the default machines file.

Were you already aware of this?



Gary







Atsuko Miyashita <[email protected]> on 01/13/2004 01:27:15 AM

To:   "Dr. Markus Fischer" <[email protected]>
cc:   [email protected] (bcc: Gary Hannon/CSP)
Subject:  Re: [Myrinet] gm_board_info only shows the node itself




Hi Markus,

Thank you for the reply.
These are the output of "gm start" and "gm stop".

>[root@hgt001 root]# service gm stop
>Stopping gm... /etc/init.d/gm: kill: (13599) - No such process
>done.
>[root@hgt001 root]# service gm start
>Starting gm... active mapper... done.

And on the last lines of dmesg, I saw these.

>GM: WARNING:
drivers/linux/gm/gm_arch.c:1368:gm_arch_lock_user_buffer_page():ker
>nel
>GM: trying to register a page with count 0

Sometimes I also see this message in dmesg too.

>GM: Application closed file descriptor while mappings still alive: port
destruct
> delayed

If I should attach some other message, please let me know.
Thank you very much in advance.

(Additional odd thing)
BTW... while I did gm stop/gm start so many times, I suddenly see that the
mapper
process properly worked.

>Port: Status  PID
>   0:   BUSY 14644  (this process [gm_board_info])
>   1:   BUSY 14558
>Route table for this node follows:
>gmID MAC Address                                 gmName Route
>---- ----------------- --------------------------------
---------------------
>   1 00:60:dd:7f:61:be                           hgt001 (this node)
>   2 00:60:dd:7f:61:fc                           hgt002 84 (mapper)

But I couldn't make mpirun work because of the message below.

>[root@hgt001 basic]# mpirun -np 4 ./cpi
>Process 2 of 4 on hgt001.linux.coc
>[3]: alloc failed, not enough memory (Fatal Error)
>Context: <(gmpi_init) gmpi_dma_alloc: dma send buffers>
>[1]: alloc failed, not enough memory (Fatal Error)
>Context: <(gmpi_init) gmpi_dma_alloc: dma send buffers>

I confirmed the memory is not full...
After I rebooted the system, again I see the mapper process doesn't work
anymore.

>[root@hgt001 basic]# cat /proc/meminfo
>        total:    used:    free:  shared: buffers:  cached:
>Mem:  4219338752 400519168 3818819584        0 63873024 178229248
>Swap: 2089209856        0 2089209856
>MemTotal:      4120448 kB
>MemFree:       3729316 kB
>MemShared:           0 kB
(Cont...)

Do you have any idea why my environment is such unstable ?????


----------------------------------------------------
ATSUKO MIYASHITA
Linux Support Center, SWSC, Technical Support - IBM Japan
Tel : 03-3808-8993 ( ext. 1712-8993 )  /  Fax : 03-3664-4893
e-mail : ATSUKOM

@jp.ibm.com
** Moved to Hakozaki - mind that tel # has changed ! **




"Dr. Markus Fischer" <[email protected]>
2004/01/11 07:08


        To:     Atsuko Miyashita/Japan/IBM@IBMJP
        cc:     [email protected]
        Subject:        Re: [Myrinet] gm_board_info only shows the node
itself



I guess you should provide the output of

/etc/init.d/gm start

and provide the output of the last lines of dmesg

Markus

Atsuko Miyashita wrote:

>
> Hi,
>
> Thank you for the reply.
> Yes, I run the mapper. I did "gm stop" and "gm start" many times.
>
> Actually, I noticed mapper process hasn't been started - when I
> executed "gm stop", it says there's nothing to kill.
> Attached is my /tmp/gm_log.root, a log of "gm start".
> According to this log, it seems that "gm start" tried to run the
> mapper and writes the pid into /var/run/gm_mapper/pid.0 file, but
> mapper died soon.
>
> I powered off the Myrinet switch and leave it for a while(15min or
> so), then power it on again, but it remains the same.
>
> I am at a loss what to do next... has anyone had similar experience ?
>
> Sat Jan 10 17:28:12 JST 2004
> /etc/init.d/gm start
> root@hpc001> test xgm != x
> Starting gm...
> root@hpc001> /opt/gm/bin/gm_board_info
> GM build ID is "2.0.6_Linux_rc20030908173009PDT
> root@hpc001:/usr/local/gm-2.0.6_Linux Wed Dec 24 16:34:42 JST 2003."
> No boards found
> root@hpc001> rmmod --verbose gm
> Checking gm for persistent data
> rmmod: module gm is not loaded
> root@hpc001> insmod --verbose /lib/modules/2.4.18-3bigmem/gm/gm.o
> Warning: loading /lib/modules/2.4.18-3bigmem/gm/gm.o will taint the
> kernel: non-GPL license - Myricom
> Using /lib/modules/2.4.18-3bigmem/gm/gm.o
> Symbol version prefix 'smp_'
> root@hpc001> /opt/gm/bin/gm_board_info
> GM build ID is "2.0.6_Linux_rc20030908173009PDT
> root@hpc001:/usr/local/gm-2.0.6_Linux Wed Dec 24 16:34:42 JST 2003."
>
>
> Board number 0:
> lanai_clockval = 0x082082a0
> lanai_cpu_version = 0x0900 (LANai9.0)
> lanai_board_id = 00:60:dd:7f:61:be
> lanai_sram_size = 0x00200000 (2048K bytes)
> max_lanai_speed = 134 MHz
> product_code = 109
> serial_number = 89434
> (should be labeled: "M3S-PCI64B-2-89434")
> LANai time is 0x000072c15 ticks, or about 0 minutes since reset.
> Mapper is 00:00:00:00:00:00.
> Map version is 0.
> 0 hosts.
> Network is NOT fully configured.
> This node is "hpc001"
> Board has room for 16 ports, 1600 nodes/routes, 16384 cache entries
> Port token cnt: send=61, recv=254
> Port: Status PID
> 0: BUSY 2742 (this process [gm_board_info])
> Route table for this node follows:
> gmID MAC Address gmName Route
> ---- ----------------- --------------------------------
> ---------------------
> 1 00:60:dd:7f:61:be hpc001 (this node)
> root@hpc001> /opt/gm/bin/gm_set_name --board=0 --host-name=hpc001
> Could not open board 0, port 1 : out of memory
> root@hpc001> /opt/gm/sbin/gm_mapper --unit=0
> --daemon-pid-file=/var/run/gm_mapper/pid.0
> --verbose-file=/var/run/gm_mapper/verbose.0
> --map-file=/var/run/gm_mapper/map.0
> Storing daemon PID in file "/var/run/gm_mapper/pid.0"
> done.
>
> Thank you very much in advance !
>
>
> ----------------------------------------------------
> ATSUKO MIYASHITA
> Linux Support Center, SWSC, Technical Support - IBM Japan
> Tel : 03-3808-8993 ( ext. 1712-8993 ) / Fax : 03-3664-4893
> e-mail : ATSUKOM@jp.ibm.com
> ** Moved to Hakozaki - mind that tel # has changed ! **
>
>
>
>                "Dr. Markus Fischer" <[email protected]>
>
> 2004/01/10 05:32
>
>
> To: Atsuko Miyashita/Japan/IBM@IBMJP
> cc: [email protected]
> Subject: Re: [Myrinet] gm_board_info only shows the node itself
>
>
>
> Have you run the mapper ?
> (active mappers will shop up using port 1)
>
> /etc/init.d/gm start
>
> will start mappers.
>
> Markus
>
> Atsuko Miyashita wrote:
>
> >
> > Hi,
> >
> > I've got a problem that GM mapper only sees the node itself, and
> > doesn't communicate each other.
> > I appreciate if anybody has a clue/suggestion to this... Thank you in
> > advance !
> >
> > I am currently using GM 2.0.6 and Myrinet "B" card with serial cable.
> > I have two nodes(hpc001/hpc002) with Linux 2.4.18-3bigmem(RedHat7.3)
> > installed.
> > I know the combination of driver/card is not the good one, but this is
> > a test environment for some software.
> >
> > I noticed that I couldn't use MPICH-GM properly, and this was probably
> > because of driver layer.
> > gm_board_info only shows the node itself, and doesn't show the another
> > node.
> > I did this several times, but the situation remains the same.
> > service gm stop
> > service gm start
> >
> > Also, I changed the physical port(of the switch), but it didn't solve
> > the problem.
> >
> > The odd thing is that I could use the environment without any problem
> > for several days.
> > I tested the MPICH-GM with its sample program, cpi, and saw it works
> > without problem.
> > I left the machines for a couple of days and then returned, and saw
> > this has happend.
> >
> > Does anyone have any idea or suggestion, at least how to diag this ?
> >
> > These are the output of gm_board_info on the each nodes.
> >
> > [root@hpc001 root]# gm_board_info
> > GM build ID is "2.0.6_Linux_rc20030908173009PDT
> > root@hpc001:/usr/local/gm-2.0.6_
> > Linux Wed Dec 24 16:34:42 JST 2003."
> > Board number 0:
> > lanai_clockval = 0x082082a0
> > lanai_cpu_version = 0x0900 (LANai9.0)
> > lanai_board_id = 00:60:dd:7f:61:be
> > lanai_sram_size = 0x00200000 (2048K bytes)
> > max_lanai_speed = 134 MHz
> > product_code = 109
> > serial_number = 89434
> > (should be labeled: "M3S-PCI64B-2-89434")
> > LANai time is 0x058a585f0 ticks, or about 11 minutes since reset.
> > Mapper is 00:00:00:00:00:00.
> > Map version is 0.
> > 0 hosts.
> > Network is NOT fully configured.
> > This node is "hpc001"
> > Board has room for 16 ports, 1600 nodes/routes, 16384 cache entries
> > Port token cnt: send=61, recv=254
> > Port: Status PID
> > 0: BUSY 1503 (this process [gm_board_info])
> > Route table for this node follows:
> > gmID MAC Address gmName Route
> > ---- ----------------- --------------------------------
> > ---------------------
> > 1 00:60:dd:7f:61:be hpc001 (this node)
> >
> > [root@hpc002 root]# gm_board_info
> > GM build ID is "2.0.6_Linux_rc20030908173009PDT
> > root@hpc002:/usr/local/gm-2.0.6_
> > Linux Wed Dec 24 17:48:24 JST 2003."
> > Board number 0:
> > lanai_clockval = 0x082082a0
> > lanai_cpu_version = 0x0900 (LANai9.0)
> > lanai_board_id = 00:60:dd:7f:61:fc
> > lanai_sram_size = 0x00200000 (2048K bytes)
> > max_lanai_speed = 134 MHz
> > product_code = 109
> > serial_number = 89372
> > (should be labeled: "M3S-PCI64B-2-89372")
> > LANai time is 0x08e6f211a ticks, or about 17 minutes since reset.
> > Mapper is 00:00:00:00:00:00.
> > Map version is 0.
> > 0 hosts.
> > Network is NOT fully configured.
> > This node is "hpc002"
> > Board has room for 16 ports, 1600 nodes/routes, 16384 cache entries
> > Port token cnt: send=61, recv=254
> > Port: Status PID
> > 0: BUSY 1692 (this process [gm_board_info])
> > Route table for this node follows:
> > gmID MAC Address gmName Route
> > ---- ----------------- --------------------------------
> > ---------------------
> > 1 00:60:dd:7f:61:fc hpc002 (this node)
> >
> > Thank you in advance for your help !!!
> >
> > ----------------------------------------------------
> > ATSUKO MIYASHITA
> > Linux Support Center, SWSC, Technical Support - IBM Japan
> > Tel : 03-3808-8993 ( ext. 1712-8993 ) / Fax : 03-3664-4893
> > e-mail : ATSUKOM@jp.ibm.com
> > ** Moved to Hakozaki - mind that tel # has changed ! **
> >
>
>------------------------------------------------------------------------
> >
> >_______________________________________________
> >Myrinet mailing list
> >[email protected]
> >http://email.osc.edu/mailman/listinfo/myrinet
> >
> >
>
>
>
>



_______________________________________________
Myrinet mailing list
[email protected]
http://email.osc.edu/mailman/listinfo/myrinet

_______________________________________________
Myrinet mailing list
[email protected]
http://email.osc.edu/mailman/listinfo/myrinet
att1.htm (text/html, 16.1 KB)
<br><font size=2 face="sans-serif">Hi Markus,</font>
<br>
<br><font size=2 face="sans-serif">Thank you for the reply. </font>
<br><font size=2 face="sans-serif">These are the output of &quot;gm start&quot; and &quot;gm stop&quot;.</font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;[root@hgt001 root]# service gm stop</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Stopping gm... /etc/init.d/gm: kill: (13599) - No such process</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;done.</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;[root@hgt001 root]# service gm start</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Starting gm... active mapper... done.</font>
<br>
<br><font size=2 face="sans-serif">And on the last lines of dmesg, I saw these.</font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;GM: WARNING: drivers/linux/gm/gm_arch.c:1368:gm_arch_lock_user_buffer_page():ker</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;nel</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;GM: trying to register a page with count 0</font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">Sometimes I also see this message in dmesg too. </font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;GM: Application closed file descriptor while mappings still alive: port destruct</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt; delayed</font>
<br>
<br><font size=2 face="sans-serif">If I should attach some other message, please let me know. </font>
<br><font size=2 face="sans-serif">Thank you very much in advance.</font>
<br>
<br><font size=2 face="sans-serif">(Additional odd thing)</font>
<br><font size=2 face="sans-serif">BTW... while I did gm stop/gm start so many times, I suddenly see that the mapper </font>
<br><font size=2 face="sans-serif">process properly worked. </font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Port: Status &nbsp;PID</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt; &nbsp; 0: &nbsp; BUSY 14644 &nbsp;(this process [gm_board_info])</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt; &nbsp; 1: &nbsp; BUSY 14558</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Route table for this node follows:</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;gmID MAC Address &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; gmName Route</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;---- ----------------- -------------------------------- ---------------------</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt; &nbsp; 1 00:60:dd:7f:61:be &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; hgt001 (this node)</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt; &nbsp; 2 00:60:dd:7f:61:fc &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; hgt002 84 (mapper)</font>
<br>
<br><font size=2 face="sans-serif">But I couldn't make mpirun work because of the message below.</font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;[root@hgt001 basic]# mpirun -np 4 ./cpi</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Process 2 of 4 on hgt001.linux.coc</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;[3]: alloc failed, not enough memory (Fatal Error)</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Context: &lt;(gmpi_init) gmpi_dma_alloc: dma send buffers&gt;</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;[1]: alloc failed, not enough memory (Fatal Error)</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Context: &lt;(gmpi_init) gmpi_dma_alloc: dma send buffers&gt;</font>
<br>
<br><font size=2 face="sans-serif">I confirmed the memory is not full... </font>
<br><font size=2 face="sans-serif">After I rebooted the system, again I see the mapper process doesn't work anymore.</font>
<br>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;[root@hgt001 basic]# cat /proc/meminfo</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt; &nbsp; &nbsp; &nbsp; &nbsp;total: &nbsp; &nbsp;used: &nbsp; &nbsp;free: &nbsp;shared: buffers: &nbsp;cached:</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Mem: &nbsp;4219338752 400519168 3818819584 &nbsp; &nbsp; &nbsp; &nbsp;0 63873024 178229248</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;Swap: 2089209856 &nbsp; &nbsp; &nbsp; &nbsp;0 2089209856</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;MemTotal: &nbsp; &nbsp; &nbsp;4120448 kB</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;MemFree: &nbsp; &nbsp; &nbsp; 3729316 kB</font>
<br><font size=2 face="$B#M#S(B $B#P%4%7%C%/(B">&gt;MemShared: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 0 kB</font>
<br><font size=2 face="sans-serif">(Cont...)</font>
<br>
<br><font size=2 face="sans-serif">Do you have any idea why my environment is such unstable ????? </font>
<br>
<br><font size=2 face="sans-serif"><br>
----------------------------------------------------<br>
ATSUKO MIYASHITA<br>
Linux Support Center, SWSC, Technical Support - IBM Japan<br>
Tel : 03-3808-8993 ( ext. 1712-8993 ) &nbsp;/ &nbsp;Fax : 03-3664-4893<br>
e-mail : ATSUKOM$B!w(Bjp.ibm.com<br>
** Moved to Hakozaki - mind that tel # has changed ! ** </font>
<br>
<br>
<br>
<table width=100%>
<tr valign=top>
<td>
<td><font size=1 face="sans-serif"><b>&quot;Dr. Markus Fischer&quot; &lt;[email protected]&gt;</b></font>
<p><font size=1 face="sans-serif">2004/01/11 07:08</font>
<br>
<td><font size=1 face="Arial">&nbsp; &nbsp; &nbsp; &nbsp; </font>
<br><font size=1 face="sans-serif">&nbsp; &nbsp; &nbsp; &nbsp; To: &nbsp; &nbsp; &nbsp; &nbsp;Atsuko Miyashita/Japan/IBM@IBMJP</font>
<br><font size=1 face="sans-serif">&nbsp; &nbsp; &nbsp; &nbsp; cc: &nbsp; &nbsp; &nbsp; &nbsp;[email protected]</font>
<br><font size=1 face="sans-serif">&nbsp; &nbsp; &nbsp; &nbsp; Subject: &nbsp; &nbsp; &nbsp; &nbsp;Re: [Myrinet] gm_board_info only shows the node itself</font>
<br>
<br><font size=1 face="Arial">&nbsp; &nbsp; &nbsp; &nbsp;</font></table>
<br>
<br><font size=2><tt>I guess you should provide the output of<br>
<br>
/etc/init.d/gm start<br>
<br>
and provide the output of the last lines of dmesg<br>
<br>
Markus<br>
<br>
Atsuko Miyashita wrote:<br>
<br>
&gt;<br>
&gt; Hi,<br>
&gt;<br>
&gt; Thank you for the reply.<br>
&gt; Yes, I run the mapper. I did &quot;gm stop&quot; and &quot;gm start&quot; many times.<br>
&gt;<br>
&gt; Actually, I noticed mapper process hasn't been started - when I<br>
&gt; executed &quot;gm stop&quot;, it says there's nothing to kill.<br>
&gt; Attached is my /tmp/gm_log.root, a log of &quot;gm start&quot;.<br>
&gt; According to this log, it seems that &quot;gm start&quot; tried to run the<br>
&gt; mapper and writes the pid into /var/run/gm_mapper/pid.0 file, but<br>
&gt; mapper died soon.<br>
&gt;<br>
&gt; I powered off the Myrinet switch and leave it for a while(15min or<br>
&gt; so), then power it on again, but it remains the same.<br>
&gt;<br>
&gt; I am at a loss what to do next... has anyone had similar experience ?<br>
&gt;<br>
&gt; Sat Jan 10 17:28:12 JST 2004<br>
&gt; /etc/init.d/gm start<br>
&gt; root@hpc001&gt; test xgm != x<br>
&gt; Starting gm...<br>
&gt; root@hpc001&gt; /opt/gm/bin/gm_board_info<br>
&gt; GM build ID is &quot;2.0.6_Linux_rc20030908173009PDT<br>
&gt; root@hpc001:/usr/local/gm-2.0.6_Linux Wed Dec 24 16:34:42 JST 2003.&quot;<br>
&gt; No boards found<br>
&gt; root@hpc001&gt; rmmod --verbose gm<br>
&gt; Checking gm for persistent data<br>
&gt; rmmod: module gm is not loaded<br>
&gt; root@hpc001&gt; insmod --verbose /lib/modules/2.4.18-3bigmem/gm/gm.o<br>
&gt; Warning: loading /lib/modules/2.4.18-3bigmem/gm/gm.o will taint the<br>
&gt; kernel: non-GPL license - Myricom<br>
&gt; Using /lib/modules/2.4.18-3bigmem/gm/gm.o<br>
&gt; Symbol version prefix 'smp_'<br>
&gt; root@hpc001&gt; /opt/gm/bin/gm_board_info<br>
&gt; GM build ID is &quot;2.0.6_Linux_rc20030908173009PDT<br>
&gt; root@hpc001:/usr/local/gm-2.0.6_Linux Wed Dec 24 16:34:42 JST 2003.&quot;<br>
&gt;<br>
&gt;<br>
&gt; Board number 0:<br>
&gt; lanai_clockval = 0x082082a0<br>
&gt; lanai_cpu_version = 0x0900 (LANai9.0)<br>
&gt; lanai_board_id = 00:60:dd:7f:61:be<br>
&gt; lanai_sram_size = 0x00200000 (2048K bytes)<br>
&gt; max_lanai_speed = 134 MHz<br>
&gt; product_code = 109<br>
&gt; serial_number = 89434<br>
&gt; (should be labeled: &quot;M3S-PCI64B-2-89434&quot;)<br>
&gt; LANai time is 0x000072c15 ticks, or about 0 minutes since reset.<br>
&gt; Mapper is 00:00:00:00:00:00.<br>
&gt; Map version is 0.<br>
&gt; 0 hosts.<br>
&gt; Network is NOT fully configured.<br>
&gt; This node is &quot;hpc001&quot;<br>
&gt; Board has room for 16 ports, 1600 nodes/routes, 16384 cache entries<br>
&gt; Port token cnt: send=61, recv=254<br>
&gt; Port: Status PID<br>
&gt; 0: BUSY 2742 (this process [gm_board_info])<br>
&gt; Route table for this node follows:<br>
&gt; gmID MAC Address gmName Route<br>
&gt; ---- ----------------- --------------------------------<br>
&gt; ---------------------<br>
&gt; 1 00:60:dd:7f:61:be hpc001 (this node)<br>
&gt; root@hpc001&gt; /opt/gm/bin/gm_set_name --board=0 --host-name=hpc001<br>
&gt; Could not open board 0, port 1 : out of memory<br>
&gt; root@hpc001&gt; /opt/gm/sbin/gm_mapper --unit=0<br>
&gt; --daemon-pid-file=/var/run/gm_mapper/pid.0<br>
&gt; --verbose-file=/var/run/gm_mapper/verbose.0<br>
&gt; --map-file=/var/run/gm_mapper/map.0<br>
&gt; Storing daemon PID in file &quot;/var/run/gm_mapper/pid.0&quot;<br>
&gt; done.<br>
&gt;<br>
&gt; Thank you very much in advance !<br>
&gt;<br>
&gt;<br>
&gt; ----------------------------------------------------<br>
&gt; ATSUKO MIYASHITA<br>
&gt; Linux Support Center, SWSC, Technical Support - IBM Japan<br>
&gt; Tel : 03-3808-8993 ( ext. 1712-8993 ) / Fax : 03-3664-4893<br>
&gt; e-mail : ATSUKOM$B!w(Bjp.ibm.com<br>
&gt; ** Moved to Hakozaki - mind that tel # has changed ! **<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&quot;Dr. Markus Fischer&quot; &lt;[email protected]&gt;<br>
&gt;<br>
&gt; 2004/01/10 05:32<br>
&gt;<br>
&gt; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<br>
&gt; To: Atsuko Miyashita/Japan/IBM@IBMJP<br>
&gt; cc: [email protected]<br>
&gt; Subject: Re: [Myrinet] gm_board_info only shows the node itself<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt; Have you run the mapper ?<br>
&gt; (active mappers will shop up using port 1)<br>
&gt;<br>
&gt; /etc/init.d/gm start<br>
&gt;<br>
&gt; will start mappers.<br>
&gt;<br>
&gt; Markus<br>
&gt;<br>
&gt; Atsuko Miyashita wrote:<br>
&gt;<br>
&gt; &gt;<br>
&gt; &gt; Hi,<br>
&gt; &gt;<br>
&gt; &gt; I've got a problem that GM mapper only sees the node itself, and<br>
&gt; &gt; doesn't communicate each other.<br>
&gt; &gt; I appreciate if anybody has a clue/suggestion to this... Thank you in<br>
&gt; &gt; advance !</tt></font>
<br><font size=2><tt>&gt; &gt;<br>
&gt; &gt; I am currently using GM 2.0.6 and Myrinet &quot;B&quot; card with serial cable.<br>
&gt; &gt; I have two nodes(hpc001/hpc002) with Linux 2.4.18-3bigmem(RedHat7.3)<br>
&gt; &gt; installed.<br>
&gt; &gt; I know the combination of driver/card is not the good one, but this is<br>
&gt; &gt; a test environment for some software.<br>
&gt; &gt;<br>
&gt; &gt; I noticed that I couldn't use MPICH-GM properly, and this was probably<br>
&gt; &gt; because of driver layer.<br>
&gt; &gt; gm_board_info only shows the node itself, and doesn't show the another<br>
&gt; &gt; node.<br>
&gt; &gt; I did this several times, but the situation remains the same.<br>
&gt; &gt; service gm stop<br>
&gt; &gt; service gm start<br>
&gt; &gt;<br>
&gt; &gt; Also, I changed the physical port(of the switch), but it didn't solve<br>
&gt; &gt; the problem.<br>
&gt; &gt;<br>
&gt; &gt; The odd thing is that I could use the environment without any problem<br>
&gt; &gt; for several days.<br>
&gt; &gt; I tested the MPICH-GM with its sample program, cpi, and saw it works<br>
&gt; &gt; without problem.<br>
&gt; &gt; I left the machines for a couple of days and then returned, and saw<br>
&gt; &gt; this has happend.<br>
&gt; &gt;<br>
&gt; &gt; Does anyone have any idea or suggestion, at least how to diag this ?<br>
&gt; &gt;<br>
&gt; &gt; These are the output of gm_board_info on the each nodes.<br>
&gt; &gt;<br>
&gt; &gt; [root@hpc001 root]# gm_board_info<br>
&gt; &gt; GM build ID is &quot;2.0.6_Linux_rc20030908173009PDT<br>
&gt; &gt; root@hpc001:/usr/local/gm-2.0.6_<br>
&gt; &gt; Linux Wed Dec 24 16:34:42 JST 2003.&quot;<br>
&gt; &gt; Board number 0:<br>
&gt; &gt; lanai_clockval = 0x082082a0<br>
&gt; &gt; lanai_cpu_version = 0x0900 (LANai9.0)<br>
&gt; &gt; lanai_board_id = 00:60:dd:7f:61:be<br>
&gt; &gt; lanai_sram_size = 0x00200000 (2048K bytes)<br>
&gt; &gt; max_lanai_speed = 134 MHz<br>
&gt; &gt; product_code = 109<br>
&gt; &gt; serial_number = 89434<br>
&gt; &gt; (should be labeled: &quot;M3S-PCI64B-2-89434&quot;)<br>
&gt; &gt; LANai time is 0x058a585f0 ticks, or about 11 minutes since reset.<br>
&gt; &gt; Mapper is 00:00:00:00:00:00.<br>
&gt; &gt; Map version is 0.<br>
&gt; &gt; 0 hosts.<br>
&gt; &gt; Network is NOT fully configured.<br>
&gt; &gt; This node is &quot;hpc001&quot;<br>
&gt; &gt; Board has room for 16 ports, 1600 nodes/routes, 16384 cache entries<br>
&gt; &gt; Port token cnt: send=61, recv=254<br>
&gt; &gt; Port: Status PID<br>
&gt; &gt; 0: BUSY 1503 (this process [gm_board_info])<br>
&gt; &gt; Route table for this node follows:<br>
&gt; &gt; gmID MAC Address gmName Route<br>
&gt; &gt; ---- ----------------- --------------------------------<br>
&gt; &gt; ---------------------<br>
&gt; &gt; 1 00:60:dd:7f:61:be hpc001 (this node)<br>
&gt; &gt;<br>
&gt; &gt; [root@hpc002 root]# gm_board_info<br>
&gt; &gt; GM build ID is &quot;2.0.6_Linux_rc20030908173009PDT<br>
&gt; &gt; root@hpc002:/usr/local/gm-2.0.6_<br>
&gt; &gt; Linux Wed Dec 24 17:48:24 JST 2003.&quot;<br>
&gt; &gt; Board number 0:<br>
&gt; &gt; lanai_clockval = 0x082082a0<br>
&gt; &gt; lanai_cpu_version = 0x0900 (LANai9.0)<br>
&gt; &gt; lanai_board_id = 00:60:dd:7f:61:fc<br>
&gt; &gt; lanai_sram_size = 0x00200000 (2048K bytes)<br>
&gt; &gt; max_lanai_speed = 134 MHz<br>
&gt; &gt; product_code = 109<br>
&gt; &gt; serial_number = 89372<br>
&gt; &gt; (should be labeled: &quot;M3S-PCI64B-2-89372&quot;)<br>
&gt; &gt; LANai time is 0x08e6f211a ticks, or about 17 minutes since reset.<br>
&gt; &gt; Mapper is 00:00:00:00:00:00.<br>
&gt; &gt; Map version is 0.<br>
&gt; &gt; 0 hosts.<br>
&gt; &gt; Network is NOT fully configured.<br>
&gt; &gt; This node is &quot;hpc002&quot;<br>
&gt; &gt; Board has room for 16 ports, 1600 nodes/routes, 16384 cache entries<br>
&gt; &gt; Port token cnt: send=61, recv=254<br>
&gt; &gt; Port: Status PID<br>
&gt; &gt; 0: BUSY 1692 (this process [gm_board_info])<br>
&gt; &gt; Route table for this node follows:<br>
&gt; &gt; gmID MAC Address gmName Route<br>
&gt; &gt; ---- ----------------- --------------------------------<br>
&gt; &gt; ---------------------<br>
&gt; &gt; 1 00:60:dd:7f:61:fc hpc002 (this node)<br>
&gt; &gt;<br>
&gt; &gt; Thank you in advance for your help !!!<br>
&gt; &gt;<br>
&gt; &gt; ----------------------------------------------------<br>
&gt; &gt; ATSUKO MIYASHITA<br>
&gt; &gt; Linux Support Center, SWSC, Technical Support - IBM Japan<br>
&gt; &gt; Tel : 03-3808-8993 ( ext. 1712-8993 ) / Fax : 03-3664-4893<br>
&gt; &gt; e-mail : ATSUKOM$B!w(Bjp.ibm.com<br>
&gt; &gt; ** Moved to Hakozaki - mind that tel # has changed ! **<br>
&gt; &gt;<br>
&gt; &gt;------------------------------------------------------------------------<br>
&gt; &gt;<br>
&gt; &gt;_______________________________________________<br>
&gt; &gt;Myrinet mailing list<br>
&gt; &gt;[email protected]<br>
&gt; &gt;http://email.osc.edu/mailman/listinfo/myrinet<br>
&gt; &gt;<br>
&gt; &gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
&gt;<br>
<br>
<br>
</tt></font>
<br>
<br>