FW: Odd Problems on Cluster......
"Lowe, David (Xontech)" <[email protected]> Fri, 9 Jan 2004 08:43:54 -0800
| Newsgroups | gmane.network.myrinet.general |
|---|---|
| Message-ID | <[email protected]> |
Listed below is a problem a user has seen several times. I was wondering
if this error is documented and if so what I might be able to do to
correct the problem.
Thanks,
Dave
I've been seeing some odd problems on the cluster recently. I am only
using saturn5 - saturn8. Every once in a while I get the following
error and the job hangs:
Program passed bad send message to GM.
Port 2 disabled.
PANIC: libgm/gm_unknown.c:335:gm_unknown():userland
Sometimes I will get this error and hang:
shandle is 8d5e028
shandle cookie is 0
shandle at 8d5e028
cookie = 0
is_complete = 1
start = 0
bytes_as_contig = 0
[3] MPI internal Aborting program Bad address in Rendezvous send
[3] Bad address in Rendezvous send
This have happened in the last week or so. The first error I've seen
before, but maybe once a month. The last week, I've seen it several
times and about 5 times today. Today is the first time I've ever seen
the second error. All jobs run fine if I just re-run them.
ALSO, some jobs will hang on a MPI_BARRIER command. This commands halts
a process until all nodes have checked in. Recently I've seen that the
barrier is hit and all jobs sync on the barrier, but once they all sync,
one process never starts up again. If we are dropping message packets
(like the about errors suggest) then this could be causing sync errors
in many areas (like the Barrier command).
-Todd
_______________________________________________
Myrinet mailing list
[email protected]
http://email.osc.edu/mailman/listinfo/myrinet