Re: [myrinet] RE: Discarded Data

Bob Felderman <[email protected]>
Newsgroups gmane.network.myrinet.general
Message-ID <[email protected]>
=> From: "Davis, Harvey" <[email protected]>
=> Hello
=> 
=> I am running red hat linux 6.2, gm-1.2pre16 acroos a number of pentium class
=> machines.
=> 
=> I am sending out some data from the master box, receiving it at a number of
=> worker nodes, doing some processing and then returning the results to the
=> master.
=> 
=> Individually the results come back fine. However when all the machines fire
=> back their results at the same time, the master receives them but only yhte
=> first chunk of data is valid. The rest is filled with zeros.
=> 
=> I have allocated receive buffers equal in number to the number of workers
=> sending messages.
=> 
=> Do you have any ideas why this is occuring.
=> 


=> From: "Greg Lindahl" <[email protected]>
=> To: "Davis, Harvey" <[email protected]>, <[email protected]>
=> Subject: [myrinet] RE:  Discarded Data
=> Date: Mon, 8 Jan 2001 11:59:50 -0500
=> Status: R
=> 
=> > Individually the results come back fine. However when all the
=> > machines fire
=> > back their results at the same time, the master receives them but
=> > only yhte
=> > first chunk of data is valid. The rest is filled with zeros.
=> 
=> You've run out of memory. GM-1.2prewhatever doesn't test that error
=> condition correctly.

Most likely this is correct. gm_register_memory was NOT returning proper error
codes until gm-1.2. I would suggest upgrading to
gm-1.4pre45 (or whatever is current on the ftp site).

	ftp://[email protected]
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.