RE: benchmarking i/o throughput
"Moshe Bar" <moshe-ay74M1d3r6RWk0Htik3J/[email protected]>
| Newsgroups | gmane.linux.cluster.openmosix.general |
|---|---|
| Message-ID | <[email protected]> |
Just tried it and it worked! I think this is a good temporary work-around and long-term we need to re-write that area from scratch. Well, done! It must have been quite a lot of forensic work! Moshe -----Original Message----- From: openmosix-general-admin-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org [mailto:openmosix-general-admin-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org] On Behalf Of Andreas Schäfer Sent: Saturday, February 25, 2006 9:35 PM To: openmosix-general-5NWGOfrQmneRv+LV9MX5uipxlwaOVQ5f@public.gmane.org Subject: Re: [openMosix-general] benchmarking i/o throughput Short story: ============ I think I've tracked down the weird write slowdowns I've seen and the attached patch increases the write performance from ~31kb/s to ~2705kb/s (so it's a ~87x speedup). Long story: =========== Originally I intended to leave this alone. It seemed too strange and no one of you guys seemed to be able to reproduce this error (you do perform some write accesses in your programms, don't you? ;-) ). But after being able to reproduce the abnormality on different machines I went a bit deeper into this... On 17:45 Tue 21 Feb , Andreas Schäfer wrote: > I've modified my initial setup and when writing directly to /dev/zero > I get the expected rate net bandwith. With Ethereal at the ready I sniffed the traffic between deputy and remote when piping into /dev/zero and compared this to the traffic I got when piping into a file. It was really striking to notice that nearly no data at all (except from small control packages) was transmitted. Admittedly, this is quite obvious, when you know how the write() syscall is implemented (I did not know but in summary no copy_from_user() is performed, thus no data from userspace needs to be read or transmitted through the network), and it led me to some thoughts I'll post separately on this list. In essence it was only a coincidence that some strange delays added up to making my program /think/ it was transmitting with a rate of ~11mb/s. As said, when piping into a real file the data rate was as low as ~31kb/s: > gentryx@wintermute ~ $ ./writeTest 4 1000 >/tmp/foo > status: 0 blksize: 131072 > sleeping 1000 seconds > dumping 4 mega bytes > write took 129175 milli seconds > 0.0317089 MB/s Ethereal showed that /now/ there was indeed traffic on the line, but it looked rather crumpled. I had seen before the traffic caused by the (fine performing) read(), which is basically one long stream of max. sized packets from remote to deputy, but the write() traffic did consist of many small transmissions. It copied only one page at a time, preceded by a number of control packages being exchanged back and forth. Still, this didn't cause the real performance flaw. Ethereal showed that there was now and then an ACK which should have been sent from the deputy to the remote but was delayed for half an eternity (exactly 40ms). Without the ACK, the remote refused to send another packed to the deputy. Exactly 3 delayed/timed-out ACKs could be observed during each page transmission. Further investigations proved that this timeout was exactly the one controlled by TCP_DELACK_MIN in include/net/tcp.h. By changing it's default value from "HZ/25" (which equals 40ms) to "HZ/200" I was able to reduce the delays/timeouts to 5ms. This again led to roughly 8x faster write transfers (~250kb/s), but it nevertheless remains a questionable--- and above all---dirty hack, so I refrained from it (notice that there is no connection to the actual openMosix code so far ;-) ). Again, I went deeper into the kernel's TCP stack to discover the cause of the delays (a surprisingly pleasant perusal at 4am). After spicing the better part the code of net/ipv4/tcp_input.c on the remote machine with debug statements it came apparent that the ACK was not send due to the socket not in quick-ack-mode (tcp_in_quickack_mode()). With some more debug statements I could prove that the socket started off alright with a value of 5 quick-acks. During the transmissions between remote and deputy, this value was magically decreased. When finally reaching 0 the timeouts set in. This could be fixed by a) resetting the quick-acks from time to time, b) by not letting them be decreased at all or... wait a moment... I silently implied that the performance degradation was the deputy-machine's fault just because the timeouts happened there. In deed the correct solution to fixing the quickack problem is c) telling the remote host not wait for an ACK in this case but send the next packet anyway. The cause for the remote socket refusing to do so was quickly found (or at least assumed) in hpc/comm.c. As said the error doesn't occur symmetrically but only for ACKs from deputy to remote. The assumption was that this was caused by an asymmetric configuration of the sockets. It turned out that comm_setup_tcp() from the afore mentioned file sets the specified socket (besides other stuff) to TCP_NODELAY mode. Strangely, it is only called for the deputy after accepting a connection. My patch adds a call to comm_setup_tcp() for the remote after connecting to the deputy. This causes all timeouts to disappear and the datarate goes up from the said ~31kb/s to ~2.7mb/s (roughly 87x faster), as can be seen below: > gentryx@wintermute ~ $ ./writeTest 1024 1000 >/tmp/foo > status: 0 blksize: 131072 > sleeping 1000 seconds > dumping 1024 mega bytes > write took 387639 milli seconds > 2.70503 MB/s The read-performance remains unaffected (and flawless): > gentryx@wintermute ~ $ dd if=/dev/zero bs=1M count=1024 | ./readTest > 1000 sleeping 1000 seconds > 1024+0 records in > 1024+0 records out > 1073741824 bytes (1,1 GB) copied, 96,3279 seconds, 11,1 MB/s write > took 94449 milli seconds > 11.102 MB/s Ok, sorry for this mammoth posting, but I thought it would be good for documentation/explanation as well as (perhaps) being interesting to read/learn from (for some kernel noobs like me (-; ). As I'm far from being a proficient kernel hacker I'd be glad if some of you guys could a) try to reproduce the performance flaw and b) check if the patch fixes it and works stable. For me it works like a charm. In consequence I'd of course be glad if it could be merged into upstream. 8-) Cheers! -Andreas ps: thanks for the recent introduction of debugfs. It made debugging a much easier, painless job. ------------------------------------------------------- This SF.Net email is sponsored by xPML, a groundbreaking scripting language that extends applications into web and mobile media. Attend the live webcast and join the prime developer group breaking into this new coding territory! http://sel.as-us.falkag.net/sel?cmd=lnk&kid0944&bid$1720&dat1642