Re: Mean vs Median
"MORTON, ALFRED C (AL)" <[email protected]>
| Newsgroups | gmane.ietf.bmwg |
|---|---|
| Message-ID | <4AF73AA205019A4C8A1DDD32C034631D0BB6ADB7AF@NJFPSRVEXG0.research.att.com> |
Hi Marius, Paul, and all who have contributed so far. a quick reply/differing opinion below. > -----Original Message----- > From: bmwg [mailto:[email protected]] On Behalf Of Marius Georgescu ... > > On Nov 10, 2015, at 02:40, Paul Emmerich <[email protected]> > wrote: > > > > Hi, > > > > On 03.11.15 09:45, Stenio Fernandes wrote: > >> a word of caution here... a number of phenomena in computer networks > >> follows a heavy-tailed probability distribution function, which means > >> that there is a non-negligible probability that a random variable > >> will take huge values. these values might be erroneously considered > as outliers. > > > > this is a really important point. I have benchmarked software where > the 99th percentile of the latency is twice the average/median and the > 99.9th percentile ten times the average/median. > > Can you give us more context (test setup; physical/virtualized > tester/DUT; one tester/sender_receiver tester ... ) on these > measurements? [ACM] My understanding (and I've seen some results, but I've had trouble re-locating them) is that both outliers and bimodal distributions are more common in the world of virtual DUTs than they were in the physical/past. Not only does this affect analysis, but the threshold waiting time for packet arrival must be chosen carefully to even measure such outliers. > > > This is an important performance characteristic for latency-sensitive > applications that isn't captured by taking just 20 measurements. So I'd > really like to see a standard that calls for thousands of latency > measurements to capture this properly. > > > > I think we should keep practicality in mind here. If we follow > RFC2544.latency measurement, the frame stream has to be 2 min long. 2000 > min ~ 33h of testing for just one test sounds unreasonable to me. I > would agree to have a lower bound for the sample size as RFC2544 > actually recommends (n > 20). [ACM] Latency (delay) and delay variation need many single delay measurements to be meaningful. One way to view the variation is for a single flow of packets with spacing that might come from an application, say 20ms spacing for VoIP. Collecting a few thousand of such packets should not take so long. > > > You can also get interesting insights into a black-box device by > > looking at histograms/probability density functions. For example, you > > can figure out if the device processes packets in batches, estimate > > the batch size, figure out at which rates interrupt moderation > > algorithms change etc. (This is, of course, not really a performance > > metric, just an interesting insight.) > > > > I agree this is an interesting insight. It can also be the base for a > decision between summarizing functions. However, in the light of > consistency and simplicity of the methodology, I think we would need to > recommend one function. We could do that depending on the metric/DUT > characteristics, previous testing behavior … [ACM] I agree the right summary statistics can only be chosen after an examination of the raw distribution for a particular scenario. If Bi-modal, the central statistics of the sample could be meaningless. Without this examination, I don't think one recommendation can always be right. my 2 cents Al (as a participant) > > > > > Paul > > > > -- > > Paul Emmerich > > Technical University of Munich (TUM) > > Department of Informatics > > Chair for Network Architectures and Services > > > > _______________________________________________ > > bmwg mailing list > > [email protected] > > https://www.ietf.org/mailman/listinfo/bmwg > > _______________________________________________ > bmwg mailing list > [email protected] > https://www.ietf.org/mailman/listinfo/bmwg _______________________________________________ bmwg mailing list [email protected] https://www.ietf.org/mailman/listinfo/bmwg