Re: Mean vs Median
Stenio Fernandes <[email protected]>
| Newsgroups | gmane.ietf.bmwg |
|---|---|
| Message-ID | <CAPrseCo-E82O+tSvRC=4x-yXYTMEHUW6UjeQK6HBRZwXey=sKg@mail.gmail.com> |
my two cents on this... see inline comments this is my first interaction with the wg... so, bear with me if i'm a bit wordy :-) stenio On Tu e, Nov 3, 2015 at 3:57 AM, GEORGESCU LIVIU MARIUS <[email protected] <https://mail.google.com/mail/?view=cm&fs=1&tf=1&[email protected]> > wrote: > Hello BMWG, > > Following some of the discussion we had in IETF93 about using either mean > or median as a summarizing function for the results of multiple test > iterations, I added the following section in > http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00 > . > > 10 <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-10>. Summarizing function and repeatability > > To ensure the stability of the benchmarking scores obtained using > the tests presented in Sections 6 <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-6>-9 <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-9>, multiple test iterations are > recommended. Following the recommendations of RFC2544 <http://tools.ietf.org/html/rfc2544>, the average > was chosen to be the summarizing function for the reported values. > While median can be an alternative summarizing function, a rationale > for using one or the other is needed. > > average is a colloquial term. although, in this context, there might be nothing wrong with that, imho precise terms are preferred. measures of central tendency could be used as the general term, where mean and median fit in. > The median can be useful for summarizing especially when outliers > are not a desired quantity. However, in the overall performance of a > network device the outliers can represent a malfunction or > misconfiguration in the DUT, which should be taken into account. > > a word of caution here... a number of phenomena in computer networks follows a heavy-tailed probability distribution function, which means that there is a non-negligible probability that a random variable will take huge values. these values might be erroneously considered as outliers. > The average is a more inclusive summarizing function. Moreover, as > underlined in [DeNijs <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#ref-DeNijs>], the average is less exposed to statistical > uncertainty. These reasons make it the RECOMMENDED summarizing > function for the results of different test iterations, unless stated > otherwise. > > i'm having a hard time to understand this paragraph... i) "inclusive" is very vague; ii) "less exposed to uncertainty" is confusing. the mean is just a measure of centrality, whereas measures of dispersions (e.g., sd, variance) can be used to assess the degree of uncertainty around that measure (think of confidence intervals). i know this is not the objective of the document, but the recommendation could be very simple. for example, stating that *one should assess statistical significance of results *would be enough, and then pointing out to a strong reference. le boudec's book seems appropriate here (cf. chap 2). free pdf at http://perfeval.epfl.ch/ @book{leboudec2010performance, title={Performance Evaluation of Computer and Communication Systems}, author={Le Boudec, Jean-Yves}, year={2010}, publisher={EPFL Press, Lausanne, Switzerland} } > To express the repeatability of the benchmarking tests through a > number, the Margin of error (MoE) can be used. Of course, other > functions, such as standard error could be employed as well. The > advantage the MoE has is expressing an associated confidence > interval by using the alpha parameter. > > The recommended formula for calculating the MoE is presented in > > Section 6.3.1 > <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-6.3.1> > . > > After discussing this rationale with Al (Morton) and Kostas (Pentakousis), > I am tending to lean towards using median. One of the reasons is non-normal > probability distribution cases (e.g. bimodal distribution), where the Mean > might not mean much (trying to paraphrase Al). One could add a step in the > procedure like "analyze the probability distribution of the 20 measurements > after deciding the summarizing function", but this might be an undesired > over-complication. In any case, I think a measure of variance should be > provided with the summarized results, in order to express the > stability/repeatability of the results. > > if the document will not give detailed approaches for summarizing performance data (and i think it shouldn't), it should provied the simplest recommendation as possible. otherwise, in order to be scientifically correct, lots of assumptions must be made and provided, like iid, which might not hold true for all cases. > Since the rationale for using Mean or Median (or ...) could be reused in > other documents produced by this WG, I would like to ask for more feedback > on this subject. > > Best regards, > Marius > > > > > > > > > _______________________________________________ > bmwg mailing list > [email protected] > <https://mail.google.com/mail/?view=cm&fs=1&tf=1&[email protected]> > https://www.ietf.org/mailman/listinfo/bmwg > > -- Prof. Stenio Fernandes CIn/UFPE http://www.steniofernandes.com _______________________________________________ bmwg mailing list [email protected] https://www.ietf.org/mailman/listinfo/bmwg