Re: Mean vs Median

Stenio Fernandes <[email protected]>
Newsgroups gmane.ietf.bmwg
Message-ID <CAPrseCo-E82O+tSvRC=4x-yXYTMEHUW6UjeQK6HBRZwXey=sKg@mail.gmail.com>
my two cents on this... see inline comments

this is my first interaction with the wg... so, bear with me if i'm a bit
wordy :-)

stenio

On Tu
e, Nov 3, 2015 at 3:57 AM, GEORGESCU LIVIU MARIUS <[email protected]
<https://mail.google.com/mail/?view=cm&fs=1&tf=1&[email protected]>
> wrote:

> Hello BMWG,
>
> Following some of the discussion we had in IETF93 about using either mean
> or median as a summarizing function for the results of multiple test
> iterations, I added the following section in
> http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00
> .
>
> 10 <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-10>. Summarizing function and repeatability
>
>    To ensure the stability of the benchmarking scores obtained using
>    the tests presented in Sections 6 <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-6>-9 <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-9>, multiple test iterations are
>    recommended. Following the recommendations of RFC2544 <http://tools.ietf.org/html/rfc2544>, the average
>    was chosen to be the summarizing function for the reported values.
>    While median can be an alternative summarizing function, a rationale
>    for using one or the other is needed.
>
>

average is a colloquial term. although, in this context, there might be
nothing wrong with that, imho precise terms are preferred. measures of
central tendency could be used as the general term, where mean and median
fit in.


>    The median can be useful for summarizing especially when outliers
>    are not a desired quantity. However, in the overall performance of a
>    network device the outliers can represent a malfunction or
>    misconfiguration in the DUT, which should be taken into account.
>
>
a word of caution here... a number of phenomena in computer networks
follows a heavy-tailed probability distribution function, which means that
there is a non-negligible probability that a random variable will take huge
values. these values might be erroneously considered as outliers.


>    The average is a more inclusive summarizing function. Moreover, as
>    underlined in [DeNijs <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#ref-DeNijs>], the average is less exposed to statistical
>    uncertainty. These reasons make it the RECOMMENDED summarizing
>    function for the results of different test iterations, unless stated
>    otherwise.
>
>
i'm having a hard time to understand this paragraph... i) "inclusive" is
very vague; ii) "less exposed to uncertainty" is confusing. the mean is
just a measure of centrality, whereas measures of dispersions (e.g., sd,
variance) can be used to assess the degree of uncertainty around that
measure (think of confidence intervals).

i know this is not the objective of the document, but the recommendation
could be very simple. for example, stating that *one should assess
statistical significance of results *would be enough, and then pointing out
to a strong reference. le boudec's book seems appropriate here (cf. chap 2).

free pdf at http://perfeval.epfl.ch/

@book{leboudec2010performance,
   title={Performance Evaluation of Computer and Communication Systems},
   author={Le Boudec, Jean-Yves},
   year={2010},
   publisher={EPFL Press, Lausanne, Switzerland}
}



>    To express the repeatability of the benchmarking tests through a
>    number, the Margin of error (MoE) can be used. Of course, other
>    functions, such as standard error could be employed as well. The
>    advantage the MoE has is expressing an associated confidence
>    interval by using the alpha parameter.
>
>    The recommended formula for calculating the MoE is presented in
>
> Section 6.3.1
> <http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-6.3.1>
> .
>
> After discussing this rationale with Al (Morton) and Kostas (Pentakousis),
> I am tending to lean towards using median. One of the reasons is non-normal
> probability distribution cases (e.g. bimodal distribution), where the Mean
> might not mean much (trying to paraphrase Al). One could add a step in the
> procedure like "analyze the probability distribution of the 20 measurements
> after deciding the summarizing function", but this might be an undesired
> over-complication. In any case, I think a measure of variance should be
> provided with the summarized results, in order to express the
> stability/repeatability of the results.
>
>
if the document will not give detailed approaches for summarizing
performance data (and i think it shouldn't), it should provied the simplest
recommendation as possible. otherwise, in order to be scientifically
correct, lots of assumptions must be made and provided, like iid, which
might not hold true for all cases.


> Since the rationale for using Mean or Median (or ...) could be reused in
> other documents produced by this WG, I would like to ask for more feedback
> on this subject.
>
> Best regards,
> Marius
>
>
>
>
>
>
>
>
> _______________________________________________
> bmwg mailing list
> [email protected]
> <https://mail.google.com/mail/?view=cm&fs=1&tf=1&[email protected]>
> https://www.ietf.org/mailman/listinfo/bmwg
>
>


-- 
Prof. Stenio Fernandes
CIn/UFPE
http://www.steniofernandes.com

_______________________________________________
bmwg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/bmwg
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.