Re: Mean vs Median

"GEORGESCU LIVIU MARIUS" <[email protected]>
Newsgroups gmane.ietf.bmwg
Message-ID <[email protected]>
> 
> 
> 
> 
> Hello Stenio,
>  
>  
>  
> Thanks for your comments. Please see my comments inline.
>  
>  
>  
> From: [email protected]
> [mailto:[email protected] <[email protected]>]
> On Behalf Of Stenio Fernandes
> 
> Sent: Tuesday, November 3, 2015 5:45 PM
> 
> To: GEORGESCU LIVIU MARIUS <[email protected]>
> 
> Cc: [email protected]; [email protected]
> 
> Subject: Re: [bmwg] Mean vs Median
>  
>  
>  
> my two cents on this... see inline comments
>  
>  
>  
> this is my first interaction with the wg... so, bear
> with me if i'm a bit wordy :-)
>  
>  
>  
> stenio
>  
>  
>  
>  
>  
> Thanks for joining the discussion.
>  
>  
>  
> On Tu
>  
> e, Nov 3, 2015 at 3:57 AM, GEORGESCU LIVIU MARIUS <[email protected](https://mail.google.com/mail/?view=cm&fs=1&tf=1&[email protected])> wrote:
>  
> Hello BMWG,
>  
>  
>  
> Following some of the discussion we had in IETF93 about
> using either mean or median as a summarizing function for the results of
> multiple test iterations, I added the following section in http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00 
>  
> .
>  10. Summarizing function
> and repeatability 
>  
> 
>  
> 
>  To ensure the stability of the benchmarking scores obtained using
> 
>  the tests presented in Sections 6(http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-6)-9(http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#section-9), multiple test iterations are
> 
>  recommended. Following the recommendations of RFC2544(http://tools.ietf.org/html/rfc2544), the average
> 
>  was chosen to be the summarizing function for the reported values.
> 
>  While median can be an alternative summarizing function, a rationale
> 
>  for using one or the other is needed.
>  
>  
>  
>  
>  
> average is a colloquial term. although, in this context,
> there might be nothing wrong with that, imho precise terms are preferred.
> measures of central tendency could be used as the general term, where mean and
> median fit in. 
>  
>  
>  
> Average seems to be a term accepted and used by industry, and
> not just a “colloquial” term. We could quibble about definitions and what
> we need to follow just as well. I prefer to go with RFC2544.
>  
>  
>  
>  The median can be useful for summarizing especially when outliers
> 
>  are not a desired quantity. However, in the overall performance of a
> 
>  network device the outliers can represent a malfunction or
> 
>  misconfiguration in the DUT, which should be taken into account.
> 
>  The average is a more inclusive summarizing function. Moreover, as
> 
>  underlined in [DeNijs(http://tools.ietf.org/html/draft-ietf-bmwg-ipv6-tran-tech-benchmarking-00#ref-DeNijs)], the average is less exposed to statistical
> 
>  uncertainty. These reasons make it the RECOMMENDED summarizing
> 
>  function for the results of different test iterations, unless stated
> 
>  otherwise.
>  
>  
>  
> i'm having a hard time to understand this paragraph... i)
> "inclusive" is very vague; ii) "less exposed to uncertainty"
> is confusing. the mean is just a measure of centrality, whereas measures of
> dispersions (e.g., sd, variance) can be used to assess the degree of
> uncertainty around that measure (think of confidence intervals).
>  
> i know this is not the objective of the document, but the
> recommendation could be very simple. for example, stating that one should
> assess statistical significance of results would be enough, and then
> pointing out to a strong reference. le boudec's book seems appropriate
> here (cf. chap 2).
>  
>  
>  
> Just assessing statistical significance would be great. I would
> recommend the study of this book: “Jain, Raj. The art of computer systems
> performance analysis. John Wiley & Sons, 2008.”
>  
> I wonder if text like this in an RFC would see the print.
>  
> I am sure this would be perfect for certain types of academic
> papers. I would rather have a clearer recommendation.
>  
>  
>  
>  To express the repeatability of the benchmarking tests through a
> 
>  number, the Margin of error (MoE) can be used. Of course, other
> 
>  functions, such as standard error could be employed as well. The
> 
>  advantage the MoE has is expressing an associated confidence
> 
>  interval by using the alpha parameter.
> 
>  
> 
>  The recommended formula for calculating the MoE is presented in 
>  
> Section 6.3.1.
>  
>  
>  
> if the document will not give detailed approaches for
> summarizing performance data (and i think it shouldn't), it should provied the
> simplest recommendation as possible. otherwise, in order to be scientifically
> correct, lots of assumptions must be made and provided, like iid, which might
> not hold true for all cases.
>  
>  
>  
> In order to be scientifically correct, the most important
> part would be to assess the probability distribution of the data. I am trying
> to find a solution where that wouldn’t be necessary. Of course it would
> not hold for all cases, but the goal is to find a pareto-optimal solution where
> the summarized result would be representative enough for the test sample and
> simple enough to obtain. Or at least that’s how I see things.
>  
>  
>  
> Marius
>  
>  
>

_______________________________________________
bmwg mailing list
[email protected]
https://www.ietf.org/mailman/listinfo/bmwg
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.