Re: Mean vs Median
Marius Georgescu <[email protected]>
| Newsgroups | gmane.ietf.bmwg |
|---|---|
| Message-ID | <[email protected]> |
Hello Stenio, > On Nov 12, 2015, at 03:46, Stenio Fernandes <[email protected]> wrote: > > Very interesting results, Paul. I'd like to have a look at the results in the journal paper. > > In the past, statisticians struggled with a shortage of samples to do their stuff. This is not the case for the computer networking people as we are constantly flooded with samples. > > The discussions so far have led me to conclude that in the context here, there is no need to make any assumptions on the sample set. Any specific measure of centrality or dispersion might not be precise enough in some cases. As stated by others, mean/median would work for well-behaved (e.g., normally distributed) data, but would not work for multi-modal or heavy-tailed ones. Recall that heavy-tailed distributions are usually characterized by the shape and location parameters instead of mean and variance. Regarding the number of samples, it is really tough to characterize heavy-tailed or multi-modal distributions with a few samples, even using advanced algorithms for maximum-likelihood estimation. As stated in the email to Paul, I don’t think increasing the sample size would be a problem, as long as the (minimum) test time is within reasonable limits. I agree that the stability of the dataset is important. However, a more stable dataset would still need to be summarized/reported. If not average (arithmetic mean), median … how can we meaningfully and consistently express a test result ? From the Al and Paul’s examples/replies as well as other discussions I had in IETF94 about this, a “fast&hard” solution could be the Median + a measure of variance. Marius _______________________________________________ bmwg mailing list [email protected] https://www.ietf.org/mailman/listinfo/bmwg