Skip to main content

Module statistics

Module statistics 

Source
Expand description

Summaries of a sample of numbers: the spread of a Monte Carlo run’s apogees, say.

A Distribution keeps every value, sorted, and the number of samples that were tried, so a sample that failed or never gave a value (a flight with no apogee) is counted, not dropped: its share is Distribution::missing of Distribution::attempted, and a probability is reported as the bounds those unknowns allow (Distribution::share_at_least).

  • Mean and standard deviation are taken on the values shifted by the smallest one, the standard deviation by the two-pass formula with n − 1 (T. F. Chan, G. H. Golub and R. J. LeVeque, “Algorithms for computing the sample variance: analysis and recommendations”, The American Statistician 37(3), 242–247, 1983, https://doi.org/10.2307/2683386). Shifting by a value of the sample keeps the sums small, and makes a sample of equal values give that value and a deviation of exactly zero.
  • Quantiles are Hyndman and Fan’s definition 7, linear between order statistics, the default of R and NumPy: with the values sorted x₀ ≤ … ≤ xₙ₋₁ and h = (n − 1) p, Q(p) = x⌊h⌋ + (h − ⌊h⌋)(x⌊h⌋₊₁ − x⌊h⌋) (R. J. Hyndman and Y. Fan, “Sample quantiles in statistical packages”, The American Statistician 50(4), 361–365, 1996, https://doi.org/10.2307/2684934).

Every sum runs over the sorted values in order, so a summary is bit-for-bit the same however the values were computed, in parallel or not.

Structs§

Distribution
The values a sample of runs gave, sorted, and how many runs were tried. It serializes as those two, and reads back through Distribution::new’s checks.
Share
Bounds on the share of all the runs tried whose value passed a test: low counts a run with no value as failing it, high as passing.
Summary
The usual numbers of a Distribution, for a report.