Skip to main content

Module sobol

Module sobol 

Source
Expand description

Sobol’ indices: the share of the output’s variance each factor causes, alone and in all.

Guide: Sensitivity analysis’s Sobol’ indices section.

§The indices

With the factors independent, the output’s variance V = Var(Y) splits into the parts each factor causes alone, each pair together, and so on. Factor i’s first-order index is the share it causes alone, Sᵢ = Var(E[Y | Xᵢ]) / V; its total index is the share it has any part in, S_Tᵢ = E[Var(Y | X₋ᵢ)] / V = 1 − Var(E[Y | X₋ᵢ]) / V, with X₋ᵢ every factor but i. Sᵢ ≤ S_Tᵢ, and the gap is the share of i‘s interactions with the others. A factor whose total index is near zero can be left at its nominal value. I. M. Sobol’, “Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates”, Mathematics and Computers in Simulation 55, 271–280, 2001, https://doi.org/10.1016/S0378-4754(00)00270-6, defines both.

§The estimates

Two independent samples of N rows, A and B, each row every factor drawn uniformly over its range, and for each factor i a third, A_B⁽ⁱ⁾: A with its column i taken from B. The model runs at each row of all k + 2, N (k + 2) runs in all, and with f its output

  • Vᵢ ≈ (1/N) Σⱼ f(B)ⱼ (f(A_B⁽ⁱ⁾)ⱼ − f(A)ⱼ), and
  • V_Tᵢ ≈ (1/(2N)) Σⱼ (f(A)ⱼ − f(A_B⁽ⁱ⁾)ⱼ)² (Jansen’s),

with V the variance of the 2N outputs of A and B together, so Sᵢ = Vᵢ/V and S_Tᵢ = V_Tᵢ/V. Taking V from both samples, not A’s alone, is more accurate (A. Saltelli and others, Global Sensitivity Analysis: The Primer, Wiley, 2008, p. 166).

A. Saltelli, P. Annoni, I. Azzini, F. Campolongo, M. Ratto and S. Tarantola, “Variance based sensitivity analysis of model output. Design and estimator for the total sensitivity index”, Computer Physics Communications 181, 259–270, 2010, https://doi.org/10.1016/j.cpc.2009.09.018, give both (their Table 2, p. 262, rows (b) and (f)). They call Jansen’s the best practice so far for the total index (p. 262), and recommend (b) for the first order, as its design holds more useful points (p. 263).

The outputs are first shifted by their mean over A and B. That changes no index. Sobol’ (2001, p. 277, remark 3) advises it against a loss of accuracy when the mean is large; it also keeps the first-order estimate’s own variance, which has a term in the mean squared, from growing with the output’s mean, as an apogee’s would.

§Sampling error

Each row j is drawn independently, so every estimate is a smooth function of means over rows, and its standard error follows from the central limit theorem by the delta method: with ψⱼ row j’s first-order change to the estimate (its influence), the standard error is √(Σⱼ ψⱼ² / (N (N − 1))). For Sᵢ = P/V, with pⱼ row j‘s term of Vᵢ, qⱼ and mⱼ the means of its two outputs’ squares and of the two outputs, and D the mean of f(A_B⁽ⁱ⁾) − f(A) (the shift’s effect),

ψⱼ = ((pⱼ − P) − D (mⱼ − M)) / V − Sᵢ ((qⱼ − Q) − 2 M (mⱼ − M)) / V,

and the same for S_Tᵢ with Jansen’s terms and no D. A unit test checks it against the gradient of each index over the raw row means, and the tests check, over 1,000 seeds of Ishigami’s function and 500 of the g function, that it is the spread the estimates really have.

Structs§

Sobol
A Sobol’ analysis: the factors and the number of rows N of each of the samples A and B. It serializes as its fields, and reads back through Sobol::new’s checks.
SobolDesign
The points of a Sobol’ analysis, row after row: A’s row, B’s, then A_B⁽ⁱ⁾’s for each factor i in order, k + 2 points a row. It is not serialized: Sobol::design rebuilds it, bit for bit, from the analysis and its seed.
SobolIndex
A factor’s indices and their standard errors.
SobolIndices
What a Sobol’ analysis found.