Expand description
Sobol’ indices: the share of the output’s variance each factor causes, alone and in all.
Guide: Sensitivity analysis’s Sobol’ indices section.
§The indices
With the factors independent, the output’s variance V = Var(Y) splits into the parts each
factor causes alone, each pair together, and so on. Factor i’s first-order index is the
share it causes alone, Sᵢ = Var(E[Y | Xᵢ]) / V; its total index is the share it has any
part in, S_Tᵢ = E[Var(Y | X₋ᵢ)] / V = 1 − Var(E[Y | X₋ᵢ]) / V, with X₋ᵢ every factor but
i. Sᵢ ≤ S_Tᵢ, and the gap is the share of i‘s interactions with the others. A factor
whose total index is near zero can be left at its nominal value. I. M. Sobol’, “Global
sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates”,
Mathematics and Computers in Simulation 55, 271–280, 2001,
https://doi.org/10.1016/S0378-4754(00)00270-6, defines both.
§The estimates
Two independent samples of N rows, A and B, each row every factor drawn uniformly over
its range, and for each factor i a third, A_B⁽ⁱ⁾: A with its column i taken from B.
The model runs at each row of all k + 2, N (k + 2) runs in all, and with f its output
Vᵢ ≈ (1/N) Σⱼ f(B)ⱼ (f(A_B⁽ⁱ⁾)ⱼ − f(A)ⱼ), andV_Tᵢ ≈ (1/(2N)) Σⱼ (f(A)ⱼ − f(A_B⁽ⁱ⁾)ⱼ)²(Jansen’s),
with V the variance of the 2N outputs of A and B together, so Sᵢ = Vᵢ/V and
S_Tᵢ = V_Tᵢ/V. Taking V from both samples, not A’s alone, is more accurate (A. Saltelli
and others, Global Sensitivity Analysis: The Primer, Wiley, 2008, p. 166).
A. Saltelli, P. Annoni, I. Azzini, F. Campolongo, M. Ratto and S. Tarantola, “Variance based sensitivity analysis of model output. Design and estimator for the total sensitivity index”, Computer Physics Communications 181, 259–270, 2010, https://doi.org/10.1016/j.cpc.2009.09.018, give both (their Table 2, p. 262, rows (b) and (f)). They call Jansen’s the best practice so far for the total index (p. 262), and recommend (b) for the first order, as its design holds more useful points (p. 263).
The outputs are first shifted by their mean over A and B. That changes no index. Sobol’
(2001, p. 277, remark 3) advises it against a loss of accuracy when the mean is large; it also
keeps the first-order estimate’s own variance, which has a term in the mean squared, from
growing with the output’s mean, as an apogee’s would.
§Sampling error
Each row j is drawn independently, so every estimate is a smooth function of means over
rows, and its standard error follows from the central limit theorem by the delta method: with
ψⱼ row j’s first-order change to the estimate (its influence), the standard error is
√(Σⱼ ψⱼ² / (N (N − 1))). For Sᵢ = P/V, with pⱼ row j‘s term of Vᵢ, qⱼ and mⱼ the
means of its two outputs’ squares and of the two outputs, and D the mean of
f(A_B⁽ⁱ⁾) − f(A) (the shift’s effect),
ψⱼ = ((pⱼ − P) − D (mⱼ − M)) / V − Sᵢ ((qⱼ − Q) − 2 M (mⱼ − M)) / V,
and the same for S_Tᵢ with Jansen’s terms and no D. A unit test checks it against the
gradient of each index over the raw row means, and the tests check, over 1,000 seeds of
Ishigami’s function and 500 of the g function, that it is the spread the estimates really
have.
Structs§
- Sobol
- A Sobol’ analysis: the factors and the number of rows
Nof each of the samplesAandB. It serializes as its fields, and reads back throughSobol::new’s checks. - Sobol
Design - The points of a Sobol’ analysis, row after row:
A’s row,B’s, thenA_B⁽ⁱ⁾’s for each factoriin order,k + 2points a row. It is not serialized:Sobol::designrebuilds it, bit for bit, from the analysis and its seed. - Sobol
Index - A factor’s indices and their standard errors.
- Sobol
Indices - What a Sobol’ analysis found.