REVIEW 2 major objections 5 minor 16 references
Implementing Errors on Errors: Bayesian vs Frequentist
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Bayesian and frequentist prescriptions for uncertain errors are the same statistical model, matched parameter by parameter.
desk verdict A correct short note connecting D'Agostini and Cowan on errors on errors; the parameter mapping is real, but the claimed structural equivalence outruns what is actually shown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the gamma-distributed auxiliary ratio: $\omega_i=s_i^2/\hat\sigma_i^2$ on the Bayesian side and $\hat w_i/v_i$ on the frequentist side, joined by the correspondence in Eq. (44). The derivation uses the scaling property of the gamma distribution to rewrite Cowan's chi-squared probability element in exactly the same canonical form as D'Agostini's prior, so that equal shape and rate parameters follow immediately. This machinery shows that D'Agostini's symmetric prior choice $k=\lambda$ is not an ad hoc recommendation but the frequentist statement that each quoted variance is a sample variance from $n=2k+1$ virtual repetitions.
What would settle it
Simulate many experiments under Cowan's model with a fixed integer $n$ (for example $n=3$, so $\varepsilon=0.5$ and $k=\lambda=1$), form D'Agostini's posterior interval for $\mu$, and measure its frequentist coverage; if coverage deviates from the nominal rate, the claimed structural equivalence fails. A sharper observation internal to the paper is that D'Agostini's recommended $k=1.29$ gives $n=2k+1=3.58$, which cannot be a literal count of virtual repetitions, so the equivalence is a formal parameter identity rather than a realizable sampling scheme unless $n$ is treated as effective and fractional.
Extended reading notes
Core claim
On the paper's own terms, the central finding is a parameter-by-parameter mapping between D'Agostini's Bayesian treatment and Cowan's frequentist treatment of uncertainty in quoted variances. D'Agostini models the inverse-square ratio $\hat\omega_i=s_i^2/\hat\sigma_i^2$ as a gamma random variable with shape $k$ and rate $\lambda$; Cowan models each measured variance $\hat w_i$ as the sample variance of $n$ virtual normal draws around the true variance $v_i$, which makes $\hat w_i/v_i$ gamma-distributed with shape and rate $(n-1)/2$. The paper asserts on physical grounds that these two ratios are the same quantity (Eq. (44)); matching the probability elements then gives $k=\lambda=(n-1)/2=1/(4\varepsilon^2)$, where $\varepsilon$ is Cowan's errors-on-errors parameter and $n=1+1/(2\varepsilon^2)$. The authors state explicitly that they are not proposing a new model, but unifying two existing ones by showing their structural equivalence.
Load-bearing premise
The whole comparison rests on the physical identification that the Bayesian ratio of the quoted variance to the theoretical variance, $s_i^2/\hat\sigma_i^2$, is the same quantity as the frequentist ratio of the estimated variance to the true variance, $\hat w_i/v_i$; if this mapping is rejected, the equality $k=\lambda=(n-1)/2$ does not follow.
Editorial extensions
If this is right
- The heuristic Bayesian prior $k=\lambda$ is replaced by a principled frequentist interpretation: it says the quoted variance is an average over $n=2k+1$ virtual replications, with relative uncertainty $\varepsilon=1/\sqrt{2(n-1)}$.
- D'Agostini's 'democratic skepticism'—identical gamma parameters for all experiments—corresponds to assuming the same number of virtual repetitions for every experiment, so one scalar $\varepsilon$ controls the skepticism.
- The marginal likelihood becomes a Student's $t$ with $\nu=2k$ degrees of freedom, so the correspondence fixes the heavy-tailed behavior of the combined result in terms of the relative error-on-error.
- The PDG scale factor $S=\max(1,\sqrt{\chi^2/(N-1)})$ is not a standalone fudge factor; it is the practical face of a gamma-variance averaging model with uncertainty in quoted variances.
- Setting $k=\lambda=1$ (i.e., $\varepsilon=0.5$, $n=3$) is a defensible default because it makes the variance of the variance ratio equal to its squared mean.
Reading between the lines
- An implication the authors leave implicit is a calibration rule: any Bayesian choice of $k$ can be translated into an effective number of repetitions $n=2k+1$, so priors with small $k$ describe variance estimates based on less than one effective replication—a useful diagnostic for excessively heavy-tailed combinations.
- A testable extension would be to estimate $\varepsilon$ directly from the scatter of published systematic uncertainties across a set of comparable measurements, then compare the implied $k=1/(4\varepsilon^2)$ with the value that best fits the combined posterior; agreement would validate the correspondence on real data rather than algebraically.
- If the equivalence holds for independent errors, a natural multivariate extension is to let the $n$ virtual repetitions be correlated across experiments, turning Cowan's scalar model into a correlated gamma-variance model; this is an extrapolation, not a claim the paper makes.
- The failure of $n=2k+1$ to be an integer for typical fitted $k$ suggests the 'virtual repetitions' should be understood as an effective statistical device, not a literal experimental design; that reading preserves the mathematical equivalence while abandoning any claim about actual replication counts.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares two approaches to implementing 'errors on errors' in the combination of inconsistent experimental results: D'Agostini's Bayesian formulation, which introduces a gamma prior on ω_i = s_i²/σ̂_i², and Cowan's frequentist formulation, which introduces a gamma sampling distribution for ŵ_i/v_i via n virtual repetitions. The central claim is that these two formulations admit a parameter-by-parameter correspondence and are structurally equivalent. The technical core (Section 4) shows that identifying ω_i with ŵ_i/v_i makes D'Agostini's gamma prior g(ω_i|k,λ) equal to Cowan's sampling gamma g(w_i/v_i|(n-1)/2,(n-1)/2), yielding k = λ = (n-1)/2 = 1/(4ε²). The paper then argues that this correspondence provides a frequentist rationale for the Bayesian prior choice k=λ and for 'democratic skepticism.' Minor caveats are noted, including the fact that Eq. (44) is asserted on physical grounds rather than derived.
Significance. If the claimed correspondence is taken as a statement about the auxiliary gamma variable models, the paper is a useful and elegantly presented contribution. The distributional algebra in Section 4 is correct and clearly explained, and the identification of k=λ=(n-1)/2=1/(4ε²) is a neat result that connects two previously disparate treatments. The paper also explicitly clarifies that its aim is not to introduce a new model but to unify existing ones, and it cites relevant literature. However, the abstract's stronger claim of 'structural equivalence' of the two formulations is not fully supported by the comparison of the auxiliary densities alone, and the load-bearing mapping in Eq. (44) is an asserted modeling assumption. These issues affect the central contribution and require attention before publication.
major comments (2)
- [Section 4, Eq. (44)] The identification s_i²/σ̂_i² ↔ ŵ_i/v_i, stated 'on physical grounds,' is the bridge that makes the parameter identity k=λ=(n-1)/2 follow. This identification is not derived; it is a modeling assumption. In particular, quoted systematic uncertainties s_i are not generally sample variances from n repeated, identical measurements, so the correspondence is not a mathematical consequence but a chosen convention. Since the paper's central claim depends on this step, the authors should either provide a more explicit justification (for instance, by relating s_i² to the standard error of the mean in Cowan's virtual-repetition picture, so that s_i² = ŵ_i/n and σ̂_i² = v_i/n) or clearly state that this is a modeling assumption whose scope and limits (e.g., for non-normal or systematic uncertainties) should be discussed.
- [Abstract and Section 4] The paper concludes that the two formulations are 'structurally equivalent,' but the comparison in Section 4 establishes equality only of the auxiliary gamma densities. In D'Agostini's model, ω_i is a prior on the theoretical variance and is marginalized over the likelihood (Eq. (29) and (30)), while in Cowan's model, ŵ_i/v_i is a sampling distribution for an observed statistic in a joint likelihood for (d_i, w_i) given μ and v_i. The paper does not show that the resulting likelihood for μ or the posterior for μ coincide under the mapping. Without such a demonstration, the abstract's claim of structural equivalence of the full formulations is an overstatement. I recommend either extending the analysis to compare the full inference procedures or qualifying the claim to 'equivalence of the auxiliary variable models,' which would still be a valuable contribution.
minor comments (5)
- [Section 2.4, Eq. (30)] The PDF in Eq. (30) is the kernel of a Student-t distribution, but the degrees of freedom and scale are not explicitly stated. For clarity, it would be helpful to note that this is a t-distribution with ν = 2k degrees of freedom and scale s_i √(λ/k), which becomes simply s_i when k=λ.
- [Section 3.1, Eq. (36)] The notation χ²_{n-1}((n-1)/v_i w_i) d((n-1)/v_i w_i) may confuse readers because the argument and the differential are written in a nonstandard way. Consider introducing an explicit variable, e.g., z_i = (n-1) w_i / v_i, and then applying the scaling property to obtain the gamma density for w_i.
- [Section 3.2, Eq. (40)] The discussion around Eq. (40) and the footnote about the naive treatment leading to V[s_i] ~ 0 is somewhat cryptic. A brief explanation of why this approximation fails would improve readability.
- [Figure 1 caption] The caption states that ε=0.5 corresponds to 'our version of D'Agostini's criterion V[ω]=E²[ω]' but does not point to the equation where this criterion is introduced. Adding a reference to Eq. (26) and Eq. (27) would be helpful.
- [Section 2.2, footnote 4] The term 'democratic skepticism' is introduced in a footnote but not defined in the main text. Since the paper later refers to this concept in the summary, a brief definition in the main text would make the argument more self-contained.
Circularity Check
No circularity: the parameter correspondence follows from an explicitly stated dictionary between two published gamma distributions.
full rationale
The paper's central claim is that D'Agostini's Bayesian and Cowan's frequentist gamma-distributed auxiliary variables are structurally equivalent. The derivation chain is transparent: Eq. (15) is D'Agostini's published gamma prior for omega_i, Eq. (36)-(37) give Cowan's chi-squared/gamma sampling distribution for w_i/v_i, and Eq. (44) explicitly states the physical dictionary omega_i = w_i/v_i. Under this dictionary, Eq. (45) becomes a gamma distribution with shape and rate both equal to (n-1)/2, matching Eq. (15) with k=lambda=(n-1)/2. No parameter is fitted to data, no subset is used to predict a closely related quantity, and no load-bearing self-citation appears; the cited works are D'Agostini and Cowan, not the present authors. The mapping in Eq. (44) is an assumption stated 'on physical grounds,' not a derived result, and the comparison is limited to the auxiliary gamma densities rather than the full marginalized likelihoods. These are limitations on the strength and completeness of the claimed equivalence, but they do not constitute circularity: the paper does not use its conclusion as an input, and the parameter identity is a direct algebraic consequence of the stated dictionary and the two published formulas. The derivation is therefore self-contained and non-circular.
Assumptions & free parameters
free parameters (3)
- D'Agostini's nuisance parameters (k, lambda) =
k approx 1.29, lambda approx 0.589
- Cowan's errors-on-errors parameter epsilon (or virtual repetitions n) =
examples epsilon = 0.2, 0.5, 1; n = 13.5, 3, 1.5
- Prior criterion V[omega_i] = E^2[omega_i] leading to k = lambda = 1 =
k = lambda = 1
assumptions (7)
- domain assumption Each experimental estimator d_i is normal with mean mu and variance sigma_i^2 (Eq. (4)).
- domain assumption The inverse square ratio omega_i = s_i^2/sigma_i^2 follows a gamma distribution g(omega_i | k, lambda) (Eq. (15)).
- domain assumption All experiments share the same gamma parameters, called democratic skepticism, with k_i = k and lambda_i = lambda (footnote 4; also n_i = n in Sec. 3).
- standard math Cowan's virtual repetitions produce an unbiased sample variance whose scaled distribution is chi-squared, hence gamma (Eqs. (35)-(37)).
- standard math A uniform improper prior on mu is used in the Bayesian update (Eqs. (2), (7)).
- standard math The gamma scaling property alpha g(alpha omega | k, lambda) = g(omega | k, alpha lambda) (Eq. (16)).
- ad hoc to paper The correspondence s_i^2/sigma_i^2 corresponds to w_i/v_i (Eq. (44)).
Cite this review
Pith. "Pith review of Implementing Errors on Errors: Bayesian vs Frequentist." pith.science (2026). https://pith.science/paper/WT5OHWDT
@misc{pith2026250506521,
author = {Pith},
title = {Pith review of: Implementing Errors on Errors: Bayesian vs Frequentist},
year = {2026},
howpublished = {\url{https://pith.science/paper/WT5OHWDT}},
note = {Machine review of arXiv:2505.06521}
}
read the original abstract
When combining apparently inconsistent experimental results, one often implements errors on errors. The Particle Data Group's phenomenological prescription offers a practical solution but lacks a firm theoretical foundation. To address this, D'Agostini and Cowan have proposed Bayesian and frequentist approaches, respectively, both introducing gamma-distributed auxiliary variables to model uncertainty in quoted errors. In this Letter, we show that these two formulations admit a parameter-by-parameter correspondence, and are structurally equivalent. This identification clarifies how Bayesian prior choices can be interpreted in terms of frequentist sampling assumptions, providing a unified probabilistic framework for modeling uncertainty in quoted variances.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
G. D’Agostini, Sceptical combination of experimental results: General considerations and application to epsilon-prime/epsilon , hep-ex/9910036
-
[3]
Cowan, Statistical Models with Uncertain Error Parameters , Eur
G. Cowan, Statistical Models with Uncertain Error Parameters , Eur. Phys. J. C 79 (2019) 133 [ 1809.05778]
arXiv 2019
-
[4]
Aoyama et al., The anomalous magnetic moment of the muon in the Standard Model, Phys
T. Aoyama et al., The anomalous magnetic moment of the muon in the Standard Model, Phys. Rept. 887 (2020) 1 [ 2006.04822]
arXiv 2020
-
[5]
Effect of Systematic Uncertainty Estimation on the Muon $g-2$ Anomaly
G. Cowan, Effect of Systematic Uncertainty Estimation on the Muon g− 2 Anomaly, EPJ Web Conf. 258 (2022) 09002 [ 2107.02652]
work page Pith review arXiv 2022
-
[6]
E. Canonero and G. Cowan, Correlated systematic uncertainties and errors-on-errors in measurement combinations with an application to the 7–8 TeV ATLAS–CMS top quark mass combination, Eur. Phys. J. C 85 (2025) 156 [ 2407.05322]
work page Pith review arXiv 2025
-
[7]
Higher-order asymptotic corrections and their application to the Gamma Variance Model
E. Canonero, A. R. Brazzale and G. Cowan, Higher-order asymptotic corrections and their application to the Gamma Variance Model , Eur. Phys. J. C 83 (2023) 1100 [2304.10574]
work page Pith review arXiv 2023
-
[8]
Statistical models with uncertain error parameters (PHYSTAT- ν workshop at CERN)
G. Cowan, “Statistical models with uncertain error parameters (PHYSTAT- ν workshop at CERN).” Available at https://indico.cern.ch/event/735431/contributions/ 3268138/attachments/1779397/2894151/cowan_phystatnu_ErrOnErr.pdf (slides) and https://cds.cern.ch/record/2655408 (movie), 2019
Show all 16 references
-
[9]
Erler and R
J. Erler and R. Ferro-Hern´ andez,Alternative to the application of PDG scale factors , Eur. Phys. J. C 80 (2020) 541 [ 2004.01219]
2020 arXiv
-
[10]
D’Agostini, On a curious bias arising when the p χ2/ν scaling prescription is first applied to a sub-sample of the individual results , 2001.07562
G. D’Agostini, On a curious bias arising when the p χ2/ν scaling prescription is first applied to a sub-sample of the individual results , 2001.07562
2001 arXiv
- [11]
-
[12]
Rallapalli and S
A. Rallapalli and S. Desai, Bayesian inference of W-boson mass , Eur. Phys. J. C 83 (2023) 580 [ 2301.09557]
2023 arXiv
-
[13]
Aliberti et al., The anomalous magnetic moment of the muon in the Standard Model: an update , 2505.21476
R. Aliberti et al., The anomalous magnetic moment of the muon in the Standard Model: an update , 2505.21476
-
[14]
Muon g-2 collaboration, D. P. Aguillard et al., Measurement of the Positive Muon Anomalous Magnetic Moment to 127 ppb , 2506.03069
-
[15]
Fuwa et al., Improved measurements of neutron lifetime with cold neutron beam at J-PARC, 2412.19519
Y. Fuwa et al., Improved measurements of neutron lifetime with cold neutron beam at J-PARC, 2412.19519
-
[16]
T. Karim et al., Measuringσ8 using DESI Legacy Imaging Surveys Emission-Line galaxies and Planck CMB lensing, and the impact of dust on parameter inference , JCAP 02 (2025) 045 [ 2408.15909]. 11
2025 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.