Pith. sign in

REVIEW 2 major objections 5 minor 16 references

Implementing Errors on Errors: Bayesian vs Frequentist

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Bayesian and frequentist prescriptions for uncertain errors are the same statistical model, matched parameter by parameter.

desk verdict A correct short note connecting D'Agostini and Cowan on errors on errors; the parameter mapping is real, but the claimed structural equivalence outruns what is actually shown. read the letter →

arxiv 2505.06521 v2 pith:WT5OHWDT submitted 2025-05-10 hep-ph astro-ph.IMhep-exstat.ME

classification hep-phastro-ph.IMhep-exstat.ME
keywords errorsongammadistributionBayesianinferencefrequentistuncertaintycombinationvirtualrepetitionsPDGscalefactorstructuralequivalence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to prove that the two standard remedies for 'errors on errors'—the Bayesian gamma-prior prescription and the frequentist virtual-repetition sampling model—are structurally identical, not merely analogous. The identification runs through the ratio $\omega_i=s_i^2/\hat\sigma_i^2$ in the Bayesian picture and $\hat w_i/v_i$ in the frequentist picture; equating their probability elements forces $k=\lambda=(n-1)/2=1/(4\varepsilon^2)$. That means the symmetric gamma prior that the Bayesian approach advocated as a heuristic is exactly the distribution the frequentist approach obtains by imagining each quoted variance as the sample variance of $n$ virtual repetitions. If true, the Particle Data Group's practical $S$-factor rescaling becomes a special case of one coherent probabilistic model, and Bayesian prior choices acquire a concrete frequentist meaning. A reader should care because the result converts a philosophical divide into a testable equivalence.

What carries the argument

The load-bearing object is the gamma-distributed auxiliary ratio: $\omega_i=s_i^2/\hat\sigma_i^2$ on the Bayesian side and $\hat w_i/v_i$ on the frequentist side, joined by the correspondence in Eq. (44). The derivation uses the scaling property of the gamma distribution to rewrite Cowan's chi-squared probability element in exactly the same canonical form as D'Agostini's prior, so that equal shape and rate parameters follow immediately. This machinery shows that D'Agostini's symmetric prior choice $k=\lambda$ is not an ad hoc recommendation but the frequentist statement that each quoted variance is a sample variance from $n=2k+1$ virtual repetitions.

What would settle it

Simulate many experiments under Cowan's model with a fixed integer $n$ (for example $n=3$, so $\varepsilon=0.5$ and $k=\lambda=1$), form D'Agostini's posterior interval for $\mu$, and measure its frequentist coverage; if coverage deviates from the nominal rate, the claimed structural equivalence fails. A sharper observation internal to the paper is that D'Agostini's recommended $k=1.29$ gives $n=2k+1=3.58$, which cannot be a literal count of virtual repetitions, so the equivalence is a formal parameter identity rather than a realizable sampling scheme unless $n$ is treated as effective and fractional.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central finding is a parameter-by-parameter mapping between D'Agostini's Bayesian treatment and Cowan's frequentist treatment of uncertainty in quoted variances. D'Agostini models the inverse-square ratio $\hat\omega_i=s_i^2/\hat\sigma_i^2$ as a gamma random variable with shape $k$ and rate $\lambda$; Cowan models each measured variance $\hat w_i$ as the sample variance of $n$ virtual normal draws around the true variance $v_i$, which makes $\hat w_i/v_i$ gamma-distributed with shape and rate $(n-1)/2$. The paper asserts on physical grounds that these two ratios are the same quantity (Eq. (44)); matching the probability elements then gives $k=\lambda=(n-1)/2=1/(4\varepsilon^2)$, where $\varepsilon$ is Cowan's errors-on-errors parameter and $n=1+1/(2\varepsilon^2)$. The authors state explicitly that they are not proposing a new model, but unifying two existing ones by showing their structural equivalence.

Load-bearing premise

The whole comparison rests on the physical identification that the Bayesian ratio of the quoted variance to the theoretical variance, $s_i^2/\hat\sigma_i^2$, is the same quantity as the frequentist ratio of the estimated variance to the true variance, $\hat w_i/v_i$; if this mapping is rejected, the equality $k=\lambda=(n-1)/2$ does not follow.

Editorial extensions

If this is right

  • The heuristic Bayesian prior $k=\lambda$ is replaced by a principled frequentist interpretation: it says the quoted variance is an average over $n=2k+1$ virtual replications, with relative uncertainty $\varepsilon=1/\sqrt{2(n-1)}$.
  • D'Agostini's 'democratic skepticism'—identical gamma parameters for all experiments—corresponds to assuming the same number of virtual repetitions for every experiment, so one scalar $\varepsilon$ controls the skepticism.
  • The marginal likelihood becomes a Student's $t$ with $\nu=2k$ degrees of freedom, so the correspondence fixes the heavy-tailed behavior of the combined result in terms of the relative error-on-error.
  • The PDG scale factor $S=\max(1,\sqrt{\chi^2/(N-1)})$ is not a standalone fudge factor; it is the practical face of a gamma-variance averaging model with uncertainty in quoted variances.
  • Setting $k=\lambda=1$ (i.e., $\varepsilon=0.5$, $n=3$) is a defensible default because it makes the variance of the variance ratio equal to its squared mean.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the authors leave implicit is a calibration rule: any Bayesian choice of $k$ can be translated into an effective number of repetitions $n=2k+1$, so priors with small $k$ describe variance estimates based on less than one effective replication—a useful diagnostic for excessively heavy-tailed combinations.
  • A testable extension would be to estimate $\varepsilon$ directly from the scatter of published systematic uncertainties across a set of comparable measurements, then compare the implied $k=1/(4\varepsilon^2)$ with the value that best fits the combined posterior; agreement would validate the correspondence on real data rather than algebraically.
  • If the equivalence holds for independent errors, a natural multivariate extension is to let the $n$ virtual repetitions be correlated across experiments, turning Cowan's scalar model into a correlated gamma-variance model; this is an extrapolation, not a claim the paper makes.
  • The failure of $n=2k+1$ to be an integer for typical fitted $k$ suggests the 'virtual repetitions' should be understood as an effective statistical device, not a literal experimental design; that reading preserves the mathematical equivalence while abandoning any claim about actual replication counts.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper compares two approaches to implementing 'errors on errors' in the combination of inconsistent experimental results: D'Agostini's Bayesian formulation, which introduces a gamma prior on ω_i = s_i²/σ̂_i², and Cowan's frequentist formulation, which introduces a gamma sampling distribution for ŵ_i/v_i via n virtual repetitions. The central claim is that these two formulations admit a parameter-by-parameter correspondence and are structurally equivalent. The technical core (Section 4) shows that identifying ω_i with ŵ_i/v_i makes D'Agostini's gamma prior g(ω_i|k,λ) equal to Cowan's sampling gamma g(w_i/v_i|(n-1)/2,(n-1)/2), yielding k = λ = (n-1)/2 = 1/(4ε²). The paper then argues that this correspondence provides a frequentist rationale for the Bayesian prior choice k=λ and for 'democratic skepticism.' Minor caveats are noted, including the fact that Eq. (44) is asserted on physical grounds rather than derived.

Significance. If the claimed correspondence is taken as a statement about the auxiliary gamma variable models, the paper is a useful and elegantly presented contribution. The distributional algebra in Section 4 is correct and clearly explained, and the identification of k=λ=(n-1)/2=1/(4ε²) is a neat result that connects two previously disparate treatments. The paper also explicitly clarifies that its aim is not to introduce a new model but to unify existing ones, and it cites relevant literature. However, the abstract's stronger claim of 'structural equivalence' of the two formulations is not fully supported by the comparison of the auxiliary densities alone, and the load-bearing mapping in Eq. (44) is an asserted modeling assumption. These issues affect the central contribution and require attention before publication.

major comments (2)
  1. [Section 4, Eq. (44)] The identification s_i²/σ̂_i² ↔ ŵ_i/v_i, stated 'on physical grounds,' is the bridge that makes the parameter identity k=λ=(n-1)/2 follow. This identification is not derived; it is a modeling assumption. In particular, quoted systematic uncertainties s_i are not generally sample variances from n repeated, identical measurements, so the correspondence is not a mathematical consequence but a chosen convention. Since the paper's central claim depends on this step, the authors should either provide a more explicit justification (for instance, by relating s_i² to the standard error of the mean in Cowan's virtual-repetition picture, so that s_i² = ŵ_i/n and σ̂_i² = v_i/n) or clearly state that this is a modeling assumption whose scope and limits (e.g., for non-normal or systematic uncertainties) should be discussed.
  2. [Abstract and Section 4] The paper concludes that the two formulations are 'structurally equivalent,' but the comparison in Section 4 establishes equality only of the auxiliary gamma densities. In D'Agostini's model, ω_i is a prior on the theoretical variance and is marginalized over the likelihood (Eq. (29) and (30)), while in Cowan's model, ŵ_i/v_i is a sampling distribution for an observed statistic in a joint likelihood for (d_i, w_i) given μ and v_i. The paper does not show that the resulting likelihood for μ or the posterior for μ coincide under the mapping. Without such a demonstration, the abstract's claim of structural equivalence of the full formulations is an overstatement. I recommend either extending the analysis to compare the full inference procedures or qualifying the claim to 'equivalence of the auxiliary variable models,' which would still be a valuable contribution.
minor comments (5)
  1. [Section 2.4, Eq. (30)] The PDF in Eq. (30) is the kernel of a Student-t distribution, but the degrees of freedom and scale are not explicitly stated. For clarity, it would be helpful to note that this is a t-distribution with ν = 2k degrees of freedom and scale s_i √(λ/k), which becomes simply s_i when k=λ.
  2. [Section 3.1, Eq. (36)] The notation χ²_{n-1}((n-1)/v_i w_i) d((n-1)/v_i w_i) may confuse readers because the argument and the differential are written in a nonstandard way. Consider introducing an explicit variable, e.g., z_i = (n-1) w_i / v_i, and then applying the scaling property to obtain the gamma density for w_i.
  3. [Section 3.2, Eq. (40)] The discussion around Eq. (40) and the footnote about the naive treatment leading to V[s_i] ~ 0 is somewhat cryptic. A brief explanation of why this approximation fails would improve readability.
  4. [Figure 1 caption] The caption states that ε=0.5 corresponds to 'our version of D'Agostini's criterion V[ω]=E²[ω]' but does not point to the equation where this criterion is introduced. Adding a reference to Eq. (26) and Eq. (27) would be helpful.
  5. [Section 2.2, footnote 4] The term 'democratic skepticism' is introduced in a footnote but not defined in the main text. Since the paper later refers to this concept in the summary, a brief definition in the main text would make the argument more self-contained.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the parameter correspondence follows from an explicitly stated dictionary between two published gamma distributions.

full rationale

The paper's central claim is that D'Agostini's Bayesian and Cowan's frequentist gamma-distributed auxiliary variables are structurally equivalent. The derivation chain is transparent: Eq. (15) is D'Agostini's published gamma prior for omega_i, Eq. (36)-(37) give Cowan's chi-squared/gamma sampling distribution for w_i/v_i, and Eq. (44) explicitly states the physical dictionary omega_i = w_i/v_i. Under this dictionary, Eq. (45) becomes a gamma distribution with shape and rate both equal to (n-1)/2, matching Eq. (15) with k=lambda=(n-1)/2. No parameter is fitted to data, no subset is used to predict a closely related quantity, and no load-bearing self-citation appears; the cited works are D'Agostini and Cowan, not the present authors. The mapping in Eq. (44) is an assumption stated 'on physical grounds,' not a derived result, and the comparison is limited to the auxiliary gamma densities rather than the full marginalized likelihoods. These are limitations on the strength and completeness of the claimed equivalence, but they do not constitute circularity: the paper does not use its conclusion as an input, and the parameter identity is a direct algebraic consequence of the stated dictionary and the two published formulas. The derivation is therefore self-contained and non-circular.

Assumptions & free parameters 3 free parameters · 7 assumptions · 0 invented entities

The central claim is a dictionary between two published models. The only new ingredient is the interpretive mapping Eq. (44), which is not a derived theorem. The numeric hyperparameters (k, lambda, epsilon) are illustrative choices, not fitted constants. No new entities are introduced.

free parameters (3)
  • D'Agostini's nuisance parameters (k, lambda) = k approx 1.29, lambda approx 0.589
    Chosen in Eqs. (22)-(23) to satisfy E[r_i] = 1 and V[r_i] approx 1. They are examples, not load-bearing for the correspondence; the correspondence holds for symbolic k and lambda.
  • Cowan's errors-on-errors parameter epsilon (or virtual repetitions n) = examples epsilon = 0.2, 0.5, 1; n = 13.5, 3, 1.5
    Input of Cowan's sampling model, Eqs. (39)-(41). It is not fitted, but it is a free choice in the frequentist formulation; the paper uses illustrative values.
  • Prior criterion V[omega_i] = E^2[omega_i] leading to k = lambda = 1 = k = lambda = 1
    Eq. (27), adopted as the paper's version of D'Agostini's criterion. It is a hand-chosen conservative condition, not uniquely determined.
assumptions (7)
  • domain assumption Each experimental estimator d_i is normal with mean mu and variance sigma_i^2 (Eq. (4)).
    Standard measurement model in both approaches; used to write the likelihood in Eq. (5).
  • domain assumption The inverse square ratio omega_i = s_i^2/sigma_i^2 follows a gamma distribution g(omega_i | k, lambda) (Eq. (15)).
    D'Agostini's prior assumption; the paper states that providing a theoretical foundation for it is one purpose of the Letter.
  • domain assumption All experiments share the same gamma parameters, called democratic skepticism, with k_i = k and lambda_i = lambda (footnote 4; also n_i = n in Sec. 3).
    Simplifies the model; the paper notes that i-dependence can be trivially recovered.
  • standard math Cowan's virtual repetitions produce an unbiased sample variance whose scaled distribution is chi-squared, hence gamma (Eqs. (35)-(37)).
    Standard result for normal samples; used to write the sampling distribution of w_i/v_i.
  • standard math A uniform improper prior on mu is used in the Bayesian update (Eqs. (2), (7)).
    Standard uninformative prior; the paper notes the improper prior does not affect the analysis as long as the likelihood decays fast.
  • standard math The gamma scaling property alpha g(alpha omega | k, lambda) = g(omega | k, alpha lambda) (Eq. (16)).
    Used to rewrite Cowan's chi-squared density in ratio form in Eq. (45).
  • ad hoc to paper The correspondence s_i^2/sigma_i^2 corresponds to w_i/v_i (Eq. (44)).
    This is the interpretative bridge introduced to unify the two models; it is stated on physical grounds, not derived from either formulation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Implementing Errors on Errors: Bayesian vs Frequentist." pith.science (2026). https://pith.science/paper/WT5OHWDT

@misc{pith2026250506521,
  author       = {Pith},
  title        = {Pith review of: Implementing Errors on Errors: Bayesian vs Frequentist},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WT5OHWDT}},
  note         = {Machine review of arXiv:2505.06521}
}
read the original abstract

When combining apparently inconsistent experimental results, one often implements errors on errors. The Particle Data Group's phenomenological prescription offers a practical solution but lacks a firm theoretical foundation. To address this, D'Agostini and Cowan have proposed Bayesian and frequentist approaches, respectively, both introducing gamma-distributed auxiliary variables to model uncertainty in quoted errors. In this Letter, we show that these two formulations admit a parameter-by-parameter correspondence, and are structurally equivalent. This identification clarifies how Bayesian prior choices can be interpreted in terms of frequentist sampling assumptions, providing a unified probabilistic framework for modeling uncertainty in quoted variances.

Figures

Figures reproduced from arXiv: 2505.06521 by the authors.

Figure 1
Figure 1. PDFs g(ω | k, λ) and f(r | k, λ) are shown in the left and right panels, respectively. For simplicity, the i-dependence of ωi and ri is omitted in this figure and caption. The nuisance parameters are set to D’Agostini’s recommended values k = 1.29, λ = 0.589 (solid), and to Cowan’s k = λ = 1/4ε 2 with ε = 0.2 (dashed), 0.5 (dotted), and 1 (dot-dashed), corresponding to k = λ = 6.25, 1, and 0.25, respectively. The as… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 8 canonical work pages

  1. [1]

    Navas, C

    Particle Data Group collaboration, S. Navas, C. Amsler, T. Gutsche, C. Hanhart, J. J. Hern´ andez-Rey, C. Louren¸ co et al.,Review of particle physics , Phys. Rev. D 110 (2024) 030001

  2. [2]

    Sceptical combination of experimental results: General considerations and application to epsilon-prime/epsilon

    G. D’Agostini, Sceptical combination of experimental results: General considerations and application to epsilon-prime/epsilon , hep-ex/9910036

  3. [3]

    Cowan, Statistical Models with Uncertain Error Parameters , Eur

    G. Cowan, Statistical Models with Uncertain Error Parameters , Eur. Phys. J. C 79 (2019) 133 [ 1809.05778]

  4. [4]

    Aoyama et al., The anomalous magnetic moment of the muon in the Standard Model, Phys

    T. Aoyama et al., The anomalous magnetic moment of the muon in the Standard Model, Phys. Rept. 887 (2020) 1 [ 2006.04822]

  5. [5]

    Effect of Systematic Uncertainty Estimation on the Muon $g-2$ Anomaly

    G. Cowan, Effect of Systematic Uncertainty Estimation on the Muon g− 2 Anomaly, EPJ Web Conf. 258 (2022) 09002 [ 2107.02652]

  6. [6]

    Correlated Systematic Uncertainties and Errors-on-Errors in Measurement Combinations: Methodology and Application to the 7-8 TeV ATLAS-CMS Top Quark Mass Combination

    E. Canonero and G. Cowan, Correlated systematic uncertainties and errors-on-errors in measurement combinations with an application to the 7–8 TeV ATLAS–CMS top quark mass combination, Eur. Phys. J. C 85 (2025) 156 [ 2407.05322]

  7. [7]

    Higher-order asymptotic corrections and their application to the Gamma Variance Model

    E. Canonero, A. R. Brazzale and G. Cowan, Higher-order asymptotic corrections and their application to the Gamma Variance Model , Eur. Phys. J. C 83 (2023) 1100 [2304.10574]

  8. [8]

    Statistical models with uncertain error parameters (PHYSTAT- ν workshop at CERN)

    G. Cowan, “Statistical models with uncertain error parameters (PHYSTAT- ν workshop at CERN).” Available at https://indico.cern.ch/event/735431/contributions/ 3268138/attachments/1779397/2894151/cowan_phystatnu_ErrOnErr.pdf (slides) and https://cds.cern.ch/record/2655408 (movie), 2019

Show all 16 references
  1. [9]

    Erler and R

    J. Erler and R. Ferro-Hern´ andez,Alternative to the application of PDG scale factors , Eur. Phys. J. C 80 (2020) 541 [ 2004.01219]

  2. [10]

    D’Agostini, On a curious bias arising when the p χ2/ν scaling prescription is first applied to a sub-sample of the individual results , 2001.07562

    G. D’Agostini, On a curious bias arising when the p χ2/ν scaling prescription is first applied to a sub-sample of the individual results , 2001.07562

  3. [11]

    Crivellin and B

    A. Crivellin and B. Mellado, Anomalies in Particle Physics , 2309.03870

  4. [12]

    Rallapalli and S

    A. Rallapalli and S. Desai, Bayesian inference of W-boson mass , Eur. Phys. J. C 83 (2023) 580 [ 2301.09557]

  5. [13]

    Aliberti et al., The anomalous magnetic moment of the muon in the Standard Model: an update , 2505.21476

    R. Aliberti et al., The anomalous magnetic moment of the muon in the Standard Model: an update , 2505.21476

  6. [14]

    Muon g-2 collaboration, D. P. Aguillard et al., Measurement of the Positive Muon Anomalous Magnetic Moment to 127 ppb , 2506.03069

  7. [15]

    Fuwa et al., Improved measurements of neutron lifetime with cold neutron beam at J-PARC, 2412.19519

    Y. Fuwa et al., Improved measurements of neutron lifetime with cold neutron beam at J-PARC, 2412.19519

  8. [16]

    T. Karim et al., Measuringσ8 using DESI Legacy Imaging Surveys Emission-Line galaxies and Planck CMB lensing, and the impact of dust on parameter inference , JCAP 02 (2025) 045 [ 2408.15909]. 11

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.