{"id":"18f214b1-a9be-42c9-937c-dec5e308143a","arxiv_id":"2505.06521","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"D'Agostini's Bayesian gamma prior and Cowan's frequentist gamma variance model are the same gamma law under the parameter map k = lambda = (n-1)/2 = 1/(4 epsilon^2).","lead":"When combining conflicting measurements, scientists often inflate uncertainties. This paper shows that two standard statistical recipes for doing that, one Bayesian and one frequentist, are mathematically the same recipe.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (44) is asserted, not derived, and the paper compares only auxiliary gamma densities, not the full inference procedures; the claimed structural equivalence is therefore not established.","rationale":"The reader's weakest_assumption correctly identifies Eq. (44) as the bridge on which the entire correspondence rests. My stress-test confirms that this step is genuinely load-bearing: the algebraic manipulations in Eqs. (36)-(47) are internally consistent, but they only match the functional form of two gamma densities. The deeper claim of 'structural equivalence' between a Bayesian prior and a frequentist sampling model is not established by matching densities, because the two frameworks use the auxiliary variable in different roles: D'Agostini marginalizes over it as a prior, while Cowan treats it as an observed statistic whose sampling distribution is part of the likelihood. The paper itself flags the heuristic nature of Eq. (44) with the phrase 'on physical grounds', and this is the point where a skeptical reader can stop the argument. I do not see a mathematical error in the paper, and the conditional verdict is appropriate: the equivalence is plausible but not forced, and the full-model comparison has not been provided. A concrete check comparing the full Cowan likelihood with D'Agostini's marginal likelihood would settle whether the claimed structural equivalence extends beyond the auxiliary gamma density. Such a check could be done analytically with profiling or with a small numerical study; it is inexpensive and directly tests the strongest claim.","tokens_in":7712,"tokens_out":10147,"duration_ms":96117,"concrete_test":"Build Cowan's full model with n virtual repetitions: let the reported central value be the sample mean \\bar{x}_i and the quoted variance be \\hat{w}_i/n. The joint sampling density is f(d_i, s_i^2 | \\mu, v_i) = N(d_i | \\mu, v_i/n) \\times Gamma(n s_i^2/v_i | (n-1)/2, (n-1)/2) times the Jacobian. Eliminate v_i either by profiling (maximize over v_i) or by integrating with a reference prior, and compare the resulting likelihood for \\mu with D'Agostini's marginal likelihood Eq. (30) with k=\\lambda=(n-1)/2. If the two are not proportional, Eq. (44) establishes only a density-level duality rather than structural equivalence of the full Bayesian and frequentist formulations, and the headline claim should be weakened accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the parameter-by-parameter equivalence of D'Agostini's Bayesian and Cowan's frequentist gamma models. The load-bearing step is Eq. (44), which asserts 'on physical grounds' that D'Agostini's ratio s_i^2/\\hat{\\sigma}_i^2 (quoted variance over random theoretical variance) corresponds to Cowan's \\hat{w}_i/v_i (random sample variance over fixed true variance). This identification is not derived: it requires treating the quoted variance s_i^2 as if it were a sample variance \\hat{w}_i from n virtual repetitions, with \\hat{\\sigma}_i^2 identified with the variance of the sample mean v_i/n. Real quoted errors, especially systematic uncertainties, are not generally sample variances from n repeated identical measurements, so the mapping is a modeling assumption rather than a mathematical consequence. Moreover, the paper compares only the gamma densities (15) and (45); it does not compare the full inference procedures. In D'Agostini's model, \\omega_i is a prior on the true variance and is marginalized in the likelihood (Eq. 29), whereas in Cowan's model, \\hat{w}_i/v_i is a sampling distribution for an observed statistic. Without showing that the resulting likelihood/posterior for \\mu coincide, 'structurally equivalent' is established only for the auxiliary gamma density, not for the two formulations as whole. All subsequent results (Eq. 47, k=\\lambda, the democratic-skepticism rationale) inherit this fragility.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares two approaches to implementing 'errors on errors' in the combination of inconsistent experimental results: D'Agostini's Bayesian formulation, which introduces a gamma prior on ω_i = s_i²/σ̂_i², and Cowan's frequentist formulation, which introduces a gamma sampling distribution for ŵ_i/v_i via n virtual repetitions. The central claim is that these two formulations admit a parameter-by-parameter correspondence and are structurally equivalent. The technical core (Section 4) shows that identifying ω_i with ŵ_i/v_i makes D'Agostini's gamma prior g(ω_i|k,λ) equal to Cowan's sampling gamma g(w_i/v_i|(n-1)/2,(n-1)/2), yielding k = λ = (n-1)/2 = 1/(4ε²). The paper then argues that this correspondence provides a frequentist rationale for the Bayesian prior choice k=λ and for 'democratic skepticism.' Minor caveats are noted, including the fact that Eq. (44) is asserted on physical grounds rather than derived.","tokens_in":8131,"tokens_out":6290,"duration_ms":61236,"significance":"If the claimed correspondence is taken as a statement about the auxiliary gamma variable models, the paper is a useful and elegantly presented contribution. The distributional algebra in Section 4 is correct and clearly explained, and the identification of k=λ=(n-1)/2=1/(4ε²) is a neat result that connects two previously disparate treatments. The paper also explicitly clarifies that its aim is not to introduce a new model but to unify existing ones, and it cites relevant literature. However, the abstract's stronger claim of 'structural equivalence' of the two formulations is not fully supported by the comparison of the auxiliary densities alone, and the load-bearing mapping in Eq. (44) is an asserted modeling assumption. These issues affect the central contribution and require attention before publication.","major_comments":[{"comment":"The identification s_i²/σ̂_i² ↔ ŵ_i/v_i, stated 'on physical grounds,' is the bridge that makes the parameter identity k=λ=(n-1)/2 follow. This identification is not derived; it is a modeling assumption. In particular, quoted systematic uncertainties s_i are not generally sample variances from n repeated, identical measurements, so the correspondence is not a mathematical consequence but a chosen convention. Since the paper's central claim depends on this step, the authors should either provide a more explicit justification (for instance, by relating s_i² to the standard error of the mean in Cowan's virtual-repetition picture, so that s_i² = ŵ_i/n and σ̂_i² = v_i/n) or clearly state that this is a modeling assumption whose scope and limits (e.g., for non-normal or systematic uncertainties) should be discussed.","section":"Section 4, Eq. (44)"},{"comment":"The paper concludes that the two formulations are 'structurally equivalent,' but the comparison in Section 4 establishes equality only of the auxiliary gamma densities. In D'Agostini's model, ω_i is a prior on the theoretical variance and is marginalized over the likelihood (Eq. (29) and (30)), while in Cowan's model, ŵ_i/v_i is a sampling distribution for an observed statistic in a joint likelihood for (d_i, w_i) given μ and v_i. The paper does not show that the resulting likelihood for μ or the posterior for μ coincide under the mapping. Without such a demonstration, the abstract's claim of structural equivalence of the full formulations is an overstatement. I recommend either extending the analysis to compare the full inference procedures or qualifying the claim to 'equivalence of the auxiliary variable models,' which would still be a valuable contribution.","section":"Abstract and Section 4"}],"minor_comments":[{"comment":"The PDF in Eq. (30) is the kernel of a Student-t distribution, but the degrees of freedom and scale are not explicitly stated. For clarity, it would be helpful to note that this is a t-distribution with ν = 2k degrees of freedom and scale s_i √(λ/k), which becomes simply s_i when k=λ.","section":"Section 2.4, Eq. (30)"},{"comment":"The notation χ²_{n-1}((n-1)/v_i w_i) d((n-1)/v_i w_i) may confuse readers because the argument and the differential are written in a nonstandard way. Consider introducing an explicit variable, e.g., z_i = (n-1) w_i / v_i, and then applying the scaling property to obtain the gamma density for w_i.","section":"Section 3.1, Eq. (36)"},{"comment":"The discussion around Eq. (40) and the footnote about the naive treatment leading to V[s_i] ~ 0 is somewhat cryptic. A brief explanation of why this approximation fails would improve readability.","section":"Section 3.2, Eq. (40)"},{"comment":"The caption states that ε=0.5 corresponds to 'our version of D'Agostini's criterion V[ω]=E²[ω]' but does not point to the equation where this criterion is introduced. Adding a reference to Eq. (26) and Eq. (27) would be helpful.","section":"Figure 1 caption"},{"comment":"The term 'democratic skepticism' is introduced in a footnote but not defined in the main text. Since the paper later refers to this concept in the summary, a brief definition in the main text would make the argument more self-contained.","section":"Section 2.2, footnote 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a hep-ph journal and addresses a practical problem in data combination. The mathematical core is sound, but the central claim of 'structural equivalence' is stronger than what the comparison of auxiliary gamma densities demonstrates. The authors should either provide a full comparison of the likelihood/posterior or carefully qualify their claim. I do not see any citation or novelty concerns. The paper is likely to be a useful contribution after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a short, honest paper that connects two known gamma-based recipes for handling errors on errors. The explicit dictionary k = lambda = (n-1)/2 = 1/(4 epsilon^2) is, as far as I can tell, not stated in the cited literature, and the paper is transparent that it is not proposing a new statistical model. The algebra in Sec. 4 is correct, and the writing is clear.\n\nThe main soft spot is Eq. (44), the bridge between D'Agostini's ratio s_i^2/sigma_i^2 and Cowan's hatw_i/v_i. It is asserted \"on physical grounds\" rather than derived. For systematic uncertainties, treating the quoted error as if it were a sample variance from n virtual repetitions is a modeling assumption, not a mathematical consequence. The paper should say this more carefully and state where the analogy breaks down.\n\nThere is a second, related overreach: the paper compares the auxiliary gamma densities, not the full inference procedures. D'Agostini marginalizes over the true variance, while Cowan uses a sampling distribution for an observed variance. The authors do not show that the resulting likelihoods or posteriors for mu coincide, so \"structurally equivalent\" is stronger than what is demonstrated. If that phrase were softened to \"the auxiliary variance models match parameter by parameter,\" the paper would be solid.\n\nThe citation pattern is clean: the relevant D'Agostini and Cowan papers are cited, and there is no circular reasoning or self-citation loop. The practical relevance is modest but real: it gives practitioners a rationale for choosing k = lambda when setting gamma priors in data combinations.\n\nWho is this for? HEP physicists who combine inconsistent measurements and want to understand where the gamma prior comes from. It is a useful clarifying note, not a technical breakthrough. I would send it to peer review; a serious referee can request a toned-down equivalence claim and a clearer statement of the modeling assumptions behind Eq. (44). With those revisions, it would be a reasonable contribution to the literature.","headline":"A correct short note connecting D'Agostini and Cowan on errors on errors; the parameter mapping is real, but the claimed structural equivalence outruns what is actually shown.","tokens_in":8610,"tokens_out":4864,"would_cite":true,"duration_ms":48327,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Bayesian and frequentist prescriptions for uncertain errors are the same statistical model, matched parameter by parameter.","keywords":["errors on errors","gamma distribution","Bayesian inference","frequentist inference","uncertainty combination","virtual repetitions","PDG scale factor","structural equivalence"],"falsifier":"Simulate many experiments under Cowan's model with a fixed integer $n$ (for example $n=3$, so $\\varepsilon=0.5$ and $k=\\lambda=1$), form D'Agostini's posterior interval for $\\mu$, and measure its frequentist coverage; if coverage deviates from the nominal rate, the claimed structural equivalence fails. A sharper observation internal to the paper is that D'Agostini's recommended $k=1.29$ gives $n=2k+1=3.58$, which cannot be a literal count of virtual repetitions, so the equivalence is a formal parameter identity rather than a realizable sampling scheme unless $n$ is treated as effective and fractional.","tokens_in":7553,"feed_emoji":"⚛️","tokens_out":11325,"duration_ms":105172,"temperature":0.7,"pith_summary":"This paper sets out to prove that the two standard remedies for 'errors on errors'—the Bayesian gamma-prior prescription and the frequentist virtual-repetition sampling model—are structurally identical, not merely analogous. The identification runs through the ratio $\\omega_i=s_i^2/\\hat\\sigma_i^2$ in the Bayesian picture and $\\hat w_i/v_i$ in the frequentist picture; equating their probability elements forces $k=\\lambda=(n-1)/2=1/(4\\varepsilon^2)$. That means the symmetric gamma prior that the Bayesian approach advocated as a heuristic is exactly the distribution the frequentist approach obtains by imagining each quoted variance as the sample variance of $n$ virtual repetitions. If true, the Particle Data Group's practical $S$-factor rescaling becomes a special case of one coherent probabilistic model, and Bayesian prior choices acquire a concrete frequentist meaning. A reader should care because the result converts a philosophical divide into a testable equivalence.","feed_headline":"Bayesian and frequentist 'errors on errors' are the same model","feed_subtitle":"It ties D'Agostini's gamma prior to Cowan's virtual repetitions and grounds the PDG rescale.","key_machinery":"The load-bearing object is the gamma-distributed auxiliary ratio: $\\omega_i=s_i^2/\\hat\\sigma_i^2$ on the Bayesian side and $\\hat w_i/v_i$ on the frequentist side, joined by the correspondence in Eq. (44). The derivation uses the scaling property of the gamma distribution to rewrite Cowan's chi-squared probability element in exactly the same canonical form as D'Agostini's prior, so that equal shape and rate parameters follow immediately. This machinery shows that D'Agostini's symmetric prior choice $k=\\lambda$ is not an ad hoc recommendation but the frequentist statement that each quoted variance is a sample variance from $n=2k+1$ virtual repetitions.","core_discovery":"On the paper's own terms, the central finding is a parameter-by-parameter mapping between D'Agostini's Bayesian treatment and Cowan's frequentist treatment of uncertainty in quoted variances. D'Agostini models the inverse-square ratio $\\hat\\omega_i=s_i^2/\\hat\\sigma_i^2$ as a gamma random variable with shape $k$ and rate $\\lambda$; Cowan models each measured variance $\\hat w_i$ as the sample variance of $n$ virtual normal draws around the true variance $v_i$, which makes $\\hat w_i/v_i$ gamma-distributed with shape and rate $(n-1)/2$. The paper asserts on physical grounds that these two ratios are the same quantity (Eq. (44)); matching the probability elements then gives $k=\\lambda=(n-1)/2=1/(4\\varepsilon^2)$, where $\\varepsilon$ is Cowan's errors-on-errors parameter and $n=1+1/(2\\varepsilon^2)$. The authors state explicitly that they are not proposing a new model, but unifying two existing ones by showing their structural equivalence.","pith_inferences":["An implication the authors leave implicit is a calibration rule: any Bayesian choice of $k$ can be translated into an effective number of repetitions $n=2k+1$, so priors with small $k$ describe variance estimates based on less than one effective replication—a useful diagnostic for excessively heavy-tailed combinations.","A testable extension would be to estimate $\\varepsilon$ directly from the scatter of published systematic uncertainties across a set of comparable measurements, then compare the implied $k=1/(4\\varepsilon^2)$ with the value that best fits the combined posterior; agreement would validate the correspondence on real data rather than algebraically.","If the equivalence holds for independent errors, a natural multivariate extension is to let the $n$ virtual repetitions be correlated across experiments, turning Cowan's scalar model into a correlated gamma-variance model; this is an extrapolation, not a claim the paper makes.","The failure of $n=2k+1$ to be an integer for typical fitted $k$ suggests the 'virtual repetitions' should be understood as an effective statistical device, not a literal experimental design; that reading preserves the mathematical equivalence while abandoning any claim about actual replication counts."],"forward_implications":["The heuristic Bayesian prior $k=\\lambda$ is replaced by a principled frequentist interpretation: it says the quoted variance is an average over $n=2k+1$ virtual replications, with relative uncertainty $\\varepsilon=1/\\sqrt{2(n-1)}$.","D'Agostini's 'democratic skepticism'—identical gamma parameters for all experiments—corresponds to assuming the same number of virtual repetitions for every experiment, so one scalar $\\varepsilon$ controls the skepticism.","The marginal likelihood becomes a Student's $t$ with $\\nu=2k$ degrees of freedom, so the correspondence fixes the heavy-tailed behavior of the combined result in terms of the relative error-on-error.","The PDG scale factor $S=\\max(1,\\sqrt{\\chi^2/(N-1)})$ is not a standalone fudge factor; it is the practical face of a gamma-variance averaging model with uncertainty in quoted variances.","Setting $k=\\lambda=1$ (i.e., $\\varepsilon=0.5$, $n=3$) is a defensible default because it makes the variance of the variance ratio equal to its squared mean."],"supporting_citations":[{"why":"Supplies the PDG scale-factor prescription $S=\\max(1,\\sqrt{\\chi^2/(N-1)})$ that the paper seeks to put on a theoretical footing.","marker":"[1]"},{"why":"Provides D'Agostini's Bayesian gamma-prior formulation with $\\omega_i=s_i^2/\\hat\\sigma_i^2$ and the recommended $k=1.29$, $\\lambda=0.589$.","marker":"[2]"},{"why":"Provides Cowan's frequentist model of $n$ virtual repetitions and the errors-on-errors parameter $\\varepsilon$, whose sampling gamma the Bayesian prior is matched to.","marker":"[3]"}],"fun_headline_variants":["Same gamma, two names: errors-on-errors unified","Bayesian and frequentist error rescale are identical","D'Agostini's prior = Cowan's virtual repeats","One gamma model for all reported variance errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on the physical identification that the Bayesian ratio of the quoted variance to the theoretical variance, $s_i^2/\\hat\\sigma_i^2$, is the same quantity as the frequentist ratio of the estimated variance to the true variance, $\\hat w_i/v_i$; if this mapping is rejected, the equality $k=\\lambda=(n-1)/2$ does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Same gamma, two names: errors-on-errors unified","Bayesian and frequentist error rescale are identical","D'Agostini's prior = Cowan's virtual repeats","One gamma model for all reported variance errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000546,"raw_usage":{"total_tokens":2568,"prompt_tokens":860,"completion_tokens":1708,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":1643}},"tokens_in":476,"tokens_out":1708,"duration_ms":12215,"temperature":1.0,"reasoning_tokens":1643,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:40:30.514672+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate many experiments under Cowan's model with a fixed integer $n$ (for example $n=3$, so $\\varepsilon=0.5$ and $k=\\lambda=1$), form D'Agostini's posterior interval for $\\mu$, and measure its frequentist coverage; if coverage deviates from the nominal rate, the claimed structural equivalence fails. A sharper observation internal to the paper is that D'Agostini's recommended $k=1.29$ gives $n=2k+1=3.58$, which cannot be a literal count of virtual repetitions, so the equivalence is a formal parameter identity rather than a realizable sampling scheme unless $n$ is treated as effective and fractional.","supporting_citations":[{"cited_title":"Navas, C","cited_arxiv_id":null,"evidence_quote":"Supplies the PDG scale-factor prescription $S=\\max(1,\\sqrt{\\chi^2/(N-1)})$ that the paper seeks to put on a theoretical footing."},{"cited_title":"Sceptical combination of experimental results: General considerations and application to epsilon-prime/epsilon","cited_arxiv_id":"hep-ex/9910036","evidence_quote":"Provides D'Agostini's Bayesian gamma-prior formulation with $\\omega_i=s_i^2/\\hat\\sigma_i^2$ and the recommended $k=1.29$, $\\lambda=0.589$."}],"review_version":1}