REVIEW 4 major objections 3 minor 17 references
Measuring Statistical Evidence: A Short Report
T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This report argues that statistical evidence is the change in belief and that the relative belief ratio, the posterior-to-prior probability ratio, is the primary measure of that evidence.
desk verdict A readable survey of Evans's Relative Belief program, but as a standalone preprint it misstates its own central ratio in a way that must be fixed before the formal parts can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the relative belief ratio, $RB(A|C)=P(A|C)/P(A)$, the change in belief expressed as a ratio. In the continuous case the paper defines a generalized version as a limit of ratios on 'nice' neighbourhoods, yielding the density ratio $f(\xi_1|\xi_2)/f_{\psi_1}(\xi_1)$, and in Bayesian inference the posterior-to-prior density ratio. Its key properties are invariance under smooth 1-1 reparameterization, a general additivity identity, the Savage-Dickey identity relating it to prior predictive densities, and a posterior-based strength measure that calibrates how strongly the data support a value. These properties are what let the paper separate evidence from strength and diagnose prior bias.
What would settle it
Compute the generalized relative belief ratio for the normal example using two different natural sequences of shrinking neighbourhoods, such as intervals versus balls of different aspect ratios, with the same data and prior. If the limiting ratios differ, or if one sequence fails to converge, the core definition of generalized evidence is not well-defined and the central claim does not have a continuous-space formulation.
Extended reading notes
Core claim
The central claim is the Principle of Evidence: for an event $A$ and obtained information $C$, if $P(A|C)>P(A)$ there is evidence in favor of $A$, if $P(A|C)<P(A)$ evidence against, and equality means no evidence. This makes evidence a ratio-scale concept, measured by the relative belief ratio $RB(A|C)=P(A|C)/P(A)$, which in continuous Bayesian settings becomes the posterior-to-prior density ratio. The report argues this ratio, not the Bayes factor, is the primary characterization of statistical evidence, because it is invariant under 1-1 reparameterization, satisfies additivity and a Savage-Dickey identity, and separates the existence of evidence from its strength. Strength is assessed by posterior tail probabilities, so a large relative belief ratio can still be weak evidence when the posterior mass concentrates elsewhere. The analysis of the Jeffreys-Lindley paradox then shows that a diffuse prior can inflate the ratio while the strength measure remains small, and that a prior-bias calculation can flag when weak evidence is an artifact of the prior.
Load-bearing premise
The argument depends on the assumption that the generalized relative belief ratio, defined as a limit over specially chosen shrinking neighbourhoods, always exists and does not depend on which neighbourhoods are chosen; if that fails, the continuous-space examples and the Jeffreys-Lindley analysis are undefined.
Editorial extensions
If this is right
- If evidence is change in belief, then likelihood ratios, p-values, and confidence intervals are at best indirect evidence measures, since none is defined in terms of prior-to-posterior change.
- Evidence strength must be reported with evidence, and strength is context-dependent: the same relative belief ratio can be strong evidence in one posterior and weak in another.
- Large Bayes factors should not be read as strong evidence; the Jeffreys-Lindley paradox is resolved by separating the relative belief ratio from its posterior-calibrated strength.
- Prior bias is quantifiable before seeing data, so a user can check whether a chosen prior artificially favors or disfavors a hypothesis.
- The maximum relative belief estimate and the $ \gamma$-relative belief regions supply accuracy statements naturally from the same evidence measure.
Reading between the lines
- Editorial extension: if evidence is defined by change in belief, then the same data can yield different evidence under different priors, and the method is only as objective as the prior; the report's stance is that prior elicitation is a separate, falsifiable step.
- Testable extension: one could discretize a continuous parameter space in several natural ways and compare the discrete relative belief ratios with the continuous limiting ratio, checking whether the neighbourhood choice actually matters in applications.
- Editorial extension: if the account is right, scientific reporting would shift from reporting p-values or Bayes factors alone to reporting posterior-to-prior ratios plus their posterior-calibrated strengths.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This short report argues that statistical evidence should be defined as the change in belief induced by conditioning on observed data, and it proposes the Relative Belief Ratio as the primary measure of evidence. It surveys pure likelihood inference, Birnbaum's theorem, p-values, confidence regions, and Bayes factors, then develops relative belief inference: the ratio, strength of evidence, relative belief regions, plausibility, bias, and a revisitation of the Jeffreys-Lindley paradox. The exposition relies heavily on Evans (2015) and related work, and most formal results are stated as restatements of that program rather than derived here.
Significance. If the formal corrections below are made, the paper would be a compact pedagogical introduction to the relative belief program, with a useful separation of evidence from its strength and an intuitive demonstration of the Jeffreys-Lindley resolution. It does not claim to prove new optimality results, and the proposed modification in Definition 21 is explicitly unproven. The main value of the manuscript is synthesis and classroom accessibility rather than novel theory; its significance therefore depends on the accuracy of the exposition, which currently contains load-bearing internal inconsistencies.
major comments (4)
- [Section 3.1, Definition 18] As printed, RB(A|C)=P(A|C)/P(C) contradicts Definition 17 and the paper's own examples: evidence is declared by comparing P(A|C) with P(A), which corresponds to P(A|C)/P(A), not division by P(C). Lemma 4 (Savage-Dickey) and Example 3.1.1 both use P(A) and P(B) in the denominator; for A=C the printed formula gives RB(C|C)=1/P(C), which is not generally equal to 1. The denominator of the central ratio must be corrected to P(A) throughout, or the formal foundation of the report is internally inconsistent.
- [Section 3.2, Example 3.2.2] The sentence "our estimation using relative belief preference ordering is ψM RBE(x)=20.72" conflates the evidence ratio with the parameter estimate. In the N(θ,1) model with n=50 and x̄√n=1.96, the MRBE is approximately x̄≈0.277, whereas 20.72 is the computed value of RB(0|x)=BF(H0|x). The subsequent bias calculation at ψ′=ψM RBE therefore evaluates bias at the wrong point, and the illustrative conclusion about prior bias is not supported as written.
- [Section 3.2, Example 3.2.1] The second case repeats the condition "RB(ψ1|x)>1 and ΠΨ(ψ1|x) is small"; it should read RB(ψ1|x)<1 (with ΠΨ(ψ2|x) large). As printed, the example does not actually illustrate evidence against ψ1, which is the motivation for Definition 21 and the subsequent discussion of strength of evidence.
- [Section 3.1, Definition 19; Section 3.2, Definition 21] The generalized relative belief ratio is defined only as a limit over "nice" neighbourhoods with regularity conditions deferred to Evans (2015), yet the continuous-space examples and the Jeffreys-Lindley calculation use the limiting formula without stating those conditions. Similarly, Definition 21 is introduced with the admission that no formal justification is available. Because these two definitions carry the continuous-case and strength-of-evidence claims, the paper should either state the required conditions explicitly or clearly mark the results as conditional on unproved regularity and unproved choice of strength measure.
minor comments (3)
- [Throughout] There are numerous typos and misspellings, including "frequntist", "sufficent", "densify", "Jefferys-Lindely", "Savage-Dicky" (should be Savage-Dickey), and "Chambernowne" (should be Champernowne).
- [Section 2.3, Example 2.3.1] The text states T(X)∼N(θ,1/n^2); for a sample mean of n i.i.d. N(θ,1) observations the variance should be 1/n, so the displayed distribution needs correction.
- [Section 3.1, Lemma 4] After correcting Definition 18, the statement of Lemma 4 should be rechecked: the identity RB(A|C)=RB(C|A) holds for the ratio P(A|C)/P(A), not for the printed P(A|C)/P(C).
Circularity Check
The 'primary method' conclusion leans on Evans's own relative-belief program via load-bearing self-citation, but no fitted-input or equation-level circularity appears.
-
self citation load bearing
[Section 3.1, after Definition 18; also Section 3.2 and Conclusion]
"There is an axiomatic construction of relative belief which can be found in (Evans, 2015), but like any other axiomatization, concerns can be raised. However, it's really the power that the theory gives us that demonstrates its appropriateness."
The report's central conclusion—that the Relative Belief Ratio is 'the primary method of characterizing statistical evidence'—is supported by asserting that the theory is appropriate because of its axiomatic construction and because of its power/optimality properties, both deferred to Evans (2015). Evans is the supervisor and originator of the relative-belief framework, so the cited support is not an external, independently verified source but the same research program that already asserts the centrality of this ratio. The 'motivation' for the primary-method claim therefore reduces to an appeal to that program's own authority.
full rationale
No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors to force the choice, and no equation-level derivation is equivalent to its own input by construction. The Principle of Evidence (Definition 17) is explicitly taken as an axiom, and the Relative Belief Ratio is proposed as a natural ratio measure of change in belief; that is a philosophical/conventional step, not a circular derivation. The worked examples and the Jeffreys-Lindley analysis are genuine demonstrations of the framework rather than predictions forced by fitted values. The main circularity concern is the paper's heavy reliance on Evans (2015) for the axiomatic construction, regularity conditions in Definition 19, and 'numerous optimality properties'; the 'primary method' conclusion is inherited from the supervisor's research program rather than independently proved here. A separate correctness issue, not itself circularity, is that Definition 18 prints RB(A|C) = P(A|C)/P(C), while the Principle of Evidence and Example 3.1.1 use P(A|C)/P(A); Lemma 4 and the example are only consistent with the latter denominator. That formal inconsistency should be repaired, but it does not constitute a circular reduction.
Assumptions & free parameters
assumptions (5)
- domain assumption Principle of Conditional Probability
- standard math Kolmogorov probability axioms with countable additivity
- domain assumption Existence of a valid Information Generator Υ
- domain assumption Finite parameter spaces are fundamental; continuity is an approximation requiring a user-supplied discretization δ
- ad hoc to paper Regularity conditions for the generalized Relative Belief limit
Cite this review
Pith. "Pith review of Measuring Statistical Evidence: A Short Report." pith.science (2026). https://pith.science/paper/NQM6ROII
@misc{pith2026241116831,
author = {Pith},
title = {Pith review of: Measuring Statistical Evidence: A Short Report},
year = {2026},
howpublished = {\url{https://pith.science/paper/NQM6ROII}},
note = {Machine review of arXiv:2411.16831}
}
read the original abstract
This short text tried to establish a big picture of what evidential statistics is about and how an ideal inference method should behave. Moreover, by examining shortcomings of some of the currently used methods for measuring evidence and utilizing some intuitive principles, we motivated the Relative Belief Ratio as the primary method of characterizing statistical evidence. Number of topics has been omitted for the interest of this text and the reader is strongly advised to refer to (Evans, 2015) as the primary source for further readings of the subject.
Figures
Reference graph
Works this paper leans on
-
[1]
Al-Labadi, L., Alzaatreh, A., & Evans, M. (2024). How to Measure Evidence and Its Strength: Bayes Factors or Relative Belief Ratios? [arXiv:2301.08994]. https://doi.org/10.48550/arXiv.2301. 08994
work page Pith review arXiv doi:10.48550/arxiv.2301.08994 2024
-
[2]
Birnbaum, A. (1962). On the Foundations of Statistical Inference [Publisher: ASA Website eprint: https://www.tandfonline.com/doi/pdf/10.1080/01621459.1962.10480660]. Journal of the Amer- ican Statistical Association, 57 (298), 269–306. https://doi.org/10.1080/01621459.1962.10480660
arXiv 1962
-
[3]
Carnap, R. (1950). Logical Foundations of Probability. Chicago University of Chicago Press
work page 1950
-
[4]
Cox, D. R., & Hinkley, D. V. (1979). Theoretical Statistics. Chapman; Hall/CRC. https://doi.org/10. 1201/b14832
work page 1979
-
[5]
Evans, M. (1989). An example concerning the likelihood function. Statistics & Probability Letters , 7 (5), 417–418. https://doi.org/10.1016/0167-7152(89)90097-7
-
[6]
Evans, M. (2013). What does the proof of birnbaum’s theorem prove? https://arxiv.org/abs/1302.5468
work page Pith review arXiv 2013
-
[7]
Evans, M. (2015). Measuring Statistical Evidence Using Relative Belief . Chapman; Hall/CRC. https : //doi.org/10.1201/b18587
-
[8]
Evans, M. (2024). The concept of statistical evidence: Historical roots and current developments. https: //arxiv.org/abs/2406.05843
work page Pith review arXiv 2024
Show all 17 references
-
[9]
Fisher, R. A. (1922). On the Mathematical Foundations of Theoretical Statistics [Publisher: Royal So- ciety]. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character , 222, 309–368. Retrieved November 25, 2...
1922
-
[10]
Glymour, C. N. (1980). Theory and Evidence . Princeton University Press
1980
-
[11]
Hume, D. (1739). A Treatise of Human Nature. https://oll.libertyfund.org/titles/bigge- a- treatise- of- human-nature Jeffreys, t. l. H. (1998). Theory of Probability (Third Edition, Third Edition). Oxford University Press
1998
-
[12]
(2008).An Introduction to Kolmogorov Complexity and Its Applications
Li, M., & Vit´ anyi, P. (2008).An Introduction to Kolmogorov Complexity and Its Applications . Springer. https://doi.org/10.1007/978-0-387-49820-1
2008 doi
-
[13]
G., & Spanos, A
Mayo, D. G., & Spanos, A. (2006). Severe Testing as a Basic Concept in a Neyman?Pearson Philosophy of Induction [Publisher: University of Chicago Press]. British Journal for the Philosophy of Science , 57 (2), 323–357. https://doi.org/10.1093/bjps/axl003
2006 doi
-
[14]
Peterson, M. (2017). An introduction to decision theory (2nd ed.). Cambridge University Press
2017
-
[15]
Popper, K. R. (2002). The Logic of Scientific Discovery [Google-Books-ID: Yq6xeupNStMC]. Psychology Press
2002
-
[16]
Ramdas, A., & Wang, R. (2024). Hypothesis testing with e-values. https://arxiv.org/abs/2410.23614
2024 arXiv
-
[17]
Royall, R. (2017). Statistical Evidence: A Likelihood Paradigm . Routledge. https://doi.org/10.1201/ 9780203738665 19
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.