Pith. sign in

REVIEW 4 major objections 3 minor 17 references

Measuring Statistical Evidence: A Short Report

T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This report argues that statistical evidence is the change in belief and that the relative belief ratio, the posterior-to-prior probability ratio, is the primary measure of that evidence.

desk verdict A readable survey of Evans's Relative Belief program, but as a standalone preprint it misstates its own central ratio in a way that must be fixed before the formal parts can be trusted. read the letter →

arxiv 2411.16831 v2 pith:NQM6ROII submitted 2024-11-25 stat.ME math.STstat.TH

classification stat.MEmath.STstat.TH MSC 62A0162F15
keywords statisticalevidencerelativebeliefratiochangeinBayesfactorJeffreys-LindleyparadoxstrengthofBayesianinferencep-valuecritique
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This report argues that statistical evidence should be defined as the change in belief produced by data, not by the data's likelihood or a fixed threshold. On that definition, the relative belief ratio $RB(A|C)=P(A|C)/P(A)$ is the natural way to measure evidence. The paper contrasts this with likelihood ratios, p-values, confidence intervals, and Bayes factors, claiming each either needs an arbitrary cutoff or confuses evidence with its strength. It separates the measurement of evidence from the measurement of its strength, using the posterior distribution to calibrate whether evidence for or against a hypothesis is strong or weak. If the account is right, evidence becomes context-dependent and can detect prior-induced bias, and the Jeffreys-Lindley paradox is resolved by noting that large Bayes factors can still be weak evidence.

What carries the argument

The load-bearing object is the relative belief ratio, $RB(A|C)=P(A|C)/P(A)$, the change in belief expressed as a ratio. In the continuous case the paper defines a generalized version as a limit of ratios on 'nice' neighbourhoods, yielding the density ratio $f(\xi_1|\xi_2)/f_{\psi_1}(\xi_1)$, and in Bayesian inference the posterior-to-prior density ratio. Its key properties are invariance under smooth 1-1 reparameterization, a general additivity identity, the Savage-Dickey identity relating it to prior predictive densities, and a posterior-based strength measure that calibrates how strongly the data support a value. These properties are what let the paper separate evidence from strength and diagnose prior bias.

What would settle it

Compute the generalized relative belief ratio for the normal example using two different natural sequences of shrinking neighbourhoods, such as intervals versus balls of different aspect ratios, with the same data and prior. If the limiting ratios differ, or if one sequence fails to converge, the core definition of generalized evidence is not well-defined and the central claim does not have a continuous-space formulation.

Watch

Extended reading notes

Core claim

The central claim is the Principle of Evidence: for an event $A$ and obtained information $C$, if $P(A|C)>P(A)$ there is evidence in favor of $A$, if $P(A|C)<P(A)$ evidence against, and equality means no evidence. This makes evidence a ratio-scale concept, measured by the relative belief ratio $RB(A|C)=P(A|C)/P(A)$, which in continuous Bayesian settings becomes the posterior-to-prior density ratio. The report argues this ratio, not the Bayes factor, is the primary characterization of statistical evidence, because it is invariant under 1-1 reparameterization, satisfies additivity and a Savage-Dickey identity, and separates the existence of evidence from its strength. Strength is assessed by posterior tail probabilities, so a large relative belief ratio can still be weak evidence when the posterior mass concentrates elsewhere. The analysis of the Jeffreys-Lindley paradox then shows that a diffuse prior can inflate the ratio while the strength measure remains small, and that a prior-bias calculation can flag when weak evidence is an artifact of the prior.

Load-bearing premise

The argument depends on the assumption that the generalized relative belief ratio, defined as a limit over specially chosen shrinking neighbourhoods, always exists and does not depend on which neighbourhoods are chosen; if that fails, the continuous-space examples and the Jeffreys-Lindley analysis are undefined.

Editorial extensions

If this is right

  • If evidence is change in belief, then likelihood ratios, p-values, and confidence intervals are at best indirect evidence measures, since none is defined in terms of prior-to-posterior change.
  • Evidence strength must be reported with evidence, and strength is context-dependent: the same relative belief ratio can be strong evidence in one posterior and weak in another.
  • Large Bayes factors should not be read as strong evidence; the Jeffreys-Lindley paradox is resolved by separating the relative belief ratio from its posterior-calibrated strength.
  • Prior bias is quantifiable before seeing data, so a user can check whether a chosen prior artificially favors or disfavors a hypothesis.
  • The maximum relative belief estimate and the $ \gamma$-relative belief regions supply accuracy statements naturally from the same evidence measure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: if evidence is defined by change in belief, then the same data can yield different evidence under different priors, and the method is only as objective as the prior; the report's stance is that prior elicitation is a separate, falsifiable step.
  • Testable extension: one could discretize a continuous parameter space in several natural ways and compare the discrete relative belief ratios with the continuous limiting ratio, checking whether the neighbourhood choice actually matters in applications.
  • Editorial extension: if the account is right, scientific reporting would shift from reporting p-values or Bayes factors alone to reporting posterior-to-prior ratios plus their posterior-calibrated strengths.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This short report argues that statistical evidence should be defined as the change in belief induced by conditioning on observed data, and it proposes the Relative Belief Ratio as the primary measure of evidence. It surveys pure likelihood inference, Birnbaum's theorem, p-values, confidence regions, and Bayes factors, then develops relative belief inference: the ratio, strength of evidence, relative belief regions, plausibility, bias, and a revisitation of the Jeffreys-Lindley paradox. The exposition relies heavily on Evans (2015) and related work, and most formal results are stated as restatements of that program rather than derived here.

Significance. If the formal corrections below are made, the paper would be a compact pedagogical introduction to the relative belief program, with a useful separation of evidence from its strength and an intuitive demonstration of the Jeffreys-Lindley resolution. It does not claim to prove new optimality results, and the proposed modification in Definition 21 is explicitly unproven. The main value of the manuscript is synthesis and classroom accessibility rather than novel theory; its significance therefore depends on the accuracy of the exposition, which currently contains load-bearing internal inconsistencies.

major comments (4)
  1. [Section 3.1, Definition 18] As printed, RB(A|C)=P(A|C)/P(C) contradicts Definition 17 and the paper's own examples: evidence is declared by comparing P(A|C) with P(A), which corresponds to P(A|C)/P(A), not division by P(C). Lemma 4 (Savage-Dickey) and Example 3.1.1 both use P(A) and P(B) in the denominator; for A=C the printed formula gives RB(C|C)=1/P(C), which is not generally equal to 1. The denominator of the central ratio must be corrected to P(A) throughout, or the formal foundation of the report is internally inconsistent.
  2. [Section 3.2, Example 3.2.2] The sentence "our estimation using relative belief preference ordering is ψM RBE(x)=20.72" conflates the evidence ratio with the parameter estimate. In the N(θ,1) model with n=50 and x̄√n=1.96, the MRBE is approximately x̄≈0.277, whereas 20.72 is the computed value of RB(0|x)=BF(H0|x). The subsequent bias calculation at ψ′=ψM RBE therefore evaluates bias at the wrong point, and the illustrative conclusion about prior bias is not supported as written.
  3. [Section 3.2, Example 3.2.1] The second case repeats the condition "RB(ψ1|x)>1 and ΠΨ(ψ1|x) is small"; it should read RB(ψ1|x)<1 (with ΠΨ(ψ2|x) large). As printed, the example does not actually illustrate evidence against ψ1, which is the motivation for Definition 21 and the subsequent discussion of strength of evidence.
  4. [Section 3.1, Definition 19; Section 3.2, Definition 21] The generalized relative belief ratio is defined only as a limit over "nice" neighbourhoods with regularity conditions deferred to Evans (2015), yet the continuous-space examples and the Jeffreys-Lindley calculation use the limiting formula without stating those conditions. Similarly, Definition 21 is introduced with the admission that no formal justification is available. Because these two definitions carry the continuous-case and strength-of-evidence claims, the paper should either state the required conditions explicitly or clearly mark the results as conditional on unproved regularity and unproved choice of strength measure.
minor comments (3)
  1. [Throughout] There are numerous typos and misspellings, including "frequntist", "sufficent", "densify", "Jefferys-Lindely", "Savage-Dicky" (should be Savage-Dickey), and "Chambernowne" (should be Champernowne).
  2. [Section 2.3, Example 2.3.1] The text states T(X)∼N(θ,1/n^2); for a sample mean of n i.i.d. N(θ,1) observations the variance should be 1/n, so the displayed distribution needs correction.
  3. [Section 3.1, Lemma 4] After correcting Definition 18, the statement of Lemma 4 should be rechecked: the identity RB(A|C)=RB(C|A) holds for the ratio P(A|C)/P(A), not for the printed P(A|C)/P(C).

Circularity Check

1 steps flagged · score 4.0 of 10

The 'primary method' conclusion leans on Evans's own relative-belief program via load-bearing self-citation, but no fitted-input or equation-level circularity appears.

  1. self citation load bearing [Section 3.1, after Definition 18; also Section 3.2 and Conclusion]
    "There is an axiomatic construction of relative belief which can be found in (Evans, 2015), but like any other axiomatization, concerns can be raised. However, it's really the power that the theory gives us that demonstrates its appropriateness."

    The report's central conclusion—that the Relative Belief Ratio is 'the primary method of characterizing statistical evidence'—is supported by asserting that the theory is appropriate because of its axiomatic construction and because of its power/optimality properties, both deferred to Evans (2015). Evans is the supervisor and originator of the relative-belief framework, so the cited support is not an external, independently verified source but the same research program that already asserts the centrality of this ratio. The 'motivation' for the primary-method claim therefore reduces to an appeal to that program's own authority.

full rationale

No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors to force the choice, and no equation-level derivation is equivalent to its own input by construction. The Principle of Evidence (Definition 17) is explicitly taken as an axiom, and the Relative Belief Ratio is proposed as a natural ratio measure of change in belief; that is a philosophical/conventional step, not a circular derivation. The worked examples and the Jeffreys-Lindley analysis are genuine demonstrations of the framework rather than predictions forced by fitted values. The main circularity concern is the paper's heavy reliance on Evans (2015) for the axiomatic construction, regularity conditions in Definition 19, and 'numerous optimality properties'; the 'primary method' conclusion is inherited from the supervisor's research program rather than independently proved here. A separate correctness issue, not itself circularity, is that Definition 18 prints RB(A|C) = P(A|C)/P(C), while the Principle of Evidence and Example 3.1.1 use P(A|C)/P(A); Lemma 4 and the example are only consistent with the latter denominator. That formal inconsistency should be repaired, but it does not constitute a circular reduction.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No data are fitted and no numbers are estimated; user-supplied quantities (discretization δ, region threshold γ, prior π) are part of the proposed methodology rather than parameters fitted to the central claim. The axioms listed are the background assumptions the exposition depends on, especially the conditional-probability axiom and the regularity of the generalized ratio. No new physical or mathematical entities are introduced; the Relative Belief Ratio is a defined measure from prior literature.

assumptions (5)
  • domain assumption Principle of Conditional Probability
    Section 1.4 states it as an axiom: 'there is no mathematical justification of why we should do so, but this seems the most plausible way to modify beliefs.'
  • standard math Kolmogorov probability axioms with countable additivity
    Section 1.3 assumes the probability triple (Ω, A, P) with countable additivity, which implies continuity needed for conditional probability.
  • domain assumption Existence of a valid Information Generator Υ
    Section 1.4 requires Υ to specify the obtained information B = Υ^{-1}{ξ0}; misapplications are blamed on its absence, and its existence and correctness are assumed.
  • domain assumption Finite parameter spaces are fundamental; continuity is an approximation requiring a user-supplied discretization δ
    Sections 1.1.3 and 1.1.4 argue that real-world inferences are for finite spaces and a discretization must be supplied; the generalized Relative Belief ratio relies on this view.
  • ad hoc to paper Regularity conditions for the generalized Relative Belief limit
    Definition 19 says 'under regularity conditions' the ratio equals a density ratio, and 'nice' neighbourhoods are required; conditions are not stated, only cited to Evans (2015) Appendix.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Measuring Statistical Evidence: A Short Report." pith.science (2026). https://pith.science/paper/NQM6ROII

@misc{pith2026241116831,
  author       = {Pith},
  title        = {Pith review of: Measuring Statistical Evidence: A Short Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NQM6ROII}},
  note         = {Machine review of arXiv:2411.16831}
}
read the original abstract

This short text tried to establish a big picture of what evidential statistics is about and how an ideal inference method should behave. Moreover, by examining shortcomings of some of the currently used methods for measuring evidence and utilizing some intuitive principles, we motivated the Relative Belief Ratio as the primary method of characterizing statistical evidence. Number of topics has been omitted for the interest of this text and the reader is strongly advised to refer to (Evans, 2015) as the primary source for further readings of the subject.

Figures

Figures reproduced from arXiv: 2411.16831 by the authors.

Figure 1
Figure 1. H0 : θ = 0 9 [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

17 extracted references · 13 canonical work pages

  1. [1]

    Al-Labadi, L., Alzaatreh, A., & Evans, M. (2024). How to Measure Evidence and Its Strength: Bayes Factors or Relative Belief Ratios? [arXiv:2301.08994]. https://doi.org/10.48550/arXiv.2301. 08994

  2. [2]

    Birnbaum, A. (1962). On the Foundations of Statistical Inference [Publisher: ASA Website eprint: https://www.tandfonline.com/doi/pdf/10.1080/01621459.1962.10480660]. Journal of the Amer- ican Statistical Association, 57 (298), 269–306. https://doi.org/10.1080/01621459.1962.10480660

  3. [3]

    Carnap, R. (1950). Logical Foundations of Probability. Chicago University of Chicago Press

  4. [4]

    R., & Hinkley, D

    Cox, D. R., & Hinkley, D. V. (1979). Theoretical Statistics. Chapman; Hall/CRC. https://doi.org/10. 1201/b14832

  5. [5]

    Evans, M. (1989). An example concerning the likelihood function. Statistics & Probability Letters , 7 (5), 417–418. https://doi.org/10.1016/0167-7152(89)90097-7

  6. [6]

    Evans, M. (2013). What does the proof of birnbaum’s theorem prove? https://arxiv.org/abs/1302.5468

  7. [7]

    Evans, M. (2015). Measuring Statistical Evidence Using Relative Belief . Chapman; Hall/CRC. https : //doi.org/10.1201/b18587

  8. [8]

    Evans, M. (2024). The concept of statistical evidence: Historical roots and current developments. https: //arxiv.org/abs/2406.05843

Show all 17 references
  1. [9]

    Fisher, R. A. (1922). On the Mathematical Foundations of Theoretical Statistics [Publisher: Royal So- ciety]. Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character , 222, 309–368. Retrieved November 25, 2...

  2. [10]

    Glymour, C. N. (1980). Theory and Evidence . Princeton University Press

  3. [11]

    Hume, D. (1739). A Treatise of Human Nature. https://oll.libertyfund.org/titles/bigge- a- treatise- of- human-nature Jeffreys, t. l. H. (1998). Theory of Probability (Third Edition, Third Edition). Oxford University Press

  4. [12]

    (2008).An Introduction to Kolmogorov Complexity and Its Applications

    Li, M., & Vit´ anyi, P. (2008).An Introduction to Kolmogorov Complexity and Its Applications . Springer. https://doi.org/10.1007/978-0-387-49820-1

  5. [13]

    G., & Spanos, A

    Mayo, D. G., & Spanos, A. (2006). Severe Testing as a Basic Concept in a Neyman?Pearson Philosophy of Induction [Publisher: University of Chicago Press]. British Journal for the Philosophy of Science , 57 (2), 323–357. https://doi.org/10.1093/bjps/axl003

  6. [14]

    Peterson, M. (2017). An introduction to decision theory (2nd ed.). Cambridge University Press

  7. [15]

    Popper, K. R. (2002). The Logic of Scientific Discovery [Google-Books-ID: Yq6xeupNStMC]. Psychology Press

  8. [16]

    Ramdas, A., & Wang, R. (2024). Hypothesis testing with e-values. https://arxiv.org/abs/2410.23614

  9. [17]

    Royall, R. (2017). Statistical Evidence: A Likelihood Paradigm . Routledge. https://doi.org/10.1201/ 9780203738665 19

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.