{"id":"94a3e7ac-f804-4d5b-81a1-c99cd96779d6","arxiv_id":"2411.16831","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A report summarizing and motivating the Relative Belief Ratio framework for statistical evidence, adding one unproven modification to strength-of-evidence measurement.","lead":"This short report argues that statistical evidence should be measured by the change in belief it produces, and promotes the Relative Belief Ratio as the primary tool for that job. It is a survey of existing ideas, mostly drawn from Michael Evans's 2015 book, with a modified strength-of-evidence measure added without proof.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Definition 18 prints the Relative Belief Ratio as P(A|C)/P(C); the Principle of Evidence, Lemma 4, and Example 3.1.1 require P(A|C)/P(A), so the formal central object is internally inconsistent as written.","rationale":"The reader's weakest assumption concerns Definition 19's limit over 'nice' neighbourhoods and the deferred regularity conditions. That concern is real but secondary: the report explicitly argues that continuous spaces are approximations to finite spaces and that correct behaviour in the finite case is sufficient, so the central claim could in principle stand without the generalized limit. The same cannot be said for Definition 18: it is the formal statement of the object the entire report recommends. As printed, RB(A|C) = P(A|C)/P(C) is inconsistent with the Principle of Evidence (Definition 17), with Lemma 4, and with Example 3.1.1. If taken literally, the central ratio is not a measure of change in belief at all. The concrete test settles whether this is a typo. If it is a typo, the report can be accepted after a one-line correction; if it is not, the central claim is formally unsupported. I therefore move the verdict from UNVERDICTED to CONDITIONAL: the paper needs a verifiable, internally consistent formal definition before its central claim can be assessed. This is an internal consistency issue, not a disagreement with the Relative Belief framework or a critique of Evans (2015); the report as submitted, however, does not state its own central object correctly.","tokens_in":16104,"tokens_out":9331,"duration_ms":86344,"concrete_test":"Re-derive the simple case from Definition 17 and Lemma 4. Take A = C with P(C) = 0.5: the printed definition yields RB(C|C) = 2, while the Principle of Evidence requires RB(C|C) = 1. Then take a nontrivial example, e.g. P(A) = 0.4, P(C) = 0.5, P(A∩C) = 0.2: the printed definition gives RB(A|C) = 0.4 and RB(C|A) = 0.5, violating Savage-Dickey; the corrected ratio P(A|C)/P(A) gives RB(A|C) = RB(C|A) = 1.0. Finally compare with the definition of Relative Belief Ratio in Evans (2015). If the corrected denominator reproduces all Section 3.1 formulas and the printed denominator does not, the definition is a typo requiring a one-line revision; if not, the central claim is formally unsupported as written.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.1, Definition 18 defines RB(A|C) = P(A|C)/P(C). Taken literally, this does not measure change in belief: Definition 17 declares evidence by comparing P(A|C) with P(A), which corresponds to P(A|C)/P(A), not division by P(C). The inconsistency is provable internally. For A = C, Definition 18 gives RB(C|C) = 1/P(C), which is not generally 1 even though conditioning on C should leave belief in C unchanged. Lemma 4 asserts RB(A|C) = RB(C|A); under the printed definition this becomes P(A|C)/P(C) = P(C|A)/P(A), which fails in general. Example 3.1.1 also silently uses denominators P(A) and P(B), not P(C), to obtain its quoted values RB(A|C) = m/m1 and RB(B|C) = n1m/(m1n). Thus the formal foundation of the report is mis-stated. Unless this is a typographical error, the central claim lacks a consistent formal definition. The generalized limit in Definition 19 and all continuous examples inherit this ambiguity, since they are motivated as limits of the simple case. This is the most load-bearing issue: the denominator of the central ratio must be fixed and verified before assessing the neighbourhood-limit and regularity conditions.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This short report argues that statistical evidence should be defined as the change in belief induced by conditioning on observed data, and it proposes the Relative Belief Ratio as the primary measure of evidence. It surveys pure likelihood inference, Birnbaum's theorem, p-values, confidence regions, and Bayes factors, then develops relative belief inference: the ratio, strength of evidence, relative belief regions, plausibility, bias, and a revisitation of the Jeffreys-Lindley paradox. The exposition relies heavily on Evans (2015) and related work, and most formal results are stated as restatements of that program rather than derived here.","tokens_in":16390,"tokens_out":5826,"duration_ms":53574,"significance":"If the formal corrections below are made, the paper would be a compact pedagogical introduction to the relative belief program, with a useful separation of evidence from its strength and an intuitive demonstration of the Jeffreys-Lindley resolution. It does not claim to prove new optimality results, and the proposed modification in Definition 21 is explicitly unproven. The main value of the manuscript is synthesis and classroom accessibility rather than novel theory; its significance therefore depends on the accuracy of the exposition, which currently contains load-bearing internal inconsistencies.","major_comments":[{"comment":"As printed, RB(A|C)=P(A|C)/P(C) contradicts Definition 17 and the paper's own examples: evidence is declared by comparing P(A|C) with P(A), which corresponds to P(A|C)/P(A), not division by P(C). Lemma 4 (Savage-Dickey) and Example 3.1.1 both use P(A) and P(B) in the denominator; for A=C the printed formula gives RB(C|C)=1/P(C), which is not generally equal to 1. The denominator of the central ratio must be corrected to P(A) throughout, or the formal foundation of the report is internally inconsistent.","section":"Section 3.1, Definition 18"},{"comment":"The sentence \"our estimation using relative belief preference ordering is ψM RBE(x)=20.72\" conflates the evidence ratio with the parameter estimate. In the N(θ,1) model with n=50 and x̄√n=1.96, the MRBE is approximately x̄≈0.277, whereas 20.72 is the computed value of RB(0|x)=BF(H0|x). The subsequent bias calculation at ψ′=ψM RBE therefore evaluates bias at the wrong point, and the illustrative conclusion about prior bias is not supported as written.","section":"Section 3.2, Example 3.2.2"},{"comment":"The second case repeats the condition \"RB(ψ1|x)>1 and ΠΨ(ψ1|x) is small\"; it should read RB(ψ1|x)<1 (with ΠΨ(ψ2|x) large). As printed, the example does not actually illustrate evidence against ψ1, which is the motivation for Definition 21 and the subsequent discussion of strength of evidence.","section":"Section 3.2, Example 3.2.1"},{"comment":"The generalized relative belief ratio is defined only as a limit over \"nice\" neighbourhoods with regularity conditions deferred to Evans (2015), yet the continuous-space examples and the Jeffreys-Lindley calculation use the limiting formula without stating those conditions. Similarly, Definition 21 is introduced with the admission that no formal justification is available. Because these two definitions carry the continuous-case and strength-of-evidence claims, the paper should either state the required conditions explicitly or clearly mark the results as conditional on unproved regularity and unproved choice of strength measure.","section":"Section 3.1, Definition 19; Section 3.2, Definition 21"}],"minor_comments":[{"comment":"There are numerous typos and misspellings, including \"frequntist\", \"sufficent\", \"densify\", \"Jefferys-Lindely\", \"Savage-Dicky\" (should be Savage-Dickey), and \"Chambernowne\" (should be Champernowne).","section":"Throughout"},{"comment":"The text states T(X)∼N(θ,1/n^2); for a sample mean of n i.i.d. N(θ,1) observations the variance should be 1/n, so the displayed distribution needs correction.","section":"Section 2.3, Example 2.3.1"},{"comment":"After correcting Definition 18, the statement of Lemma 4 should be rechecked: the identity RB(A|C)=RB(C|A) holds for the ratio P(A|C)/P(A), not for the printed P(A|C)/P(C).","section":"Section 3.1, Lemma 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads as a supervised student report summarizing Professor Evans's research program. Its central formal object is mis-stated in Definition 18, and Example 3.2.2 contains a numerical/interpretive error. These are fixable within the manuscript's scope. If the author corrects them and clarifies the paper's contribution relative to Evans (2015) and Al-Labadi et al. (2024), the report could be a useful pedagogical introduction, but novelty is limited."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What should you know? This is a survey-style report, not a new research contribution. Its value is pedagogical: it walks through likelihood, Birnbaum, p-values, confidence intervals, and Bayes factors, then motivates the Relative Belief Ratio from Evans (2015). The writing is clear, and the criticisms of p-values and diffuse-prior Bayes factors are assembled thoughtfully. The one genuinely new element, Definition 21 for strength of evidence, is openly admitted by the author to lack formal justification, so there is no new theorem to evaluate.\n\nWhere the paper does well: the examples are well chosen—the crime/neighbourhood case, the location normal, Jeffreys–Lindley—and the discrete-to-continuous motivation is honest. The paper also flags its own limitations, repeatedly pointing to Evans (2015) for proofs and regularity conditions. That makes it a decent entry point for a graduate student.\n\nThe soft spots: first and most important, Definition 18 defines RB(A|C) = P(A|C)/P(C), but the Principle of Evidence compares P(A|C) to P(A), and all subsequent uses—Lemma 4, Example 3.1.1, and the generalized limit—implicitly use P(A|C)/P(A). Example 3.1.1 literally prints RB(A|C) = P(A|C)/P(A) = m/m1, contradicting the printed definition. As written, the central formal object is wrong; this is not cosmetic, because every continuous example inherits the limit definition. I assume it is a typo, but it has to be fixed before the formal parts can be assessed.\n\nSecond, Definition 19's limit over 'nice' neighbourhoods is deferred to Evans (2015). That is fine for an expository report, but it means the continuous case is not independently established here. Third, Example 3.2.2 has a category slip: RB(H0|x) = 20.72, which is a ratio value, is later reported as ψ_MRBE(x) = 20.72, as though it were a parameter estimate. Also, Example 3.2.1 repeats 'RB > 1 and posterior small' in both cases, when the second case should presumably be RB < 1. These are minor but indicate careless copyediting.\n\nThe citation pattern is heavily concentrated on Evans's own program. That is not disqualifying for a report that is explicitly a summary of that book, but it does mean the 'primary method' conclusion is inherited rather than independently argued.\n\nBottom line: this is not a paper with a defensible new claim. But as a short report it has real expository value, and the formal issue is fixable. I would not publish it as a research article; if it is being considered as a pedagogical or overview piece, it needs correction and a serious referee mainly to catch the Definition 18 error and the numerical slips.","headline":"A readable survey of Evans's Relative Belief program, but as a standalone preprint it misstates its own central ratio in a way that must be fixed before the formal parts can be trusted.","tokens_in":16901,"tokens_out":2229,"would_cite":false,"duration_ms":20598,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62A01","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This report argues that statistical evidence is the change in belief and that the relative belief ratio, the posterior-to-prior probability ratio, is the primary measure of that evidence.","keywords":["statistical evidence","relative belief ratio","change in belief","Bayes factor","Jeffreys-Lindley paradox","strength of evidence","Bayesian inference","p-value critique"],"falsifier":"Compute the generalized relative belief ratio for the normal example using two different natural sequences of shrinking neighbourhoods, such as intervals versus balls of different aspect ratios, with the same data and prior. If the limiting ratios differ, or if one sequence fails to converge, the core definition of generalized evidence is not well-defined and the central claim does not have a continuous-space formulation.","tokens_in":15879,"feed_emoji":"⚖️","tokens_out":7530,"duration_ms":63368,"temperature":0.7,"pith_summary":"This report argues that statistical evidence should be defined as the change in belief produced by data, not by the data's likelihood or a fixed threshold. On that definition, the relative belief ratio $RB(A|C)=P(A|C)/P(A)$ is the natural way to measure evidence. The paper contrasts this with likelihood ratios, p-values, confidence intervals, and Bayes factors, claiming each either needs an arbitrary cutoff or confuses evidence with its strength. It separates the measurement of evidence from the measurement of its strength, using the posterior distribution to calibrate whether evidence for or against a hypothesis is strong or weak. If the account is right, evidence becomes context-dependent and can detect prior-induced bias, and the Jeffreys-Lindley paradox is resolved by noting that large Bayes factors can still be weak evidence.","feed_headline":"Relative belief ratio is the primary measure of statistical evidence","feed_subtitle":"Evidence means changed belief; this report says strength is posterior-calibrated, resolving the Jeffreys-Lindley paradox.","key_machinery":"The load-bearing object is the relative belief ratio, $RB(A|C)=P(A|C)/P(A)$, the change in belief expressed as a ratio. In the continuous case the paper defines a generalized version as a limit of ratios on 'nice' neighbourhoods, yielding the density ratio $f(\\xi_1|\\xi_2)/f_{\\psi_1}(\\xi_1)$, and in Bayesian inference the posterior-to-prior density ratio. Its key properties are invariance under smooth 1-1 reparameterization, a general additivity identity, the Savage-Dickey identity relating it to prior predictive densities, and a posterior-based strength measure that calibrates how strongly the data support a value. These properties are what let the paper separate evidence from strength and diagnose prior bias.","core_discovery":"The central claim is the Principle of Evidence: for an event $A$ and obtained information $C$, if $P(A|C)>P(A)$ there is evidence in favor of $A$, if $P(A|C)<P(A)$ evidence against, and equality means no evidence. This makes evidence a ratio-scale concept, measured by the relative belief ratio $RB(A|C)=P(A|C)/P(A)$, which in continuous Bayesian settings becomes the posterior-to-prior density ratio. The report argues this ratio, not the Bayes factor, is the primary characterization of statistical evidence, because it is invariant under 1-1 reparameterization, satisfies additivity and a Savage-Dickey identity, and separates the existence of evidence from its strength. Strength is assessed by posterior tail probabilities, so a large relative belief ratio can still be weak evidence when the posterior mass concentrates elsewhere. The analysis of the Jeffreys-Lindley paradox then shows that a diffuse prior can inflate the ratio while the strength measure remains small, and that a prior-bias calculation can flag when weak evidence is an artifact of the prior.","pith_inferences":["Editorial extension: if evidence is defined by change in belief, then the same data can yield different evidence under different priors, and the method is only as objective as the prior; the report's stance is that prior elicitation is a separate, falsifiable step.","Testable extension: one could discretize a continuous parameter space in several natural ways and compare the discrete relative belief ratios with the continuous limiting ratio, checking whether the neighbourhood choice actually matters in applications.","Editorial extension: if the account is right, scientific reporting would shift from reporting p-values or Bayes factors alone to reporting posterior-to-prior ratios plus their posterior-calibrated strengths."],"forward_implications":["If evidence is change in belief, then likelihood ratios, p-values, and confidence intervals are at best indirect evidence measures, since none is defined in terms of prior-to-posterior change.","Evidence strength must be reported with evidence, and strength is context-dependent: the same relative belief ratio can be strong evidence in one posterior and weak in another.","Large Bayes factors should not be read as strong evidence; the Jeffreys-Lindley paradox is resolved by separating the relative belief ratio from its posterior-calibrated strength.","Prior bias is quantifiable before seeing data, so a user can check whether a chosen prior artificially favors or disfavors a hypothesis.","The maximum relative belief estimate and the $\n\\gamma$-relative belief regions supply accuracy statements naturally from the same evidence measure."],"supporting_citations":[{"why":"Supplies the formal definition of the generalized relative belief ratio, the regularity conditions for the limit, and the optimality properties the paper relies on.","marker":"(Evans, 2015)"},{"why":"Provides the theorem connecting sufficiency, conditionality, and the likelihood principle that the paper reinterprets as not supporting the likelihood principle.","marker":"(Birnbaum, 1962)"},{"why":"Corrects the statement of Birnbaum's theorem, a key step in the report's survey of why likelihood-based evidence measures fail.","marker":"(Evans, 2013)"},{"why":"Supplies the Bayes factor scale and the point-null prior construction that the report critiques in the Jeffreys-Lindley paradox.","marker":"(Jeffreys, 1998)"},{"why":"Gives the definition of p-values used in the report's demonstration that p-values are not sensitive to sample size and cannot indicate evidence for the null.","marker":"(Cox & Hinkley, 1979)"},{"why":"Provides the general Savage-Dickey ratio identity that connects the relative belief ratio to prior predictive densities.","marker":"(Dickey 1971)"}],"fun_headline_variants":["Evidence is a ratio: relative belief beats Bayes factor","Why relative belief ratio is the metric for statistical evidence","Posterior-to-prior ratio: the core measure of statistical evidence","Relative belief resolves the Jeffreys-Lindley paradox"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the assumption that the generalized relative belief ratio, defined as a limit over specially chosen shrinking neighbourhoods, always exists and does not depend on which neighbourhoods are chosen; if that fails, the continuous-space examples and the Jeffreys-Lindley analysis are undefined.","fun_headline_variants_meta":{"raw":{"variants":["Evidence is a ratio: relative belief beats Bayes factor","Why relative belief ratio is the metric for statistical evidence","Posterior-to-prior ratio: the core measure of statistical evidence","Relative belief resolves the Jeffreys-Lindley paradox"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000707,"raw_usage":{"total_tokens":3129,"prompt_tokens":833,"completion_tokens":2296,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":2230}},"tokens_in":449,"tokens_out":2296,"duration_ms":14859,"temperature":1.0,"reasoning_tokens":2230,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:49:09.390558+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the generalized relative belief ratio for the normal example using two different natural sequences of shrinking neighbourhoods, such as intervals versus balls of different aspect ratios, with the same data and prior. If the limiting ratios differ, or if one sequence fails to converge, the core definition of generalized evidence is not well-defined and the central claim does not have a continuous-space formulation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the formal definition of the generalized relative belief ratio, the regularity conditions for the limit, and the optimality properties the paper relies on."},{"cited_title":"What does the proof of Birnbaum's theorem prove?","cited_arxiv_id":"1302.5468","evidence_quote":"Corrects the statement of Birnbaum's theorem, a key step in the report's survey of why likelihood-based evidence measures fail."},{"cited_title":"R., & Hinkley, D","cited_arxiv_id":null,"evidence_quote":"Gives the definition of p-values used in the report's demonstration that p-values are not sensitive to sample size and cannot indicate evidence for the null."}],"review_version":1}