Pith. sign in

REVIEW 3 major objections 4 minor 10 references

A lower bound on Intrinsic Bayes Factors is the exact Bayes factor under a proper Least Favorable Intrinsic Prior built from the worst training sample.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 00:51 UTC pith:63AV4ZEE

load-bearing objection Useful conservative IBF bound and proper LF prior for Gaussian cases; the general linear-model claim rests on an unproved independence step and a design-replicate approximation. the 3 major comments →

arxiv 2607.10035 v1 pith:63AV4ZEE submitted 2026-07-10 math.ST stat.TH

Bounds on Intrinsic Bayes Factors and Least Favorable Intrinsic Priors for General Statistical Hypothesis Testing

classification math.ST stat.TH MSC 62F1562F0362J05
keywords Intrinsic Bayes FactorsLeast Favorable Intrinsic Priorsdynamic bounds on Bayes factorshypothesis testingtraining samplesrobust Bayesian inferencemodel selection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

P-values discard null hypotheses too readily as samples grow, while ordinary Bayes factors depend on arbitrary choices of how to average over training samples. This paper replaces every such average by its most extreme value over the full space of possible minimal training samples, producing a lower bound on the Intrinsic Bayes Factor that tightens automatically with sample size. The extremal training sample itself defines a proper prior—the Least Favorable Intrinsic Prior—under which the bound becomes an exact Bayes factor. The construction therefore supplies a single, prior-based number that is more conservative than any particular Intrinsic Bayes Factor yet still improves with information, and it recovers the classical imaginary-training-sample idea of Smith and Spiegelhalter as a special case.

Core claim

The Intrinsic Bayes Factor is bounded from below by replacing the average of the training-sample correction factors with their infimum (or the supremum of the reciprocal) over the entire theoretical training-sample space; that same extremal sample defines a proper Least Favorable Intrinsic Prior under which the bound is realized exactly as a Bayes factor, and under conditional independence the Least Favorable Bayes factor equals the expanded-sample bound obtained by treating the imaginary training sample as additional data.

What carries the argument

The Basic Lemma identity B10(y(−ℓ)|y(ℓ)) = BN10(y) · BN01(y(ℓ)), optimized by taking the infimum of BN01 over all theoretical minimal training samples rather than any average; the optimizing sample ys(ℓ) then supplies the Least Favorable Intrinsic Prior πLFk(θk) := πNk(θk | ys(ℓ)).

Load-bearing premise

In linear models the design-matrix determinants cancel only if the full data set can be treated as an approximate replicate of one minimal training design; unbalanced or non-replicable designs break that cancellation.

What would settle it

Compute both the theoretical lower bound LBT and the exact Least Favorable Bayes factor on a deliberately unbalanced ANOVA or regression design; if they diverge systematically while the sample-size approximation holds for balanced designs, the claimed general closed form fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper constructs lower (and upper) bounds on Intrinsic Bayes Factors by replacing the usual average of minimal-training-sample correction factors BN_01(y(ℓ)) with their infimum (or supremum of the reciprocal) over the full theoretical training-sample space D*. The resulting theoretical bound LBT is claimed to be realized exactly as the Bayes factor under a proper Least Favorable Intrinsic Prior π^LF_k(θ_k) := π^N_k(θ_k | y_s(ℓ)) obtained from the extremal training sample. Under an additional conditional-independence assumption the LF Bayes factor equals the expanded-sample bound B_10(y | y*(ℓ)). Explicit closed forms are derived for normal precision, normal mean (known and unknown variance), and nested normal-linear models including one-way ANOVA; the bounds are compared with the Sellke–Bayarri–Berger robust bound, with ordinary IBFs, and with BIC, and are illustrated on Student’s sleep data.

Significance. If the claimed bridge holds, the work supplies a single, sample-size-adaptive lower bound that is free of the choice of averaging method (arithmetic, geometric, median, …) and that is simultaneously the exact Bayes factor under a proper, non-degenerate prior. That would give a concrete, operational link between the Intrinsic-Bayes-Factor programme of Berger–Pericchi and the imaginary-training-sample device of Spiegelhalter–Smith, and would furnish a practical tool for robust null-hypothesis assessment that automatically improves with n. The explicit algebra for the classical Gaussian and ANOVA cases, the matching asymptotic relations with BIC and ordinary IBFs (Props. 1–3), and the real-data illustration are concrete contributions that can be checked and used even if the most general claim requires further work.

major comments (3)
  1. Proposition 4 (Section 4) asserts that B^LF_10(y) equals the expanded-sample bound B_10(y | y*_s(ℓ)). The short proof explicitly invokes “assume conditional independence” of the observed sample y and the extremal training sample y*(ℓ). In the i.i.d. Gaussian location/precision examples the assumption holds by construction, but in the general nested normal-linear model of §2.4.1 the design matrices A_i and A_i(ℓ) share rows whenever the imaginary training design is taken from the same experimental layout; the joint likelihood therefore does not factor. Consequently the passage from the posterior-prior construction (12)–(13) to the expanded-sample identity (20) is not justified precisely where the paper claims greatest generality. Either a rigorous justification for non-i.i.d. designs or an explicit restriction of the claim to i.i.d. sampling is required.
  2. In §2.4.1 the general theoretical lower bound is simplified to expression (8) by treating the full design as an approximate (n/n_01)-fold replicate of a minimal training design, i.e. A_i^T A_i ≈ (n/n_01) A_i(ℓ)^T A_i(ℓ). The approximation cancels the design-matrix determinants and yields a closed form that depends only on residual sums of squares. For unbalanced or non-replicable designs the cancellation fails and the reported LBT is not justified. The manuscript should either state the precise design class for which (8) holds, or replace the approximation by an exact (possibly design-dependent) expression.
  3. Proofs of Propositions 1–3 and several key marginal calculations (e.g., m^LF_0 in the unknown-variance case) are deferred to “supplemental material” that is not supplied with the manuscript. These results underwrite the asymptotic comparisons with ordinary IBFs and BIC that are used to argue that the LF construction is competitive. Without the proofs the central claims cannot be fully verified.
minor comments (4)
  1. Numerous typographical and orthographic errors appear throughout (e.g., “authomatically”, “favourable”, “suplemental”, “avaible”, “desi” truncation). A careful copy-edit is needed.
  2. Notation for the training-sample space oscillates between D, D*, y(ℓ) and y*(ℓ); a single consistent convention would improve readability.
  3. Figures 1–3 are described but the actual graphics are not embedded in the supplied text; axis labels, legends and numerical values should be checked for legibility once the figures are restored.
  4. The abstract and introduction repeatedly contrast “which average?” with the new infimum bound; a short explicit statement that the bound is simultaneously a lower bound for every conventional IBF average would make the contribution clearer.

Circularity Check

0 steps flagged

No significant circularity: LBT and LF priors are explicit constructions (extrema of IBF training corrections; posteriors from those samples), with Prop. 4 a short derivation under stated independence; self-citations are foundational prior art, not a forcing loop.

full rationale

The paper begins from the Berger–Pericchi Basic Lemma (already standard) and defines empirical/theoretical bounds simply by replacing the arithmetic average of training-sample correction factors with their max/min or sup/inf (Eqs. 4–6). The Least Favorable Intrinsic Prior is then defined, by construction, as the ordinary posterior under the improper prior given the extremal training sample that attains that sup/inf (Eq. 12); the LF Bayes factor is the ordinary Bayes factor under those proper priors. Proposition 4 shows, under an explicitly stated conditional-independence assumption, that this BF equals the expanded-sample quantity B10(y|y*(ℓ)); the short algebraic proof is just the definition of the posterior-prior construction and does not reduce a claimed prediction to an input by tautology. The design-matrix approximation used for nested linear models (A_iᵀA_i ≈ (n/n01)A_i(ℓ)ᵀA_i(ℓ)) is an explicit modelling assumption, not a circular fit. Self-citations to Berger–Pericchi (1996) and Spiegelhalter–Smith are ordinary references to the IBF and imaginary-training-sample literature that the paper extends; they supply the starting lemma and the idea of imaginary samples, not an unverified uniqueness theorem that forces the numerical bound. No free parameter is estimated from data and then re-presented as a prediction; no ansatz is smuggled via citation; the concrete expressions for normal mean/precision and ANOVA are obtained by direct calculus or residual-sum-of-squares algebra. The derivation chain is therefore self-contained and non-circular.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 2 invented entities

The work is theoretical Bayesian methodology. It inherits the Intrinsic Bayes Factor program (minimal training samples, Basic Lemma, improper default priors) and Smith–Spiegelhalter imaginary samples. No numerical constants are fitted to data. The main invented objects are the Least Favorable Intrinsic Prior and the associated LF Bayes factor. Load-bearing modeling choices are the default prior family (especially modified Jeffreys vs reference) and the design-replicate approximation used to cancel determinants in linear models.

axioms (6)
  • domain assumption Berger–Pericchi Basic Lemma: B_10(y(−ℓ)|y(ℓ)) = B^N_10(y) × B^N_01(y(ℓ)) for a minimal training sample.
    Invoked at the start of §1 as the starting identity for all bounds; taken from Berger & Pericchi (1996) without re-proof.
  • domain assumption Improper default priors (Jeffreys / modified Jeffreys / reference) may be used for common parameters; indeterminacy is resolved only via training samples.
    Throughout §2; standard in objective Bayesian testing but essential—without it the BN factors are undefined.
  • domain assumption Minimal training sample size k = max{dim(θ_i), dim(θ_j)} (or n_01 = max{p_0+1, p_1+1} in linear models) yields proper posteriors for all models compared.
    Definition of D and D* in §1 and §2.4; if the sample is not truly minimal/properizing, the correction factors are invalid.
  • ad hoc to paper Full design matrices satisfy A_i^T A_i ≈ (n/n_01) A_i(ℓ)^T A_i(ℓ) so determinants cancel in the general LBT.
    §2.4.1 leading to equation (8); an approximation not generally true for arbitrary designs.
  • standard math Conditional independence of observations given parameters when equating B^LF to the expanded-sample bound (Prop. 4).
    Stated explicitly in the proof of Prop. 4; standard i.i.d./conditional-independence modeling assumption.
  • domain assumption Modified Jeffreys prior (q_j = p_j − p_i, q_i = 0) is preferred because reference prior makes the bound uninformative.
    §2.4.5 conclusions; choice affects whether the bound is useful, and is justified by appeal to Berger–Pericchi rather than a uniqueness theorem.
invented entities (2)
  • Least Favorable Intrinsic Prior (LF prior) no independent evidence
    purpose: Proper prior defined as the posterior under the default improper prior given the extremal (least favorable) training sample y_s(ℓ); used to define an exact LF Bayes factor.
    Introduced in §3 as π^LF_k(θ_k) := π^N_k(θ_k | y_s(ℓ)). Independent evidence is conceptual (equals expanded bound under stated assumptions), not an external empirical prediction.
  • Least Favorable Bayes Factor / theoretical IBF lower bound (LBT) no independent evidence
    purpose: Conservative, sample-size-dependent lower bound on evidence for the null obtained by inf/sup over theoretical training samples rather than an average.
    Defined in §1 eqs. (6) and developed through examples; the main methodological output of the paper.

pith-pipeline@v1.1.0-grok45 · 20284 in / 3898 out tokens · 35706 ms · 2026-07-14T00:51:05.001911+00:00 · methodology

0 comments
read the original abstract

Hypothesis Testing is the most contentious procedure in statistical Methodology. P values rejects Null Hypotheses far too easily, specially for large samples. On the other hand, Bayes Factors depends on assumptions, for example regarding Intrinsic Bayes Factors, which average? Arithmetic, Geometric, Median? Our bound is the infimum over all the averages. We develop a lower bound on Intrinsic Bayes Factors that adjust authomatically with the sample size. Furthermore, we introduce the new idea of {\it{\textbf{Least Favorable Intrinsic Prior}}}, which corresponds to the least favourable possible training samples. The bound sets a bridge between Intrinsic Bayes Factors and Adrian Smith and David Spiegelhalter methodology.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

10 extracted references

  1. [1]

    and Pericchi, Luis R

    Berger, James O. and Pericchi, Luis R. , title =. Journal of the American Statistical Association , year =

  2. [2]

    , title =

    Pericchi, Luis R. , title =. Handbook of Statistics , editor =. 2005 , volume =

  3. [3]

    Smith, A. F. M. and Spiegelhalter, D. J. , title =. Journal of the Royal Statistical Society: Series B (Methodological) , year =

  4. [4]

    Spiegelhalter, D. J. and Smith, A. F. M. , title =. Journal of the Royal Statistical Society: Series B (Methodological) , year =

  5. [5]

    Sellke, Thomas and Bayarri, M. J. and Berger, James O. , title =. The American Statistician , year =

  6. [6]

    Expected-Posterior Prior Distributions for Model Selection , journal =

    P. Expected-Posterior Prior Distributions for Model Selection , journal =. 2002 , volume =

  7. [7]

    and Liu, G

    Pericchi, Luis R. and Liu, G. and Torres, D. , title =. Bayesian Evaluation of Informative Hypotheses , editor =. 2008 , chapter =

  8. [8]

    and Pericchi, Luis R

    Berger, James O. and Pericchi, Luis R. and Varshavsky, Julia A. , title =. Sankhya: The Indian Journal of Statistics, Series A (1961-2002) , year =

  9. [9]

    and Peebles, A

    Cushny, Arthur R. and Peebles, A. R. , title =. The Journal of Physiology , year =

  10. [10]

    The Annals of Statistics , year =

    Schwarz, Gideon , title =. The Annals of Statistics , year =