{"id":"2b9bf0e2-7144-4c6e-8ac7-c3051168d7a7","arxiv_id":"2607.10035","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Infima over training-sample corrections yield automatic lower bounds on Intrinsic Bayes Factors and define proper Least Favorable Intrinsic Priors that bridge IBF and Smith–Spiegelhalter methods.","lead":"The paper builds lower bounds on Intrinsic Bayes Factors by taking the worst-case training sample, and defines Least Favorable Intrinsic Priors from those samples. This gives sample-size-aware, more conservative Bayesian evidence against nulls without choosing an average of training samples.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The LF-prior / expanded-sample equality (Prop. 4) is the linchpin of the central claim, yet it is proved only under an unstated conditional-independence assumption that is not verified for the general nested linear model.","rationale":"The Reader correctly isolates a genuine technical weakness—the unquantified design-replicate approximation that produces the simplified LBT (8)—and rightly withholds full acceptance until the missing supplemental derivations appear. That concern, however, is not the most load-bearing one for the paper’s central claim. The claim is not merely that a lower bound exists, but that the bound is realized as an exact Bayes factor under a proper prior (the LF prior). That identification rests on Proposition 4, whose proof inserts an independence assumption never verified outside pure i.i.d. sampling. Because the general nested linear model is the setting in which the paper claims its greatest novelty, the independence gap is more fundamental than the design-approximation gap. The two concerns are related (both arise from the non-i.i.d. structure of linear models), so agreement with the Reader is only partial. The concrete ANOVA check proposed above would settle the issue with a single numerical comparison and does not require the missing supplements. Until that check (or an analytic proof that removes the independence hypothesis) is supplied, the verdict remains CONDITIONAL.","tokens_in":16348,"tokens_out":735,"duration_ms":6677,"concrete_test":"Take the one-way ANOVA layout of §2.4.2 with m=3 unbalanced groups (n1=2,n2=3,n3=5). Construct an explicit extremal training sample y*_s of size n01=4 that attains F=0, form the joint design matrix for the expanded data (y,y*_s), compute both the ordinary expanded Bayes factor B_10(y|y*_s) under the modified Jeffreys prior and the LF Bayes factor obtained by plugging the LF prior (12) into the original sample y alone. If the two numerical values differ by more than machine precision, Prop. 4 does not hold for this design and the central identification fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s strongest claim is that the theoretical lower bound LBT is realized exactly as the Bayes factor under a proper Least Favorable Intrinsic Prior, i.e. B^LF_10(y) = B_10(y | y*_s(ℓ)). Proposition 4 asserts this equality, but the short proof explicitly invokes “assume conditional independence” of the observed sample y and the extremal training sample y*_s(ℓ). In the Gaussian location and precision examples the assumption holds by construction (i.i.d. sampling). In the general nested normal-linear model of §2.4.1, however, the design matrices A_i and A_i(ℓ) share rows whenever the imaginary training design is taken from the same experimental layout; the likelihood factors f(y,y*(ℓ)|θ) therefore do not factor into independent marginals. Consequently the passage from the posterior-prior construction (12)–(13) to the expanded-sample identity (20) is not justified, and the claimed bridge between the bound and a proper prior fails precisely where the paper claims greatest generality. The design-replicate approximation already flagged by the Reader is secondary; even if that approximation were exact, the independence gap would still leave Prop. 4 unproved for non-i.i.d. designs.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper constructs lower (and upper) bounds on Intrinsic Bayes Factors by replacing the usual average of minimal-training-sample correction factors BN_01(y(ℓ)) with their infimum (or supremum of the reciprocal) over the full theoretical training-sample space D*. The resulting theoretical bound LBT is claimed to be realized exactly as the Bayes factor under a proper Least Favorable Intrinsic Prior π^LF_k(θ_k) := π^N_k(θ_k | y_s(ℓ)) obtained from the extremal training sample. Under an additional conditional-independence assumption the LF Bayes factor equals the expanded-sample bound B_10(y | y*(ℓ)). Explicit closed forms are derived for normal precision, normal mean (known and unknown variance), and nested normal-linear models including one-way ANOVA; the bounds are compared with the Sellke–Bayarri–Berger robust bound, with ordinary IBFs, and with BIC, and are illustrated on Student’s sleep data.","tokens_in":16654,"tokens_out":1173,"duration_ms":8526,"significance":"If the claimed bridge holds, the work supplies a single, sample-size-adaptive lower bound that is free of the choice of averaging method (arithmetic, geometric, median, …) and that is simultaneously the exact Bayes factor under a proper, non-degenerate prior. That would give a concrete, operational link between the Intrinsic-Bayes-Factor programme of Berger–Pericchi and the imaginary-training-sample device of Spiegelhalter–Smith, and would furnish a practical tool for robust null-hypothesis assessment that automatically improves with n. The explicit algebra for the classical Gaussian and ANOVA cases, the matching asymptotic relations with BIC and ordinary IBFs (Props. 1–3), and the real-data illustration are concrete contributions that can be checked and used even if the most general claim requires further work.","major_comments":[{"comment":"Proposition 4 (Section 4) asserts that B^LF_10(y) equals the expanded-sample bound B_10(y | y*_s(ℓ)). The short proof explicitly invokes “assume conditional independence” of the observed sample y and the extremal training sample y*(ℓ). In the i.i.d. Gaussian location/precision examples the assumption holds by construction, but in the general nested normal-linear model of §2.4.1 the design matrices A_i and A_i(ℓ) share rows whenever the imaginary training design is taken from the same experimental layout; the joint likelihood therefore does not factor. Consequently the passage from the posterior-prior construction (12)–(13) to the expanded-sample identity (20) is not justified precisely where the paper claims greatest generality. Either a rigorous justification for non-i.i.d. designs or an explicit restriction of the claim to i.i.d. sampling is required.","section":null},{"comment":"In §2.4.1 the general theoretical lower bound is simplified to expression (8) by treating the full design as an approximate (n/n_01)-fold replicate of a minimal training design, i.e. A_i^T A_i ≈ (n/n_01) A_i(ℓ)^T A_i(ℓ). The approximation cancels the design-matrix determinants and yields a closed form that depends only on residual sums of squares. For unbalanced or non-replicable designs the cancellation fails and the reported LBT is not justified. The manuscript should either state the precise design class for which (8) holds, or replace the approximation by an exact (possibly design-dependent) expression.","section":null},{"comment":"Proofs of Propositions 1–3 and several key marginal calculations (e.g., m^LF_0 in the unknown-variance case) are deferred to “supplemental material” that is not supplied with the manuscript. These results underwrite the asymptotic comparisons with ordinary IBFs and BIC that are used to argue that the LF construction is competitive. Without the proofs the central claims cannot be fully verified.","section":null}],"minor_comments":[{"comment":"Numerous typographical and orthographic errors appear throughout (e.g., “authomatically”, “favourable”, “suplemental”, “avaible”, “desi” truncation). A careful copy-edit is needed.","section":null},{"comment":"Notation for the training-sample space oscillates between D, D*, y(ℓ) and y*(ℓ); a single consistent convention would improve readability.","section":null},{"comment":"Figures 1–3 are described but the actual graphics are not embedded in the supplied text; axis labels, legends and numerical values should be checked for legibility once the figures are restored.","section":null},{"comment":"The abstract and introduction repeatedly contrast “which average?” with the new infimum bound; a short explicit statement that the bound is simultaneously a lower bound for every conventional IBF average would make the contribution clearer.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core idea is interesting and the special-case algebra is largely correct, but the two load-bearing gaps (conditional independence for Prop. 4 and the design-replicate approximation) sit exactly at the claimed generality. I would be prepared to recommend acceptance after a revision that either supplies the missing arguments or clearly delimits the scope to i.i.d./balanced designs. The missing supplemental material should be required before any final decision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is simple: replace the contested average of training corrections in Intrinsic Bayes Factors by the extremum over theoretical training samples, get a sample-size-aware lower bound on B01, and realize that extremum as a proper Least Favorable Intrinsic Prior whose Bayes factor matches an expanded-sample construction. That is new relative to Berger–Pericchi and Smith–Spiegelhalter, and it is the right kind of object for people who want a conservative, automatic calibration that still improves with n.\n\nThe Gaussian location and precision examples work cleanly. The algebra is explicit, the LF priors come out as Gamma/Cauchy or Normal, the bounds track F/t/BIC asymptotics in the stated propositions, and the sleep-data illustration is sensible. The bridge to Spiegelhalter–Smith imaginary samples is real, and the comparison with the Sellke–Bayarri–Berger –e p log p bound is fair: their bound is fixed while this one tightens with n. For i.i.d. normal problems the construction is solid enough to use.\n\nThe soft spots are real but localized. Proposition 4 (the linchpin that B^LF equals the expanded bound) is proved only under an explicit conditional-independence assumption. That holds for the i.i.d. examples; it does not automatically hold for the general nested linear model once the imaginary training design shares rows with the observed design. The design-replicate approximation that cancels the determinant terms is likewise unquantified for unbalanced layouts. Several marginals and the proofs of Props. 1–3 are deferred to missing supplemental material, and the writing has the usual notation slips. None of this kills the idea; it just means the generality claim is not yet established.\n\nThis is for people who already work with objective Bayes factors or who need a proper prior that is not an arbitrary point mass. It deserves a serious referee once the independence step is cleaned up for non-i.i.d. designs and the missing derivations appear. I would engage with the Gaussian cases now and wait for the general linear-model version to be tightened.","headline":"Useful conservative IBF bound and proper LF prior for Gaussian cases; the general linear-model claim rests on an unproved independence step and a design-replicate approximation.","tokens_in":17300,"tokens_out":526,"would_cite":false,"duration_ms":5137,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62F03","62J05"],"pacs":[],"model":"grok-4.5","headline":"A lower bound on Intrinsic Bayes Factors is the exact Bayes factor under a proper Least Favorable Intrinsic Prior built from the worst training sample.","keywords":["Intrinsic Bayes Factors","Least Favorable Intrinsic Priors","dynamic bounds on Bayes factors","hypothesis testing","training samples","robust Bayesian inference","model selection"],"falsifier":"Compute both the theoretical lower bound LBT and the exact Least Favorable Bayes factor on a deliberately unbalanced ANOVA or regression design; if they diverge systematically while the sample-size approximation holds for balanced designs, the claimed general closed form fails.","tokens_in":17166,"feed_emoji":"📐","tokens_out":631,"duration_ms":5287,"temperature":0.7,"pith_summary":"P-values discard null hypotheses too readily as samples grow, while ordinary Bayes factors depend on arbitrary choices of how to average over training samples. This paper replaces every such average by its most extreme value over the full space of possible minimal training samples, producing a lower bound on the Intrinsic Bayes Factor that tightens automatically with sample size. The extremal training sample itself defines a proper prior—the Least Favorable Intrinsic Prior—under which the bound becomes an exact Bayes factor. The construction therefore supplies a single, prior-based number that is more conservative than any particular Intrinsic Bayes Factor yet still improves with information, and it recovers the classical imaginary-training-sample idea of Smith and Spiegelhalter as a special case.","feed_headline":"Bayes-factor lower bound becomes exact under worst-case prior","feed_subtitle":"The extremal training sample yields a proper prior that automatically tightens with sample size","key_machinery":"The Basic Lemma identity B10(y(−ℓ)|y(ℓ)) = BN10(y) · BN01(y(ℓ)), optimized by taking the infimum of BN01 over all theoretical minimal training samples rather than any average; the optimizing sample ys(ℓ) then supplies the Least Favorable Intrinsic Prior πLFk(θk) := πNk(θk | ys(ℓ)).","core_discovery":"The Intrinsic Bayes Factor is bounded from below by replacing the average of the training-sample correction factors with their infimum (or the supremum of the reciprocal) over the entire theoretical training-sample space; that same extremal sample defines a proper Least Favorable Intrinsic Prior under which the bound is realized exactly as a Bayes factor, and under conditional independence the Least Favorable Bayes factor equals the expanded-sample bound obtained by treating the imaginary training sample as additional data.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Infimum over training samples yields exact least-favorable Bayes bound","Least-favorable intrinsic prior realizes lower bound as true Bayes factor","Extremal training sample makes intrinsic Bayes bound exact and adaptive","Worst-case prior turns intrinsic Bayes lower bound into exact factor","Intrinsic Bayes bound equals expanded-sample factor under independence"],"cache_read_input_tokens":128,"weakest_assumption_plain":"In linear models the design-matrix determinants cancel only if the full data set can be treated as an approximate replicate of one minimal training design; unbalanced or non-replicable designs break that cancellation.","fun_headline_variants_meta":{"raw":{"variants":["Infimum over training samples yields exact least-favorable Bayes bound","Least-favorable intrinsic prior realizes lower bound as true Bayes factor","Extremal training sample makes intrinsic Bayes bound exact and adaptive","Worst-case prior turns intrinsic Bayes lower bound into exact factor","Intrinsic Bayes bound equals expanded-sample factor under independence"]},"model":"grok-4.5","effort":"low","cost_usd":0.0028,"raw_usage":{"total_tokens":978,"prompt_tokens":671,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":28000000,"prompt_tokens_details":{"text_tokens":671,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":241,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":671,"tokens_out":66,"duration_ms":2756,"temperature":1.0,"reasoning_tokens":241,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T00:51:05.001911+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Compute both the theoretical lower bound LBT and the exact Least Favorable Bayes factor on a deliberately unbalanced ANOVA or regression design; if they diverge systematically while the sample-size approximation holds for balanced designs, the claimed general closed form fails.","supporting_citations":[],"review_version":1}