{"id":"46e6844d-b981-42ef-8c5d-86df230b0b89","arxiv_id":"2508.11814","paper_version":4,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Applying simulation-based calibration and binary prediction calibration to the Bayesian model averaging supermodel reliably validates Bayes factor computation, catching errors that data-averaged posterior checks and the Good check miss.","lead":"This paper introduces two statistical checks that test whether Bayes factors, a common tool for comparing scientific models, are computed correctly. The checks catch errors that earlier methods miss, and the authors recommend them for anyone building or using Bayes factor software.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"§3.2 asserts SBC and binary calibration are the same for the model index, but Theorem 6's converse requires Dφ,f non-singular; in the paper's discrete-data examples D is atomic, so the recommended substitution of binary calibration for SBC is not proven in that regime.","rationale":"The reader's weakest assumption, the non-singularity condition in Theorem 6, is indeed the most load-bearing technical point. I agree that it is the weakest link in the central theoretical argument. My stress-test sharpens this into a concrete gap: the paper's §3.2 and abstract state the equivalence unconditionally, and the practical recommendation to prefer binary calibration over SBC for the model index relies on that unconditional reading. However, the theorem itself does state the condition, and in the discrete cases where the condition fails there is reason to suspect the conclusion may still hold (atoms appear to force calibration under SBC), so the issue is best characterized as an overstatement or proof-gap rather than a demonstrated failure of the methodology. The empirical work is extensive, the code is available, and the central claims about sensitivity are supported by the simulations. Therefore the reader's ACCEPT verdict stands, but the manuscript should qualify the equivalence claim in §3.2 and the abstract to match Theorem 6's hypotheses.","tokens_in":29744,"tokens_out":32047,"duration_ms":370493,"concrete_test":"For the §5.1 single-binary-observation model, enumerate all candidate posteriors (b0,b1) on a fine grid over [0,1]^2 (or symbolically), and for each candidate check whether continuous SBC w.r.t. the test quantity f(i,y)=i passes, i.e. whether E[r(x)] = x for all x in [0,1], and whether binary prediction calibration Pr(i=1|D=d)=d holds at every atom of D. If any candidate passes SBC while failing calibration, or vice versa, the unconditional equivalence claimed in §3.2 fails in the discrete regime and the §7.1 recommendation needs qualification; if no discrepancy is found, the non-singularity condition is likely unnecessary for atomic cases and the concern reduces to a proof-gap rather than a substantive defect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 states that for K=2, SBC for the model index in the BMA supermodel 'is in fact the same as binary prediction calibration.' The proof in Appendix A (Theorem 6) establishes only one direction unconditionally: binary calibration implies continuous SBC. The converse requires that the distribution of Dφ,f = Pr_phi(f=1|y) be non-singular, i.e. a mixture of an absolutely continuous part and point masses. In the paper's own discrete-data examples (§5.1 single binary observation; §5.2 Poisson vs. NB counts), D takes countably or finitely many values, so the converse is not covered by the theorem as stated. The unqualified equivalence is then used in §7.1 to justify the recommendation to prefer binary calibration over SBC for the model index: 'For checking the calibration of the model index, we would currently recommend binary prediction calibration over SBC.' If the two checks can disagree in the atomic/counting-data regime, that recommendation is not supported by the proved result, and the abstract's 'in theory ... equivalent' is stronger than what is proven. The paper does state the non-singularity condition in the theorem, but the narrative in §3.2 and the abstract does not carry the caveat, so the central theoretical claim is overbroad as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops and evaluates simulation-based checks for Bayes factor computation. It introduces an improved variant of simulation-based calibration checking (SBC) applied to the Bayesian model averaging (BMA) supermodel, and proposes binary prediction calibration of the model index as a complementary check. The main theoretical result (Theorem 6 in Appendix A) establishes that, for binary test quantities, binary prediction calibration implies continuous SBC, and the converse holds under a non-singularity condition on the distribution of binary predictions. The authors also analyse data-averaged posterior checking and the Good check, showing limitations of the latter, and they introduce posterior SBC for settings with improper priors. Extensive simulation studies on toy models and realistic examples validate the bridgesampling and BayesFactor R packages. The paper includes reproducible code.","tokens_in":30009,"tokens_out":10742,"duration_ms":108149,"significance":"The paper provides a practical and theoretically grounded toolkit for validating Bayes factor computations, a topic of increasing importance. The empirical comparison is thorough, the code is publicly available, and the demonstration that the Good check lacks error control and that data-averaged posterior checking can be blind to important failure modes is valuable. The posterior SBC extension for improper priors is a useful contribution. The main theoretical equivalence result is elegant, but the narrative overstates its scope (see major comment below), so the central claim needs qualification before the paper can be accepted.","major_comments":[{"comment":"The paper states in Section 3.2 that for K=2, SBC for the model index in the BMA supermodel is \"in fact the same as binary prediction calibration\", and the abstract claims that \"in theory, binary prediction calibration is equivalent to a special case of SBC\". This is stronger than what is proved. Theorem 6 in Appendix A proves that binary prediction calibration implies continuous SBC unconditionally, but the converse requires that the distribution of D_{φ,f} is not singular. In the paper's own discrete-data examples (Section 5.1 with a binary observation, Section 5.2 with Poisson/NB counts), D takes finitely or countably many values, so the distribution is singular, and the converse is not covered by the theorem as stated. The unqualified equivalence is subsequently used in Section 7.1 to justify preferring binary prediction calibration over SBC for the model index. Please qualify the equivalence in the abstract and Section 3.2 by stating the non-singularity condition, and either prove the converse for the atomic case or explicitly note that the two checks may in principle disagree when D is discrete.","section":"Section 3.2 and Abstract"},{"comment":"The claim that \"with well-designed test quantities, SBC can however detect all possible problems in computation\" is a strong theoretical statement. It is credible given the arguments in Modrák et al. (2025), but the paper should make explicit that this is an in-principle statement requiring the user to choose suitable test quantities, and that finite simulation budgets and practical choice of test quantities will limit sensitivity. A brief clarification would prevent overinterpretation by practitioners.","section":"Section 3.2 and Abstract"}],"minor_comments":[{"comment":"The main text refers to \"Theorem 1 in Appendix A\" for the SBC–binary calibration equivalence, but the theorem is numbered Theorem 6 in the appendix. Similarly, \"Theorem 3 in Appendix A\" in Section 3.4 appears to refer to Theorem 9. Please fix these cross-references.","section":"Section 3.2 and Appendix A"},{"comment":"The captions state that the vertical orange line \"marks when the check first attains 80% power\", but the procedure for computing this point from the 100 resampled histories is not described. Please clarify how the 80% power point is defined and how the shared simulation pool is used.","section":"Figures 2 and 3"},{"comment":"The recommendation that \"at least several hundred simulations per validation scenario should be run\" is reasonable, but given that some of the paper's own scenarios require 20,000 simulations to detect mild miscalibration (Appendix E), consider adding a more explicit warning that the required number depends on the severity and type of the potential problem.","section":"Section 7.3"},{"comment":"There are several places where \"bridgesamplingandBayesFactorR\" appears without proper spacing, and the phrase \"or a warning on inaccuracy\" is unclear; please specify what kind of warning was observed.","section":"Abstract and Section 7.3"},{"comment":"The sentence \"Modrák et al. (2025) show that checking continuous SBC is equivalent to checking M-sample SBC for all M\" could be misread as claiming equivalence for a single finite M; rephrase to clarify that passing continuous SBC is equivalent to passing M-sample SBC for every M simultaneously.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper builds heavily on the authors' own prior work (Modrák et al., 2025), but the new theorem is proved independently and the empirical validation of external R packages is a genuine contribution. The overstatement of the equivalence between binary prediction calibration and SBC is the main obstacle; it can be fixed by qualifying the claims in the abstract and Section 3.2 and by discussing the atomic-distribution case. Once that is done, the paper is a solid methodological contribution suitable for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new contribution is using SBC on the BMA supermodel to validate Bayes factor computation, together with a real theorem connecting continuous SBC for binary test quantities to binary prediction calibration. The proof in Appendix A is careful, the empirical studies are extensive, and the code is on GitHub. This is a useful paper.\n\nWhat it does well: it shows data-averaged posterior checking misses plausible bugs (flipped model indices, partially ignored data, added noise) that SBC or binary calibration catch; it shows the Good check, as originally proposed, has no error control and can be nearly useless because BF moments can have infinite or huge variance; and it extends the checks to improper priors via posterior SBC. The validation of bridgesampling and BayesFactor is a practical bonus.\n\nThe soft spot is a mismatch between the theorem and how it is described. Theorem 6 proves binary calibration implies continuous SBC unconditionally, but the converse requires the distribution of Dφ,f to be non-singular. The abstract and §3.2 say the two are \"equivalent\" for K=2 without carrying that caveat, and §7.1 recommends binary calibration over SBC for the model index. In their own discrete-data examples (§5.1, §5.2), D is atomic, so the converse is not proven in exactly the regime where they apply the language of equivalence. This doesn't sink the methods—the recommendation about finite-sample power stands on its own—but the theoretical claim should be qualified in the abstract and main text, and the authors should say explicitly whether the converse can fail with atomic D or whether it can be recovered by a different argument. As written, the paper proves less than it claims.\n\nMinor: the theorem numbering jumps (Theorem 1 in the text, Theorem 6 in the appendix), and the \"detect all problems\" sentence is a statement of principle, not a guarantee; they do qualify it with \"well-designed test quantities,\" so that's acceptable.\n\nWho this is for: anyone implementing or evaluating Bayes factor software, and people doing Bayesian workflow diagnostics. The paper deserves a serious referee. I would send it to review; the fix is a rewrite of a few passages, not a change in the substance.","headline":"SBC for the BMA supermodel is a real advance for Bayes factor validation, and the binary-calibration equivalence is close to right—but the paper states it more broadly than the theorem proves.","tokens_in":30528,"tokens_out":3994,"would_cite":true,"duration_ms":44710,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62C10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that validating Bayes factor computation reduces to checking one binary prediction: whether the computed model probability is calibrated.","keywords":["Bayes factor","simulation-based calibration","binary prediction calibration","Bayesian model averaging","marginal likelihood","improper priors","model validation","calibration"],"falsifier":"Simulate a deliberately broken Bayes factor computation in which the reported posterior model probabilities take only two values, such as 0.5 and 1, by e.g. ignoring half the data in a Poisson-versus-negative-binomial comparison, then run both checks; binary prediction calibration should report near-perfect calibration while the SBC gamma statistic on a data-dependent quantity rejects, confirming the non-singularity condition is doing real work.","tokens_in":29520,"feed_emoji":"📊","tokens_out":6910,"duration_ms":69192,"temperature":0.7,"pith_summary":"The paper's central claim is that Bayes factor computation can be validated by treating the competing models as one Bayesian model-averaging supermodel and checking the calibration of the model index: when the computed posterior probability for a model is $p$, that model should be the true one in $p$ of simulated cases. It proves a theoretical equivalence: for binary test quantities, continuous simulation-based calibration (SBC) and binary prediction calibration are the same check, provided the distribution of predicted probabilities is not singular. Empirically, the paper shows that with finite simulations binary calibration is often the more sensitive of the two, while SBC with data-dependent test quantities catches problems such as ignored data, noisy Bayes factors, and biased normalization constants that both binary calibration and data-averaged posterior checks miss. It further shows that the Good check, as originally described, does not control error rates, and that a posterior version of SBC is needed when models have improper priors. For the standard scenarios tested, the paper finds that the Bayes factor implementations it examined satisfy all available checks.","feed_headline":"Bayes factor checks: model index calibration catches what older checks miss","feed_subtitle":"Binary prediction calibration equals SBC for model indices, and standard Bayes factor software passes all checks.","key_machinery":"The machinery is the BMA supermodel: a hierarchical model in which the model index $i$ is drawn from a categorical prior, parameters are drawn from the chosen submodel's prior, and data are drawn from that submodel's likelihood. The posterior probability of the model index is a one-to-one transform of the Bayes factor via Bayes' theorem, so checking correctness of one checks the other. Two test statistics carry the argument: continuous SBC, which ranks the prior draw within the candidate posterior and demands uniformity, and binary prediction calibration, which demands $\\Pr(i=k \\mid D_k=p)=p$ for the reported probability $D_k$. Theorem 6 is the pivot: for binary test quantities these two conditions coincide whenever the distribution of $D_k$ has no singular part, so the practical choice between them is a matter of sensitivity and finite-sample power.","core_discovery":"The discovery is a direct reduction: a correct Bayes factor computation is equivalent to calibrated predictions about which model generated the data. Working inside the BMA supermodel, the paper defines binary prediction calibration as the condition that, among all simulated datasets for which an algorithm reports posterior model probability $p$, the model it favors is indeed the data-generating model in exactly a proportion $p$ of them. Theorem 6 states that this condition holds if and only if the candidate posterior passes continuous SBC with respect to the binary test quantity $f(\\cdot,y) = \\text{model index}$, with the caveat that the distribution of predicted probabilities must not be singular. The paper also shows that passing calibration forces the correct data-averaged posterior, so data-averaged checking is a strictly weaker test in the infinite-simulation limit; conversely, SBC with suitably chosen test quantities can detect any mismatch, including ones invisible to binary calibration. On the practical side, the paper demonstrates that posterior SBC resolves apparent miscalibration caused by improper priors, and it validates widely used Bayes factor software on several real models.","pith_inferences":["Editorial inference: the non-singularity caveat in Theorem 6 means practitioners should not choose between SBC and binary calibration solely by theoretical equivalence; running both and reporting disagreements is a cheap diagnostic for atom-heavy prediction distributions.","Editorial inference: the same binary-calibration logic could be applied directly to other reported quantities like posterior predictive $p$-values or confidence-interval coverage, since any binary event can play the role of the model index.","Editorial inference: the sensitivity contrast between data-dependent SBC and model-index checks suggests that future validation workflows should include at least one test quantity that is a function of the data itself, not just parameters.","Editorial inference: the slow convergence of the Good check implies moment-based validation is generally fragile for heavy-tailed Bayes factor distributions; rank-based or calibration-based checks are better suited to a validation standard."],"forward_implications":["Bayes factor software can be validated end-to-end by checking one binary prediction, requiring only simulations from the prior and the models, with several hundred simulations as a practical minimum.","Data-averaged posterior checks should be used only as a supplement: they add no information in the infinite-simulation limit and miss flipped model indices, ignored data, and noisy computations.","The Good check as originally implemented should not be relied on for error control, because Bayes factor distributions can have infinite or very large variance.","When models contain improper priors, estimating a posterior SBC prior from a fixed dataset is necessary; using fixed or surrogate values for shared parameters produces false miscalibration.","For the standard scenarios tested, the Bayes factor packages examined pass all available checks, suggesting they are safe for routine low-stakes use."],"supporting_citations":[{"why":"Supplies the improved SBC methodology, including the role of test quantities, that the paper applies to the BMA supermodel.","marker":"Modrák et al. 2025"},{"why":"Origin of the prior/posterior exchangeability that underlies simulation-based calibration.","marker":"Talts et al. 2020"},{"why":"Introduces posterior SBC, which the paper shows is necessary for Bayes factor checks under improper priors.","marker":"Säilynoja et al. 2025"},{"why":"Provides the bootstrap miscalibration test used to implement binary prediction calibration empirically.","marker":"Dimitriadis et al. 2021"},{"why":"Earlier workflow for Bayes factor validation based on data-averaged posterior checking, which the paper shows is weaker than SBC.","marker":"Schad et al. 2023"},{"why":"The Good check method whose error-rate control and convergence the paper analyses and finds lacking.","marker":"Sekulovski et al. 2024"},{"why":"Introduces data-averaged posterior checking, the family of methods compared throughout.","marker":"Geweke 2004"},{"why":"Presents the bridge sampling implementation used in the realistic validation examples.","marker":"Gronau et al. 2020"},{"why":"The Bayes factor package validated on several real and toy models in the paper.","marker":"Morey and Rouder 2024"}],"fun_headline_variants":["Bayes factor validation: calibration equals SBC for model indices","Simulation-based checks expose Bayes factor flaws","Binary calibration: the sharper Bayes factor check","Bayes factor software passes new simulation validation","SBC and binary calibration: new gold standard for Bayes factors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The equivalence between the two checks assumes the algorithm's predicted model probabilities have a distribution without singular atoms; if predictions concentrate on a few isolated values, binary calibration can pass while SBC with a data-dependent test quantity flags a problem.","fun_headline_variants_meta":{"raw":{"variants":["Bayes factor validation: calibration equals SBC for model indices","Simulation-based checks expose Bayes factor flaws","Binary calibration: the sharper Bayes factor check","Bayes factor software passes new simulation validation","SBC and binary calibration: new gold standard for Bayes factors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000866,"raw_usage":{"total_tokens":3785,"prompt_tokens":1008,"completion_tokens":2777,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":2703}},"tokens_in":624,"tokens_out":2777,"duration_ms":19807,"temperature":1.0,"reasoning_tokens":2703,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:27:11.702282+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a deliberately broken Bayes factor computation in which the reported posterior model probabilities take only two values, such as 0.5 and 1, by e.g. ignoring half the data in a Poisson-versus-negative-binomial comparison, then run both checks; binary prediction calibration should report near-perfect calibration while the SBC gamma statistic on a data-dependent quantity rejects, confirming the non-singularity condition is doing real work.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Origin of the prior/posterior exchangeability that underlies simulation-based calibration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Good check method whose error-rate control and convergence the paper analyses and finds lacking."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces data-averaged posterior checking, the family of methods compared throughout."},{"cited_title":"F., Singmann, H., and Wagenmakers, E.-J","cited_arxiv_id":null,"evidence_quote":"Presents the bridge sampling implementation used in the realistic validation examples."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Bayes factor package validated on several real and toy models in the paper."}],"review_version":1}