{"id":"df501c3d-6a87-4338-93e9-ba47d7397475","arxiv_id":"2508.01403","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An empirical Bayes factor based on a Savage-Dickey density ratio tests whether all random effects are zero without fitting multiple models.","lead":"This paper proposes an empirical Bayes factor for testing whether random effects exist in mixed models. The method aims to avoid manual prior choices and to test all random effects from a single model fit.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The EBF's Savage-Dickey ratio uses the data both to form the random-effects prior and to evaluate the likelihood, so it is not a calibrated Bayes factor unless the authors prove otherwise.","rationale":"The reader's UNVERDICTED verdict is cautious and appropriate given that only the abstract is available. My stress-test identifies the same weakest assumption: the empirical-Bayes double-use of the data in the Savage-Dickey ratio. This concern is load-bearing because the central claim is that the EBF provides a valid test of variance components without prior specification; if the data used to form the prior is the same data used to evaluate the likelihood, the resulting ratio is not a Bayes factor and has an unknown null distribution. The proposed concrete test is a direct simulation check of calibration under the null hypothesis, which would settle whether the double-use actually inflates false positives. I do not change the verdict because the full text may address this point; the concern is a reason to remain unverified rather than a demonstrated refutation. No ad hominem is intended; the issue is purely methodological.","tokens_in":687,"tokens_out":5658,"duration_ms":76203,"concrete_test":"Simulate 1,000 datasets from a one-way random-intercepts model with the variance component set to 0, residual variance known, and 50 groups of size 10. For each replicate, compute the proposed EBF from the full model fit. Record the median log EBF and the proportion of replicates for which log EBF exceeds log(3). Repeat the same simulation using a properly specified Bayes factor with a fixed prior on the variance component, for example half-Cauchy(0,1). If the EBF has a substantially higher median log BF or a higher rate of exceeding the log(3) threshold than the proper BF, or if the EBF is undefined for any replicate because the variance estimate is exactly zero, then the double-use of the data invalidates the EBF as a calibrated evidence measure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the Savage-Dickey density ratio computed after estimating the random-effects distribution from the data is a valid Bayes factor. The abstract states that the random-effects distribution is 'part of the model's lower level and estimated from the data,' which means the prior density in the denominator of the Savage-Dickey ratio is a function of the observed data rather than a prior fixed before seeing y. Consequently, the proposed EBF is not a ratio of marginal likelihoods under two fixed hypotheses. In an empirical-Bayes construction with a normal random-effects prior, the prior density at zero is p(0 | sigma_hat^2) = (2 pi sigma_hat^2)^(-d/2). If the estimated variance is positive but small, this density is large; if the estimated variance is large, this density is small, and the ratio can be arbitrarily inflated in either direction depending on the estimate. When the estimated variance is zero, the prior density is degenerate and the ratio is undefined. The finite-sample null distribution of a data-dependent ratio is not the distribution of a calibrated Bayes factor, so hypothesis tests based on the EBF can be anti-conservative. The manuscript must supply either an analytical demonstration that this data-dependent ratio is calibrated or a simulation study that shows the null distribution behaves correctly; otherwise the strongest claim that the EBF 'tests' random effects without prior specification is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an empirical Bayes factor (EBF) for testing the presence of random effects, reformulated as a test of whether all random effects are zero. The method uses a Savage-Dickey density ratio that, according to the abstract, can be computed from a single full-model fit, avoiding manual prior specification because the random-effects distribution is estimated from the data. The paper claims broad applicability across generalized linear crossed mixed models, spatial random effects models, dynamic structural equation models, random intercept cross-lagged panel models, and nonlinear mixed effects models, and reports that simulations on synthetic data evaluate the method's general behavior. This report is based solely on the abstract, as the full text was not available; consequently, all assessment here refers to the abstract's claims and the evidence provided therein.","tokens_in":952,"tokens_out":3995,"duration_ms":50564,"significance":"If the central claim holds, the EBF would be a practically valuable tool: it promises to test variance components in many mixed-model families using only the full model fit, with no subjective prior specification. That would eliminate the need to fit multiple models and would side-step the boundary problem in variance-component testing. However, the significance hinges on whether the EBF is a calibrated Bayesian procedure. The abstract explicitly states that the random-effects distribution is estimated from the data and used in the Savage-Dickey ratio, which raises the concern of double use of data. Without a proof of calibration or null-distribution simulations, the method's usefulness as a hypothesis test remains unestablished. The paper does not seem to provide machine-checked proofs or reproducible code in the available text, and the abstract gives no error quantification, so the strength of the empirical claims cannot be assessed.","major_comments":[{"comment":"The abstract states that the random-effects distribution is 'estimated from the data' and used in the Savage-Dickey density ratio. Because the same data are then used both to form the prior and to evaluate the likelihood, the resulting ratio is a data-dependent statistic, not a ratio of marginal likelihoods under two fixed hypotheses. For a normal random-effects prior with estimated variance σ̂², the prior density at zero is p(0 | σ̂²) = (2πσ̂²)^(-d/2), which is a function of the data. The paper provides no analytical argument or simulation evidence that this statistic is calibrated under the null hypothesis (e.g., uniform p-values or correct Type I error rates). This is the load-bearing issue for the paper's central claim that the EBF 'tests' random effects, and it must be resolved.","section":"Abstract"},{"comment":"The abstract does not address the boundary case where the estimated variance component is zero. In that case the prior density in the Savage-Dickey ratio is degenerate and the ratio is undefined. Since the paper motivates itself by the non-negativity-constrained variance component boundary problem, the behavior of the EBF at and near this boundary is central. The abstract supplies no statement about how the method handles this case or how the constrained parameter space affects the null distribution of the test statistic. This missing point is directly relevant to the method's validity in the very setting it aims to address.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'equivalent hypothesis that all random effects are zero' should be made formally precise, especially for nonlinear mixed-effects models, where random effects can enter nonlinearly and the definition of 'all random effects zero' may not be equivalent to a variance component being zero.","section":"Abstract"},{"comment":"The abstract does not cite prior work on empirical Bayes model comparison or on the Savage-Dickey density ratio; situating the proposal in that literature would clarify the intended contribution.","section":"Abstract"},{"comment":"The list of application families is long, but the abstract does not explain how the EBF is computed in each; for reproducibility, the manuscript should indicate whether a shared computational procedure is used or whether family-specific adjustments are required.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The review was conducted on the abstract alone because the full text was not present in the submission package. The referee could not verify whether the full text already contains the calibration proofs or simulations that would address the double-use-of-data concern. I recommend sending the manuscript back to the authors with the request to provide either an analytical demonstration of EBF calibration or a thorough simulation study of its null distribution, and to clarify the boundary behavior. If the full text already covers these, the editor may disregard this remark and consider the manuscript on its merits."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"One thing to know: this is an abstract-only read, so my verdict is provisional. The authors propose an empirical Bayes factor for testing variance components, using a Savage-Dickey density ratio computed from the full model fit. That is genuinely attractive: applied users would avoid fitting multiple reduced models and avoid subjective prior choice. The applications list is broad (crossed mixed models, spatial, DSEM, RI-CLPM, NLMEM) and the simulation study is at least in the abstract.\n\nWhat the abstract does not do is address the obvious double-use problem. The random-effects distribution is said to be 'part of the model's lower level and estimated from the data,' then that same estimated distribution is used as the prior density in the Savage-Dickey ratio. That means the denominator of the ratio is data-dependent. The value p(0 | \\hat\\sigma^2) can be arbitrarily large when the estimated variance is small, so the Bayes factor can be inflated in either direction. When \\hat\\sigma^2=0 the ratio is degenerate. The finite-sample null distribution of such a data-dependent ratio is not that of a calibrated Bayes factor; hypothesis tests based on it may be anti-conservative. The stress-test note makes this point precisely.\n\nThis is not a nitpick; it is load-bearing. If the authors can prove calibration analytically or show in simulations that the null distribution matches nominal error rates, the method is valuable. If not, the 'EBF' is a heuristic model-selection criterion, not a Bayes factor. The abstract gives no hint of such a defense.\n\nI want to be fair: I cannot see the full text. The derivation may already address the issue, perhaps via a proper prior on the variance component or a correction term. But from the abstract alone, the central claim is unsupported. The novelty and scope are real; the soundness is unverified.\n\nFor a colleague: I would bring this to reading group only after a full text is available. I would not cite it yet. I would, however, send it to peer review: the topic is important, the method is novel, and the potential flaw is testable. A good referee can check the calibration argument, ask for null-distribution simulations, and request code/data. If the authors can resolve the double-use issue, this could be a useful contribution to the mixed-models literature. If they cannot, the paper's main selling point collapses.","headline":"Useful single-fit empirical Bayes factor for variance components, but the abstract gives no evidence that the data-dependent Savage-Dickey ratio is calibrated.","tokens_in":1422,"tokens_out":2316,"would_cite":false,"duration_ms":29784,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes an empirical Bayes factor that tests for random effects using only the full mixed model fit, replacing the boundary problem of testing a variance component at zero with a Savage-Dickey density ratio for all random…","keywords":["empirical Bayes factor","random effects","variance components","Savage-Dickey density ratio","boundary problem","mixed models","hypothesis testing"],"falsifier":"Generate many datasets under the null model where all random effects are zero, fit only the full model, and compare the distribution of the EBF with a benchmark Bayes factor obtained from proper priors; if the EBF is not calibrated under the null or systematically favors the full model, the claim that it is a valid Bayes factor fails.","tokens_in":508,"feed_emoji":"⚖️","tokens_out":3129,"duration_ms":39565,"temperature":0.7,"pith_summary":"Random effects are standard tools for capturing heterogeneity, but deciding whether they are present is awkward because the variance component sits at a boundary under the null. This paper proposes an empirical Bayes factor (EBF) that instead tests whether all random effects are zero, using only the full model fit and avoiding the need to set priors by hand. The random effects distribution is estimated from the data, which is what makes the procedure empirical. If it works, researchers can test random-effect terms across many mixed-model families from a single fit, with no separate null-model fits and no subjective prior elicitation. Simulations and applications to several model classes illustrate the claimed flexibility.","feed_headline":"One mixed-model fit tests all random effects, no priors needed","feed_subtitle":"An empirical Bayes factor via Savage-Dickey density ratio skips separate null models and subjective priors.","key_machinery":"The engine is the Savage-Dickey density ratio: for a point-null hypothesis, a Bayes factor can be written as the ratio of the posterior density to the prior density of the parameter evaluated at the null value. The EBF uses this identity with all random effects set to zero, so the full model fit supplies everything needed; no reduced model needs to be fitted. Because the random effects distribution is estimated from the data, the resulting quantity is an empirical Bayes factor.","core_discovery":"The paper's central claim is that testing a variance component for absence can be recast as testing that all random effects are zero, and that the resulting Bayes factor can be computed from the full model alone via a Savage-Dickey density ratio. The random effects distribution is part of the lower level of the model and is estimated from the data, so the method is empirical and requires no external prior knowledge. The EBF therefore sidesteps the non-negativity boundary problem and, in principle, applies uniformly to linear and generalized linear crossed mixed models, spatial random effects models, dynamic structural equation models, random intercept cross-lagged panel models, and nonlinear mixed effects models.","pith_inferences":["A natural extension, not pursued here, would be ranking or selecting among several random-effect structures using one full-model EBF, since the full fit already contains all candidate random effects.","The empirical construction raises a calibration question for future work: if estimating the prior from the same data biases the density ratio under the null, a correction or a calibration check would be needed before the EBF is used as a decisive Bayes factor.","One could test the method's frequentist properties under the null, for instance whether repeated EBF values behave like a valid Bayes factor, which would connect it to the broader empirical Bayes model selection literature."],"forward_implications":["Only the full model needs to be fitted to test any random-effect term; null and reduced models become unnecessary.","The boundary problem of testing $\\sigma^2=0$ is bypassed by testing the equivalent statement that the entire random-effects vector is zero.","The procedure applies across mixed-model families, including crossed, spatial, dynamic, panel, and nonlinear specifications, from one estimation framework.","No subjective prior for the variance component is required, eliminating a common barrier to routine Bayes factor testing."],"supporting_citations":[],"fun_headline_variants":["Test all random effects with one model fit, no priors","Empirical Bayes factor: one fit, no priors, all random effects","One fit, no priors: empirical Bayes factor tests all random effects","Empirical Bayes factor skips priors and separate null fits","Empirical Bayes factor tests random effects across many model families"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that using the same data to estimate the random-effects distribution and then to evaluate the likelihood still yields a valid Bayes factor; this empirical shortcut is assumed not to amount to double use of the data.","fun_headline_variants_meta":{"raw":{"variants":["Test all random effects with one model fit, no priors","Empirical Bayes factor: one fit, no priors, all random effects","One fit, no priors: empirical Bayes factor tests all random effects","Empirical Bayes factor skips priors and separate null fits","Empirical Bayes factor tests random effects across many model families"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00133,"raw_usage":{"total_tokens":5377,"prompt_tokens":876,"completion_tokens":4501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":4411}},"tokens_in":492,"tokens_out":4501,"duration_ms":32555,"temperature":1.0,"reasoning_tokens":4411,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:36:18.331925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate many datasets under the null model where all random effects are zero, fit only the full model, and compare the distribution of the EBF with a benchmark Bayes factor obtained from proper priors; if the EBF is not calibrated under the null or systematically favors the full model, the claim that it is a valid Bayes factor fails.","supporting_citations":[],"review_version":1}