{"id":"019297d7-2542-4892-8933-2ca61374053d","arxiv_id":"2509.05832","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new BRAIDS utility interpolates between risk-seeking and risk-averse subgroup detection, and regularized Bayesian models can keep nominal coverage for subgroup effects without sample splitting.","lead":"This paper presents a Bayesian rule for picking patient subgroups in clinical trials that can be tuned from aggressive to cautious. It argues that with strongly regularized statistical models, researchers can find subgroups and measure their effects using the same data without losing statistical validity.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage of adaptive Bayesian subgroup intervals is demonstrated only for DGPs that are draws from the same model class; misspecified-prior simulations are needed to support the general claim.","rationale":"The reader's weakest assumption identifies exactly the same issue: Section 4.2 makes the key premise true by construction, and Theorem 3 uses a modified prior that does not cover the priors actually used. I agree and would locate the decisive gap in the absence of any misspecified-DGP coverage check. The paper's own limitation paragraph in Section 5 concedes that Bayesian post-selection validity depends on well-calibrated priors and that diffuse priors degrade performance, which raises the question of whether the empirical coverage results are an artifact of the simulation design rather than a robust property of regularization. The proposed test would directly distinguish these possibilities: if coverage collapses under a step-function DGP with similar heterogeneity, then the central claim should be qualified as holding only when the prior family is approximately correct; if coverage survives, the claim is substantially strengthened. The BRAIDS utility contribution and the risk-seeking/risk-averse analysis are not affected by this concern, and the paper is honest about its limitations. No code or data are provided, which compounds the difficulty of assessing the simulations, but the prior-misspecification gap is the more load-bearing scientific issue. Since the reader's conditional verdict already reflects this uncertainty, no verdict change is needed.","tokens_in":20451,"tokens_out":4953,"duration_ms":60870,"concrete_test":"Re-run the Section 4.2 design (N=1000, σ=0.1, 200 replications) with a deliberately misspecified DGP: e.g., τ(x)=c·1{x1>median(x1), x2>median(x2)}+c·1{x3<q25(x3)} with c chosen so the empirical heterogeneity H is near 0.3 (matching the MEPS-based prior scale), and main effects/prognostic function as in the ridge DGP. Apply the identical BCF and ridge BRAIDS procedures (same priors, same subgroup search, same nominal 95% intervals) and compute empirical coverage of the data-adaptive subgroup ATE. If coverage drops below ~90%, the central claim fails to generalize beyond the well-specified setting; if coverage stays near 95%, the regularization story is substantially more robust. A useful second arm: set the true τ(x) to the MEPS-fitted BCF function but deliberately misspecify the prior scale (sτ=0.05 and sτ=10) to test sensitivity to the 'appropriate prior' choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that regularization priors allow fully Bayesian subgroup inference to keep nominal frequentist coverage despite data-adaptive subgroup selection (Section 5, Figure 4). The argument in Section 3 is an identity averaged over the prior: for any procedure with posterior probability 1−α, Pr(E)=1−α when θ0∼π(θ). This gives no uniform or conditional guarantee for a fixed θ0; the subsequent assertion that real θ0s are 'typical draws' is informal. The Section 4.2 simulations instantiate 'typical draws' by construction: DGPs are generated from BCF and ridge fits to MEPS, i.e., from the same regularized model families whose priors are then evaluated. Thus the empirical coverage numbers estimate well-specified prior-averaged coverage, not coverage under the kind of model misspecification that motivates regularization. The flat-prior case in Figure 4 (coverage as low as 0.59) shows that a fixed θ0 can be very atypical under an unregularized prior, and Section 5 concedes that performance deteriorates when priors are diffuse. Since 'appropriate prior' is only defined ex post by matching the true heterogeneity, the claim that regularization safeguards double dipping is not yet established outside a narrow, prior-generated simulation regime. Theorem 3 concerns E(H^2) under a modified BART prior (Poisson depth, continuous covariates, uniform random cutpoints) and does not directly support frequentist coverage for the BCF priors actually used. No theorem or robustness study addresses what happens when the true τ(x) has interactions, sparsity, or effect sizes outside the Exp(1) prior's support. This is the load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Bayesian decision-theoretic framework, BRAIDS, for detecting subgroups with heterogeneous treatment effects. It shows that a naive utility (maximizing between-group CATE variance) is risk-seeking because it rewards posterior variance in subgroup effects, and it proposes a family of utilities indexed by a risk parameter λ that interpolates between risk-seeking, risk-neutral (virtual-twins-like), and risk-averse behavior. The authors then argue that, to avoid winner's-curse bias when the same data are used for subgroup discovery and post-selection inference, the treatment-effect heterogeneity must be strongly regularized. They provide empirical evidence that, with appropriately chosen shrinkage priors, Bayesian credible intervals for adaptive subgroup effects can achieve near-nominal frequentist coverage without data splitting. The methods are illustrated on a canagliflozin type-2 diabetes trial, and the supplementary material proves Theorems 1–3, including a result on the prior mean of treatment-effect heterogeneity under a modified BART prior.","tokens_in":20854,"tokens_out":6248,"duration_ms":75235,"significance":"If the central claim held as stated, the paper would be an important contribution: it would give practitioners a principled, efficient alternative to sample splitting for subgroup discovery and inference, backed by a decision-theoretic justification and a flexible class of utilities. The posterior-expectation decompositions in Theorems 1 and 2 are clean and the risk-seeking/risk-averse interpolation is conceptually useful. The simulation design is thoughtful in using fitted models on real MEPS data to generate plausible DGPs, and the application to a real clinical trial is valuable. However, the load-bearing claim—that regularized Bayesian inference maintains nominal frequentist coverage after adaptive subgroup selection—is established only in a narrow, prior-aligned simulation regime. The formal argument in Section 3 is a prior-averaged identity, not a coverage theorem for fixed θ0, and the empirical design in Section 4.2 constructs the true τ(x) from the same model families whose priors are subsequently evaluated. The paper's own Section 5 acknowledges that coverage deteriorates under diffuse priors. Thus the significance of the central finding is real but considerably more condi","major_comments":[{"comment":"The central claim—that regularization priors safeguard fully Bayesian subgroup inference from the winner's curse—is supported only by a prior-averaged identity, not by a coverage theorem for fixed θ0. The argument in Section 3 correctly shows that Pr(E) = 1−α when θ0 is drawn from the prior, but this does not provide a uniform or conditional guarantee for a fixed, data-generating θ0; the text's move from 'there exist θ0 with coverage ≥ 1−α' to 'typical draws' is informal. Section 4.2 makes the typical-draw condition true by construction: the DGPs are generated by fitting BCF and ridge regression to MEPS data, i.e., from the same regularized model families whose posterior intervals are then evaluated. The flat-prior case in Figure 4 (coverage as low as 0.59) shows that a fixed θ0 can be very atypical under an unregularized prior, and Section 5 concedes that 'performance deteriorates when","section":"Section 3 and Section 4.2"},{"comment":"Theorem 3 is proved for a modified BART prior—Poisson leaf depth, continuous covariates, and cutpoints sampled from conditional distributions—and the text explicitly acknowledges: 'Strictly speaking, Theorem 3 does not cover the BART priors used in practice' (Section 3.2). The paper nevertheless lists this as contribution 4 and uses it to argue that the prior on heterogeneity is invariant to covariate distribution. As stated, the theorem applies to a nearby prior, not to the implemented BCF priors; the transfer is an approximation. This is not a fatal flaw, but the claim should be reframed as an approximation or a heuristic property, or a version of Theorem 3 should be proved for the actual default BART prior. This is load-bearing for the paper's theoretical contribution, even if it does not directly drive the coverage simulations.","section":"Section 3.2 and Supplement S.1"},{"comment":"The empirical coverage rates in Figure 4 are reported as point estimates from 200 replications. For nominal 95% coverage, the Monte Carlo standard error is approximately sqrt(0.95×0.05/200) ≈ 1.5%, so reported coverages of 0.93–0.97 are not statistically distinguished from each other or from 0.95 with high confidence. The claim that Bayesian methods 'generally work well' and the comparison with honest methods would be strengthened by adding standard errors or confidence intervals for the coverage rates, or by increasing the number of replications.","section":"Section 4.2, Figure 4"},{"comment":"The lasso double-dipping variant is reported to 'perform well even when double dipping, producing relatively narrow intervals and conservative inferences in all settings we examined.' This is an interesting but unexplained finding that partially undercuts the narrative that Bayesian regularization is special or that 'Bayesian logic does not naturally lead to any direct correction for post-selection inference.' The paper should discuss why the lasso, a frequentist regularized estimator, also appears to mitigate the winner's curse in these experiments, and clarify whether the phenomenon is about shrinkage generally rather than Bayesian inference specifically.","section":"Section 4.2"}],"minor_comments":[{"comment":"The BRAIDS utility is illustrated with λ ∈ {0,1,2} in Table 1, but Figure 5 and the subsequent subgroup analysis are only for the risk-neutral λ = 1 setting. The paper claims to cover risk-seeking and risk-averse behavior in the application, but adaptive subgroup discovery is demonstrated only in the risk-neutral case. This should be stated explicitly.","section":"Section 4.3"},{"comment":"The figure caption and text layout make it hard to map the empirical coverage labels to the corresponding boxplots, especially with the two noise levels σ ∈ {1/10, 1/3}. A cleaner label layout or a table of coverages with standard errors would improve readability.","section":"Figure 4"},{"comment":"The notation τ(X) is used both for the population-level average treatment effect and later for the sample average; Eq. (2) and the text around it could benefit from a notational distinction (e.g., τ̄ vs τ(X)) to avoid ambiguity.","section":"Section 1.1"},{"comment":"There is no code or data availability statement. Since the simulations use MEPS data and a YODA-accessed clinical trial, the reproducibility of the empirical results would be greatly enhanced by releasing code and, where possible, the exact data-processing steps.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a solid methods paper with clean decision-theoretic results and a useful empirical comparison, but the central frequentist-coverage claim is currently supported only under a prior-generated simulation design. The authors' own limitation statement in Section 5 acknowledges the dependence on well-calibrated priors, yet the abstract and discussion state the result more broadly. I would recommend major revision: the paper needs either a misspecification-robustness study or a carefully narrowed claim, plus a clearer separation between theorems that hold exactly and those that hold only approximately. The paper fits the journal's scope well; the issues are about the strength of the evidence, not the novelty of the framework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the BRAIDS utility and the risk-seeking/risk-averse interpolation. Theorem 1 cleanly shows why the naive posterior-expected utility rewards uncertainty in subgroup effects, and Theorem 2 gives the exact trade-off. Recovering virtual twins at lambda=1 as the risk-neutral midpoint is a nice unification, not just a restatement. Theorem 3, despite being proved for a modified BART prior, does give real intuition about why tree priors stabilize prior heterogeneity across covariate distributions. The clinical application is careful, and the authors openly admit the central limitation: post-selection coverage depends on choosing a well-calibrated prior.\n\nThe soft spots are real but narrower than the stress-test makes them sound. The Section 3 identity is prior-averaged—true on average over theta ~ pi, not a fixed-theta guarantee. The authors know this and say the real question is whether true theta_0 is a typical draw. The problem is that Section 4.2 makes that true by construction: the DGPs are generated by fitting the very same Bayesian models being evaluated, so the simulations are checking well-specified prior-averaged coverage, not misspecification. The flat-prior collapse to 0.59 coverage shows what happens when theta_0 is atypical, which does undermine the general claim that regularization safeguards against double dipping. The lack of code and data also prevents independent checking of the BCF approximation. These are proportionate concerns: the abstract slightly oversells, but the Discussion explicitly flags the prior-dependence. This is not a load-bearing hidden flaw; it is a stated limitation that needs stronger evidence.\n\nWho gets value: statisticians working on subgroup detection, Bayesian causal inference, and post-selection inference. The framework is a useful conceptual contribution even if the coverage guarantee remains conditional on prior adequacy. I would bring it to a reading group and would cite the BRAIDS utility if I worked in this area. It deserves a serious referee: the methods are new, the math in Theorems 1 and 2 is clean, and the empirical claim is interesting but needs a robustness section with misspecified priors and non-model-generated DGPs, plus release of code and data. I would recommend engaging, with revision.","headline":"Genuine new utility formulation and a useful warning about risk-seeking Bayesian subgroup selection, but the headline coverage claim is supported only by prior-generated simulations—worth a serious referee with a request for robustness checks and code.","tokens_in":21304,"tokens_out":1195,"would_cite":true,"duration_ms":17714,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62C10","62G05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that adaptive subgroup discovery and estimation can share one dataset when the prior is regularized: credible intervals keep nominal Frequentist coverage, and the new BRAIDS utility exposes the hidden risk-seeking bias in n","keywords":["heterogeneous treatment effects","Bayesian decision theory","subgroup identification","Bayesian additive regression trees","post-selection inference","regularization priors","winner's curse","clinical trials"],"falsifier":"Run 500 simulations with a true treatment-effect function whose heterogeneity is large relative to the prior scale (say ten times the prior's typical H), apply BRAIDS subgroup selection with the linear ridge prior, and count the coverage of the nominal 95% credible intervals for the selected subgroup effects: the paper's own account predicts coverage will fall well below 95% in this misspecified-prior regime, while the honest data-splitting estimator keeps nominal coverage.","tokens_in":20374,"feed_emoji":"🎯","tokens_out":17353,"duration_ms":168423,"temperature":0.7,"pith_summary":"This paper takes on a routine but fraught problem in clinical trials: after the trial is over, which subgroups of patients responded differently to the treatment, and how much did the treatment help them? The paper first shows that the most natural Bayesian decision rule for this task is self-defeating: the utility that rewards heterogeneous subgroups also rewards subgroups whose treatment effects are least precisely estimated (a \"risk-seeking\" preference), which lowers the chance that the finding will replicate. To fix this, it introduces the BRAIDS utility, whose tuning parameter lets the analyst slide from risk-seeking through risk-neutral to risk-averse subgroup selection; the risk-neutral setting turns out to be a variant of the well-known virtual-twins algorithm, giving that heuristic a decision-theoretic foundation. The paper's central claim is that when treatment-effect heterogeneity is modeled with a regularization prior that makes the true effect function a typical draw from the prior, credible intervals for the effects of data-identified subgroups keep nominal Frequentist coverage — so analysts can use the full dataset for both finding subgroups and estimating their effects, avoiding the efficiency loss of data splitting. This is demonstrated in simulations built from real survey data and illustrated on the canagliflozin diabetes trial.","feed_headline":"95% coverage survives subgroup selection without data splitting","feed_subtitle":"Regularized priors let analysts find and estimate subgroups from one dataset: valid intervals, no efficiency loss.","key_machinery":"The central object is the BRAIDS utility, a multi-stage decision rule that scores a subgroup partition G and reported subgroup means t by how far the subgroup effects deviate from the overall effect, minus a penalty λ for how far the reported means are from the truth. The tuning parameter λ interpolates between risk-seeking (λ<1), risk-neutral (λ=1), and risk-averse (λ>1) selection: Theorem 2 reduces the posterior expected utility to a weighted combination of the posterior variance of each subgroup effect and the within-subgroup scatter of estimated individual effects. The supporting machinery is the regularization prior: Bayesian ridge with βτ ~ Normal(0, σ_τ²I) and σ_τ ~ Exp(1), and BCF wh","core_discovery":"Three results carry the paper. (1) The natural heterogeneity utility's posterior expectation (Theorem 1) contains the variance Var{τ(Gk)|D} of each subgroup effect, so maximizing it prefers subgroups whose effects are hardest to estimate — risk-seeking behavior. (2) The BRAIDS utility weights that variance by (1−λ); at λ=1 it vanishes, yielding a variant of virtual twins (Theorem 2). (3) Bayesian credible sets are calibrated for prior draws, so full-data subgroup discovery keeps nominal Frequentist coverage when the true τ(x) is a typical prior draw. Hierarchical shrinkage priors (Bayesian ridge, BCF with tuned scales) achieve this in simulation; flat priors undercover — the winner's curse.","pith_inferences":["Beyond the paper: the coverage rationale implies a pre-registration recipe — fix the prior's heterogeneity scale before seeing the data (e.g., by drawing the prior of H and M as the paper suggests) so that validity rests on a public commitment rather than on tuning after subgroup discovery.","Beyond the paper: the observation that the lasso also survived double-dipping suggests explicit shrinkage, not Bayesian updating itself, is the active ingredient; a frequentist version of BRAIDS with an ℓ1 or ℓ2 penalty on subgroup selection might achieve comparable coverage without posterior computation.","Beyond the paper: the \"typical draw of the prior\" principle transfers to other adaptive settings, such as high-dimensional variable selection followed by effect estimation, and could be tested by checking interval coverage from sparsity priors calibrated to the true signal strength."],"forward_implications":["Trial analysts can report subgroup effects with nominal-coverage intervals without holding out data for subgroup discovery, so power is no longer lost to sample splitting.","The λ=1 BRAIDS choice gives the virtual-twins heuristic a strict Bayesian decision-theoretic justification, anchoring a widely used clinical-trial workflow in expected-utility theory.","Risk-averse settings (λ>1) deliberately pool dissimilar patients to stabilize subgroup estimates — a defensible choice when power is the binding constraint — while risk-seeking settings (λ<1) prioritize covariate-homogeneous subgroups for descriptive purposes.","Theorem 3's invariance means Bayesian causal forest priors can be specified without reference to the covariate distribution, so adding or transforming covariates does not inflate the prior's expected heterogeneity; this is not true of Bayesian linear models.","The regularize-and-infer recipe is portable: the same logic suggests that any post-selection Bayesian analysis — variable selection, policy learning, mediation — keeps nominal coverage when its prior makes the truth a typical draw."],"supporting_citations":[{"why":"Introduces the virtual-twins algorithm that the risk-neutral BRAIDS utility (λ=1) exactly recovers, giving the heuristic a Bayesian foundation.","marker":"Foster et al. (2011)"},{"why":"Supplies the Bayesian causal forest model and default prior that the paper extends with a separable heterogeneity function and a tuned στ scale.","marker":"Hahn et al. (2020)"},{"why":"Defines the BART tree prior that the paper's nonparametric models use and that Theorem 3 modifies to compute the prior heterogeneity E(H²).","marker":"Chipman et al. (2010)"},{"why":"Formalizes inference on winners, the Frequentist post-selection problem that the paper argues regularization priors avoid.","marker":"Andrews et al. (2024)"},{"why":"Provides the data-splitting algorithm for valid inference on adaptively selected subgroups that the paper's full-data Bayesian approach is meant to outperform.","marker":"Chernozhukov et al. (2018)"},{"why":"Earlier Bayesian decision-theoretic utility for subgroup finding that the paper extends and generalizes with the risk-aware family.","marker":"Morita and Müller (2017)"},{"why":"States the position that Bayesian selection need not be corrected for double use of data, which the paper tests empirically.","marker":"Woody et al. (2021)"},{"why":"Review of post-selection inference strategies used to frame data splitting as the standard but inefficient alternative.","marker":"Kuchibhotla et al. (2022)"}],"fun_headline_variants":["Bayesian subgroup selection: valid intervals without data splitting","Regularized priors tame winner's curse in subgroup discovery","New Bayesian utility balances risk in subgroup detection","Full-data subgroup inference keeps 95% coverage with right priors","BRAIDS: risk-aware subgroup detection with valid uncertainty"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The coverage guarantee holds only when the prior is essentially the truth about how much treatment effects vary: if the real effect function is not a typical draw from the prior — the paper's own flat-prior experiments show the failure mode — the credible intervals for selected subgroups lose their nominal coverage.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian subgroup selection: valid intervals without data splitting","Regularized priors tame winner's curse in subgroup discovery","New Bayesian utility balances risk in subgroup detection","Full-data subgroup inference keeps 95% coverage with right priors","BRAIDS: risk-aware subgroup detection with valid uncertainty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000148,"raw_usage":{"total_tokens":1012,"prompt_tokens":713,"completion_tokens":299,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":221}},"tokens_in":457,"tokens_out":299,"duration_ms":3636,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:57:36.372971+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run 500 simulations with a true treatment-effect function whose heterogeneity is large relative to the prior scale (say ten times the prior's typical H), apply BRAIDS subgroup selection with the linear ridge prior, and count the coverage of the nominal 95% credible intervals for the selected subgroup effects: the paper's own account predicts coverage will fall well below 95% in this misspecified-prior regime, while the honest data-splitting estimator keeps nominal coverage.","supporting_citations":[{"cited_title":"C., Taylor, J","cited_arxiv_id":null,"evidence_quote":"Introduces the virtual-twins algorithm that the risk-neutral BRAIDS utility (λ=1) exactly recovers, giving the heuristic a Bayesian foundation."},{"cited_title":"R., Murray, J","cited_arxiv_id":null,"evidence_quote":"Supplies the Bayesian causal forest model and default prior that the paper extends with a separable heterogeneity function and a tuned στ scale."},{"cited_title":"A., George, E","cited_arxiv_id":null,"evidence_quote":"Defines the BART tree prior that the paper's nonparametric models use and that Theorem 3 modifies to compute the prior heterogeneity E(H²)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes inference on winners, the Frequentist post-selection problem that the paper argues regularization priors avoid."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the data-splitting algorithm for valid inference on adaptively selected subgroups that the paper's full-data Bayesian approach is meant to outperform."},{"cited_title":"and M\\\" u ller, P","cited_arxiv_id":null,"evidence_quote":"Earlier Bayesian decision-theoretic utility for subgroup finding that the paper extends and generalizes with the risk-aware family."},{"cited_title":"M., and Murray, J","cited_arxiv_id":null,"evidence_quote":"States the position that Bayesian selection need not be corrected for double use of data, which the paper tests empirically."}],"review_version":1}