Pith. sign in

REVIEW 4 major objections 4 minor 49 references

Decision Theoretic Subgroup Detection With Bayesian Machine Learning

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper shows that adaptive subgroup discovery and estimation can share one dataset when the prior is regularized: credible intervals keep nominal Frequentist coverage, and the new BRAIDS utility exposes the hidden risk-seeking bias in n

desk verdict Genuine new utility formulation and a useful warning about risk-seeking Bayesian subgroup selection, but the headline coverage claim is supported only by prior-generated simulations—worth a serious referee with a request for robustness checks and code. read the letter →

arxiv 2509.05832 v1 pith:FJV4KPM7 submitted 2025-09-06 stat.ME

classification stat.ME MSC 62F1562C1062G05
keywords heterogeneoustreatmenteffectsBayesiandecisiontheorysubgroupidentificationadditiveregressiontreespost-selectioninferenceregularizationpriorswinner'scurseclinicaltrials
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper takes on a routine but fraught problem in clinical trials: after the trial is over, which subgroups of patients responded differently to the treatment, and how much did the treatment help them? The paper first shows that the most natural Bayesian decision rule for this task is self-defeating: the utility that rewards heterogeneous subgroups also rewards subgroups whose treatment effects are least precisely estimated (a "risk-seeking" preference), which lowers the chance that the finding will replicate. To fix this, it introduces the BRAIDS utility, whose tuning parameter lets the analyst slide from risk-seeking through risk-neutral to risk-averse subgroup selection; the risk-neutral setting turns out to be a variant of the well-known virtual-twins algorithm, giving that heuristic a decision-theoretic foundation. The paper's central claim is that when treatment-effect heterogeneity is modeled with a regularization prior that makes the true effect function a typical draw from the prior, credible intervals for the effects of data-identified subgroups keep nominal Frequentist coverage — so analysts can use the full dataset for both finding subgroups and estimating their effects, avoiding the efficiency loss of data splitting. This is demonstrated in simulations built from real survey data and illustrated on the canagliflozin diabetes trial.

What carries the argument

The central object is the BRAIDS utility, a multi-stage decision rule that scores a subgroup partition G and reported subgroup means t by how far the subgroup effects deviate from the overall effect, minus a penalty λ for how far the reported means are from the truth. The tuning parameter λ interpolates between risk-seeking (λ<1), risk-neutral (λ=1), and risk-averse (λ>1) selection: Theorem 2 reduces the posterior expected utility to a weighted combination of the posterior variance of each subgroup effect and the within-subgroup scatter of estimated individual effects. The supporting machinery is the regularization prior: Bayesian ridge with βτ ~ Normal(0, σ_τ²I) and σ_τ ~ Exp(1), and BCF wh

What would settle it

Run 500 simulations with a true treatment-effect function whose heterogeneity is large relative to the prior scale (say ten times the prior's typical H), apply BRAIDS subgroup selection with the linear ridge prior, and count the coverage of the nominal 95% credible intervals for the selected subgroup effects: the paper's own account predicts coverage will fall well below 95% in this misspecified-prior regime, while the honest data-splitting estimator keeps nominal coverage.

Watch

Extended reading notes

Core claim

Three results carry the paper. (1) The natural heterogeneity utility's posterior expectation (Theorem 1) contains the variance Var{τ(Gk)|D} of each subgroup effect, so maximizing it prefers subgroups whose effects are hardest to estimate — risk-seeking behavior. (2) The BRAIDS utility weights that variance by (1−λ); at λ=1 it vanishes, yielding a variant of virtual twins (Theorem 2). (3) Bayesian credible sets are calibrated for prior draws, so full-data subgroup discovery keeps nominal Frequentist coverage when the true τ(x) is a typical prior draw. Hierarchical shrinkage priors (Bayesian ridge, BCF with tuned scales) achieve this in simulation; flat priors undercover — the winner's curse.

Load-bearing premise

The coverage guarantee holds only when the prior is essentially the truth about how much treatment effects vary: if the real effect function is not a typical draw from the prior — the paper's own flat-prior experiments show the failure mode — the credible intervals for selected subgroups lose their nominal coverage.

Editorial extensions

If this is right

  • Trial analysts can report subgroup effects with nominal-coverage intervals without holding out data for subgroup discovery, so power is no longer lost to sample splitting.
  • The λ=1 BRAIDS choice gives the virtual-twins heuristic a strict Bayesian decision-theoretic justification, anchoring a widely used clinical-trial workflow in expected-utility theory.
  • Risk-averse settings (λ>1) deliberately pool dissimilar patients to stabilize subgroup estimates — a defensible choice when power is the binding constraint — while risk-seeking settings (λ<1) prioritize covariate-homogeneous subgroups for descriptive purposes.
  • Theorem 3's invariance means Bayesian causal forest priors can be specified without reference to the covariate distribution, so adding or transforming covariates does not inflate the prior's expected heterogeneity; this is not true of Bayesian linear models.
  • The regularize-and-infer recipe is portable: the same logic suggests that any post-selection Bayesian analysis — variable selection, policy learning, mediation — keeps nominal coverage when its prior makes the truth a typical draw.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the coverage rationale implies a pre-registration recipe — fix the prior's heterogeneity scale before seeing the data (e.g., by drawing the prior of H and M as the paper suggests) so that validity rests on a public commitment rather than on tuning after subgroup discovery.
  • Beyond the paper: the observation that the lasso also survived double-dipping suggests explicit shrinkage, not Bayesian updating itself, is the active ingredient; a frequentist version of BRAIDS with an ℓ1 or ℓ2 penalty on subgroup selection might achieve comparable coverage without posterior computation.
  • Beyond the paper: the "typical draw of the prior" principle transfers to other adaptive settings, such as high-dimensional variable selection followed by effect estimation, and could be tested by checking interval coverage from sparsity priors calibrated to the true signal strength.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper develops a Bayesian decision-theoretic framework, BRAIDS, for detecting subgroups with heterogeneous treatment effects. It shows that a naive utility (maximizing between-group CATE variance) is risk-seeking because it rewards posterior variance in subgroup effects, and it proposes a family of utilities indexed by a risk parameter λ that interpolates between risk-seeking, risk-neutral (virtual-twins-like), and risk-averse behavior. The authors then argue that, to avoid winner's-curse bias when the same data are used for subgroup discovery and post-selection inference, the treatment-effect heterogeneity must be strongly regularized. They provide empirical evidence that, with appropriately chosen shrinkage priors, Bayesian credible intervals for adaptive subgroup effects can achieve near-nominal frequentist coverage without data splitting. The methods are illustrated on a canagliflozin type-2 diabetes trial, and the supplementary material proves Theorems 1–3, including a result on the prior mean of treatment-effect heterogeneity under a modified BART prior.

Significance. If the central claim held as stated, the paper would be an important contribution: it would give practitioners a principled, efficient alternative to sample splitting for subgroup discovery and inference, backed by a decision-theoretic justification and a flexible class of utilities. The posterior-expectation decompositions in Theorems 1 and 2 are clean and the risk-seeking/risk-averse interpolation is conceptually useful. The simulation design is thoughtful in using fitted models on real MEPS data to generate plausible DGPs, and the application to a real clinical trial is valuable. However, the load-bearing claim—that regularized Bayesian inference maintains nominal frequentist coverage after adaptive subgroup selection—is established only in a narrow, prior-aligned simulation regime. The formal argument in Section 3 is a prior-averaged identity, not a coverage theorem for fixed θ0, and the empirical design in Section 4.2 constructs the true τ(x) from the same model families whose priors are subsequently evaluated. The paper's own Section 5 acknowledges that coverage deteriorates under diffuse priors. Thus the significance of the central finding is real but considerably more condi

major comments (4)
  1. [Section 3 and Section 4.2] The central claim—that regularization priors safeguard fully Bayesian subgroup inference from the winner's curse—is supported only by a prior-averaged identity, not by a coverage theorem for fixed θ0. The argument in Section 3 correctly shows that Pr(E) = 1−α when θ0 is drawn from the prior, but this does not provide a uniform or conditional guarantee for a fixed, data-generating θ0; the text's move from 'there exist θ0 with coverage ≥ 1−α' to 'typical draws' is informal. Section 4.2 makes the typical-draw condition true by construction: the DGPs are generated by fitting BCF and ridge regression to MEPS data, i.e., from the same regularized model families whose posterior intervals are then evaluated. The flat-prior case in Figure 4 (coverage as low as 0.59) shows that a fixed θ0 can be very atypical under an unregularized prior, and Section 5 concedes that 'performance deteriorates when
  2. [Section 3.2 and Supplement S.1] Theorem 3 is proved for a modified BART prior—Poisson leaf depth, continuous covariates, and cutpoints sampled from conditional distributions—and the text explicitly acknowledges: 'Strictly speaking, Theorem 3 does not cover the BART priors used in practice' (Section 3.2). The paper nevertheless lists this as contribution 4 and uses it to argue that the prior on heterogeneity is invariant to covariate distribution. As stated, the theorem applies to a nearby prior, not to the implemented BCF priors; the transfer is an approximation. This is not a fatal flaw, but the claim should be reframed as an approximation or a heuristic property, or a version of Theorem 3 should be proved for the actual default BART prior. This is load-bearing for the paper's theoretical contribution, even if it does not directly drive the coverage simulations.
  3. [Section 4.2, Figure 4] The empirical coverage rates in Figure 4 are reported as point estimates from 200 replications. For nominal 95% coverage, the Monte Carlo standard error is approximately sqrt(0.95×0.05/200) ≈ 1.5%, so reported coverages of 0.93–0.97 are not statistically distinguished from each other or from 0.95 with high confidence. The claim that Bayesian methods 'generally work well' and the comparison with honest methods would be strengthened by adding standard errors or confidence intervals for the coverage rates, or by increasing the number of replications.
  4. [Section 4.2] The lasso double-dipping variant is reported to 'perform well even when double dipping, producing relatively narrow intervals and conservative inferences in all settings we examined.' This is an interesting but unexplained finding that partially undercuts the narrative that Bayesian regularization is special or that 'Bayesian logic does not naturally lead to any direct correction for post-selection inference.' The paper should discuss why the lasso, a frequentist regularized estimator, also appears to mitigate the winner's curse in these experiments, and clarify whether the phenomenon is about shrinkage generally rather than Bayesian inference specifically.
minor comments (4)
  1. [Section 4.3] The BRAIDS utility is illustrated with λ ∈ {0,1,2} in Table 1, but Figure 5 and the subsequent subgroup analysis are only for the risk-neutral λ = 1 setting. The paper claims to cover risk-seeking and risk-averse behavior in the application, but adaptive subgroup discovery is demonstrated only in the risk-neutral case. This should be stated explicitly.
  2. [Figure 4] The figure caption and text layout make it hard to map the empirical coverage labels to the corresponding boxplots, especially with the two noise levels σ ∈ {1/10, 1/3}. A cleaner label layout or a table of coverages with standard errors would improve readability.
  3. [Section 1.1] The notation τ(X) is used both for the population-level average treatment effect and later for the sample average; Eq. (2) and the text around it could benefit from a notational distinction (e.g., τ̄ vs τ(X)) to avoid ambiguity.
  4. [General] There is no code or data availability statement. Since the simulations use MEPS data and a YODA-accessed clinical trial, the reproducibility of the empirical results would be greatly enhanced by releasing code and, where possible, the exact data-processing steps.

Circularity Check

1 steps flagged · score 4.0 of 10

Coverage validation is partly self-referential: Section 4.2 DGPs are generated from the same BCF/ridge models being evaluated, so the 'typical draw' premise of the Section 3 coverage identity is satisfied by construction.

  1. self definitional [Section 3 (coverage identity) and Section 4.2 (Data generation)]
    "Bayesian intervals are guaranteed to attain nominal coverage levels marginally for θ’s sampled from the prior... This suggests that if the true θ0 looks like a “typical” draw of θ ∼ π(θ) then we should expect the Bayesian approach to have Frequentist coverage at or above the nominal level... Data generation We generate plausible data generating mechanisms in the same fashion as Section 4.1, fitting the BCF and ridge regression models to this data."

    The Section 3 identity Pr(E) = 1 − α holds by construction whenever θ0 is averaged over the prior. The Section 4.2 simulation generates the true τ(x) by fitting BCF and ridge regression to MEPS data—the same model classes whose priors are then used for inference. This makes the 'θ0 is a typical draw' condition true by construction. The near-nominal coverage reported in Figure 4 therefore estimates the prior-averaged guarantee, not coverage under the kind of prior misspecification or model uncertainty that motivates the paper's regularization recommendation. The flat-prior failure only illustrates that a fixed θ0 from a regularized fit is atypical under a diffuse prior, which is again the same identity in reverse.

full rationale

The BRAIDS utility derivation (Theorems 1 and 2) is self-contained and not circular: it computes posterior expected utilities from the stated utility functions. The policy-estimation discussion and the connection to virtual twins are also independent identifications, not circular renaming. No load-bearing self-citation chain appears; citations to the authors' prior work (e.g., Linero and Antonelli 2023; Hahn et al. 2020) provide context and models, but the central coverage argument does not reduce to a self-citation. The main circularity concern is the empirical support for the headline claim that 'fully Bayesian subgroup inference can maintain nominal Frequentist coverage' under adaptive subgroup selection. Section 3 proves a prior-averaged identity: if θ0 is drawn from the prior, posterior credible intervals have nominal coverage. Section 4.2 then constructs DGPs by fitting BCF and ridge regression to real data, i.e., from the same model families whose priors are evaluated. This makes the 'typical draw' premise true by construction, so the coverage numbers in Figure 4 are an illustration of the identity rather than independent evidence under prior misspecification. The paper is honest about the conditionality: Section 4.2 says the Bayesian approaches 'generally work well provided that the regression models are well-specified,' and Section 5 concedes that 'performance deteriorates when diffuse priors are used.' These limitations weaken the general claim but do not themselves constitute circularity. Theorem 3, which gives E(H^2) = στ^2(1 − e^{−λ/3}) for a modified BART prior, is not used as the basis for the coverage claim and is explicitly noted not to cover the priors used in the applications. This is a support gap, not circularity. Overall, the paper's central methodological contribution—the BRAIDS utility—is independent and non-circular, but the coverage validation is partly self-referential, warranting a score of 4 rather than 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method introduces no new physical or causal entities. The main free inputs are the analyst-chosen utility tuning parameter lambda and the prior shrinkage scale for treatment effect heterogeneity. The key unproven premise is that real treatment effect heterogeneity is small enough to be typical under the regularized prior.

free parameters (3)
  • BRAIDS risk-tuning lambda = lambda in {0, 1, 2} in the application
    Interpolates between risk-seeking, risk-neutral, and risk-averse subgroup selection. It is chosen by the analyst, not estimated from data, and is central to the utility family.
  • Heterogeneity scale s_tau / prior on sigma_tau = s_tau = 1; sigma_tau ~ Exp(1)
    Controls shrinkage of treatment effect heterogeneity. Section 3.1 fixes s_tau = 1 in the canagliflozin analysis; this is load-bearing for the coverage results.
  • Decision tree depth constraint or complexity penalty = not specified
    The utility is maximized over trees of bounded depth or with a complexity penalty, and the chosen depth changes the subgroup structure. The paper does not state a specific default value.
assumptions (3)
  • domain assumption SUTVA, consistency, strong ignorability, and positivity of treatment assignment
    Section 1.1 assumes these standard causal conditions to interpret tau(x) as the conditional average treatment effect.
  • domain assumption The true treatment effect function is a typical draw from the regularization prior, meaning real heterogeneity is small
    Section 3 argues that Bayesian intervals have frequentist coverage when theta_0 is a typical draw from pi(theta). This is the load-bearing premise for the post-selection validity claim.
  • ad hoc to paper Theorem 3 uses a modified BART prior: Poisson leaf depth, continuous covariates, and cutpoints sampled from conditional distributions
    Supplementary S.1 states the theorem does not cover the BART priors used in practice but works as an approximation, limiting the theoretical anchoring.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Decision Theoretic Subgroup Detection With Bayesian Machine Learning." pith.science (2026). https://pith.science/paper/FJV4KPM7

@misc{pith2026250905832,
  author       = {Pith},
  title        = {Pith review of: Decision Theoretic Subgroup Detection With Bayesian Machine Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FJV4KPM7}},
  note         = {Machine review of arXiv:2509.05832}
}
read the original abstract

We consider the problem of identifying promising subpopulations in terms of treatment effectiveness or treatment effect heterogeneity, from a Bayesian decision theoretic perspective. We first show that a straight-forward application of Bayesian decision theory to subgroup detection leads to a counter-intuitive risk-seeking (RS) behavior. Motivated by this observation, we introduce the Bayesian Risk-Aware Inference and Detection of Subgroups (BRAIDS) utility and use it to perform subgroup selection and post selection inference. The BRAIDS utility interpolates between risk-seeking (RS) and risk-averse (RA) identifications of subgroups, with a variant of the virtual twins algorithm as its risk-neutral midpoint. We also argue that effective subgroup estimation and inference requires the use of regularization priors to safeguard inferences from the winner's curse. We provide empirical evidence that posterior credible intervals for subgroup effects can still obtain nominal coverage levels, provided that an appropriate prior distribution is chosen. The proposed framework is illustrated on data from clinical trial assessing the efficacy of canagliflozin as a treatment for type 2 diabetes.

Figures

Figures reproduced from arXiv: 2509.05832 by the authors.

Figure 1
Figure 1. Prior distribution of the root mean squared heterog [PITH_FULL_IMAGE:figures/full_fig_p019_1.png] view at source ↗
Figure 2
Figure 2. True values of τ (X) for the simulation in Section 4.1 for each of the data generating mechanisms we consider. treatment effect heterogeneity; this is displayed in [PITH_FULL_IMAGE:figures/full_fig_p021_2.png] view at source ↗
Figure 3
Figure 3. Results of the simulation in Section 4.1. The MSE plots (left) display the average value of {τ0(Xi)−τb(Xi)} 2 across datasets and observations. The utility plots (right) display the average utility of the discovered subgroups across the datasets. Different panels give the results for different simulation ground truths. error for the CATE. Among the non-linear methods, the horserule BCF consistently yields lower MSE … view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Results for the simulation experiment in Section [PITH_FULL_IMAGE:figures/full_fig_p024_4.png]
Figure 5
Figure 5. Figure 5: Posterior summarization of the treatment effect het [PITH_FULL_IMAGE:figures/full_fig_p027_5.png]
Figure 6
Figure 6. Figure 6: Posterior density for the deviation ∆ from overall/ [PITH_FULL_IMAGE:figures/full_fig_p028_6.png]
Figure 7
Figure 7. Figure 7: Posterior for the deviation ∆ from the overall/aver [PITH_FULL_IMAGE:figures/full_fig_p029_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 40 canonical work pages

  1. [1]

    Andrews, I., Kitagawa, T., and McCloskey, A. (2024). Inference on winners. The Quarterly Journal of Economics , 139(1):305--358

  2. [2]

    and Imbens, G

    Athey, S. and Imbens, G. (2016). Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences , 113(27):7353--7360

  3. [3]

    and Wager, S

    Athey, S. and Wager, S. (2018). Policy learning with observational data. Econometrica , 89(1):133--161

  4. [4]

    and Dunn, J

    Bertsimas, D. and Dunn, J. (2017). Optimal classification trees. Machine Learning , 106:1039--1082

  5. [5]

    Chernozhukov, V., Demirer, M., Duflo, E., and Fernandez-Val, I. (2018). Generic machine learning inference on heterogeneous treatment effects in randomized experiments, with an application to immunization in I ndia. Technical report, National Bureau of Economic Research

  6. [6]

    A., George, E

    Chipman, H. A., George, E. I., and McCulloch, R. E. (1998). Bayesian CART model search. Journal of the American Statistical Association , 93(443):935--948

  7. [7]

    A., George, E

    Chipman, H. A., George, E. I., and McCulloch, R. E. (2010). BART : B ayesian additive regression trees. The Annals of Applied Statistics , 4(1):266

  8. [8]

    W., Cohen, S

    Cohen, J. W., Cohen, S. B., and Banthin, J. S. (2009). The medical expenditure panel survey: a national information resource to support healthcare cost research and inform policy and practice. Medical care , 47(7\_Supplement\_1):S44--S50

Show all 49 references
  1. [9]

    DiTraglia, F. J. and Liu, L. (2025). Bayesian double machine learning for causal inference. arXiv preprint arXiv:2508.12688

  2. [10]

    Dorie, V., Hill, J., Shalit, U., Scott, M., and Cervone, D. (2019). Automated versus do-it-yourself methods for causal inference: Lessons learned from a data analysis competition. Statistical science , 34(1):43--68

  3. [11]

    Fayyad, U. M. and Irani, K. B. (1992). On the handling of continuous-valued attributes in decision tree generation. Machine learning , 8(1):87--102

  4. [12]

    C., Taylor, J

    Foster, J. C., Taylor, J. M., and Ruberg, S. J. (2011). Subgroup identification from randomized clinical trial data. Statistics in medicine , 30(24):2867--2880

  5. [13]

    Gelman, A., Hill, J., and Yajima, M. (2012). Why we (usually) don't have to worry about multiple comparisons. Journal of research on educational effectiveness , 5(2):189--211

  6. [14]

    Grubinger, T., Zeileis, A., and Pfeiffer, K.-P. (2014). evtree: Evolutionary learning of globally optimal classification and regression trees in r. Journal of statistical software , 61:1--29

  7. [15]

    R., Carvalho, C

    Hahn, P. R., Carvalho, C. M., Puelz, D., et al. (2018). Regularization and confounding in linear regression for treatment effect estimation. Bayesian Analysis , 13(1)

  8. [16]

    R., Murray, J

    Hahn, P. R., Murray, J. S., and Carvalho, C. M. (2020). Bayesian regression tree models for causal inference: Regularization, confounding, and heterogeneous effects (with discussion). Bayesian Analysis , 15(3):965--1056

  9. [17]

    Hill, J. L. (2011). Bayesian nonparametric modeling for causal inference. Journal of Computational and Graphical Statistics , 20(1):217--240

  10. [18]

    J., Misra, S., and Zhang, W

    Hitsch, G. J., Misra, S., and Zhang, W. W. (2024). Heterogeneous treatment effects and optimal targeting policy evaluation. Quantitative Marketing and Economics , 22(2):115--168

  11. [19]

    M., and Kenney, A

    Huang, M., Tang, T. M., and Kenney, A. M. (2025). Distilling heterogeneous treatment effects: Stable subgroup estimation in causal inference. arXiv preprint arXiv:2502.07275

  12. [20]

    and Rivest, R

    Hyafil, L. and Rivest, R. L. (1976). Constructing optimal binary decision trees is NP -complete. Information Processing Letters , 5(1):15--17

  13. [21]

    and Strauss, A

    Imai, K. and Strauss, A. (2011). Estimation of heterogeneous treatment effects from randomized experiments, with application to the optimal planning of the get-out-the-vote campaign. Political Analysis , 19(1):1--19

  14. [22]

    E., Ohlssen, D

    Jones, H. E., Ohlssen, D. I., Neuenschwander, B., Racine, A., and Branson, M. (2011). Bayesian models for subgroup analysis in clinical trials. Clinical Trials , 8(2):129--143

  15. [23]

    Kennedy, E. H. (2023). Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics , 17(2):3008--3049

  16. [24]

    and Tetenov, A

    Kitagawa, T. and Tetenov, A. (2018). Who should be treated? empirical welfare maximization methods for treatment choice. Econometrica , 86(2):591--616

  17. [25]

    K., Kolassa, J

    Kuchibhotla, A. K., Kolassa, J. E., and Kuffner, T. A. (2022). Post-selection inference. Annual Review of Statistics and Its Application , 9:505--527

  18. [26]

    H., and Madigan, D

    Letham, B., Rudin, C., McCormick, T. H., and Madigan, D. (2015). Interpretable classifiers using rules and B ayesian analysis: Building a better stroke prediction model. The Annals of Applied Statistics , 9(3):1350 -- 1371

  19. [27]

    and Belford, G

    Li, R.-H. and Belford, G. G. (2002). Instability of decision tree classification algorithms. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining , pages 570--575

  20. [28]

    Linero, A. R. (2024). In nonparametric and high-dimensional models, bayesian ignorability is an informative prior. Journal of the American Statistical Association , 119(548):2785--2798

  21. [29]

    Linero, A. R. and Antonelli, J. L. (2023). The how and why of bayesian nonparametric causal inference. Wiley Interdisciplinary Reviews: Computational Statistics , 15(1):e1583

  22. [30]

    J., and Tibshirani, R

    Lockhart, R., Taylor, J., Tibshirani, R. J., and Tibshirani, R. (2014). A significance test for the lasso. The Annals of statistics , 42(2):413

  23. [31]

    Manski, C. F. (2004). Statistical treatment rules for heterogeneous populations. Econometrica , 72(4):1221--1246

  24. [32]

    and M\" u ller, P

    Morita, S. and M\" u ller, P. (2017). Bayesian population finding with biomarkers in a randomized clinical trial. Biometrics , 73:1355--1365

  25. [33]

    and Villani, M

    Nalenz, M. and Villani, M. (2018). Tree ensembles with rule structured horseshoe regularization. The Annals of Applied Statistics , 12(4):2379--2408

  26. [34]

    W., De Zeeuw, D., Fulcher, G., Erondu, N., Shaw, W., Law, G., Desai, M., and Matthews, D

    Neal, B., Perkovic, V., Mahaffey, K. W., De Zeeuw, D., Fulcher, G., Erondu, N., Shaw, W., Law, G., Desai, M., and Matthews, D. R. (2017). Canagliflozin and cardiovascular and renal events in type 2 diabetes . New England Journal of Medicine , 377(7):644--657

  27. [35]

    and Wager, S

    Nie, X. and Wager, S. (2021). Quasi-oracle estimation of heterogeneous treatment effects. Biometrika , 108(2):299--319

  28. [36]

    Nugent, C., Guo, W., M \"u ller, P., and Ji, Y. (2019). Bayesian approaches to subgroup analysis and related adaptive clinical trial designs. JCO Precision Oncology , 3:1--9

  29. [37]

    and Linero, A

    Oganisian, A. and Linero, A. (2025). Priors and propensity scores in bayesian causal inference. Observational studies , 11(1):47--60

  30. [38]

    W., Fulcher, G., Erondu, N., Shaw, W., Barrett, T

    Perkovic, V., de Zeeuw, D., Mahaffey, K. W., Fulcher, G., Erondu, N., Shaw, W., Barrett, T. D., Weidner-Wells, M., Deng, H., Matthews, D. R., et al. (2018). Canagliflozin and renal outcomes in type 2 diabetes: results from the CANVAS Program randomised clinical trials . The la...

  31. [39]

    Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology , 66(5):688

  32. [40]

    M., Tang, Q., Offen, W

    Schnell, P. M., Tang, Q., Offen, W. W., and Carlin, B. P. (2016). A B ayesian credible subgroups approach to identifying patient subgroups with positive treatment effects. Biometrics , 72(4):1026--1036

  33. [41]

    Shin, H., Linero, A., Audirac, M., Irene, K., Braun, D., and Antonelli, J. (2024). Treatment effect heterogeneity and importance measures for multivariate continuous treatments. arXiv preprint arXiv:2404.09126

  34. [42]

    Sivaganesan, S., M \"u ller, P., and Huang, B. (2017). Subgroup finding via B ayesian additive regression trees. Statistics in Medicine , 36(15):2391--2403

  35. [43]

    Sverdrup, E., Kanodia, A., Zhou, Z., Athey, S., and Wager, S. (2020). policytree: Policy learning via doubly robust empirical welfare maximization over trees. Journal of Open Source Software , 5(50):2232

  36. [44]

    Thal, D. R. and Finucane, M. M. (2023). Causal methods madness: Lessons learned from the 2022 acic competition to estimate health policy impacts. Observational Studies , 9(3):3--27

  37. [45]

    and Linero, A

    Ting, A. and Linero, A. R. (2023). Estimating heterogeneous causal mediation effects with bayesian decision tree ensembles. arXiv preprint arXiv:2303.01620

  38. [46]

    and Athey, S

    Wager, S. and Athey, S. (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association , 113(523):1228--1242

  39. [47]

    Wager, S., Du, W., Taylor, J., and Tibshirani, R. J. (2016). High-dimensional regression adjustments in randomized experiments. Proceedings of the National Academy of Sciences , 113(45):12673--12678

  40. [48]

    M., and Murray, J

    Woody, S., Carvalho, C. M., and Murray, J. S. (2021). Model interpretation through lower-dimensional posterior summarization. Journal of Computational and Graphical Statistics , 30(1):144--161

  41. [49]

    S., Hanselman, P., Walton, G

    Yeager, D. S., Hanselman, P., Walton, G. M., Murray, J. S., Crosnoe, R., Muller, C., Tipton, E., Schneider, B., Hulleman, C. S., Hinojosa, C. P., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature , 573(7774):364--369

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.