Pith. sign in

REVIEW 2 major objections 5 minor 50 references

Nonparametric estimation of an optimal treatment rule with fused randomized trials and missing effect modifiers

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read An optimal treatment rule is still identifiable when one randomized trial omits a key effect modifier.

desk verdict Promising data-fusion idea, but the identification lemma doesn't follow from the stated assumptions and the double-robustness lemma lists the wrong condition pairs. read the letter →

arxiv 2506.10863 v1 pith:5YIFWTCC submitted 2025-06-12 stat.AP stat.ME

classification stat.APstat.ME MSC 62G0562G2062P10
keywords conditionalaveragetreatmenteffectoptimaldynamicruleproxydatafusionmissingmodifierdoublyrobustestimationcross-fittingopioidusedisorder
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that when several randomized trials study the same treatment but only one trial records a discrete patient characteristic that modifies the treatment effect, the optimal treatment rule can still be recovered from the pooled data. The central device is a conditional proxy effect (CPE), a re-expression of the conditional average treatment effect whose sign matches the sign that decides treatment but whose denominator is a nuisance term and can be discarded. Under transport assumptions on the missing modifier, the CPE is identified and can be estimated by a cross-fitted, doubly robust regression. If correct, this means a clinically important effect modifier does not have to be measured in every trial to personalize treatment. In the authors' opioid use disorder analysis, housing status emerges as the strongest measured driver of whether patients do better with buprenorphine-naloxone or extended-release naltrexone, and the housing-aware rule is estimated to cut week-12 relapse by 29 percent relative to always prescribing naltrexone.

What carries the argument

The load-bearing object is the conditional proxy effect (CPE), the numerator of the Bayes-expanded CATE: $\tilde{f}(a,v_1,v_2)=P(V_2=v_2\mid Y^a=1,V_1)P(Y^a=1\mid V_1)$, contrasted across $a=0,1$. Lemma 2.1 rewrites it as an integral over the covariate distribution, $\int b(a,v_2,v_1)m(a,v_1)P(w\mid v_1)\,dw$, which shifts the missing modifier's distribution across trials: $b$ is learned only from observations with $S=1$, while $m$ uses all observations. The estimator in Algorithm 1 regresses the uncentered efficient influence function $\xi(O,a,v_2;\eta)=\varphi_b+\varphi_m+b m$ on $V_1$; because this pseudo-outcome has conditional mean equal to $\tilde{f}(a,v_1,v_2)$ under three alternative combinations of nuisance consistency (Lemma 2.2), the CPE estimate is doubly robust. Cross-fitting supplies the oracle-style convergence guarantees stated in Lemma 2.3.

What would settle it

Estimate $P(V_2=v_2\mid Y=1,A=a,W,V_1,S=s)$ in a dataset where the missing modifier is recorded in both trial populations, or simulate the paper's data-generating mechanism with assumption A5 deliberately violated; if the cross-trial conditional distributions of $V_2$ differ enough to flip the sign of $\tilde{\tau}$ for some patient subgroup, then the proxy-effect rule is not the optimal rule for that subgroup.

Watch

Extended reading notes

Core claim

The paper's central claim is that the sign of the conditional average treatment effect $\tau(v_1,v_2)=f(1,v_1,v_2)-f(0,v_1,v_2)$ equals the sign of the conditional proxy effect $\tilde{\tau}(v_1,v_2)=\tilde{f}(1,v_1,v_2)-\tilde{f}(0,v_1,v_2)$, where $\tilde{f}(a,v_1,v_2)=P(V_2=v_2\mid Y^a=1,V_1=v_1)P(Y^a=1\mid V_1=v_1)$. Remark 1.1 establishes this sign equivalence, so the optimal dynamic treatment rule can be defined through $\tilde{\tau}$ alone; the denominator $P(V_2=v_2\mid V_1=v_1)$ drops out because it is common to both treatment arms. Lemma 2.1 identifies $\tilde{f}(a,v_1,v_2)$ as $\int b(a,v_2,v_1)m(a,v_1)P(w\mid v_1)\,dw$, where $b(a,v_2,v_1)=P(V_2=v_2\mid Y=1,A=a,W,V_1=v_1,S=1)$ is estimated in the trial that records $V_2$ and $m(a,v_1)=E[Y\mid A=a,V_1=v_1,W]$ is estimated from all trials. Algorithm 1 then learns $\tilde{\tau}$ through cross-fitted pseudo-outcome regressions built from the uncentered efficient influence function, and Lemma 2.2 shows the estimator's conditional mean remains correct when certain pairs of nuisance functions are misspecified.

Load-bearing premise

The method stands or falls on the transport assumption that, within strata of the measured covariates and the would-be outcome, the chance of having the missing modifier (such as being unhoused) is the same in every trial; if the trial populations differ in that modifier for reasons the measured covariates do not capture, the recovered treatment rule will be biased.

Editorial extensions

If this is right

  • Discrete effect modifiers that are missing in some trials can be reintroduced into the optimal treatment rule, so pooling trials no longer requires identical covariate collection across protocols.
  • Because only the numerator of the Bayes-expanded CATE is needed, the target is smoother than the CATE itself and can be estimated nonparametrically with flexible machine learning.
  • Under any of the three doubly robust conditions in Lemma 2.3, the estimated rule is oracle-efficient: if the second-stage regression converges quickly, the nuisance estimation error does not dominate the total error.
  • In the opioid use disorder application, the housing-aware rule is estimated to reduce week-12 relapse by about 29 percent relative to always prescribing XR-NTX and by about 2 percent relative to a rule that ignores housing status, although the second comparison is not statistically significant.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The sign-equivalence trick is likely more general than the paper's discrete-modifier setting: the denominator drops out for any effect modifier whose distribution does not depend on treatment within the relevant strata, so a similar proxy construction should extend to continuous missing modifiers or to several missing modifiers at once.
  • The decisive-versus-ambiguous bounds suggest an adaptive data-collection design in which the missing modifier is measured only for patients whose proxy-effect bounds straddle zero, potentially making trial fusion cheaper without sacrificing rule accuracy.
  • A reader applying these results clinically would want a sensitivity analysis over the transport assumption, for example multiplying the trial-specific housing distribution by plausible ratios, to see whether the housing-aware rule's small advantage over the no-housing rule persists when assumption A5 is violated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes a nonparametric estimator of an optimal dynamic treatment rule when a discrete effect modifier is measured in only one of several pooled randomized trials. The central idea is a Bayes-rule decomposition of the conditional average treatment effect (CATE): the sign of the CATE is shown to equal the sign of a conditional proxy effect (CPE), and the CPE is claimed to be identified from the pooled trial data under assumptions A1-A6. Estimation is carried out with a cross-fitted DR-Learner based on a doubly-robust pseudo-outcome, and the method is illustrated in a simulation study and in an application to two opioid use disorder trials (CTN0051 and CTN0030) where housing status is missing in one trial. The paper also provides bounds for treatment decisions when the missing modifier is unobserved and evaluates rules with TMLE.

Significance. If the identification result holds, the CPE decomposition is a genuinely useful contribution: it reduces the problem of learning an optimal rule with a partially missing effect modifier to a standard pseudo-outcome regression, and it does so without fitting constants or requiring the denominator of the CATE decomposition to be estimated. The cross-fitted DR-Learner formulation follows established machinery, and the authors provide code and a clinically relevant application. However, the two technical gaps described below currently prevent the central claims from being accepted as stated: Lemma 2.1 does not follow from A1-A6 as written, and Lemma 2.3's robustness conditions are inconsistent with the remainder terms in Lemma 2.2. Both issues are local and repairable, so the appropriate disposition is major revision rather than rejection.

major comments (2)
  1. [Lemma 2.1 and Supplementary §4.4] The identification proof does not follow from assumptions A4 and A5 as stated. In the proof, the step P(V2 | Y=1, A=a, W, V1) = P(V2 | Y=1, A=a, W, V1, S=1) is attributed to A5, but A5 only states V2 ⊥ S | W, V1, Y^a; it does not condition on A. Since the model allows A = f_A(S, U_A), conditioning on A can create an association between V2 and S even when A4 and A5 both hold. A concrete distribution satisfying A1-A6 is: W and V1 empty; S and A independent Bernoulli(0.5); V2 = S xor A; and P(Y^a=1 | V2) = 0.9 if V2=1, 0.1 if V2=0. Then P(V2=1 | Y=1, A=1, S=1) = 0, whereas P(V2=1 | Y^a=1) = 0.9, so the formula in Lemma 2.1 returns tilde_f(1,1)=0 instead of the true value 0.45. The lemma becomes correct if the assumptions are strengthened to V2 ⊥ (A, S) | W, V1, Y^a (which is implied by the NPSEM in Section 2.1), or if A4/A5 are revised to condition on A. As written, the central identification result is not established from A1-A6.
  2. [Lemma 2.3 versus Lemma 2.2] The three consistency conditions listed in Lemma 2.3 do not match the exact double-robustness pairs in Lemma 2.2. Lemma 2.2 shows that the remainder Rem(a,v1,v2;η') is exactly zero when (g',b') = (g,b), or (r',m') = (r,m), or (b',m') = (b,m). Lemma 2.3 condition 2, however, assumes consistency of g and m, and condition 3 assumes consistency of r and b. Under condition 2, the term C'_b = m'(r-r')/r' (b-b') need not vanish because b' and r' are unrestricted by the condition; under condition 3, the term C'_m = b'(g-g')/g' (m-m') need not vanish because g' and m' are unrestricted. As written, only condition 1 is supported by the remainder decomposition in Lemma 2.2. The conditions in Lemma 2.3 should be corrected to (g,b), (r,m), and (b,m), or the remainder analysis must be revised to show that the stated pairs indeed yield Rem = o_P(1).
minor comments (5)
  1. [Section 2.5, Equation (12)] Equation (12) uses arg min and arg max on the left- and right-hand sides of an inequality, but an arg min is an argument, not a scalar bound; the intended expressions are the minimum and maximum of tilde_tau(v1,i, v*_2) over v*_2.
  2. [Section 2.1] The NPSEM in Section 2.1 implies the stronger independence V2 ⊥ (S, A) | W, V1, which is sufficient for Lemma 2.1. The relationship between this structural assumption and the weaker stated conditions A4-A5 should be clarified, especially because the proof of Lemma 2.1 cites only A4 and A5.
  3. [Section 1] The phrase "a estimator" in the introduction should be "an estimator".
  4. [Supplementary Table 2 caption] The caption contains the typo "Conditonal" and should read "Conditional proxy effect".
  5. [Supplementary §4.3] The sentence "The red densities only be estimated among observations where S=1, while the blue density can be estimated among all observations" is ungrammatical, and the colors are not defined in the surrounding text.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the CPE identification is a self-contained Bayes-rule derivation; self-citations are not load-bearing.

full rationale

The paper's central claim is Lemma 2.1, which identifies the conditional proxy effect tilde_f(a,v1,v2) as an integral of observed-data quantities b(a,v2,v1) and m(a,v1) with respect to P(w|v1) under assumptions A1-A6. The proof in Supplementary Section 4.4 is a direct Bayes-rule and tower-rule derivation; no fitted constant enters the estimand and no equality is assumed from prior work. Remark 1.1 states that the sign of the CPE equals the sign of the CATE because the omitted denominator P(V2|V1) is positive; this is a mathematical equivalence, not a prediction forced by construction. The DR-Learner of Algorithm 1 is built on published doubly-robust estimation machinery (Kennedy et al.; Van der Laan), with nuisance parameters estimated from the data, and the simulation study evaluates the estimator against known true values. Self-citations to Rudolph, Williams, and Díaz are used to motivate the opioid-use application and select baseline covariates, but they do not supply the identification argument. A reviewer concern that Lemma 2.1's use of assumption A5 may require a stronger conditional independence condition (V2 independent of S given W, V1, Ya, A) is a correctness or assumption-strength issue, not a circularity, because the proof does not reduce to its inputs by definition. No equation in the manuscript is equivalent to a fitted value or to a self-citation chain.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The central claim rests on causal assumptions A1-A6 and the NPSEM with independent errors and no arrow V2 to S. A4 and A5 are nontrivial; in the application, A5 requires housing distribution transportability across two trials with different recruitment populations. No free parameters are introduced into the estimand, and no new entities are postulated.

assumptions (8)
  • domain assumption A1: Consistency, A=a implies Y=Y^a.
    Standard causal assumption invoked in Section 2.3 and in the proof of Lemma 2.1.
  • domain assumption A2: Positivity of treatment assignment P(A=a|W,V1) > epsilon.
    Required for the pseudo-outcome weights and for identifying m(a,v1).
  • domain assumption A3: Exchangeability Y^a independent of A given W,V1.
    Standard no-unmeasured-confounding assumption, justified by randomization within trials.
  • ad hoc to paper A4: V2 independent of A given W,V1,Y^a.
    Nonstandard assumption used to replace A by a in the conditional distribution of V2 in the identification proof.
  • ad hoc to paper A5: V2 independent of S given W,V1,Y^a.
    Key transport assumption: the distribution of the missing effect modifier must be the same across trials after conditioning on covariates and potential outcomes.
  • domain assumption A6: Joint positivity P(Y=1,A=a,S=1|W,V1) > epsilon.
    Required for the weight alpha_r used in the efficient influence function for b(a,v2,v1).
  • domain assumption NPSEM with mutually independent U and no arrow from V2 to S.
    The structural equation model in Section 2.1 encodes that trial membership does not directly cause V2 and that unmeasured factors are independent.
  • domain assumption P(tau(v1,v2)=0)=0.
    Assumed in Section 2.2 so the ODTR is the sign of the CATE without ties.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Nonparametric estimation of an optimal treatment rule with fused randomized trials and missing effect modifiers." pith.science (2026). https://pith.science/paper/5YIFWTCC

@misc{pith2026250610863,
  author       = {Pith},
  title        = {Pith review of: Nonparametric estimation of an optimal treatment rule with fused randomized trials and missing effect modifiers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5YIFWTCC}},
  note         = {Machine review of arXiv:2506.10863}
}
read the original abstract

A fundamental principle of clinical medicine is that a treatment should only be administered to those patients who would benefit from it. Treatment strategies that assign treatment to patients as a function of their individual characteristics are known as dynamic treatment rules. The dynamic treatment rule that optimizes the outcome in the population is called the optimal dynamic treatment rule. Randomized clinical trials are considered the gold standard for estimating the marginal causal effect of a treatment on an outcome; they are often not powered to detect heterogeneous treatment effects, and thus, may rarely inform more personalized treatment decisions. The availability of multiple trials studying a common set of treatments presents an opportunity for combining data, often called data-fusion, to better estimate dynamic treatment rules. However, there may be a mismatch in the set of patient covariates measured across trials. We address this problem here; we propose a nonparametric estimator for the optimal dynamic treatment rule that leverages information across the set of randomized trials. We apply the estimator to fused randomized trials of medications for the treatment of opioid use disorder to estimate a treatment rule that would match patient subgroups with the medication that would minimize risk of return to regular opioid use.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

50 extracted references · 45 canonical work pages

  1. [1]

    , Binder , Martin M

    barticle [author] Bischl , Bernd B. , Binder , Martin M. , Lang , Michel M. , Pielok , Tobias T. , Richter , Jakob J. , Coors , Stefan S. , Thomas , Janek J. , Ullmann , Theresa T. , Becker , Marc M. , Boulesteix , Anne-Laure A.-L. et al. ( 2023 ). Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges . Wiley Interdisci...

  2. [2]

    barticle [author] Brantner , Carly Lupton C. L. , Chang , Ting-Hsuan T.-H. , Nguyen , Trang Quynh T. Q. , Hong , Hwanhee H. , Di Stefano , Leon L. Stuart , Elizabeth A E. A. ( 2023 ). Methods for integrating trials and non-experimental data to examine treatment effect heterogeneity . Statistical science: a review journal of the Institute of Mathematical S...

  3. [3]

    barticle [author] Brantner , Carly Lupton C. L. , Nguyen , Trang Quynh T. Q. , Tang , Tengjie T. , Zhao , Congwen C. , Hong , Hwanhee H. Stuart , Elizabeth A E. A. ( 2024 ). Comparison of methods that combine multiple randomized trials to estimate heterogeneous treatment effects . Statistics in medicine 43 1291--1314 . barticle

  4. [4]

    ( 2024 )

    bmisc [author] Centers for Disease Control, National Center for Health Statistics, Office of Communication. ( 2024 ). U.S . overdose deaths decrease in 2023, first time since 2018 . Press release . bmisc

  5. [5]

    Guestrin , Carlos C

    binproceedings [author] Chen , Tianqi T. Guestrin , Carlos C. ( 2016 ). XGBoost: A Scalable Tree Boosting System . In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785--794 . ACM . 10.1145/2939672.2939785 binproceedings

  6. [6]

    barticle [author] Chipman , Hugh A. H. A. , George , Edward I. E. I. McCulloch , Robert E. R. E. ( 2010 ). BART : Bayesian Additive Regression Trees . The Annals of Applied Statistics 4 266--298 . 10.1214/09-AOAS285 barticle

  7. [7]

    barticle [author] Friedman , Jerome H J. H. ( 1991 ). Multivariate adaptive regression splines . The annals of statistics 19 1--67 . barticle

  8. [8]

    barticle [author] Goedel , William C W. C. , Shapiro , Aaron A. , Cerd \'a , Magdalena M. , Tsai , Jennifer W J. W. , Hadland , Scott E S. E. Marshall , Brandon DL B. D. ( 2020 ). Association of racial/ethnic segregation with treatment capacity for opioid use disorder in counties in the United States . JAMA network open 3 e203711--e203711 . barticle

Show all 50 references
  1. [9]

    Van Der Laan , Mark J M

    barticle [author] Gruber , Susan S. Van Der Laan , Mark J M. J. ( 2009 ). Targeted maximum likelihood estimation: A gentle introduction . U.C. Berkeley Division of Biostatistics working paper series . barticle

  2. [10]

    barticle [author] Hansen , Helena B H. B. , Siegel , Carole E C. E. , Case , Brady G B. G. , Bertollo , David N D. N. , DiRocco , Danae D. Galanter , Marc M. ( 2013 ). Variation in use of buprenorphine and methadone treatment by racial, ethnic, and income characteristics of re...

  3. [11]

    , Siegel , Carole C

    barticle [author] Hansen , Helena H. , Siegel , Carole C. , Wanderling , Joseph J. DiRocco , Danae D. ( 2016 ). Buprenorphine and methadone treatment for opioid dependence by income, ethnicity and race of neighborhoods in New York City . Drug and alcohol dependence 164 14--21 ...

  4. [12]

    bbook [author] Imbens , Guido W G. W. Rubin , Donald B D. B. ( 2015 ). Causal inference in statistics, social, and biomedical sciences . Cambridge university press . bbook

  5. [13]

    barticle [author] Jarvis , Brantley P B. P. , Holtyn , August F A. F. , Subramaniam , Shrinidhi S. , Tompkins , D Andrew D. A. , Oga , Emmanuel A E. A. , Bigelow , George E G. E. Silverman , Kenneth K. ( 2018 ). Extended-release injectable naltrexone for opioid use disorder: a...

  6. [14]

    barticle [author] Kennedy , Edward H E. H. ( 2023 ). Towards optimal doubly robust estimation of heterogeneous causal effects . Electronic Journal of Statistics 17 3008--3049 . barticle

  7. [15]

    barticle [author] Kennedy , Edward H E. H. , Ma , Zongming Z. , McHugh , Matthew D M. D. Small , Dylan S D. S. ( 2017 ). Non-parametric methods for doubly robust estimation of continuous treatment effects . Journal of the Royal Statistical Society Series B: Statistical Methodo...

  8. [16]

    barticle [author] Kent , David M D. M. , Rothwell , Peter M P. M. , Ioannidis , John PA J. P. , Altman , Doug G D. G. Hayward , Rodney A R. A. ( 2010 ). Assessing and reporting heterogeneity in treatment effects in clinical trials: a proposal . Trials 11 1--11 . barticle

  9. [17]

    barticle [author] Kosorok , Michael R M. R. Laber , Eric B E. B. ( 2019 ). Precision medicine . Annual review of statistics and its application 6 263--286 . barticle

  10. [18]

    bmisc [author] Kosorok , Michael R M. R. , Laber , Eric B E. B. , Small , Dylan S D. S. Zeng , Donglin D. ( 2021 ). Introduction to the theory and methods special issue on precision medicine and individualized policy discovery . bmisc

  11. [19]

    barticle [author] Lee , Joshua D J. D. , Nunes , Edward V E. V. , Novo , Patricia P. , Bachrach , Ken K. , Bailey , Genie L G. L. , Bhatt , Snehal S. , Farkas , Sarah S. , Fishman , Marc M. , Gauthier , Phoebe P. , Hodgkins , Candace C C. C. et al. ( 2018 ). Comparative effect...

  12. [20]

    , Wang , Yuanjia Y

    barticle [author] Liu , Ying Y. , Wang , Yuanjia Y. , Kosorok , Michael R M. R. , Zhao , Yingqi Y. Zeng , Donglin D. ( 2018 ). Augmented outcome-weighted learning for estimating optimal dynamic treatment regimens . Statistics in medicine 37 3776--3788 . barticle

  13. [21]

    barticle [author] Luedtke , Alexander R A. R. van der Laan , Mark J M. J. ( 2016 ). Super-learning of an optimal dynamic treatment rule . The international journal of biostatistics 12 305--332 . barticle

  14. [22]

    binproceedings [author] Lundberg , Scott M S. M. Lee , Su-In S.-I. ( 2017 ). A Unified Approach to Interpreting Model Predictions . In Advances in Neural Information Processing Systems 30 . binproceedings

  15. [23]

    barticle [author] Moodie , Erica EM E. E. , Chakraborty , Bibhas B. Kramer , Michael S M. S. ( 2012 ). Q-learning for estimating optimal dynamic treatment rules from observational data . Canadian Journal of Statistics 40 629--645 . barticle

  16. [24]

    barticle [author] Murphy , Susan A S. A. ( 2003 ). Optimal dynamic treatment regimes . Journal of the Royal Statistical Society Series B: Statistical Methodology 65 331--355 . barticle

  17. [25]

    ( 2009 )

    bbook [author] Pearl , Judea J. ( 2009 ). Causality . Cambridge university press . bbook

  18. [26]

    Bareinboim , Elias E

    binproceedings [author] Pearl , Judea J. Bareinboim , Elias E. ( 2011 ). Transportability of causal and statistical relations: A formal approach . In Proceedings of the AAAI Conference on Artificial Intelligence 25 247--254 . binproceedings

  19. [27]

    barticle [author] Richardson , Thomas S T. S. Robins , James M J. M. ( 2013 ). Single world intervention graphs (SWIGs): A unification of the counterfactual and graphical approaches to causality . Center for the Statistics and the Social Sciences, University of Washington Seri...

  20. [28]

    , Li , Lingling L

    barticle [author] Robins , James J. , Li , Lingling L. , Tchetgen , Eric E. van der Vaart , Aad W A. W. ( 2009 ). Quadratic semiparametric von mises calculus . Metrika 69 227--247 . barticle

  21. [29]

    barticle [author] Rudolph , Kara E K. E. , D \' az , Iv \'a n I. , Luo , Sean X S. X. , Rotrosen , John J. Nunes , Edward V E. V. ( 2021 ). Optimizing opioid use disorder treatment with naltrexone or buprenorphine . Drug and alcohol dependence 228 109031 . barticle

  22. [30]

    barticle [author] Rudolph , Kara E K. E. , Williams , Nicholas T N. T. , D \' az , Iv \'a n I. , Luo , Sean X S. X. , Rotrosen , John J. Nunes , Edward V E. V. ( 2023 ). Optimally choosing medication type for patients with opioid use disorder . American journal of epidemiology...

  23. [31]

    , Song , Weishan W

    barticle [author] Schader , Lindsey L. , Song , Weishan W. , Kempker , Russell R. Benkeser , David D. ( 2024 ). Don’t Let Your Analysis Go to Seed: On the Impact of Random Seed on Machine Learning-based Causal Inference . Epidemiology 35 764--778 . barticle

  24. [32]

    barticle [author] Stuart , Elizabeth A E. A. , Bradshaw , Catherine P C. P. Leaf , Philip J P. J. ( 2015 ). Assessing the generalizability of randomized trial results to target populations . Prevention Science 16 475--485 . barticle

  25. [33]

    R: A Language and Environment for Statistical Computing R Foundation for Statistical Computing , Vienna, Austria

    bmanual [author] R Core Team ( 2024 ). R: A Language and Environment for Statistical Computing R Foundation for Statistical Computing , Vienna, Austria . bmanual

  26. [34]

    ( 1996 )

    barticle [author] Tibshirani , Robert R. ( 1996 ). Regression shrinkage and selection via the lasso . Journal of the Royal Statistical Society Series B: Statistical Methodology 58 267--288 . barticle

  27. [35]

    barticle [author] Van der Laan , Mark J M. J. ( 2006 ). Statistical inference for variable importance . The International Journal of Biostatistics 2 . barticle

  28. [36]

    barticle [author] van der Laan , Mark J M. J. Gruber , Susan S. ( 2011 ). Targeted minimum loss based estimation of an intervention specific mean outcome . U.C. Berkeley Division of Biostatistics working paper series . barticle

  29. [37]

    barticle [author] van der Laan , Mark J M. J. Luedtke , Alexander R A. R. ( 2015 ). Targeted learning of the mean outcome under an optimal dynamic treatment rule . Journal of causal inference 3 61--95 . barticle

  30. [38]

    , Luedtke , Alex A

    barticle [author] van der Laan , Lars L. , Luedtke , Alex A. Carone , Marco M. ( 2024 ). Automatic doubly robust inference for linear functionals via calibrated debiased machine learning . arXiv preprint arXiv:2411.02771 . barticle

  31. [39]

    barticle [author] Van der Laan , Mark J M. J. , Polley , Eric C E. C. Hubbard , Alan E A. E. ( 2007 ). Super learner . Statistical applications in genetics and molecular biology 6 . barticle

  32. [40]

    barticle [author] Volkow , Nora D N. D. , Jones , Emily B E. B. , Einstein , Emily B E. B. Wargo , Eric M E. M. ( 2019 ). Prevention and treatment of opioid misuse and addiction: a review . JAMA psychiatry 76 208--216 . barticle

  33. [41]

    ( 1947 )

    barticle [author] von Mises , R R. ( 1947 ). On the asymptotic distribution of differentiable statistical functions . The annals of mathematical statistics 18 309--348 . barticle

  34. [42]

    barticle [author] Watkins , Christopher John Cornish Hellaby C. J. C. H. et al. ( 1989 ). Learning from delayed rewards . barticle

  35. [43]

    barticle [author] Weiss , Roger D R. D. , Potter , Jennifer Sharpe J. S. , Fiellin , David A D. A. , Byrne , Marilyn M. , Connery , Hilary S H. S. , Dickinson , William W. , Gardin , John J. , Griffin , Margaret L M. L. , Gourevitch , Marc N M. N. , Haller , Deborah L D. L. et...

  36. [44]

    barticle [author] Williams , Nicholas T N. T. , Hung , Anton A. Rudolph , Kara K. ( 2025 ). Re: Don’t let your analysis go to seed: on the impact of random seed on machine learning-based causal inference . Epidemiology . barticle

  37. [45]

    barticle [author] Williams , Arthur Robin A. R. , Nunes , Edward V E. V. , Bisaga , Adam A. , Pincus , Harold A H. A. , Johnson , Kimberly A K. A. , Campbell , Aimee N A. N. , Remien , Robert H R. H. , Crystal , Stephen S. , Friedmann , Peter D P. D. , Levin , Frances R F. R. ...

  38. [46]

    barticle [author] Williams , Nicholas T N. T. , Hoffman , Katherine L K. L. , D \' az , Iv \'a n I. Rudolph , Kara E K. E. ( 2024 ). Learning optimal dynamic treatment regimes from longitudinal data . American Journal of Epidemiology 193 1768--1775 . barticle

  39. [47]

    barticle [author] Wright , Marvin N. M. N. Ziegler , Andreas A. ( 2017 ). ranger : A Fast Implementation of Random Forests for High Dimensional Data in C++ and R . Journal of Statistical Software 77 1--17 . 10.18637/jss.v077.i01 barticle

  40. [48]

    , Zeng , Donglin D

    barticle [author] Zhao , Yingqi Y. , Zeng , Donglin D. , Rush , A John A. J. Kosorok , Michael R M. R. ( 2012 ). Estimating individualized treatment rules using outcome weighted learning . Journal of the American Statistical Association 107 1106--1118 . barticle

  41. [49]

    , Mayer-Hamblett , Nicole N

    barticle [author] Zhou , Xin X. , Mayer-Hamblett , Nicole N. , Khan , Umer U. Kosorok , Michael R M. R. ( 2017 ). Residual weighted learning for estimating individualized treatment rules . Journal of the American Statistical Association 112 169--187 . barticle

  42. [50]

    barticle [author] Zivich , Paul N P. N. Breskin , Alexander A. ( 2021 ). Machine learning for causal inference: on the use of cross-fit estimators . Epidemiology 32 393--401 . barticle

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.