Pith. sign in

REVIEW 5 major objections 4 minor 43 references

This paper argues that individualized antidepressant switching effects can be estimated from observational visit data, and that a causal-forest estimator yields modest, actionable next-visit HAMD-17 benefits—roughly 0.3 to 0.9 points—once c

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 11:03 UTC pith:HSGGCSGJ

load-bearing objection A useful benchmark undermined by untested unconfoundedness and a circular evaluation metric; CF's factual performance is plausible but the causal estimates need sensitivity analysis and the meta-learner collapse needs diagnosis. the 5 major comments →

arxiv 2607.27214 v1 pith:HSGGCSGJ submitted 2026-06-16 stat.AP cs.AI

Estimating Treatment Effects for Depression in Longitudinal Therapy Switching Settings

classification stat.AP cs.AI
keywords treatment effect estimationcausal forestantidepressant switchingHAMD-17counterfactual predictionobservational confoundingmeta-learnerspolicy value
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to show that counterfactual prediction from longitudinal depression care data can support visit-level medication-switch decisions, despite the fact that treatment changes are driven by patient severity and prior response. It frames the task as predicting next-visit HAMD-17 depression scores under six alternative treatments, then benchmarks eight estimators under a unified evaluation protocol. The central claim is that Causal Forest performs best and most consistently across all criteria, while common meta-learners fail badly in this switched, confounded setting. The paper also claims that confounding adjustment shrinks apparent switch benefits from 6–7 HAMD-17 points down to modest, plausible 0.3–0.9 point reductions, with dose intensification generally helping and a notable subset where a lower-intensity regimen beats a higher-intensity one. A sympathetic reader would care because this is a direct, prospectively testable route to clinical decision support for a common treatment dilemma.

Core claim

In a switched longitudinal MDD dataset where every record is a treatment change, the paper finds that a baseline-referenced extension of Causal Forest—pairwise honest forests comparing each treatment to a common baseline, anchored by a flexible baseline outcome model—delivers the most favorable and consistent next-visit HAMD-17 predictions across factual error, calibration, and inverse-propensity-weighted policy value. Confounding-adjusted average effects are modest: roughly 0.3 to 0.9 point reductions for duloxetine 40 mg BID, venlafaxine titration, dose escalation, and paroxetine relative to duloxetine 60 mg QD, far below the 6–7 point differences seen in crude observational comparisons. T

What carries the argument

The load-bearing mechanism is the baseline-referenced pairwise causal forest: the most frequent treatment is chosen as baseline, separate honest causal forests estimate each other treatment's effect relative to that baseline, and a flexible baseline outcome model reconstructs all potential outcomes as baseline prediction plus treatment-contrast. Honest splitting—using separate samples for tree structure and leaf-effect estimation—is what prevents adaptive overfitting and gives the method its calibration edge. The evaluation protocol pairs this with an observational calibration proxy and IPW policy value as decision-quality measures in a setting where true counterfactual outcomes are absent.

Load-bearing premise

The central claim assumes that, after controlling for the recorded covariates, nothing unmeasured that drives the decision to switch (such as side effects, prior non-response, or clinician judgment) also influences the next-visit HAMD-17 score; if that fails, every estimated effect—including the modest causal-forest estimates—is biased.

What would settle it

A randomized switching trial comparing the six regimens, or a sensitivity analysis that adds reason-for-switch and side-effect variables to the covariate set, would settle the central magnitude: if adding these variables moves the causal-forest estimates by more than roughly 0.5 HAMD-17 points, or if trial intent-to-treat effects land near zero instead of the predicted 0.3–0.9 point reductions, the paper's central claim is not robust.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim holds, clinicians can obtain per-visit rankings of six antidepressant options from observational records, with dose titration and specific switches flagged as beneficial rather than relying on crude group averages.
  • The estimated effect sizes imply that realistic gains from a data-driven switch are on the order of one HAMD-17 point, not the 6–7 points that unadjusted comparisons suggest; expectations for decision support should be calibrated accordingly.
  • The benchmark protocol—factual error, calibration, proxy-based ATE error, and IPW policy value—provides a reusable template for evaluating counterfactual predictors in any observational setting where outcomes under alternative treatments are unobserved.
  • The paper's ablation results indicate that honest splitting is essential for this class of methods; estimators that use all data for both splitting and effect estimation degrade measurably in switched settings.
  • Meta-learners as implemented here should not be deployed for switch recommendations without substantial modification, since their predictions collapse toward selection-driven outcome means.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the modest effect sizes survive prospective validation, the practical payoff is less about finding large gains and more about avoiding harmful switches or identifying which specific dose change is most likely to help a given patient.
  • The paper's counterintuitive lower-intensity exception suggests a testable hypothesis: for certain severity or comorbidity profiles, stepping down may outperform escalating; this could be probed directly in a randomized switch trial stratified by the identified patient subsets.
  • The baseline-referenced pairwise forest recipe is generic for any K-treatment observational panel and could transfer to other chronic-disease switching registries, but its reliability will hinge on capturing the actual reasons for switching—side effects, prior non-response, clinician judgment—which are absent here.
  • The observational calibration proxy assumes crude group differences are a meaningful reference; if confounding is strong, that metric could reward estimators that merely shrink predictions, so external validation remains the real test.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. This paper formulates a next-visit counterfactual prediction task for six antidepressant treatments using a proprietary longitudinal MDD dataset in which all records are flagged as switching events. It benchmarks eight estimators—Causal Forest (CF), EP-learner, KG-TREAT, R-learner, and S/T/X/DR meta-learners—under a four-criteria evaluation protocol: factual prediction error, ATE calibration error against the observational group mean difference Δ_OCP, outcome calibration, and IPW policy value. The authors report that CF performs best across all criteria, with predicted treatment effects of -0.11 to -0.92 HAMD-17 points versus crude observational differences of -6 to -7 points, and they present a baseline-referenced extension of CF for multiple treatments. Ablation studies support the importance of honest splitting and flexible baseline outcome modeling.

Significance. If the causal estimates were valid, the paper would provide a useful benchmark protocol and a practical extension of CF for visit-level switching decisions, with modest and actionable effect sizes. The manuscript is commendable for reporting bootstrap uncertainties, ablation studies, and IPW-based decision evaluation. However, the central causal claims rest on untested unconfoundedness (A2), and the headline evaluation metric E_ATE uses a confounded observational difference as its target. The implausible failure of the meta-learners also raises concerns about the benchmark's internal validity. These issues substantially limit the current support for the abstract's conclusions, though they are potentially addressable within the manuscript's scope.

major comments (5)
  1. [Section II.C (A2) and Section VII] The unconfoundedness assumption A2 is load-bearing for the abstract's claim that 'confounding-adjusted estimates yield modest, actionable magnitudes' (0.3–0.9 HAMD-17 point reductions). The dataset consists entirely of switching events; reasons for switching (side effects, prior non-response, clinician judgment) are not in the covariate list, and the authors' own descriptive statistics show paroxetine patients have 50% higher mean HAMD-17 than duloxetine 60 mg QD continuers, demonstrating severity-driven selection. The paper acknowledges A2 is untestable and postpones sensitivity analysis to future work. Without a sensitivity analysis (e.g., Cinelli–Hazlett, ref. [43]) or a clear statement that all estimates are conditional on A2, the quantitative effect sizes in Table II cannot be interpreted as causal. Please add such an analysis or substantially temper the causal language in the abstr
  2. [Section II.D and Section V.A] The ATE calibration error E_ATE is computed against Δ_OCP, the unadjusted test-set mean outcome difference between treatment groups. The authors correctly note that Δ_OCP 'is not a causal estimand,' yet they subsequently use E_ATE as one of the criteria supporting CF's 'most favorable and consistent performance across all criteria' (Table I, Section V.A). A low E_ATE only indicates that a method's predicted effects are close to the confounded group difference; it does not indicate causal accuracy. Because the paper's goal is to estimate causal effects, this metric should either be removed from the primary benchmark or explicitly labeled as a 'crude-difference reproduction' metric, not as ATE calibration. The sentence 'CF demonstrates the best ATE calibration error' should be reworded accordingly.
  3. [Section V.A, Table I] The reported meta-learner performance is implausible: S/T/X/DR-learners achieve RMSE 14.75–28.20, all worse than predicting the unconditional test mean (RMSE 8.03), with negative calibration slopes (-0.44 to -0.11). Since all baselines use XGBoost with hyperparameter tuning and are said to use the same baseline-referenced construction as CF (Section III.B), this failure pattern strongly suggests an implementation artifact in how meta-learners are converted into potential-outcome predictions, rather than a genuine property of those methods. The manuscript provides no code, pseudocode, or sanity check (e.g., a predict-mean baseline) to rule this out. Please provide a precise algorithmic description for each baseline, make code available, or run a synthetic-data validation. Without this, the headline result that CF 'substantially outperforms meta-learners' is not adequately supported.
  4. [Abstract vs. Section V] The abstract claims 'a counterintuitive exception where a lower-intensity regimen outperforms a higher-intensity alternative for specific patient subsets.' I could not locate any subgroup analysis or result in the body that supports this claim. The closest content in Section V.B discusses dose escalation being beneficial; no patient subset is identified as having a lower-intensity regimen outperform a higher-intensity one. Either add the corresponding subgroup analysis with quantitative results (e.g., a table or figure in Section V) or remove this claim from the abstract. An unsupported claim in the abstract cannot be evaluated as a contribution.
  5. [Section II.A and Section IV.A] The potential outcome framework defines Y_i(t) as the outcome under treatment t, but in a switching setting the effect of assigning treatment t at visit v depends on the treatment received before the switch (carry-over effects). The covariate list in Section IV.A mentions only demographics and HAMD-17 item scores; prior treatment is not listed. Consequently, the consistency assumption (A1) is ambiguous: patients with identical X_i and same current t could have different potential outcomes depending on their previous regimen. Please clarify whether prior treatment (or switch direction) is included as a covariate. If not, either include treatment history in the model or explicitly limit the estimand to 'effect of current treatment conditional on the observed switching context' and discuss the bias that may arise from omitting prior treatment.
minor comments (4)
  1. [Section II.D] The definition of Ω ('valid treatment pairs') is not given. Specify how pairs are selected, e.g., minimum sample size per treatment group, and whether all 15 pairs are included.
  2. [Table II caption] The caption uses 'EP' without defining it. State that 'EP' refers to EP-learner, as introduced in Section III.B.
  3. [References] Several references appear unrelated to the content: [12] (facial expression classification), [13] (large language models), [15] (volatility forecasting), [16] (text-to-image biases), and [20] (vision-language models). These citations are not relevant to treatment effect estimation and should be removed or replaced with correct citations.
  4. [Section IV.A] The statement 'All records are flagged as switched' is contradicted later in the same paragraph by 'patients continuing duloxetine 60 mg QD.' Clarify whether the dataset contains only switching events or a mix of switches and continuations, as this affects the interpretation of the treatment effects.

Circularity Check

0 steps flagged

No significant circularity: the benchmark is self-contained with held-out evaluation, and no prediction reduces to its inputs by construction.

full rationale

The paper's derivation chain is not circular. Hyperparameters are selected on validation RMSE, and all reported performance metrics (RMSE, MAE, ECE, calibration slope, IPW policy value) are computed on held-out test patients, so factual predictions are genuine out-of-sample evaluations. The ATE calibration proxy E_ATE compares model dATE estimates to the observational group difference Δ_OCP, but the paper explicitly states that Δ_OCP 'reflects selection bias and is not a causal estimand, but provides a relative consistency signal'; models are not fitted to E_ATE, so this is an internal consistency check rather than a reduction. The potential-outcome reconstruction Y(t) = m_b(x) + τ_{t,b}(x) is an additive modeling assumption, not a tautology, and the pairwise CF contrasts are estimated from training data rather than defined as the observed test differences. The paper also candidly acknowledges that unconfoundedness (A2) is 'untestable and may be violated' and defers sensitivity analysis to future work; this is a validity limitation, not circularity. The self-citations [4,5] appear only as motivational context for counterfactual prediction and are not load-bearing for any result. No step in the paper equates a prediction to its input by construction, and no fitted parameter is renamed as a prediction. Therefore the appropriate circularity score is 0.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 0 invented entities

The causal claims rest on untestable identification assumptions and an internal, confounded calibration target. Hyperparameters and clipping thresholds are data-dependent modeling choices; the paper introduces no new entities.

free parameters (3)
  • CF hyperparameters (n_trees, max_depth, honest_fraction) = 200, 10, 0.5
    Selected by grid search on validation RMSE (Section IV.B); the reported ranking of CF depends on this configuration.
  • Propensity clipping thresholds = [0.01, 0.99]
    Chosen by hand for IPW policy evaluation (Section II.D); robustness check shows stable policy value, but the specific clipping affects the metric.
  • Baseline outcome model hyperparameters (XGBoost: n_estimators, learning_rate, max_depth) = 100, 0.1, 6
    Used for m_b(x) in the potential-outcome reconstruction (Section III.A); ablation shows choice matters (RMSE 4.35 vs 4.01).
axioms (6)
  • domain assumption Consistency, unconfoundedness, and positivity (A1-A3) for causal identification from observational data
    Stated in Section II.C. Unconfoundedness is untestable; in this switched dataset with rescue-therapy prescribing, unmeasured reasons for switching likely violate it.
  • domain assumption Next-visit outcome formulation: HAMD-17 at visit v+1 depends only on covariates and treatment at visit v
    The entire estimand is defined this way (Section II.A); it ignores intermediate events, adherence, and concurrent treatments between visits.
  • ad hoc to paper Baseline-referenced additive reconstruction of potential outcomes: Ŷ(t)=m_b(x)+τ_t,b(x)
    Section III.A: pairwise forests estimate contrasts against the most frequent treatment; additivity and validity of pairwise CATEs are assumed, not tested.
  • ad hoc to paper Observational group mean difference Δ_OCP is a useful calibration reference for ATE estimates
    Section II.D defines E_ATE against Δ_OCP while also stating Δ_OCP reflects selection bias and is not causal; this is internally inconsistent and loads the evaluation toward confounded comparisons.
  • domain assumption Complete-case analysis after dropping missing values does not introduce selection bias
    Section IV.A removes records with missing values; no missingness model or sensitivity check is provided.
  • domain assumption All analysis records are switching events; findings apply to switching decisions in this trial program
    Section IV.A states all records are flagged as switched; no never-switch comparison group exists, so estimates are not generalizable to initial treatment choice.

pith-pipeline@v1.3.0-alltime-deepseek · 13095 in / 19689 out tokens · 185662 ms · 2026-08-02T11:03:45.245992+00:00 · methodology

0 comments
read the original abstract

Depression treatment often requires switching medications due to inadequate response or adverse effects. Estimating individualized treatment effects in this setting is challenging because treatment assignment is confounded by patient characteristics, switching induces time-varying selection, and counterfactual outcomes are not observed in follow-up data. Using a proprietary longitudinal major depressive disorder (MDD) clinical trial dataset, we formulate a next-visit counterfactual prediction task to estimate Hamilton Depression Rating Scale (HAMD-17) total scores under alternative treatments. We benchmark 8 estimators, including meta-learners, residual-based methods, and tree-based approaches. Causal Forest (CF) demonstrates the most favorable and consistent performance across all criteria. Our analysis shows that symptom benefits concentrate in specific switch directions, with dose intensification being generally beneficial. Notably, we identify a counterintuitive exception where a lower-intensity regimen outperforms a higher-intensity alternative for specific patient subsets. While crude observational comparisons substantially overstate gains, confounding-adjusted estimates yield modest, actionable magnitudes. These findings provide prospectively testable candidates for clinical decision support in depression care.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

43 extracted references · 2 linked inside Pith

  1. [1]

    Consider- ations when selecting an antidepressant: a narrative review for primary care providers treating adults with depression,

    C. B. Montano, W. C. Jackson, D. Vanacore, and R. Weisler, “Consider- ations when selecting an antidepressant: a narrative review for primary care providers treating adults with depression,”Postgraduate Medicine, vol. 135, no. 5, pp. 449–465, 2023

  2. [2]

    The epidemiology of major depressive disorder: results from the national comorbidity survey replication (ncs-r),

    R. C. Kessler, P. Berglund, O. Demler, R. Jin, D. Koretz, K. R. Merikan- gas, A. J. Rush, E. E. Walters, and P. S. Wang, “The epidemiology of major depressive disorder: results from the national comorbidity survey replication (ncs-r),”jama, vol. 289, no. 23, pp. 3095–3105, 2003

  3. [3]

    Estimating causal effects of treatments in randomized and nonrandomized studies,

    D. B. Rubin, “Estimating causal effects of treatments in randomized and nonrandomized studies,”Journal of Educational Psychology, vol. 66, no. 5, pp. 688–701, 1974

  4. [4]

    Enhancing counterfactual ex- planations with feasibility and diversity,

    X. Qin, S. Li, Y . Cai, and L. Wang, “Enhancing counterfactual ex- planations with feasibility and diversity,” in2025 IEEE International Conference on Data Mining Workshops (ICDMW). IEEE, 2025, pp. 2310–2319

  5. [5]

    Explainable counterfactual rea- soning in depression medication selection at multi-levels (personalized and population),

    X. Qin, M. H. Chignell, A. Greifenberger, S. Lokuge, E. Toumeh, T. Sternat, M. Katzman, and L. Wang, “Explainable counterfactual rea- soning in depression medication selection at multi-levels (personalized and population),”BMC Medical Informatics and Decision Making, 2026

  6. [6]

    Confounding by indication,

    A. M. Walker, “Confounding by indication,”Epidemiology, vol. 7, no. 4, pp. 335–336, 1996

  7. [7]

    Dealing with limited overlap in estimation of average treatment effects,

    R. K. Crump, V . J. Hotz, G. W. Imbens, and O. A. Mitnik, “Dealing with limited overlap in estimation of average treatment effects,”Biometrika, vol. 96, no. 1, pp. 187–199, 2009

  8. [8]

    Acute and longer-term outcomes in depressed outpatients requiring one or several treatment steps: a star* d report,

    A. J. Rush, M. H. Trivedi, S. R. Wisniewski, A. A. Nierenberg, J. W. Stewart, D. Warden, G. Niederehe, M. E. Thase, P. W. Lavori, B. D. Lebowitzet al., “Acute and longer-term outcomes in depressed outpatients requiring one or several treatment steps: a star* d report,” American Journal of Psychiatry, vol. 163, no. 11, pp. 1905–1917, 2006

  9. [9]

    Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: a systematic review and network meta-analysis,

    A. Cipriani, T. A. Furukawa, G. Salanti, A. Chaimani, L. Z. Atkinson, Y . Ogawa, S. Leucht, H. G. Ruhe, E. H. Turner, J. P. Higginset al., “Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: a systematic review and network meta-analysis,”The Lancet, vol. 391, no. 10128, pp. 1...

  10. [10]

    Efficient estimation of average treatment effects using the estimated propensity score,

    K. Hirano, G. W. Imbens, and G. Ridder, “Efficient estimation of average treatment effects using the estimated propensity score,”Econometrica, vol. 71, no. 4, pp. 1161–1189, 2003

  11. [11]

    Metalearners for estimating heterogeneous treatment effects using machine learning,

    S. R. K ¨unzel, J. S. Sekhon, P. J. Bickel, and B. Yu, “Metalearners for estimating heterogeneous treatment effects using machine learning,” Proceedings of the National Academy of Sciences, vol. 116, no. 10, pp. 4156–4165, 2019

  12. [12]

    Hy-facial: Hybrid feature extraction by dimensionality reduction methods for enhanced facial expression classification,

    X. Li, Y . Ma, K. Ye, J. Cao, M. Zhou, and Y . Zhou, “Hy-facial: Hybrid feature extraction by dimensionality reduction methods for enhanced facial expression classification,” inEighteenth International Conference on Machine Vision (ICMV 2025), vol. 14114. SPIE, 2026, pp. 206–213

  13. [13]

    Synergized data efficiency and compression (sec) optimization for large language models,

    X. Li, Y . Ma, Y . Huang, X. Wang, Y . Lin, and C. Zhang, “Synergized data efficiency and compression (sec) optimization for large language models,” in2024 4th International Conference on Electronic Information Engineering and Computer Science (EIECS). IEEE, 2024, pp. 586–591

  14. [14]

    Doubly robust estimation in missing data and causal inference models,

    H. Bang and J. M. Robins, “Doubly robust estimation in missing data and causal inference models,”Biometrics, vol. 61, no. 4, pp. 962–973, 2005

  15. [15]

    V olatility persistence and model choice in cross-market volatility forecasting,

    K. Cheng, X. Qi, Z. Cheng, and L. Lai, “V olatility persistence and model choice in cross-market volatility forecasting,”Available at SSRN 6610278, 2026

  16. [16]

    Bigbench: A unified benchmark for evaluating multi- dimensional social biases in text-to-image models,

    H. Luo, H. Huang, Z. Deng, X. Li, H. Wang, Y . Jin, Y . Liu, W. Xu, and Z. Liu, “Bigbench: A unified benchmark for evaluating multi- dimensional social biases in text-to-image models,”arXiv preprint arXiv:2407.15240, 2024

  17. [17]

    Estimation and inference of heterogeneous treatment effects using random forests,

    S. Wager and S. Athey, “Estimation and inference of heterogeneous treatment effects using random forests,”Journal of the American Statis- tical Association, vol. 113, no. 523, pp. 1228–1242, 2018

  18. [18]

    Generalized random forests,

    S. Athey, J. Tibshirani, and S. Wager, “Generalized random forests,”The Annals of Statistics, vol. 47, no. 2, pp. 1148–1178, 2019

  19. [19]

    Quasi-oracle estimation of heterogeneous treat- ment effects,

    X. Nie and S. Wager, “Quasi-oracle estimation of heterogeneous treat- ment effects,”Biometrika, vol. 108, no. 2, pp. 299–319, 2021

  20. [20]

    Autoneural: Co-designing vision-language models for npu inference,

    W. Chen, L. Wu, Y . Hu, Z. Li, Z. Cheng, Y . Qian, L. Zhu, Z. Hu, L. Liang, Q. Tanget al., “Autoneural: Co-designing vision-language models for npu inference,”arXiv preprint arXiv:2512.02924, 2025

  21. [21]

    Double/debiased machine learning for treatment and structural parameters,

    V . Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins, “Double/debiased machine learning for treatment and structural parameters,”The Econometrics Journal, vol. 21, no. 1, pp. C1–C68, 2018

  22. [22]

    Estimation of causal effects with multiple treatments: a review and new ideas,

    M. J. Lopez and R. Gutman, “Estimation of causal effects with multiple treatments: a review and new ideas,”Statistical Science, vol. 32, no. 3, pp. 432–454, 2017

  23. [23]

    Some methods for heterogeneous treatment effect estimation in high dimensions,

    S. Powers, J. Qian, K. Jung, A. Schuler, N. H. Shah, T. Hastie, and R. Tibshirani, “Some methods for heterogeneous treatment effect estimation in high dimensions,”Statistics in Medicine, vol. 37, no. 11, pp. 1767–1787, 2018

  24. [24]

    Estimating individual treatment effect: generalization bounds and algorithms,

    U. Shalit, F. D. Johansson, and D. Sontag, “Estimating individual treatment effect: generalization bounds and algorithms,” inProceedings of the 34th International Conference on Machine Learning, vol. 70. PMLR, 2017, pp. 3076–3085

  25. [25]

    A rating scale for depression,

    M. Hamilton, “A rating scale for depression,”Journal of neurology, neurosurgery, and psychiatry, vol. 23, no. 1, p. 56, 1960

  26. [26]

    Doubly robust policy evaluation and optimization,

    M. Dud ´ık, D. Erhan, J. Langford, and L. Li, “Doubly robust policy evaluation and optimization,”Statistical Science, vol. 29, no. 4, pp. 485– 511, 2014

  27. [27]

    Combining t-learning and dr-learning: A framework for oracle-efficient estimation of causal contrasts,

    L. van der Laan, M. Carone, and A. Luedtke, “Combining t-learning and dr-learning: A framework for oracle-efficient estimation of causal contrasts,”arXiv preprint arXiv:2402.01972, 2024

  28. [28]

    Kg-treat: pre-training for treatment effect estimation by synergizing patient data with knowledge graphs,

    R. Liu, L. Wu, and P. Zhang, “Kg-treat: pre-training for treatment effect estimation by synergizing patient data with knowledge graphs,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 8, 2024, pp. 8805–8814

  29. [29]

    G. W. Imbens and D. B. Rubin,Causal inference in statistics, social, and biomedical sciences. Cambridge university press, 2015

  30. [30]

    EconML: A Python Package for ML-Based Het- erogeneous Treatment Effects Estimation,

    K. Battocchi, E. Dillon, M. Hei, G. Lewis, P. Oka, M. Oprescu, and V . Syrgkanis, “EconML: A Python Package for ML-Based Het- erogeneous Treatment Effects Estimation,” https://github.com/py-why/ EconML, 2019, version 0.x

  31. [31]

    Multi-target strategies for the improved treatment of depressive states: conceptual foundations and neuronal substrates, drug discovery and therapeutic application,

    M. J. Millan, “Multi-target strategies for the improved treatment of depressive states: conceptual foundations and neuronal substrates, drug discovery and therapeutic application,”Pharmacology & therapeutics, vol. 110, no. 2, pp. 135–370, 2006

  32. [32]

    Monoamine neurocircuitry in depression and strategies for new treatments,

    M. Hamon and P. Blier, “Monoamine neurocircuitry in depression and strategies for new treatments,”Progress in Neuro-Psychopharmacology and Biological Psychiatry, vol. 45, pp. 54–63, Aug 2013

  33. [33]

    Snris: the pharmacology, clinical efficacy, and tolerability in comparison with other classes of antidepressants,

    S. M. Stahl, M. M. Grady, C. Moret, and M. Briley, “Snris: the pharmacology, clinical efficacy, and tolerability in comparison with other classes of antidepressants,”CNS spectrums, vol. 10, no. 9, pp. 732–747, 2005

  34. [34]

    Which factors influence psychiatrists’ selection of antidepressants?

    M. Zimmerman, M. Posternak, M. Friedman, N. Attiullah, S. Baymiller, R. Boland, S. Berlowitz, S. Rahman, K. Uy, and S. Singer, “Which factors influence psychiatrists’ selection of antidepressants?”American Journal of Psychiatry, vol. 161, no. 7, pp. 1285–1289, 2004

  35. [35]

    Effects of duloxetine, an antidepressant drug candidate, on concentrations of monoamines and their metabolites in rats and mice

    R. W. Fuller, S. K. Hemrick-Luecke, and H. D. Snoddy, “Effects of duloxetine, an antidepressant drug candidate, on concentrations of monoamines and their metabolites in rats and mice.”The Journal of pharmacology and experimental therapeutics, vol. 269, no. 1, pp. 132– 136, 1994

  36. [36]

    Duloxetine: clinical pharmacokinetics and drug interactions,

    M. P. Knadler, E. Lobo, J. Chappell, and R. Bergstrom, “Duloxetine: clinical pharmacokinetics and drug interactions,”Clinical pharmacoki- netics, vol. 50, no. 5, pp. 281–294, 2011

  37. [37]

    A systematic review of efficacy, safety, and tolerability of duloxetine,

    D. Rodrigues-Amorim, J. M. Olivares, C. Spuch, and T. Rivera-Baltanas, “A systematic review of efficacy, safety, and tolerability of duloxetine,” Frontiers in Psychiatry, vol. 11, p. 554899, 2020

  38. [38]

    Paroxetine: a review,

    M. Bourin, P. Chue, and Y . Guillon, “Paroxetine: a review,”CNS drug reviews, vol. 7, no. 1, pp. 25–47, 2001

  39. [39]

    Selective serotonin reuptake inhibitor (ssri) drugs: More risks than benefits,

    J. M. Kauffman, “Selective serotonin reuptake inhibitor (ssri) drugs: More risks than benefits,”Journal of American Physicians and Surgeons, vol. 14, no. 1, pp. 7–12, 2009

  40. [40]

    Depression subtypes in predicting antidepressant response: a report from the ispot-d trial,

    B. A. Arnow, C. Blasey, L. M. Williams, D. M. Palmer, W. Rekshan, A. F. Schatzberg, A. Etkin, J. Kulkarni, J. F. Luther, and A. J. Rush, “Depression subtypes in predicting antidepressant response: a report from the ispot-d trial,”American Journal of Psychiatry, vol. 172, no. 8, pp. 743–750, 2015

  41. [41]

    Major depressive subtypes and treatment response,

    M. Fava, L. A. Uebelacker, J. E. Alpert, A. A. Nierenberg, J. A. Pava, and J. F. Rosenbaum, “Major depressive subtypes and treatment response,”Biological psychiatry, vol. 42, no. 7, pp. 568–576, 1997

  42. [42]

    Severe and anxious depression: combining definitions of clinical sub-types to identify pa- tients differentially responsive to selective serotonin reuptake inhibitors,

    G. I. Papakostas, H. Fan, and E. Tedeschini, “Severe and anxious depression: combining definitions of clinical sub-types to identify pa- tients differentially responsive to selective serotonin reuptake inhibitors,” European Neuropsychopharmacology, vol. 22, no. 5, pp. 347–355, 2012

  43. [43]

    Making sense of sensitivity: Extending omitted variable bias,

    C. Cinelli and C. Hazlett, “Making sense of sensitivity: Extending omitted variable bias,”Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 82, no. 1, pp. 39–67, 2020