REVIEW 3 major objections 5 minor 30 references
Revolutionizing Clinical Trials: A Manifesto for AI-Driven Transformation
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This manifesto argues that AI-driven digital twins and causal inference can transform clinical trials, potentially making replication of pivotal trials obsolete.
desk verdict A polished industry manifesto on AI for clinical trials: no new methods, but a useful validation checklist; the pivotal-trial-replacement claim overreaches its cited evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's argument is carried by two technologies plus a validation procedure. A digital twin is a computational model of an individual patient or population that simulates biological and therapeutic responses over time; the paper's version is AI-driven, built from electronic health records, genomics, wearables, and longitudinal data, and can answer 'what-if' counterfactual questions. Causal inference, especially heterogeneous treatment effect (HTE) and conditional average treatment effect (CATE) estimation, supplies the counterfactual quantities needed to predict trial outcomes in silico, adjust for confounding via inverse probability weighting, targeted maximum likelihood estimation, and double/debiased machine learning, and transport findings to real-world populations. The validation framework—define context of use, data integrity, multifaceted model assessment, interpretability, benchmarking, real-world validation, ethics, feedback loops, and reporting—is the proposed safeguard that makes regulatory acceptance thinkable.
What would settle it
Take a completed placebo-controlled trial for a condition with a known unmeasured confounder (such as frailty or socioeconomic status); train a digital twin on observational records from the same population using only measured covariates; then compare the twin's simulated control-arm event rates and treatment-effect estimates with the actual trial results. Systematic divergence in this head-to-head would falsify the claim that twins can substitute for pivotal replication.
Extended reading notes
Core claim
The central claim is a programmatic one: digital twins that simulate individual patient trajectories and ML-based causal inference that estimates heterogeneous treatment effects are mature enough to be woven into every phase of clinical trials, and if validated properly, could let a single pivotal trial plus a digital-twin simulation replace the traditional requirement for two replicated pivotal trials. The paper explicitly asserts that digital twins can recreate clinical trial results from observational data alone, citing work on longitudinal counterfactual treatment-effect estimation, and that simulation of whole trials could make replication of pivotal trials obsolete. This is presented as a roadmap rather than a completed proof, with the validation framework as the condition for regulatory acceptance.
Load-bearing premise
The plan assumes that historical observational data contain all the information needed to predict what would happen to a patient under a treatment they did not receive—that is, no unmeasured confounders distort the counterfactual answers and the training population matches the trial population.
Editorial extensions
If this is right
- If accepted by regulators, a single pivotal trial could be paired with a digital-twin replication, cutting development costs and months of calendar time.
- Virtual control arms could free a larger fraction of enrolled patients to receive active treatment, improving recruitment, retention, and the ethics of placebo use.
- Causal models could identify responding subgroups early, reducing the number of failed Phase II and Phase III trials and focusing late-phase trials on the right patients and doses.
- Trial findings could be transported to older, sicker, and more diverse real-world populations using rich observational data, strengthening generalizability and equity.
Reading between the lines
- The same logic that replaces a second physical trial could also reshape post-approval surveillance: a continuously updated twin would provide a living estimate of benefit-risk as real-world data accumulate, not just a static approval snapshot.
- A decisive practical test the manifesto leaves implicit is a head-to-head, pre-registered comparison of a twin-based virtual control arm against a simultaneously enrolled real control arm in an active trial, before any regulatory decision leans on the twin.
- If twins inherit the biases of the observational data they are trained on, the approach could silently amplify inequities in who gets included in trials; the paper's diversity claims will hold only if training data themselves are population-representative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a multi-stakeholder manifesto, authored by representatives from academia, the pharmaceutical industry, consulting firms, and AI research, arguing that two AI technologies—causal inference and digital twins—can transform clinical trials. It outlines a set of goals (acceleration, improved probability of success, expanded questions, balanced confirmation and discovery), describes use cases for digital twins (safety monitoring, synthetic control arms, post-trial personalization) and causal inference (biomarker discovery, CATE estimation, transportability, discovery-driven modeling), and proposes a validation framework in Section 5. The central, strongest claim, made in Section 2 and elaborated in Appendix 1, is that digital twins capable of simulating entire clinical trials could make replication of pivotal trials obsolete. The manuscript is explicitly a roadmap rather than a report of new experimental results.
Significance. If the central claim were established, the impact on clinical development, regulation, and patient access would be substantial. The manuscript provides a useful high-level checklist of potential applications and a structured, if generic, validation framework. Its strengths include explicit acknowledgment of the need for validation and real-world testing, and the involvement of multiple stakeholders lends practical relevance. However, the paper contains no new analysis, no working system, and no quantitative evidence; its most consequential assertion—that digital twins can replace pivotal-trial replication—rests on a single citation to an effect-estimation method (Qian et al., 2021) rather than on demonstrated whole-trial simulation. The value of the paper lies in framing a research agenda, not in establishing a capability. The self-referential limitations in Appendix 1 are candid, but they are not operationalized into testable conditions.
major comments (3)
- [Section 3 and Appendix 1] The claim that digital twins 'can recreate clinical trial results from observational data alone' cites Qian et al. (2021), SyncTwin, which estimates conditional average treatment effects from longitudinal observational data. That method does not simulate an entire trial, does not predict a pre-specified primary endpoint on an independent population, and does not substitute for a second pivotal study. The subsequent assertion in Appendix 1 that such twins 'provide robust evidence that mirrors what a second physical trial would achieve' is therefore unsupported by the cited evidence. The authors should either provide a citation to a demonstrated whole-trial simulation that recreates a completed pivotal trial's primary outcome, or explicitly label this capability as a speculative aspiration rather than an established result.
- [Section 4] The prescriptions for confounding adjustment, including inverse probability weighting and targeted maximum likelihood estimation, are presented without stating the required identifiability assumptions. The discussion of combining RCT data with EHR data for subgroup analysis and transportability does not mention the conditions under which such combination yields valid causal estimates: no unmeasured confounding, positivity, consistency, no interference, and transportability of effects from the trial population to the target population. These assumptions are load-bearing; if they fail, the conclusions drawn from the proposed causal inference pipeline are not justified. The manuscript should state these conditions and discuss how they would be verified in the proposed applications.
- [Appendix 1 and Section 5] The central proposal—that digital twins could make pivotal-trial replication obsolete—is not accompanied by an operational validation criterion. Appendix 1 concedes that 'full regulatory acceptance ... will require rigorous validation', and Section 5 provides a generic list of validation components, but no specific acceptance thresholds are proposed (e.g., what level of agreement between simulated and observed trial outcomes would be considered sufficient, or how the twin should be calibrated to a target population). Without such criteria, the central claim is unfalsifiable. The authors should at minimum outline the design of a prospective validation study, such as a historical-benchmark exercise where a twin must match the result of a completed pivotal trial before it can be used in lieu of a second trial.
minor comments (5)
- [Introduction] The sentence fragment 'Even though treatments are often much more widely used [Martin et al., 2004]' is grammatically incomplete and should be merged with the preceding or following sentence.
- [Section 2] The text contains an unresolved placeholder '(add REF)' in the subsection on improving probability of success; this should be replaced with an actual citation or removed.
- [Section 2] The phrase '¡ref¿' appears in the paragraph on balancing confirmation and discovery; this formatting artifact should be cleaned up.
- [Section 5, Table 1] The table repeats the bullet-point list from the text without adding detail; consider either expanding it with concrete example metrics for each item or removing it to avoid redundancy.
- [References] The reference to the FDA discussion paper is formatted inconsistently (authors are listed as 'Department of Health, US Food Human Services, and Drug Administration'); this should be corrected.
Circularity Check
No significant circularity: the paper is a position manifesto with no derivation chain; its self-citations are external empirical works and the central claims are explicitly framed as potential, not as derived predictions.
full rationale
This manuscript is a manifesto, not a derivation. It contains no equations, no fitted parameters, and no uniqueness theorem, so the construction-based circularity patterns (self-definitional, fitted-input-as-prediction, imported uniqueness, ansatz-by-citation) do not apply. The strongest claim—that digital twins could make replication of pivotal trials obsolete (§2; Appendix 1)—is phrased as a potential ('could potentially make replications of pivotal trials obsolete') and is not derived from the cited methods. Section 3 asserts that digital twins 'can recreate clinical trial results from observational data alone' and cites Qian et al. (2021) (SyncTwin). That citation is co-authored by one of the present authors (van der Schaar), so a self-citation is load-bearing for the asserted capability. However, SyncTwin is an externally published, peer-reviewed method paper that estimates treatment effects from longitudinal observational data; the mismatch between what it actually demonstrates and the broader 'recreate clinical trial results' claim is an evidentiary overreach, not a circular reduction. The paper's own Appendix 1 concedes that 'full regulatory acceptance of digital twins as substitutes for physical trials will require rigorous validation,' and Section 5 provides an external validation framework (benchmarking against traditional methods, real-world testing, feedback loops) rather than assuming the conclusion. The only textual defect relevant to completeness is an '(add REF)' placeholder in §2, which is a missing citation, not a circular step. No step in the argument is equivalent to its input by construction, and no result is forced by a self-citation chain. The appropriate finding is therefore no significant circularity, with the SyncTwin scope gap flagged as a correctness/evidence concern rather than a circularity score driver.
Assumptions & free parameters
assumptions (3)
- domain assumption Observational data are sufficiently rich to estimate causal effects without unmeasured confounding and with adequate overlap.
- domain assumption Digital twins can accurately simulate individual patient trajectories and counterfactual outcomes.
- ad hoc to paper Regulators will eventually accept digital twin simulations as evidence in place of physical trials.
Cite this review
Pith. "Pith review of Revolutionizing Clinical Trials: A Manifesto for AI-Driven Transformation." pith.science (2026). https://pith.science/paper/3F3DYLTN
@misc{pith2026250609102,
author = {Pith},
title = {Pith review of: Revolutionizing Clinical Trials: A Manifesto for AI-Driven Transformation},
year = {2026},
howpublished = {\url{https://pith.science/paper/3F3DYLTN}},
note = {Machine review of arXiv:2506.09102}
}
read the original abstract
This manifesto represents a collaborative vision forged by leaders in pharmaceuticals, consulting firms, clinical research, and AI. It outlines a roadmap for two AI technologies - causal inference and digital twins - to transform clinical trials, delivering faster, safer, and more personalized outcomes for patients. By focusing on actionable integration within existing regulatory frameworks, we propose a way forward to revolutionize clinical research and redefine the gold standard for clinical trials using AI.
Reference graph
Works this paper leans on
-
[1]
Hermenegild J Arevalo, Fijoy Vadakkumpadan, Eliseo Guallar, Alexander Jebb, Peter Malamas, Katherine C Wu, and Natalia A Trayanova. Arrhythmia risk stratification of patients after myocardial infarction using personalized heart models. Nature communications, 7 0 (1): 0 11437, 2016
work page 2016
-
[2]
Estimating treatment effects with causal forests: An application
Susan Athey and Stefan Wager. Estimating treatment effects with causal forests: An application. Observational studies, 5 0 (2): 0 37--51, 2019
work page 2019
-
[3]
Causal inference and the data-fusion problem
Elias Bareinboim and Judea Pearl. Causal inference and the data-fusion problem. Proceedings of the National Academy of Sciences, 113 0 (27): 0 7345--7352, 2016
2016
-
[4]
From real-world patient data to individualized treatment effects using machine learning: current and future methods to address underlying challenges
Ioana Bica, Ahmed M Alaa, Craig Lambert, and Mihaela Van Der Schaar. From real-world patient data to individualized treatment effects using machine learning: current and future methods to address underlying challenges. Clinical Pharmacology & Therapeutics, 109 0 (1): 0 87--100, 2021
2021
-
[5]
Automated reverse engineering of nonlinear dynamical systems
Josh Bongard and Hod Lipson. Automated reverse engineering of nonlinear dynamical systems. Proceedings of the National Academy of Sciences, 104 0 (24): 0 9943--9948, 2007
2007
-
[6]
Peter Bower, Valerie Brueton, Carrol Gamble, Shaun Treweek, Catrin Tudur Smith, Bridget Young, and Paula Williamson. Interventions to improve recruitment and retention in clinical trials: a survey and workshop to assess current practice and future priorities. Trials, 15: 0 1--9, 2014
work page 2014
-
[7]
Discovering governing equations from data by sparse identification of nonlinear dynamical systems
Steven L Brunton, Joshua L Proctor, and J Nathan Kutz. Discovering governing equations from data by sparse identification of nonlinear dynamical systems. Proceedings of the national academy of sciences, 113 0 (15): 0 3932--3937, 2016
2016
-
[8]
Double/debiased machine learning for treatment and causal parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. Double/debiased machine learning for treatment and causal parameters. arXiv preprint arXiv:1608.00060, 2016
arXiv 2016
Show all 30 references
-
[9]
Cooper, Thomas F
Jay S. Cooper, Thomas F. Pajak, Arlene A. Forastiere, John Jacobs, Bruce H. Campbell, Scott B. Saxman, Julie A. Kish, Harold E. Kim, Anthony J. Cmelak, Marvin Rotman, Mitchell Machtay, John F. Ensley, K.S. Clifford Chao, Christopher J. Schultz, Nancy Lee, and Karen K. Fu. Post...
1937 doi
-
[10]
Exclusion rates in randomized controlled trials of treatments for physical conditions: a systematic review
Jinzhang He, Daniel R Morales, and Bruce Guthrie. Exclusion rates in randomized controlled trials of treatments for physical conditions: a systematic review. Trials, 21: 0 1--11, 2020
2020
-
[11]
Automatically learning hybrid digital twins of dynamical systems
Samuel Holt, Tennison Liu, and Mihaela van der Schaar. Automatically learning hybrid digital twins of dynamical systems. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=SOsiObSdU2
2024
-
[12]
ODE discovery for longitudinal heterogeneous treatment effects inference
Krzysztof Kacprzyk, Samuel Holt, Jeroen Berrevoets, Zhaozhi Qian, and Mihaela van der Schaar. ODE discovery for longitudinal heterogeneous treatment effects inference. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?i...
2024
-
[13]
Challenges in recruitment and retention of clinical trial subjects
Rashmi Ashish Kadam, Sanghratna Umakant Borde, Sapna Amol Madas, Sundeep Santosh Salvi, and Sneha Saurabh Limaye. Challenges in recruitment and retention of clinical trial subjects. Perspectives in clinical research, 7 0 (3): 0 137--143, 2016
2016
-
[14]
Med-real2sim: Non-invasive medical digital twins using physics-informed self-supervised learning
Keying Kuang, Frances Dean, Jack B Jedlicki, David Ouyang, Anthony Philippakis, David Sontag, and Ahmed M Alaa. Med-real2sim: Non-invasive medical digital twins using physics-informed self-supervised learning. Advances in Neural Information Processing Systems, 37: 0 5757--5788, 2024
2024
-
[15]
Using digital twins in viral infection
Reinhard Laubenbacher, James P Sluka, and James A Glazier. Using digital twins in viral infection. Science, 371 0 (6534): 0 1105--1106, 2021
2021
-
[16]
Neural-ode for pharmacokinetics modeling and its advantage to alternative machine learning models in predicting new dosing regimens
James Lu, Kaiwen Deng, Xinyuan Zhang, Gengbo Liu, and Yuanfang Guan. Neural-ode for pharmacokinetics modeling and its advantage to alternative machine learning models in predicting new dosing regimens. Iscience, 24 0 (7), 2021
2021
-
[17]
Differences between clinical trials and postmarketing use
Karin Martin, Bernard B \'e gaud, Philippe Latry, Ghada Miremont-Salam \'e , Annie Fourrier, and Nicholas Moore. Differences between clinical trials and postmarketing use. British journal of clinical pharmacology, 57 0 (1): 0 86--92, 2004
2004
-
[18]
Estimated costs of pivotal trials for novel therapeutic agents approved by the us food and drug administration, 2015-2016
Thomas J Moore, Hanzhe Zhang, Gerard Anderson, and G Caleb Alexander. Estimated costs of pivotal trials for novel therapeutic agents approved by the us food and drug administration, 2015-2016. JAMA internal medicine, 178 0 (11): 0 1451--1457, 2018
2015
-
[19]
Using artificial intelligence & machine learning in the development of drug & biological products: Discussion paper and request for feedback, 2023
Department of Health, US Food Human Services, and Drug Administration. Using artificial intelligence & machine learning in the development of drug & biological products: Discussion paper and request for feedback, 2023. URL https://www.federalregister.gov/d/2023-09985
2023
-
[20]
In silico clinical trials: concepts and early adoptions
Francesco Pappalardo, Giulia Russo, Flora Musuamba Tshinanu, and Marco Viceconti. In silico clinical trials: concepts and early adoptions. Briefings in bioinformatics, 20 0 (5): 0 1699--1708, 2019
2019
-
[21]
Economic evaluation of cost and time required for a platform trial vs conventional trials
Jay JH Park, Behnam Sharif, Ofir Harari, Louis Dron, Anna Heath, Maureen Meade, Ryan Zarychanski, Raymond Lee, Gabriel Tremblay, Edward J Mills, et al. Economic evaluation of cost and time required for a platform trial vs conventional trials. JAMA network open, 5 0 (7): 0 e222...
2022
-
[22]
Meta-analysis of chemotherapy in head and neck cancer (mach-nc): An update on 93 randomised trials and 17,346 patients
Jean-Pierre Pignon, Aurélie le Maître, Emilie Maillard, and Jean Bourhis. Meta-analysis of chemotherapy in head and neck cancer (mach-nc): An update on 93 randomised trials and 17,346 patients. Radiotherapy and Oncology, 92 0 (1): 0 4--14, 2009. ISSN 0167-8140. doi:https://doi...
2009 doi
-
[23]
Synctwin: Treatment effect estimation with longitudinal outcomes
Zhaozhi Qian, Yao Zhang, Ioana Bica, Angela Wood, and Mihaela van der Schaar. Synctwin: Treatment effect estimation with longitudinal outcomes. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Sys...
2021
-
[24]
Ai education for clinicians
Tim Schubert, Tim Oosterlinck, Robert D Stevens, Patrick H Maxwell, and Mihaela van der Schaar. Ai education for clinicians. EClinicalMedicine, 79, 2025
2025
-
[25]
Meta-learners for partially-identified treatment effects across multiple environments
Jonas Schweisthal, Dennis Frauen, Mihaela Van Der Schaar, and Stefan Feuerriegel. Meta-learners for partially-identified treatment effects across multiple environments. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and ...
2024
-
[26]
Costs of drug development and research and development intensity in the us, 2000-2018
Aylin Sertkaya, Trinidad Beleche, Amber Jessup, and Benjamin D Sommers. Costs of drug development and research and development intensity in the us, 2000-2018. JAMA network open, 7 0 (6): 0 e2415445--e2415445, 2024
2000
-
[27]
Adapting neural networks for the estimation of treatment effects
Claudia Shi, David Blei, and Victor Veitch. Adapting neural networks for the estimation of treatment effects. Advances in neural information processing systems, 32, 2019
2019
-
[28]
Optimal treatment selection in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: deep q-learning with a patient-physician digital twin dyad
Elisa Tardini, Xinhua Zhang, Guadalupe Canahuate, Andrew Wentzel, Abdallah SR Mohamed, Lisanne Van Dijk, Clifton D Fuller, and G Elisabeta Marai. Optimal treatment selection in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: deep q-learning w...
2022 doi
-
[29]
Targeted learning: causal inference for observational and experimental data, volume 4
Mark J Van der Laan, Sherri Rose, et al. Targeted learning: causal inference for observational and experimental data, volume 4. Springer, 2011
2011
-
[30]
Gpu accelerated digital twins of the human heart open new routes for cardiovascular research
Francesco Viola, Giulio Del Corso, Ruggero De Paulis, and Roberto Verzicco. Gpu accelerated digital twins of the human heart open new routes for cardiovascular research. Scientific reports, 13 0 (1): 0 8230, 2023
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.