REVIEW 3 major objections 6 minor 21 references
Estimands for Randomized Discontinuation Designs in Oncology
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Randomized discontinuation trials in oncology can be given a complete ICH E9(R1) estimand specification, with the target population defined by the open-label response criterion and intercurrent-event handling anchored at randomization.
desk verdict Useful first mapping of E9(R1) estimand attributes onto oncology RDDs, but the population pre-specification problem is left unresolved; still deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the five-attribute estimand specification of ICH E9(R1) — population, treatment, endpoint, population-level summary, and intercurrent events — applied to RDD as a structured mapping table. The load-bearing move is anchoring the estimand at the moment of randomization in the double-blind phase; this separates pre-randomization enrichment events, which define the population, from post-randomization intercurrent events, which are handled as in a standard RCT. The distinction between the enrichment endpoint, such as stable disease after 12 weeks, and the primary endpoint, such as progression status at 24 weeks or overall survival, carries the argument against circularity.
What would settle it
Re-analyze the sorafenib case with a different open-label response threshold, for example 20% tumor shrinkage instead of 25%, and check whether the five-attribute estimand and the estimated maintenance effect change; if the target population attribute cannot be stated before any open-label data exist and replicated across trials, the framework fails its own pre-specification test.
Extended reading notes
Core claim
The central claim is that a randomized discontinuation trial's estimand differs from a traditional RCT's estimand in exactly two of the five E9(R1) attributes: population and treatment. The target population is selected after the open-label phase by a response criterion, so it is a latent responder set that cannot be fully identified before data collection; the treatment contrast is continued active drug versus placebo or standard care after initial benefit, so the estimand captures a maintenance effect rather than an initiation effect. Endpoint, population-level summary, and intercurrent-event handling remain analogous to an RCT when the estimand is anchored at the randomization point, and pre-randomization events such as early progression or toxicity are design and selection features rather than intercurrent events. The case studies of the phase II sorafenib trial and the JAVELIN Gastric 100 phase III trial show how the five attributes should be filled in for a binary endpoint and a time-to-event endpoint, respectively.
Load-bearing premise
The argument assumes that a responder-defined enriched population can be specified objectively enough to count as a target population under E9(R1), even though the population is known only after the open-label phase and cannot be fixed before data collection.
Editorial extensions
If this is right
- A regulatory submission for an RDD oncology trial can state its estimand across all five E9(R1) attributes, with the population recorded as the data-driven enriched responder set.
- When the estimand is anchored at randomization, intercurrent events in the randomized phase are handled with the same strategies as in a traditional RCT, while pre-randomization events are design features, not intercurrent events.
- The treatment attribute in an RDD estimand describes maintenance of an already-initiated treatment, not initiation, so causal conclusions are limited to the conditional responder population.
- Binary phase II endpoints that closely mirror the enrichment endpoint risk circularity; time-to-event endpoints such as overall survival avoid this tautology.
- Only randomized-phase data are typically used in the analysis, and methods that use both stages of an RDD remain an open problem.
Reading between the lines
- An RDD estimand is best read as a responder-stratum conditional estimand, and its generalizability would depend on explicitly modeling the enrichment selection; the paper does not take this step.
- A natural extension is to record the enrichment criterion itself as part of the estimand definition, since the open-label response threshold determines both the population and the treatment contrast.
- One testable extension would simulate RDD trials with known treatment effects and varying responder misclassification rates to quantify how much the maintenance-effect estimand changes, informing a sensitivity analysis for the framework.
- The paper's logic that pre-randomization events are not intercurrent events implies that any estimand defined over the full treatment journey, from initial therapy onward, would require a different and more complex intercurrent-event handling than the one proposed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops a conceptual estimand framework, in the spirit of ICH E9(R1), for oncology trials using the randomized discontinuation design. It uses two case studies—the phase II sorafenib trial and the phase III JAVELIN Gastric 100 trial—and maps their objectives onto the five E9(R1) estimand attributes (population, treatment, endpoint, population-level summary, intercurrent events). The paper argues that RDD estimands differ from traditional RCT estimands chiefly in the population and treatment attributes: the target population is 'data-driven' because only responders from the open-label phase are randomized, and the treatment effect is a maintenance effect rather than an initiation effect. It also argues that, once the estimand is anchored at randomization, intercurrent-event handling is similar to that in RCTs. The paper explicitly acknowledges unresolved challenges, including the difficulty of pre-specifying a responder-defined population and the risk of tautology when the enrichment endpoint resembles the primary endpoint.
Significance. If accepted, the framework would fill a genuine gap: E9(R1) applications to enrichment designs such as RDD are rare, and regulatory communication would benefit from explicit estimand specification in these trials. The paper is balanced, well-referenced, and contains no fitted parameters or derivations, so the usual circularity concerns do not arise. It deserves credit for openly flagging the data-driven-population problem, the maintenance-versus-initiation distinction, and the enrichment-endpoint tautology risk. The central pre-specification tension and a misclassified phase III example, however, leave the paper's main claim only partially supported in its current form.
major comments (3)
- [Section 3, Table 1 (Population attribute)] The central claim that RDD has a data-driven target population is left in unresolved tension with the E9(R1) and Mütze et al. (2025) principle that an estimand should be independent of trial conduct and data. The paper concedes that the randomized population is 'retrospectively determined' and 'cannot be identified in advance' (Section 3), but the proposed remedy that the enriched population 'must be objectively defined' addresses measurement quality, not pre-specification. An objectively measured 12-week responder set is still a set that depends on observed responses during the trial. To make the Table 1 population attribute a valid estimand attribute, the authors need to define the population without reference to realized data, for example as a principal stratum defined by potential response to the open-label treatment, or by a baseline covariate rule. Without such a formalization, the Section 4 statement that RDD is distinct because its target population is data-driven supports the paper's descriptive point but not its claim that the E9(R1) framework can be applied to RDD as proposed.
- [Abstract and Section 2.2] JAVELIN Gastric 100 does not satisfy the definition of RDD given in the paper's own Abstract and Introduction: participants receive induction chemotherapy (not the investigational product avelumab) before randomization, and the randomized comparison is between initiating maintenance avelumab and continuing induction chemotherapy, not between continuing and withdrawing the same investigational product. Thus the phase III example is a maintenance-therapy or response-enrichment design rather than a randomized discontinuation design. The claim that the proposed framework covers 'phase III RDD' therefore rests on a misclassified case study. The authors should either identify a genuine phase III RDD example or explicitly reframe JAVELIN as a related design and qualify the phase III claims accordingly.
- [Section 3, Endpoint attribute] The paper notes that the sorafenib primary endpoint (progression status at 12 weeks post-randomization) is 'similar, but not identical' to the enrichment endpoint (stable disease at week 12) and cites Ghaemi and Selker on the tautology risk, but it does not explain whether the difference is sufficient to avoid conditioning on a post-baseline outcome. Since the estimand's population is defined by the same 12-week tumor-response measure, the causal interpretation of the primary-endpoint estimand in the phase II example is not established. The manuscript should discuss this explicitly, for instance by showing how the estimand can be expressed as a principal-stratum effect or by specifying a different outcome (e.g., PFS from randomization) and explaining why the enrichment-endpoint similarity does not bias the estimate.
minor comments (6)
- [Title and running header] The title and running header contain 'FORRANDOMIZEDDISCONTINUATIONDESIGNS' with missing word separators; the spacing should be corrected.
- [Section 4, first paragraph] 'orafenib' should be 'sorafenib' in the phrase 'a phase II trial of orafenib'.
- [Author affiliations] 'Advanced Quantitative Scieces' should be 'Advanced Quantitative Sciences'.
- [Table 1] The table contains the typo 'Similiar to RCT', which should be 'Similar to RCT'.
- [Throughout] The trial name is written as 'JA VELIN Gastric 100' with a space; the standard name is 'JAVELIN Gastric 100' and the spelling should be made consistent.
- [Section 4, opening sentence] The paper calls RDD 'an adaptive design', but adaptive designs are usually defined by pre-planned modifications based on interim data; the cited FDA guidance classifies RDD as an enrichment strategy. This statement should be clarified or supported.
Circularity Check
No significant circularity: the paper applies the E9(R1) estimand framework to randomized discontinuation designs using external case studies, with no fitted parameters, no derivation that reduces to its inputs, and no load-bearing self-citation chain.
full rationale
The paper makes no quantitative predictions and fits no parameters. Its central claim is that the E9(R1) estimand attributes can be applied to randomized discontinuation designs (RDD), and that the population and treatment attributes differ from traditional RCTs. These claims are supported by external trial reports (Ratain et al. 2006; Moehler et al. 2021) and external regulatory/statistical guidance (FDA 2019, 2021; Mütze et al. 2025). The population discussion is explicitly self-critical: 'in RDD, the randomized population is retrospectively determined based on individual treatment responses - information that only becomes available during the trial. This makes the estimand dependent on a latent responder population, which cannot be identified in advance and may differ across studies or real-world settings. As a result, this data-driven population definition limits generalizability and challenges the estimand framework, which assumes a clearly defined target population before data collection.' That is a stated limitation, not a circular derivation: it does not assume the conclusion, it exposes a tension between RDD practice and E9(R1)'s pre-specification principle. The observation that the RDD target population is data-driven follows from the design definition, and the paper does not present it as a fitted prediction. There is no equation that equals itself, no fitted parameter renamed as a prediction, and no author-overlapping citation invoked as load-bearing evidence. The use of the E9(R1) attributes as the organizing lens is the paper's method, not a circularity: applying a framework to a design and reporting similarities and differences is a legitimate application, not a reduction of the conclusion to the input. The skeptical objection that a responder-defined population cannot be pre-specified under E9(R1) is a validity and generalizability concern; it belongs under correctness risk, not circularity. Accordingly, no circular step is identified.
Assumptions & free parameters
assumptions (5)
- domain assumption ICH E9(R1) five-attribute estimand framework (population, treatment, endpoint, population-level summary, intercurrent events) is the correct and applicable lens for defining trial objectives in RDD oncology trials.
- domain assumption RDD is correctly classified as a predictive enrichment strategy under FDA (2019) guidance, making the enriched responder subgroup the target population.
- domain assumption The reported results of the sorafenib phase II trial (Ratain et al. 2006) and JAVELIN Gastric 100 (Moehler et al. 2021) are accurate and representative of oncology RDD practice.
- domain assumption Estimands should be independent of trial conduct and data, per Mütze et al. (2025).
- ad hoc to paper Pre-randomization events (early progression, toxicity, discontinuation during open-label treatment) are not intercurrent events when the estimand is anchored at randomization.
Cite this review
Pith. "Pith review of Estimands for Randomized Discontinuation Designs in Oncology." pith.science (2026). https://pith.science/paper/GLBZL6TA
@misc{pith2026250600556,
author = {Pith},
title = {Pith review of: Estimands for Randomized Discontinuation Designs in Oncology},
year = {2026},
howpublished = {\url{https://pith.science/paper/GLBZL6TA}},
note = {Machine review of arXiv:2506.00556}
}
read the original abstract
Randomized discontinuation design (RDD) is an enrichment strategy commonly used to address limitations of traditional placebo-controlled trials, particularly the ethical concern of prolonged placebo exposure. RDD consists of two phases: an initial open-label phase in which all eligible patients receive the investigational medicinal product (IMP), followed by a double-blind phase in which responders are randomized to continue with the IMP or switch to placebo. This design tests whether the IMP provides benefit beyond the placebo effect. The estimand framework introduced in ICH E9(R1) strengthens the dialogue among clinical research stakeholders by clarifying trial objectives and aligning them with appropriate statistical analyses. However, its application in oncology trials using RDD remains unclear. This manuscript uses the phase III JAVELIN Gastric 100 trial and the phase II trial of sorafenib (BAY 43-9006) as case studies to propose an estimand framework tailored for oncology trials employing RDD in phase III and phase II settings, respectively. We highlight some similarities and differences between RDDs and traditional randomized controlled trials in the context of ICH E9(R1). This approach aims to support more efficient regulatory decision-making.
Reference graph
Works this paper leans on
-
[6]
doi:10.1200/JCO.2011.35.3524. D. C. Smith, M. R. Smith, C. Sweeney, A. A. Elfiky, C. Logothetis, P. G. Corn, N. J. V ogelzang, E. J. Small, A. L. Harzstark, M. S. Gordon, U. N. Vaishampayan, N. B. Haas, A. I. Spira, P. N. Lara, C. C. Lin, S. Srinivas, A. Sella, P. Schöffski, C. Scheffold, A. L. Weitzman, and M. Hussain. Cabozantinib in patients with advan...
-
[13]
doi:10.1200/JCO.2006.08.102. M. J. Ratain. In reply to the letter by Sonpavde et al.Journal of Clinical Oncology, 24(28):4670–4671,
-
[14]
doi:10.1200/JCO.2006.08.4418. S. N. Ghaemi and H. P. Selker. Maintenance efficacy designs in psychiatry: Randomized discontinuation trials - enriched but not better.Journal of Clinical and Translational Science, 1(3):198–204,
-
[17]
doi:10.1136/bmj-2023-076316. M. Okwuokenye and K. E. Peace. Adaptive design and the estimand framework.Annals of Biostatistics & Biometric Applications, 1(5):1–4,
-
[19]
doi:10.1002/cpt.2575. V .V . Fedorov and T. Liu. Randomized discontinuation trials with binary outcomes.Journal of Statistical Theory and Practice, 8:30–45,
-
[21]
doi:10.1080/00031305.2012.720900. 8
arXiv 2012
-
[1975]
doi:10.1002/j.1552-4604.1975.tb05919.x. 6 arXivTemplateA PREPRINT J. A. Kopec, M. Abrahamowicz, and J. M. Esdaile. Randomized discontinuation trials: utility and efficiency.Journal of Clinical Epidemiology, 46(9):959–971,
-
[1993]
doi:10.1016/0895-4356(93)90163-u. Food and Drug Administration. Enrichment Strategies for Clinical Trials to Support Approval of Human Drugs and Biological Products. Guidance for Industry,
Show all 21 references
-
[2002]
doi:10.1200/JCO.2002.11.126. L. Trippa, G. L. Rosner, and P. Müller. Bayesian enrichment strategies for randomized discontinuation trials.Biometrics, 68(1):203–211,
2002 doi
-
[2006]
doi:10.1200/JCO.2005.03.6723. D. A. Nosov, B. Esteves, O. N. Lipatov, A. A. Lyulko, A. A. Anischenko, R. T. Chacko, D. C. Doval, A. Strahs, W. J. Slichenmyer, and P. Bhargava. Antitumor activity and safety of tivozanib (A V-951) in a phase II randomized discontinuation trial i...
2005 doi
-
[2007]
doi:10.1056/NEJMoa060655. T. Mütze, J. Bell, S. Englert, P. Hougaard, D. Jackson, V . Lanius, and H. Ravn. Principles for defining estimands in clinical trials—A proposal.Pharmaceutical Statistics, 24:e2432,
-
[2012]
doi:10.1111/j.1541-0420.2011.01623.x. M. J. Ratain, T. Eisen, W. M. Stadler, K. T. Flaherty, S. B. Kaye, G. L. Rosner, M. Gore, A. A. Desai, A. Patnaik, H. Q. Xiong, E. Rowinsky, J. L. Abbruzzese, C. Xia, R. Simantov, B. Schwartz, and P. J. O’Dwyer. Phase II placebo-controlled...
2011
-
[2013]
doi:10.1200/JCO.2012.45.0494. C. O. Harding, R. S. Amato, M. Stuy, N. Longo, B. K. Burton, J. Posner, H. H. Weng, M. Merilainen, Z. Gu, J. Jiang, J. V ockley, and PRISM-2 Investigators. Pegvaliase for the treatment of phenylketonuria: A pivotal, double- blind randomized discon...
2012 doi
-
[2014]
doi:10.1080/15598608.2014.840492. T.G. Karrison, M.J. Ratain, Stadler W.M., and G.L. Rosner. Estimation of progression-free survival for all treated patients in the randomized discontinuation trial design.The American Statistician, 66:155–162,
2014
-
[2017]
doi:10.1017/cts.2017.2. A. A. Fierenz and A. Zapf. Current developments of the estimand concept.Pharmaceutical Statistics, 23:864–869,
2017 doi
-
[2018]
doi:10.1016/j.ymgme.2018.03.003. M. Moehler, M. Dvorkin, N. Boku, M. Özgüro ˘glu, M. H. Ryu, A. S. Muntean, S. Lonardi, M. Nechaeva, A. C. Bragagnoli, H. S. Co¸ skun, A. Cubillo Gracian, T. Takano, R. Wong, H. Safran, G. M. Vaccaro, Z. A. Wainberg, M. R. Silver, H. Xiong, J. H...
2018 doi
-
[2019]
7 arXivTemplateA PREPRINT O
doi:10.33552/ABBA.2019.01.000524. 7 arXivTemplateA PREPRINT O. Collignon, A. Schiel, C. F. Burman, K. Rufibach, M. Posch, and F. Bretz. Estimands and complex innovative designs. Clinical Pharmacology & Therapeutics, 112(6):1183–1190,
2019 doi
-
[2021]
Food and Drug Administration
doi:10.1200/JCO.20.00892. Food and Drug Administration. E9 (R1) Statistical Principles for Clinical Trials: Addendum: Estimands and Sensitivity Analysis in Clinical Trials,
-
[2022]
doi:10.1007/978-3-319-52636-
-
[2024]
doi:10.1002/pst.2395. B.C. Kahan, J. Hindley, M. Edwards, S. Cro, and T.P. Morris. The estimands framework: a primer on the ICH E9(R1) addendum.BMJ, 384:e076316,
-
[2025]
doi:10.1002/pst.2432. V . V . Fedorov. Randomized discontinuation trials. In Steven Piantadosi and Curtis L. Meinert, editors,Principles and Practice of Clinical Trials, pages 1439–1453. Springer Nature Switzerland AG,
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.