{"id":"a069fcca-27a8-4915-ad6f-9d2ce229a550","arxiv_id":"2506.00556","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Maps the ICH E9(R1) estimand framework onto oncology randomized discontinuation trials, showing population and treatment attributes differ from traditional RCTs while intercurrent-event handling is similar when anchored at randomization.","lead":"This paper proposes a structured way to define the treatment question in two-stage cancer trials where all patients first get the drug and only responders are then randomized to continue it or switch to placebo. It is aimed at clinical trial statisticians and drug regulators, using two completed cancer trials to show how the official ICH 'estimand' framework applies to these enrichment designs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The population attribute in RDD cannot be pre-specified under the E9(R1) principle of estimands independent of trial data; the paper concedes this tension but offers no resolution, so the framework's central claim remains unsupported unless a principal-stratum or equivalent formalization is…","rationale":"The paper is a clear, honest framework application: it maps E9(R1) attributes onto two oncology trials and openly flags the population tension. There is no mathematical derivation, data, or code to examine, so the axes that would punish fitting or circularity are not implicated. The reader's weakest_assumption matches my own: the unresolved data-driven population attribute is the load-bearing weakness. I also considered the separate observation that JAVELIN Gastric 100 may not satisfy the paper's own definition of RDD (no investigational product during the induction phase, open-label randomization, no placebo control), but this primarily weakens the phase III example rather than the population-attribute logic, so I do not elevate it. A principal-stratum formulation could potentially rescue the paper, and therefore the appropriate verdict remains conditional: the concern is real and central, but it may be addressable with a precise definition of the responder population before data collection.","tokens_in":9305,"tokens_out":5619,"duration_ms":57180,"concrete_test":"Write a formal pre-specified estimand statement for the sorafenib example whose population attribute contains no reference to observed 12-week response, e.g., 'patients with advanced RCC who would have stable disease after 12 weeks of sorafenib' as a principal stratum. If such a statement is expressible and acceptable under E9(R1), then the paper must add it or explicitly acknowledge that its population attribute is not E9(R1)-compliant; if it is inexpressible, the Section 4 claim of a valid tailored estimand framework fails. This can be checked by a one-page analytic exercise against the E9(R1) addendum and Mütze et al. (2025) principles.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 makes the paper's central move: Table 1 maps the five E9(R1) attributes for the sorafenib and JAVELIN trials and claims that RDD populations are data-driven while the framework still applies. The load-bearing assumption is that a data-driven responder population can be an estimand population under E9(R1). The paper itself concedes that 'the randomized population is retrospectively determined based on individual treatment responses... cannot be identified in advance' and that this 'challenges the estimand framework, which assumes a clearly defined target population before data collection.' The proposed remedy—that an enriched population 'must be objectively defined'—does not resolve the tension: 'objective' describes the measurement rule, not pre-specification or independence from trial data. The actual set of randomized patients depends on observed 12-week responses; a target population fixed before data collection cannot be equivalent unless it is defined by baseline covariates or by potential response under initial treatment (principal stratification). Neither appears in the paper. Consequently, the Section 4 claim that RDD has 'unique features' including a data-driven target population is internally consistent, but it conflicts with the E9(R1)/Mütze principle the framework is supposed to apply, so the central estimand-framework claim is not yet supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a conceptual estimand framework, in the spirit of ICH E9(R1), for oncology trials using the randomized discontinuation design. It uses two case studies—the phase II sorafenib trial and the phase III JAVELIN Gastric 100 trial—and maps their objectives onto the five E9(R1) estimand attributes (population, treatment, endpoint, population-level summary, intercurrent events). The paper argues that RDD estimands differ from traditional RCT estimands chiefly in the population and treatment attributes: the target population is 'data-driven' because only responders from the open-label phase are randomized, and the treatment effect is a maintenance effect rather than an initiation effect. It also argues that, once the estimand is anchored at randomization, intercurrent-event handling is similar to that in RCTs. The paper explicitly acknowledges unresolved challenges, including the difficulty of pre-specifying a responder-defined population and the risk of tautology when the enrichment endpoint resembles the primary endpoint.","tokens_in":9416,"tokens_out":6813,"duration_ms":64538,"significance":"If accepted, the framework would fill a genuine gap: E9(R1) applications to enrichment designs such as RDD are rare, and regulatory communication would benefit from explicit estimand specification in these trials. The paper is balanced, well-referenced, and contains no fitted parameters or derivations, so the usual circularity concerns do not arise. It deserves credit for openly flagging the data-driven-population problem, the maintenance-versus-initiation distinction, and the enrichment-endpoint tautology risk. The central pre-specification tension and a misclassified phase III example, however, leave the paper's main claim only partially supported in its current form.","major_comments":[{"comment":"The central claim that RDD has a data-driven target population is left in unresolved tension with the E9(R1) and Mütze et al. (2025) principle that an estimand should be independent of trial conduct and data. The paper concedes that the randomized population is 'retrospectively determined' and 'cannot be identified in advance' (Section 3), but the proposed remedy that the enriched population 'must be objectively defined' addresses measurement quality, not pre-specification. An objectively measured 12-week responder set is still a set that depends on observed responses during the trial. To make the Table 1 population attribute a valid estimand attribute, the authors need to define the population without reference to realized data, for example as a principal stratum defined by potential response to the open-label treatment, or by a baseline covariate rule. Without such a formalization, the Section 4 statement that RDD is distinct because its target population is data-driven supports the paper's descriptive point but not its claim that the E9(R1) framework can be applied to RDD as proposed.","section":"Section 3, Table 1 (Population attribute)"},{"comment":"JAVELIN Gastric 100 does not satisfy the definition of RDD given in the paper's own Abstract and Introduction: participants receive induction chemotherapy (not the investigational product avelumab) before randomization, and the randomized comparison is between initiating maintenance avelumab and continuing induction chemotherapy, not between continuing and withdrawing the same investigational product. Thus the phase III example is a maintenance-therapy or response-enrichment design rather than a randomized discontinuation design. The claim that the proposed framework covers 'phase III RDD' therefore rests on a misclassified case study. The authors should either identify a genuine phase III RDD example or explicitly reframe JAVELIN as a related design and qualify the phase III claims accordingly.","section":"Abstract and Section 2.2"},{"comment":"The paper notes that the sorafenib primary endpoint (progression status at 12 weeks post-randomization) is 'similar, but not identical' to the enrichment endpoint (stable disease at week 12) and cites Ghaemi and Selker on the tautology risk, but it does not explain whether the difference is sufficient to avoid conditioning on a post-baseline outcome. Since the estimand's population is defined by the same 12-week tumor-response measure, the causal interpretation of the primary-endpoint estimand in the phase II example is not established. The manuscript should discuss this explicitly, for instance by showing how the estimand can be expressed as a principal-stratum effect or by specifying a different outcome (e.g., PFS from randomization) and explaining why the enrichment-endpoint similarity does not bias the estimate.","section":"Section 3, Endpoint attribute"}],"minor_comments":[{"comment":"The title and running header contain 'FORRANDOMIZEDDISCONTINUATIONDESIGNS' with missing word separators; the spacing should be corrected.","section":"Title and running header"},{"comment":"'orafenib' should be 'sorafenib' in the phrase 'a phase II trial of orafenib'.","section":"Section 4, first paragraph"},{"comment":"'Advanced Quantitative Scieces' should be 'Advanced Quantitative Sciences'.","section":"Author affiliations"},{"comment":"The table contains the typo 'Similiar to RCT', which should be 'Similar to RCT'.","section":"Table 1"},{"comment":"The trial name is written as 'JA VELIN Gastric 100' with a space; the standard name is 'JAVELIN Gastric 100' and the spelling should be made consistent.","section":"Throughout"},{"comment":"The paper calls RDD 'an adaptive design', but adaptive designs are usually defined by pre-planned modifications based on interim data; the cited FDA guidance classifies RDD as an enrichment strategy. This statement should be clarified or supported.","section":"Section 4, opening sentence"}],"recommendation":"major_revision","confidential_remarks":"This is a conceptual paper with no numerical results, and the main obstacle is the unresolved pre-specification tension, which is fixable by adding a principal-stratum formalization or by reframing the contribution as a discussion of challenges rather than a complete framework. The misclassification of JAVELIN Gastric 100 as an RDD should also be corrected. The paper is a reasonable fit for a statistical methodology or regulatory-statistics venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first explicit mapping of ICH E9(R1) estimand attributes onto randomized discontinuation designs in oncology, and it is a genuinely useful one. It is not a finished estimand framework. The paper names the central problem—the RDD population is data-driven and cannot be pre-specified—and then leaves that problem unresolved. If you work in this area, treat it as a well-scoped problem statement with a helpful table, not as a complete solution.\n\nWhat is good: Table 1 does real work. The maintenance-versus-initiation distinction is accurate and carefully argued. The point that intercurrent-event handling is conceptually the same as in an RCT once you anchor at randomization is correct and worth having in print. The paper also flags the tautology risk when the enrichment endpoint and the primary endpoint are close, which is a common practical issue. The references are appropriate—they cover the historical RDD papers, the newer estimand principles (Mütze et al.), and adjacent work on estimands in adaptive and complex designs. The authors are appropriately cautious about relying on secondary trial reports.\n\nThe soft spot is exactly the one in the stress-test note. Section 3 concedes that the randomized population is retrospectively determined and cannot be identified in advance, and that this challenges the estimand framework. The proposed remedy—an 'objectively defined' enriched population—does not help. An objective measurement rule does not make the population pre-specified or independent of trial data. To define an estimand under E9(R1), you would need something like principal stratification on the potential response to the initial treatment, or an explicit argument that the target population is a rule-based subpopulation of the trial itself. The paper hints at a latent responder population but does not formalize it. So the central claim that the framework can be tailored to RDD is not yet supported. That is not a fatal flaw for a conceptual/taxonomic paper, but the authors should either take the formalization step or scope the claim down.\n\nMinor gaps: the 'first in oncology RDD' claim rests on informal knowledge, and there is no concrete template estimand statement or a worked example of how the mapping would change a real protocol. Supplying either would make the paper much more actionable. No code or data is present, which is acceptable for a conceptual paper.\n\nWho this is for: statisticians in pharma, regulatory, and academic settings who deal with enrichment designs and the E9(R1) addendum. It deserves a serious referee—send it out. The referee should ask for the principal-stratification fix or a clear scope limitation, but this is not a desk reject.","headline":"Useful first mapping of E9(R1) estimand attributes onto oncology RDDs, but the population pre-specification problem is left unresolved; still deserves a serious referee.","tokens_in":10083,"tokens_out":5398,"would_cite":false,"duration_ms":51091,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Randomized discontinuation trials in oncology can be given a complete ICH E9(R1) estimand specification, with the target population defined by the open-label response criterion and intercurrent-event handling anchored at randomization.","keywords":["randomized discontinuation design","estimands","ICH E9(R1)","enrichment trials","oncology","intercurrent events","maintenance effect","predictive enrichment"],"falsifier":"Re-analyze the sorafenib case with a different open-label response threshold, for example 20% tumor shrinkage instead of 25%, and check whether the five-attribute estimand and the estimated maintenance effect change; if the target population attribute cannot be stated before any open-label data exist and replicated across trials, the framework fails its own pre-specification test.","tokens_in":1831,"feed_emoji":"💊","tokens_out":1829,"duration_ms":55667,"temperature":0.7,"pith_summary":"The paper argues that the ICH E9(R1) estimand framework should be applied to randomized discontinuation designs in oncology, where all patients first receive the drug openly and only responders are then randomized. It claims that the population attribute in such designs is data-driven, consisting of the enriched set of responders selected during the open-label phase, and therefore differs from the population attribute in a traditional randomized controlled trial. It also claims that the treatment attribute reflects maintenance of an already-initiated treatment rather than treatment initiation, while the endpoint, population-level summary, and intercurrent-event handling remain similar to a traditional RCT when the estimand is anchored at randomization. The proposal is worked out on two published trials: the phase II sorafenib trial and the phase III JAVELIN Gastric 100 trial. If correct, this gives drug developers and regulators a common, explicit language for specifying what an RDD trial is estimating.","feed_headline":"A five-attribute map defines RDD estimands in oncology","feed_subtitle":"The population is the open-label responders; treatment is maintenance; intercurrent events match standard trials.","key_machinery":"The machinery is the five-attribute estimand specification of ICH E9(R1) — population, treatment, endpoint, population-level summary, and intercurrent events — applied to RDD as a structured mapping table. The load-bearing move is anchoring the estimand at the moment of randomization in the double-blind phase; this separates pre-randomization enrichment events, which define the population, from post-randomization intercurrent events, which are handled as in a standard RCT. The distinction between the enrichment endpoint, such as stable disease after 12 weeks, and the primary endpoint, such as progression status at 24 weeks or overall survival, carries the argument against circularity.","core_discovery":"The central claim is that a randomized discontinuation trial's estimand differs from a traditional RCT's estimand in exactly two of the five E9(R1) attributes: population and treatment. The target population is selected after the open-label phase by a response criterion, so it is a latent responder set that cannot be fully identified before data collection; the treatment contrast is continued active drug versus placebo or standard care after initial benefit, so the estimand captures a maintenance effect rather than an initiation effect. Endpoint, population-level summary, and intercurrent-event handling remain analogous to an RCT when the estimand is anchored at the randomization point, and pre-randomization events such as early progression or toxicity are design and selection features rather than intercurrent events. The case studies of the phase II sorafenib trial and the JAVELIN Gastric 100 phase III trial show how the five attributes should be filled in for a binary endpoint and a time-to-event endpoint, respectively.","pith_inferences":["An RDD estimand is best read as a responder-stratum conditional estimand, and its generalizability would depend on explicitly modeling the enrichment selection; the paper does not take this step.","A natural extension is to record the enrichment criterion itself as part of the estimand definition, since the open-label response threshold determines both the population and the treatment contrast.","One testable extension would simulate RDD trials with known treatment effects and varying responder misclassification rates to quantify how much the maintenance-effect estimand changes, informing a sensitivity analysis for the framework.","The paper's logic that pre-randomization events are not intercurrent events implies that any estimand defined over the full treatment journey, from initial therapy onward, would require a different and more complex intercurrent-event handling than the one proposed."],"forward_implications":["A regulatory submission for an RDD oncology trial can state its estimand across all five E9(R1) attributes, with the population recorded as the data-driven enriched responder set.","When the estimand is anchored at randomization, intercurrent events in the randomized phase are handled with the same strategies as in a traditional RCT, while pre-randomization events are design features, not intercurrent events.","The treatment attribute in an RDD estimand describes maintenance of an already-initiated treatment, not initiation, so causal conclusions are limited to the conditional responder population.","Binary phase II endpoints that closely mirror the enrichment endpoint risk circularity; time-to-event endpoints such as overall survival avoid this tautology.","Only randomized-phase data are typically used in the analysis, and methods that use both stages of an RDD remain an open problem."],"supporting_citations":[{"why":"Supplies the ICH E9(R1) estimand framework that the paper applies to RDD.","marker":"[Food and Drug Administration, 2021]"},{"why":"Supplies the four principles for well-defined estimands, including independence from trial conduct, which creates the population-attribute tension the paper confronts.","marker":"[Mütze et al., 2025]"},{"why":"The phase II sorafenib RDD trial used as one of the two case studies for defining RDD estimands.","marker":"[Ratain et al., 2006]"},{"why":"The phase III JAVELIN Gastric 100 RDD trial used as the other case study for defining RDD estimands.","marker":"[Moehler et al., 2021]"},{"why":"Provides the phase II oncology RDD conceptualization that motivates the endpoint and enrichment design.","marker":"[Rosner et al., 2002]"},{"why":"Classifies RDD as a predictive enrichment strategy in FDA guidance, supporting the paper's population-attribute discussion.","marker":"[Food and Drug Administration, 2019]"},{"why":"Highlights the risk of responder misclassification in RDD, which the paper cites as a challenge to the population attribute.","marker":"[Fedorov, 2022]"},{"why":"The commentary questioning whether RDD enrichment yields a homogeneous population, cited as part of the population-attribute debate.","marker":"[Sonpavde et al., 2006]"},{"why":"Provides the argument that the predictor and outcome must differ to avoid tautology, which the paper applies to the endpoint attribute.","marker":"[Ghaemi and Selker, 2017]"}],"fun_headline_variants":["RDD estimands differ from RCTs in two attributes","Two E9(R1) attributes set RDD apart in oncology","Randomized discontinuation estimands: population and treatment differ","Maintenance effect captured by RDD estimands","Oncology RDD estimands: response-selected population, maintenance contrast"],"cache_read_input_tokens":12032,"weakest_assumption_plain":"The argument assumes that a responder-defined enriched population can be specified objectively enough to count as a target population under E9(R1), even though the population is known only after the open-label phase and cannot be fixed before data collection.","fun_headline_variants_meta":{"raw":{"variants":["RDD estimands differ from RCTs in two attributes","Two E9(R1) attributes set RDD apart in oncology","Randomized discontinuation estimands: population and treatment differ","Maintenance effect captured by RDD estimands","Oncology RDD estimands: response-selected population, maintenance contrast"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1411,"prompt_tokens":934,"completion_tokens":477,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":393}},"tokens_in":550,"tokens_out":477,"duration_ms":4233,"temperature":1.0,"reasoning_tokens":393,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:03:11.887040+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the sorafenib case with a different open-label response threshold, for example 20% tumor shrinkage instead of 25%, and check whether the five-attribute estimand and the estimated maintenance effect change; if the target population attribute cannot be stated before any open-label data exist and replicated across trials, the framework fails its own pre-specification test.","supporting_citations":[],"review_version":1}