{"id":"f8aa5351-0f66-48e5-962d-e609b6b0f954","arxiv_id":"2502.03329","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"Handling one intercurrent event by treatment policy and another by hypothetical strategy yields at least two distinct estimands, and the causal ordering of the events dictates which covariates are needed for unbiased estimation.","lead":"The paper defines how to precisely estimate treatment effects in clinical trials when two intercurrent events, such as stopping treatment and taking rescue medication, are handled with different statistical strategies. It shows that the causal order of these events determines which variables must be adjusted for, and that the same verbal description can hide two genuinely different estimands.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the central conceptual claim is internally consistent; the acknowledged need for a known causal ordering is a practical limitation, not a correctness defect.","rationale":"The reader's verdict is ACCEPT with high confidence, and I agree. The paper's contribution is conceptual: it clarifies that combining a treatment policy strategy for one intercurrent event and a hypothetical strategy for another does not yield a single unambiguous estimand, and that the causal structure between the events determines which covariates are needed for unbiased estimation. The definitions in Section 3 are careful: the hypothetical estimand (3) and the cross-world hypothetical estimand (4) are genuinely distinct, and the paper is explicit that (4) does not correspond to any single intervention and is unlikely to be useful to stakeholders. The identification section states the required assumptions, including the cross-world assumption, and the estimation sections translate the assumed ordering into concrete MI and IPW algorithms. The simulation is well designed: it verifies that the true potential-outcome effect is identical across the three structures, then shows that only the IPW estimator matching the true ordering is unbiased. This directly supports the central claim that ordering has implications for estimation. The weakest assumption, a known and uniform causal ordering of D and R, is acknowledged by the authors and is a practical limitation rather than an internal inconsistency. A heterogeneous-ordering population would indeed break the proposed algorithms, but this does not undermine the conceptual message; it reinforces the need to specify the ordering or use methods robust to such heterogeneity. The unmeasured common-cause scenario is also discussed in Section 4, so the Table 1 simplification does not mislead a careful reader. No formal verification or reproducibility issue was found; the simulation code is public, which is a point in the paper's favor. Overall, the central argument holds under scrutiny, and the identified limitations are appropriately disclosed in the manuscript.","tokens_in":14663,"tokens_out":16146,"duration_ms":162605,"concrete_test":"Simulate a population in which half the subjects follow the D-before-R data-generating process (Section 5.2) and half follow the R-before-D process (Section 5.3), with identical numerical coefficients and the same potential-outcome target. Apply the three IPW estimators from Section 5 to the pooled data. If, as expected, none of the three is unbiased within Monte Carlo error, this confirms that the uniform-ordering assumption is genuinely load-bearing whenever the trial population is heterogeneous.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No load-bearing objection identified. The paper's central claim, that distinct causal estimands arise when one ICE is handled by the treatment policy strategy and another by the hypothetical strategy, and that the required adjustment set depends on the assumed causal ordering of D and R, is internally consistent and supported by the simulation. The weakest point remains the assumption of a known and uniform temporal ordering of D and R, which the authors explicitly acknowledge in Section 3. If the ordering varies across patients, the three IPW algorithms in Sections 5.1-5.3 target different quantities and none will generally be unbiased for a common estimand. However, this is a boundary condition on application rather than a flaw in the conceptual argument. The cross-world hypothetical estimand (4) is clearly flagged as relying on a strong cross-world assumption and as being of doubtful usefulness, so its limited role does not threaten the main contribution. The only additional caveat worth noting is that Table 1's first row ('The ICEs do not affect each other') should be read as also assuming no shared cause U; Section 4 correctly explains that if an unmeasured U affects both D and R, D must still be adjusted for. This does not change the verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the definition, identification, and estimation of clinical-trial estimands when two intercurrent events are handled by different strategies: one by the treatment policy strategy and the other by the hypothetical strategy. Using potential outcomes and single-world intervention graphs, the authors define a primary 'hypothetical estimand' E[Y^{a=1,r=0} - Y^{a=0,r=0}] (Eq. 3) and a contrasting 'cross-world hypothetical estimand' (Eq. 4), and they show that the estimand and the required adjustment set depend on the assumed causal ordering of the two intercurrent events, D and R. The paper extends the discussion to a longitudinal setting with two time points, gives multiple imputation and inverse probability weighting algorithms under three causal structures, and demonstrates via simulation that only the IPW estimator whose adjustment set matches the true causal structure is unbiased. The motivating diabetes trial is revisited to illustrate the practical choice of ordering, and the paper concludes by recommending that trial protocols state the causal ordering of intercurrent events.","tokens_in":14831,"tokens_out":7013,"duration_ms":69222,"significance":"If the paper's claims hold, it makes a useful and practically important contribution to the implementation of the ICH E9(R1) estimand framework. The central conceptual point—that the phrase 'one ICE handled by treatment policy and another by the hypothetical strategy' does not define a unique causal estimand—is clearly made and is supported by formal graphical arguments. Strengths of the paper include the explicit statement of consistency, conditional exchangeability, and positivity assumptions in Section 4; the transparent treatment of the cross-world estimand, whose strong cross-world assumption (23) is stated explicitly and whose usefulness is appropriately questioned; the extension to longitudinal settings with concrete algorithms in Section 5; and the simulation study in Section 6, which is calibrated against a known generative model and whose R code is publicly available. The paper's limitation that the causal ordering of D and R may be unknown or vary across individuals is acknowledged by the authors and does not undermine the theoretical contribution, though it does constrain direct application.","major_comments":[],"minor_comments":[{"comment":"The statement that the estimands in Figure 1 j), k), and l) 'both reduce to' E[Y^{a=1,r=0} - Y^{a=0,r=0}] is concise to the point of being potentially misleading: in the column where R affects D, the D appearing in Y^{a,r=0} is D^{a,r=0}, not D^a. A sentence making this explicit would prevent misreading.","section":"§3, Eq. (3)"},{"comment":"The text says the estimation algorithms are described 'under the three causal structures depicted in Figure 5 a), b) and c)', but the relevant panels appear to be a), c), and e). Please correct the panel references.","section":"§5, first paragraph"},{"comment":"The first row, 'The ICEs do not affect each other. The treatment policy ICE can be ignored', should be read together with the Section 4 discussion of an unmeasured common cause U of D and R; as the text correctly notes, if a shared cause exists, D must still be accounted for even when there is no direct arrow between D and R. The table could state this qualification explicitly to avoid a misleading shortcut.","section":"Table 1"},{"comment":"The simulation results are presented only as box plots. Reporting numerical summaries such as Monte Carlo mean bias, empirical standard error, and coverage of confidence intervals would strengthen the reproducibility and allow readers to assess the magnitude of the biases shown.","section":"§6, Figure 6"},{"comment":"There are minor consistency issues in naming: the text refers to 'EMEA guidelines' in Section 7 while earlier sections and the abstract use EMA, and the reference list gives 'European Medicines Society' where the intended body is the European Medicines Agency. These should be harmonized.","section":"§7 and References"},{"comment":"The caption says 'unmeasured variables affecting D and U' but the figure and text indicate that U affects D and R. This typo should be corrected.","section":"Figure 2 caption"},{"comment":"The affiliation contains typographical errors: 'Enviromental Medicine' should be 'Environmental Medicine', and 'Karolinska Universitet' should be 'Karolinska Institutet'.","section":"Title page"}],"recommendation":"minor_revision","confidential_remarks":"This is a solid methodology paper whose central conceptual claim is sound and well supported. The main practical limitation—that the causal ordering of intercurrent events may not be known or uniform—is explicitly acknowledged and is a boundary condition rather than a correctness flaw. The self-citations to Olarte Parra et al. and Ocampo & Bather are appropriate as they supply the tools being extended. I support publication after the minor clarifications above."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. Olarte Parra, Daniel, and Bartlett give trialists something they can actually use: when one ICE is handled with treatment policy and another with hypothetical, there are at least two distinct estimands, and the causal ordering of D and R determines whether D enters the imputation/IPW models as an L_k or L_{k+1} covariate. That distinction – Eq. (3) versus the cross-world Eq. (4) – is not drawn in Lipkovich or the authors' own earlier work, and it matters for protocol writers. The paper is careful about the identification assumptions: Section 4 states consistency, conditional exchangeability, and positivity, and the supplementary material is explicit that Eq. (4) needs a cross-world assumption (23) that won't hold if an L2 common cause of D and Y exists. That is honest.\n\nThe simulation is well designed to isolate bias from mis-specified adjustment sets, and the code is public. The naive estimator is biased in all three structures; only the IPW variant that respects the true D-R ordering is unbiased. That gets the point across.\n\nSoft spots are real but not fatal. The biggest is the assumption that the causal ordering of D and R is known and uniform across patients. The authors flag this themselves in Section 3, but the practical guidance depends on it. If ordering varies, none of the three algorithms targets a well-defined common estimand. That is a boundary condition, not a flaw in the central argument. Minor: MI versions of the algorithms are described but not simulated; Monte Carlo standard errors are not reported; Table 1's first row should be read with the no-shared-cause caveat, though Section 4 covers it.\n\nI think the reader's summary is accurate. The paper deserves a serious referee. It fills a gap left by the E9 addendum, and the causal-ordering guidance is new practical advice. I'd cite it, and I'd bring it to a reading group that follows the estimands literature.\n\nRecommendation: send to peer review. It will need minor revisions but the core is sound and the contribution is clear.","headline":"Careful, useful paper: it separates two estimands that ICH E9 language conflates and shows the causal ordering of intercurrent events dictates the adjustment set; the main caveat is the known-ordering assumption, which the authors acknowledge.","tokens_in":15440,"tokens_out":1680,"would_cite":true,"duration_ms":15054,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The causal ordering of rescue and discontinuation determines whether and how discontinuation enters the analysis of a hypothetical estimand.","keywords":["ICH E9 addendum","causal inference","estimands","intercurrent events","treatment policy strategy","hypothetical strategy","single-world intervention graphs","time-varying confounding"],"falsifier":"Generate data under the DAG where rescue use $R$ precedes discontinuation $D$ and analyse it with the inverse-probability-weighting variant that assumes no $D$--$R$ effect; the paper's claim predicts bias and confidence intervals that miss the true effect, so observing nominal coverage in a correctly specified simulation would directly refute the claim that the ordering must enter the analysis.","tokens_in":14430,"feed_emoji":"📊","tokens_out":8900,"duration_ms":77592,"temperature":0.7,"pith_summary":"Clinical trial protocols must define estimands that say what to do with intercurrent events such as rescue medication and treatment discontinuation. This paper examines estimands that handle one event with the hypothetical strategy (set rescue use to zero) and the other with the treatment policy strategy (let discontinuation happen as it will). Using potential outcomes and causal diagrams, it shows that 'hypothetical strategy for rescue, treatment policy for discontinuation' is not a single estimand: the assumed causal ordering of the two events decides whether discontinuation is a time-varying confounder, and where it sits in the covariate set. A simulation confirms that choosing the wrong ordering biases the estimator. The practical message is that trial protocols must state the causal ordering of intercurrent events before the analysis is chosen.","feed_headline":"Right trial analysis depends on order of rescue and dropout","feed_subtitle":"Treatment discontinuation and rescue medication must have their causal order specified or the estimate misses its target.","key_machinery":"The carrying tools are DAGs and single-world intervention graphs (SWIGs). DAGs encode the assumed causal ordering of treatment, the two intercurrent events, covariates, and outcome; a SWIG represents the post-intervention world after setting $A=a$ and $R=0$, and it shows which paths remain open from $R$ to the potential outcome $Y^{a,r=0}$. That determines whether $D$ is ignored, included as a time-varying confounder at visit $k$, or included at visit $k+1$. The cross-world estimand cannot be drawn in a single SWIG, which is why it requires an extra cross-world independence assumption that does not follow from the intervention graphs alone.","core_discovery":"The central claim is that the EMA-style choice of handling rescue medication $R$ hypothetically and treatment discontinuation $D$ by treatment policy can be formalised in two genuinely different ways. The hypothetical estimand $E[Y^{a=1,r=0} - Y^{a=0,r=0}]$ lets $D$ take whatever value it would take once $R$ is intervened on; $D$ is then a post-intervention variable, and it must be adjusted for as a time-varying confounder exactly when it is a common cause of $R$ and $Y$. The cross-world hypothetical estimand $E[Y^{a=1,r=0,D^{a=1,R^{a=1}}} - Y^{a=0,r=0,D^{a=0,R^{a=0}}}]$ keeps $D$ at its natural value in the unset world; this mixture of worlds cannot be drawn in a single single-world intervention graph, has no possible trial interpretation, and is identifiable only under a cross-world assumption. For the first estimand, identification uses standard consistency, conditional exchangeability, and positivity; for the second, the paper shows that an imputation-style estimator identifies the quantity when a cross-world independence holds. The simulations show that, of three inverse-probability-weighting variants, only the one respecting the true causal ordering of $D$ and $R$ is unbiased.","pith_inferences":["Beyond the paper: when true orderings vary across patients, the DAG-based target becomes a mixture over orderings; a practical step the paper leaves implicit is a sensitivity analysis that repeats the analysis under plausible orderings and reports the spread of estimates.","Beyond the paper: the same two-interpretation ambiguity should reappear when two intercurrent events are both handled by the hypothetical strategy but one causes the other; the paper does not work out that case, but its SWIG logic suggests the estimand would again split into a single-world and a cross-world version.","Beyond the paper: in trials where event times are recorded, comparing estimates under different assumed orderings can serve as a diagnostic check on whether the protocol's ordering assumption is consistent with the data."],"forward_implications":["Trial teams that follow the EMA-style choice must decide, and state in the statistical analysis plan, whether discontinuation precedes or follows rescue use; the wrong choice biases the estimate of the treatment effect.","If discontinuation and rescue use do not affect each other, $D$ can be ignored in the imputation and weighting models; if $D$ precedes $R$, $D$ belongs in that visit's time-varying confounder set; if $R$ precedes $D$, $D$ enters the next visit's confounder set.","The cross-world version of the estimand, in which discontinuation keeps its natural value, is identifiable only under a cross-world assumption, but no conceivable trial realizes it; the paper advises that stakeholders generally should not use it.","With more than two intercurrent events handled by the treatment policy strategy, each additional event becomes part of the appropriate time-varying confounding set, so the same ordering logic generalizes.","The hypothetical estimand itself is identified by standard time-varying confounding assumptions; an unmeasured common cause of $D$ and $Y$ with a direct effect on $Y$ breaks identification even when the ordering is correctly specified."],"supporting_citations":[{"why":"Supplies the regulatory framework of intercurrent events and strategies that motivates the estimand problem.","marker":"ICH (2019)"},{"why":"Establishes that different strategies for intercurrent events produce different estimands, the starting point of the paper.","marker":"Lipkovich et al. (2020)"},{"why":"Introduces single-world intervention graphs for estimands, the graphical tool used to derive conditional exchangeability assumptions.","marker":"Ocampo & Bather (2023)"},{"why":"Defines hypothetical estimands in causal-inference terms and the multiple imputation / inverse-probability-weighting machinery that the paper extends to two intercurrent events.","marker":"Olarte Parra et al. (2023)"},{"why":"Supplies the identification assumptions for time-varying treatments that the paper adapts to the two-ICE setting.","marker":"Hernán & Robins (2020)"},{"why":"Underpins the non-parametric structural equation model interpretation used to justify the cross-world assumption.","marker":"Robins & Richardson (2010)"},{"why":"Provides the motivating diabetes trial where rescue medication is handled hypothetically and discontinuation by treatment policy.","marker":"Müller-Wieland et al. (2018)"}],"fun_headline_variants":["Rescue and dropout order decides trial estimate","Causal order of rescue vs dropout changes estimand","Trial estimand hinges on intercurrent event order","Only right causal ordering yields unbiased estimate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the causal ordering of the two intercurrent events is known, fixed for all patients, and fully captured by the assumed DAG; if the true ordering varies or is mis-specified, the analysis targets a different quantity, and if an unmeasured common cause directly affects the outcome, no adjustment set fixes it.","fun_headline_variants_meta":{"raw":{"variants":["Rescue and dropout order decides trial estimate","Causal order of rescue vs dropout changes estimand","Trial estimand hinges on intercurrent event order","Only right causal ordering yields unbiased estimate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000372,"raw_usage":{"total_tokens":1994,"prompt_tokens":952,"completion_tokens":1042,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":984}},"tokens_in":568,"tokens_out":1042,"duration_ms":7592,"temperature":1.0,"reasoning_tokens":984,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T05:06:20.298306+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data under the DAG where rescue use $R$ precedes discontinuation $D$ and analyse it with the inverse-probability-weighting variant that assumes no $D$--$R$ effect; the paper's claim predicts bias and confidence intervals that miss the true effect, so observing nominal coverage in a correctly specified simulation would directly refute the claim that the ordering must enter the analysis.","supporting_citations":[],"review_version":1}