{"id":"45f05ae2-6991-4a23-8884-babed33d084b","arxiv_id":"2411.12944","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A formal estimand framework defines treatment effects in platform trials on the entire concurrently eligible population and provides weighting and post-stratification estimators with asymptotic guarantees and efficiency gains over sub-study-only analyses.","lead":"A platform trial can test many treatments at once, but patients are not equally likely to get every treatment, which makes it hard to say what a treatment effect means. This paper defines the right population to compare two treatments, the entirely concurrently eligible patients, and provides estimators that stay valid and can use data across arms.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 1 is the load-bearing link: if observed Z omits any randomization determinant (site, block, exact time, or adaptive allocation), the ECE population and all six estimators are mis-specified and Theorem 1 does not apply.","rationale":"I read the paper and supplement with the goal of finding the weakest point in the chain from estimand to inference. The reader's weakest_assumption identifies Assumption 1, and my reading converges on the same point. The ECE population is defined through π_j(Z)>0 and π_k(Z)>0, so the estimand itself inherits every omission in Z. The IPW formula (3), the post-stratification construction, and Theorem 1 all require the conditional independence and known-propensity statement in Assumption 1. If Z is incomplete, the misspecification is not a small bias in variance estimation; it changes the target population and invalidates the consistency proof. I did not find an independent algebraic error in Theorem 1 or the supplement: the variance derivations, the AIPW/SAIPW equivalence, and the efficiency corollaries are internally coherent under the stated assumptions. The concern is therefore about applicability and the strength of the 'same minimal assumptions' claim, not about a contradiction in the mathematics. Because the paper's contribution is explicitly advertised as a general framework for platform trials, this assumption deserves a direct stress test. The proposed simulation isolates the mechanism: nominal π_j(Z) can be perfectly known while still missing randomization inputs, and the resulting coverage failure would show that Assumption 1 carries the entire construction. Since the reader already issued a CONDITIONAL verdict based on essentially this concern, my recommendation is UNCHANGED: the verdict stands, with the justification sharpened toward Assumption 1 rather than toward the reproducibility issues alone.","tokens_in":47551,"tokens_out":9741,"duration_ms":117576,"concrete_test":"Simulate a two-arm platform trial with site-level block randomization and a time trend, generating A by blocks of size 4 within site. Define Z as baseline covariates excluding site and block. Apply SIPW and SAIPW using the nominal π_j(Z) that ignores site/block to estimate θ_jk over 10,000 repetitions with n=500. If coverage of 95% confidence intervals falls materially below 95% (or bias exceeds about 10% of the standard deviation), Assumption 1's observability requirement is violated and the framework is not automatic. Control run: add site and enrollment-time block to Z; coverage should return to approximately 95%, confirming that the failure is driven by incomplete Z rather than by the estimators themselves.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim has two parts: the ECE estimand is well-defined and invariant, and IPW/SIPW/AIPW/SAIPW/PS/APS are consistent for it. Both parts are conditional on Assumption 1 (Section 3.1): A⊥(W,Y(1),...,Y(J))|Z and π_j(Z) known. This is not the 'same minimal assumption' as a traditional randomized trial. In a completely randomized trial, A is independent of potential outcomes unconditionally; here independence holds only if Z contains every variable used by the randomization procedure. In real platform trials these include enrollment order within a window, site, randomization block/sequence, or, under response-adaptive randomization, the history of previous outcomes. If any such variable is omitted or coarsened, π_j(Z) is not the true conditional assignment probability. Then the IPW identification identity (3) fails, the post-strata in Section 4.4 are mis-specified, and, more fundamentally, the ECE set {π_j(Z)>0, π_k(Z)>0} can include individuals with zero true chance of either treatment or exclude those with positive chance. The estimand itself is then not the 'entire concurrently eligible' population, and Theorem 1's conclusion no longer follows. The SIMPLIFY application does not establish the general claim because Z was available by design (enrollment window and HS/DA use); the paper offers no diagnostic or sensitivity analysis for violations of Assumption 1. This is a conditional weakness rather than an internal inconsistency: if Assumption 1 is genuinely satisfied, the algebraic and asymptotic developments in the supplement appear sound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses two foundational questions in platform trials: how to define a treatment-effect estimand that is not an artifact of randomization ratios or trial format, and how to estimate and perform inference for that estimand robustly. The authors define the ECE population as the set of individuals with positive assignment probability to both treatments under the observed randomization mechanism, propose the corresponding estimand, and develop six estimators: IPW, SIPW, AIPW, SAIPW, PS, and APS. The main theoretical contribution is Theorem 1, which gives asymptotic normality for all six estimators under Assumptions 1 and 2, along with explicit asymptotic covariance matrices and efficiency comparisons. The methodology is illustrated on the SIMPLIFY cystic fibrosis trial. Under the stated assumptions, the derivations are careful and the central claims are largely correct.","tokens_in":47881,"tokens_out":11174,"duration_ms":125574,"significance":"If the assumptions hold, this is a valuable formalization: the ECE estimand gives a clinically interpretable target that is invariant to design choices such as the randomization ratio between sub-studies, and the proposed estimators provide robust inference while allowing efficiency gains through covariate adjustment and pooling across arms. The explicit variance formulas, efficiency ordering results, simulation evidence, and the accompanying R package are strengths, and the paper goes beyond prior work by allowing distinct eligibility criteria and non-uniform assignment probabilities. The main caveats are that the identification and inference results are conditional on a strong observed-randomization-mechanism assumption, and that some practical recommendations are not covered by the stated theory.","major_comments":[{"comment":"Assumption 1 is the load-bearing condition for every result in the paper: it requires an observed baseline variable Z such that A is independent of potential outcomes given Z and such that the known π_j(Z) are the true conditional assignment probabilities. In many platform trials the actual randomization probability may depend on variables not contained in a coarse baseline Z, such as exact enrollment time within a window, site-level blocks, or, under response-adaptive randomization, the history of previous outcomes. If Z omits such a variable, then the identification formula (3) fails, the post-strata in Section 4.4 are mis-specified, and the ECE set {π_j(Z)>0, π_k(Z)>0} may not coincide with the population that could truly receive either treatment. The manuscript asserts that Z 'usually' includes these variables but offers no diagnostic, sensitivity analysis, or guidance for checking whether Assumption 1 is credible. Because this assumption defines both the estimand and the estimators, I would ask for an explicit discussion of its scope and at least a sensitivity analysis for omitted randomization variables, or a concrete balance-type diagnostic using the known π_j(Z).","section":"3.1, Assumption 1; 3.2, ECE definition"},{"comment":"The asymptotic theory in Theorem 1 rests on the assumption that (W_i,Y_i(1),...,Y_i(J),A_i) is an iid sample from a superpopulation. Platform trials unfold over calendar time, with treatment arms entering and leaving and with enrollment rates that may be non-stationary; under a fixed-sequence interpretation, the iid assumption is not automatic and the estimand itself depends on the stochastic process generating enrollment windows. The theorem conditions on n_jk but not on the realized sequence of windows, and no martingale or conditional asymptotic argument is provided for the time-dependent case. Please either state clearly that the target is a random-effects superpopulation in which enrollment windows are exchangeable, or relax the iid assumption and show how the results extend to fixed time trends and sequentially enrolled cohorts.","section":"3.1, superpopulation iid assumption"},{"comment":"The sentence 'Note that when Zi is discrete, πj(Zi) can be estimated using sample proportions and used in place of the true value πj(Z) in any of the above weighting estimators' is not covered by Theorem 1. When π_j(Z) is replaced by a sample proportion, the estimator changes its influence function; for example, the IPW estimator with estimated π_j(Z) becomes algebraically equivalent to a post-stratification estimator, whose asymptotic variance is different from the variance in Theorem 1(a). The variance estimators in Section S1.4 are derived for fixed, known π_j(Z), and plugging in estimated weights without accounting for their variability can produce invalid confidence intervals. Please either remove this remark, or provide the asymptotic theory for the estimated-weight versions and modify the variance estimators accordingly.","section":"4.3, remark on estimating π_j(Z)"}],"minor_comments":[{"comment":"The phrase 'the same minimal assumptions used in traditional randomized trials' is overstated: Assumption 1 requires conditional independence given an observed Z and a known conditional assignment mechanism, which is stronger than the unconditional independence that suffices in a completely randomized trial. Please rephrase or clarify the relationship.","section":"Abstract and Section 1.2"},{"comment":"The PS estimator is undefined if a post-stratum has no individuals assigned to treatment j, even though π_j(Z)>0 in that stratum. This finite-sample issue is acknowledged indirectly in the simulation discussion of PS(Z), but it should be stated explicitly in the main text along with a recommendation (e.g., pooling strata or using an IPW-type estimator).","section":"Section 4.4, PS estimator"},{"comment":"The application excludes 10 (1.7%) participants with missing outcomes and excludes data recorded after re-enrollment; these exclusions require additional assumptions (such as missingness independent of potential outcomes given Z and treatment, and no outcome-relevant effect of re-enrollment) for the reported estimates to be consistent for the ECE estimand. The paper should state these assumptions or explicitly label the empirical analysis as illustrative under these restrictions.","section":"Section 6, SIMPLIFY application"},{"comment":"In the text preceding the display, 'the APS estimator ˆϑ apw jk' contains a typo: the superscript should be 'aps', not 'apw'.","section":"Section 5.2, Corollary 4"},{"comment":"In Table S1, 'Baseline Age, yeas' should read 'years'.","section":"Supplement, Table S1"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a substantial contribution and the core derivations appear sound, but the manuscript needs additional work on the scope and verification of Assumption 1, on the superpopulation framing, and on the unproven recommendation to estimate known weights. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper deserves a serious referee. If you work on platform trials or estimands in master protocols, read it. The ECE population definition is the real contribution: comparing treatments j and k on everyone with positive assignment probability to both, which is invariant to randomization ratio and trial format. The formalization is new, even if FDA guidance gestured at it. The asymptotic theory for the six estimators is standard but carefully done; Theorem 1 and Corollaries 1-6 give explicit variance formulas and relative efficiency comparisons, including when covariate adjustment helps. The SIMPLIFY application is honest: point estimates align across methods, and the 5-27% variance reduction relative to sub-study-specific analysis is credible.\n\nWhat I like: the identification formula (3) is derived cleanly from Assumption 1, not assumed. The naive estimator bias formula (2) is instructive. The efficiency comparisons in the supplement are explicit, and the conditions (4)-(6) are clearly stated, with adjustment (7) to enforce (4). The paper also discusses analysis sets beyond the default, which is useful for partial blinding. This is solid work.\n\nSoft spots, in proportion:\n\n1. Assumption 1 is load-bearing and real. If the observed Z omits any variable used by the randomization procedure (site, block, exact enrollment time, adaptive allocation history), then pi_j(Z) is not the true conditional assignment probability, and the ECE set is mis-specified. In the SIMPLIFY application Z is available by design, but the paper gives no diagnostic or sensitivity analysis for violations. This is a conditional weakness, not an internal inconsistency. If Assumption 1 holds, the rest follows.\n\n2. The superpopulation iid assumption is strong for time-varying enrollment. If outcome distributions drift with calendar time beyond what Z captures, the estimand and inference are affected. Worth acknowledging more explicitly.\n\n3. Reproducibility: the R package RobinCID is named but no repository or version is provided. That is a small omission; add a link and simulation code.\n\n4. The \"first formal framework\" claim overstates relative to FDA guidance and Santacatterina et al. The paper does more, but the citation pattern should be more generous.\n\nThe central claim holds up under its assumptions. The paper is for clinical trial statisticians and regulators, and it deserves peer review. I would engage with it, and the revisions are mostly about framing, diagnostics, and artifacts, not fixing the math.\n\nBest,\n[Your name]","headline":"A careful formal treatment of the ECE estimand for platform trials with standard IPW/post-stratification machinery, conditional on a real but standard Assumption 1; worth refereeing with revisions.","tokens_in":48406,"tokens_out":2887,"would_cite":true,"duration_ms":32638,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62P10","62D05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes a framework for defining treatment effects on the entire concurrently eligible population, so estimands in platform trials no longer depend on randomization ratios or trial operation format.","keywords":["platform trials","master protocols","estimand","entire concurrently eligible population","inverse probability weighting","post-stratification","covariate adjustment","relative efficiency"],"falsifier":"Simulate a platform trial in which randomization actually depends on an unobserved enrollment-time or site variable not included in $Z$, then apply the proposed IPW and PS estimators using the protocol probabilities; if their confidence intervals fail to cover the true ECE treatment effect at the nominal rate, or the estimates depart from the truth beyond sampling error, the central claim is falsified.","tokens_in":47341,"feed_emoji":"📊","tokens_out":11331,"duration_ms":109212,"temperature":0.7,"pith_summary":"Platform trials randomize patients under constraints: some arms are unavailable to some patients, and assignment probabilities change with enrollment windows and patient characteristics. The paper argues that this flexibility makes the usual \"who is the target population?\" question ill-posed, because estimates based on patients actually assigned to a treatment or sub-study shift whenever the randomization ratio or trial format changes. Its central proposal is the entire concurrently eligible (ECE) population: all individuals who, at some point during the trial, have positive probability of being assigned to either of the two treatments being compared. Defining the estimand on this population makes the treatment effect invariant to randomization ratios and trial operation format, and the paper shows that inverse probability weighting, post-stratification, and model-assisted versions of both can estimate it consistently under the same minimal assumptions used in traditional randomized trials. Applied to the SIMPLIFY cystic fibrosis trial, the framework produces similar non-inferiority conclusions across the proposed estimators while avoiding the population drift seen in sub-study-only analyses.","feed_headline":"One estimand makes platform-trial effects independent of design","feed_subtitle":"For platform trials, effects can be defined on all concurrently eligible patients, independent of randomization ratios.","key_machinery":"The mechanism that carries the argument is Assumption 1: there is an observed baseline variable $Z$ (enrollment window, site, disease subtype, and similar) such that treatment assignment $A$ is independent of potential outcomes and baseline covariates given $Z$, and the assignment probabilities $\\pi_j(Z)$ are known, nonnegative, and sum to one. The ECE population is the support $\\{\\pi_j(Z)>0,\\ \\pi_k(Z)>0\\}$ determined by these probabilities. Estimation then works by reweighting observed outcomes with $1/\\pi_j(Z)$ (IPW and SIPW), by adding a fitted outcome-model residual (AIPW and SAIPW), or by splitting the ECE sample into post-strata on which $\\pi_j$ and $\\pi_k$ are constant and averaging within-stratum means (PS and APS). The asymptotic results hinge on the known probabilities and on Assumption 2, which requires the fitted working model to converge to a fixed limit so misspecified outcome models do not break consistency.","core_discovery":"The paper's central claim is that a treatment effect in a platform trial should be defined on the ECE population, the set of individuals with $\\pi_j(Z)>0$ and $\\pi_k(Z)>0$ for the two treatments $j$ and $k$, where $Z$ is the observed baseline variable governing the known randomization probabilities $\\pi_j(Z)$. Conditional on this population, the estimand is $\\vartheta_{jk}=(E[Y(j)\\mid \\text{ECE}], E[Y(k)\\mid \\text{ECE}])^T$, and the treatment effect is a contrast such as $\\theta_{jk}-\\theta_{kj}$. Under Assumption 1 (conditional randomization given $Z$ with known probabilities) and Assumption 2 (stability of the working outcome model, when used), all six estimators are consistent and asymptotically normal with explicit covariance matrices; the stabilized and covariate-adjusted versions are asymptotically at least as efficient, and AIPW and SAIPW attain the semiparametric efficiency bound when the working models are correctly specified. The paper also proves that post-stratification by strata in which $\\pi_j$ and $\\pi_k$ are constant is asymptotically equivalent to stabilized augmented weighting with the strata as covariates, and that adjusted post-stratification matches stabilized augmented weighting under condition (4).","pith_inferences":["One testable prediction of the invariance claim is that two trial formats with the same ECE support but different randomization ratios should produce estimates that agree up to sampling error; a simulation or reanalysis varying only the ratios could check this directly.","The reliance on known $\\pi_j(Z)$ suggests a practical sensitivity check: compare estimates using the protocol probabilities with estimates using sample-proportion probabilities, and treat systematic divergence as evidence that the recorded randomization mechanism is incomplete.","The ECE population could serve as a natural benchmark for quantifying how much nonconcurrent-control borrowing changes the estimand, giving a bias-variance trade-off for methods that use nonconcurrent data."],"forward_implications":["Changing randomization ratios over time, as in the 1:1 to 3:1 shift in the SIMPLIFY trial, no longer changes the population the estimated treatment effect refers to.","All six estimators are consistent and asymptotically normal with explicit covariance matrices and robust variance estimators, so the framework is directly usable for confirmatory analysis.","Model-assisted adjustment is safe under misspecification: AIPW and SAIPW remain asymptotically unbiased when the working outcome model is wrong, and under conditions (4)-(6) they are asymptotically at least as efficient as weighting or post-stratification alone.","Post-stratification should be built on strata in which the two assignment probabilities are constant, not on all joint levels of $Z$; redundant strata can cause small-stratum instability.","Sub-study-only analyses in umbrella and platform trials target a different, design-dependent population; in the SIMPLIFY application this produced a different point estimate for the DA comparison and wider confidence intervals than the ECE-based estimators."],"supporting_citations":[{"why":"Defines the estimand framework and the requirement to specify the target population, motivating the ECE population proposal.","marker":"ICH E9 (R1), 2019"},{"why":"Draft master-protocol guidance recommending concurrently randomized eligible populations, which the paper formalizes as the ECE population.","marker":"FDA (2023b)"},{"why":"Supplies the inverse probability weighting identification formula used by the IPW estimator.","marker":"Horvitz and Thompson (1952)"},{"why":"Gives the normalized weighting scheme used by the stabilized IPW estimator.","marker":"Hajek (1971)"},{"why":"Provides the balancing property of assignment probabilities used to justify conditioning on $\\pi_j(Z)$ and on post-strata $S$.","marker":"Rosenbaum and Rubin (1983)"},{"why":"Supplies augmented inverse probability weighting and its efficiency theory, which the paper extends to the ECE estimand.","marker":"Robins et al. (1994)"},{"why":"Provides the ANHECOVA working model whose score equations imply conditions (4)-(6) for guaranteed efficiency gains.","marker":"Ye et al. (2023)"},{"why":"Provides joint calibration, another construction under which conditions (4)-(6) hold.","marker":"Bannick et al. (2025)"},{"why":"Prior stratification by enrollment windows, which the proposed post-stratification generalizes by stratifying on the constant values of $\\pi_j$ and $\\pi_k$.","marker":"Marschner and Schou (2022)"},{"why":"The SIMPLIFY trial data and its non-inferiority conclusions, which the application section re-estimates under the ECE estimand.","marker":"Mayer-Hamblett et al. (2023)"}],"fun_headline_variants":["Define effects on ECE population to free platform trials from design quirks","ECE estimand frees platform-trial effects from randomization ratios","Design-independent estimand for robust platform-trial inference","Entire concurrent eligibility: key to robust platform-trial effects"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed baseline variable $Z$ fully captures the randomization mechanism: conditional on $Z$, treatment assignment is independent of potential outcomes and covariates, and the known probabilities $\\pi_j(Z)$ are correct; if $Z$ omits anything the randomization actually used, such as site blocks, randomization lists, or unrecorded enrollment time, the ECE support, weights, and strata are misspecified and the asymptotic results do not apply.","fun_headline_variants_meta":{"raw":{"variants":["Define effects on ECE population to free platform trials from design quirks","ECE estimand frees platform-trial effects from randomization ratios","Design-independent estimand for robust platform-trial inference","Entire concurrent eligibility: key to robust platform-trial effects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1634,"prompt_tokens":1071,"completion_tokens":563,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":687,"completion_tokens_details":{"reasoning_tokens":492}},"tokens_in":687,"tokens_out":563,"duration_ms":6188,"temperature":1.0,"reasoning_tokens":492,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:02:11.911766+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a platform trial in which randomization actually depends on an unobserved enrollment-time or site variable not included in $Z$, then apply the proposed IPW and PS estimators using the protocol probabilities; if their confidence intervals fail to cover the true ECE treatment effect at the nominal rate, or the estimates depart from the truth beyond sampling error, the central claim is falsified.","supporting_citations":[{"cited_title":"Addendum on estimands and sensitivity analysis in clinical trials to the guideline on statistical principles for clinical trials E9(R1)","cited_arxiv_id":null,"evidence_quote":"Defines the estimand framework and the requirement to specify the target population, motivating the ECE population proposal."},{"cited_title":"S., Shao, J., Liu, J., Du, Y., Yi, Y., and Ye, T","cited_arxiv_id":null,"evidence_quote":"Provides joint calibration, another construction under which conditions (4)-(6) hold."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior stratification by enrollment windows, which the proposed post-stratification generalizes by stratifying on the constant values of $\\pi_j$ and $\\pi_k$."},{"cited_title":"H., Riekert, K","cited_arxiv_id":null,"evidence_quote":"The SIMPLIFY trial data and its non-inferiority conclusions, which the application section re-estimates under the ECE estimand."}],"review_version":1}