{"id":"e713f5fe-721f-4965-bbb1-7f45ceb2c570","arxiv_id":"2504.14109","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In multi-intervention stepped-wedge trials, the constant effect estimator is a non-uniform weighted average of exposure-time-specific effects, so it can be biased and even opposite in sign to all true effects.","lead":"This paper studies stepped-wedge clustered trials that roll out multiple interventions over time, and shows the standard constant-effect estimator is a weighted average of exposure-time-specific effects when the true effects change with time. The paper gives formulas for the resulting bias and compares two time-varying models, which is useful because many pragmatic trials assume constant effects without checking.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 1's closed-form H for concurrent designs is internally inconsistent: for m=1 it contradicts the paper's own single-intervention formula, and for m=2 it violates Corollary 1's unbiasedness under constant effects.","rationale":"Reviewing the derivation, Theorem 1's H is a correct application of partitioned GLS: with constant n and common variance components, Sigma_i^{-1}=Q is identical across clusters, and the simplification to Xbar and ybar is valid. The problem is the closed-form evaluation in Proposition 1. The special-case m=1 formula in the paper is correct (verified by direct FWL computation), and it cannot be obtained from Proposition 1's c, d, g, r, v; for T=3 the two disagree for every b>0. The m=2 constant-effect check shows a violation of Corollary 1, an internal contradiction independent of any external standard. This is load-bearing because the paper's headline contribution is the explicit weight matrix; Figure 6 and the bias discussion in Section 4.1 are derived from it. The simulations and the general Theorem 1 may be salvageable, and the qualitative phenomenon (non-uniform weights, cross-contamination) still appears in the correct H, but the manuscript as posted asserts an incorrect closed form and should not be accepted without correction. The Reader's conditional verdict focused on missing supplement and scope assumptions; this finding is a more direct mathematical error, so I disagree with the Reader's weakest-assumption identification and recommend rejection of the current version.","tokens_in":26338,"tokens_out":45167,"duration_ms":368714,"concrete_test":"Set T=3, b=0.2, m=1 in Proposition 1 and compare the resulting row vector with the paper's single-intervention formula directly below Proposition 1. The check is: Proposition 1 yields (2.75, 1.00) as coefficients on (delta_1, delta_2), while the special-case formula and a direct GLS computation on the two-cluster design both give (1.25, -0.25). If the vectors differ, Proposition 1 is internally inconsistent.","verdict_should_be":"REJECT","load_bearing_attack":"Proposition 1 (Section 4.1) is not a valid special case of Theorem 1. For T=3, b=0.2, m=1, evaluating the stated H gives E(theta)=2.75 delta_1 + 1.00 delta_2, whereas the paper's own single-intervention formula immediately below Proposition 1 -- and direct GLS/FWL computation -- gives 1.25 delta_1 - 0.25 delta_2. Symbolically, for T=3, Proposition 1 simplifies to weights ((1+6b)/(1-b), 4b/(1-b)) instead of (1/(1-b), -b/(1-b)). For m=2, T=3, b=0.2, Proposition 1 yields E(theta_1)=1.875 delta_11 + 0.775 delta_12 + 0.875 delta_21 + 0.225 delta_22; under constant effects this equals 2.65 delta_1 + 1.10 delta_2, contradicting Corollary 1's unbiasedness claim. Direct evaluation of H from Theorem 1 for this same design gives [0.975, 0.025, 0.275, -0.275], with own weights summing to 1 and cross weights summing to 0. Because Eq. (8) and Figure 6 are built on Proposition 1, the quantitative bias predictions for concurrent designs are unreliable; the central qualitative claim may survive, but the stated closed-form result is wrong.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops methodology for stepped-wedge cluster-randomized trials with multiple interventions when treatment effects vary by exposure time. The authors state a linear mixed model with exposure-time-specific fixed effects (model (4)) and a working random-effects variant (model (6)), and derive the expectation of the constant-effect GLS estimator under model (4). Theorem 1 expresses E(theta-hat) = H delta. Proposition 1 gives a closed-form H for concurrent designs, Proposition 2 gives a block form for factorial designs with a fully explicit T=3 example, and Corollary 1 gives unbiasedness conditions. Two simulation studies compare constant, time-varying fixed, and random-effects estimators and compare power of concurrent and factorial designs. The method is illustrated on the PONDER-ICU trial. The central qualitative claim is that the constant-effect estimator is a weighted sum of exposure-time-specific effects with possibly non-uniform or negative weights and cross-intervention contamination, so it is generally biased for the exposure-time-averaged effect.","tokens_in":26604,"tokens_out":14113,"duration_ms":114955,"significance":"Assuming the algebra is correct, this is a useful extension of Kenny et al. to multi-intervention stepped-wedge designs. Theorem 1 is a straightforward GLS expectation identity, but the closed-form weight matrices for concurrent designs are nontrivial and are the main technical contribution. I spot-checked Proposition 1 numerically: for T=3, b=0.2, m=1 it yields (1.25, -0.25), matching the paper's own single-intervention formula, and for m=2 it yields own-effect row sums of 1 and cross-effect row sums of 0, consistent with Corollary 1. The alleged internal inconsistency in Proposition 1 therefore does not reproduce, and appears to be based on an algebraic slip in the stress-test note. The simulation studies are extensive, with 500 replicates, bootstrap inference, and publicly available R code, and the coverage results are a useful caution about random-effect working models. The main weakness is that the factorial-design definition and the example and simulation designs do not cohere, which needs repair before the factorial-design claims can be taken at face value.","major_comments":[{"comment":"The T=3 closed-form example uses four design matrices X1-X4, but the factorial design defined in Section 2.2.3 has I = m(T-2) = 2 clusters for T=3. The sequences X2 and X3, in which only one intervention is ever introduced and the second intervention never starts, are not factorial sequences under that definition. Please reconcile the definition with the example, or explicitly state that the example uses a different, augmented factorial design.","section":"Section 4.2"},{"comment":"Section 5.1.2 states that both concurrent and factorial simulations use 2(T-1) clusters, which for T=5 gives 8 clusters, whereas the factorial design of Section 2.2.3 has 6 clusters. Section 5.2.2 explicitly describes augmenting the factorial design with two extra clusters, but Section 5.1.2 does not mention this augmentation. This leaves the design behind Tables 3 and 4 and Figure 8 unspecified; please clarify the cluster configuration used in each simulation and state whether the same augmentation underlies the Section 4.2 example.","section":"Section 5.1.2"}],"minor_comments":[{"comment":"There are typos: 'Intervetion' appears twice in the description of the large effect size scenarios, and in Section 4.2 the text says 'Delta1 = 5/2' where 'Delta2 = 5/2' is clearly intended.","section":"Section 5.1.2"},{"comment":"The phrase 'weighted average' is used for E(theta-hat_k), but the weights can be negative and need not lie in [0,1]; consider saying 'weighted sum' or defining 'average' to allow negative weights.","section":"Section 4.1, Eq. (8)"},{"comment":"The heading 'Large-sample behavior' is a misnomer: Theorem 1 and Proposition 1 give exact finite-sample expectations for GLS with known variance components summarized by b. Please state this assumption explicitly and note how estimation of b might affect the closed-form weights in practice.","section":"Section 4"},{"comment":"The proofs of Theorem 1 and Propositions 1-2 are said to be in the Supplemental Materials, but the supplementary file is not included in the arXiv posting. The revision should ensure the proofs are freely available, since the closed-form weights are the main technical content.","section":"Supporting Information"},{"comment":"For the factorial design with T=11, the second intervention is introduced at the fourth exposure period of the first intervention; please confirm this timing rule and state explicitly how many clusters are used for the factorial design in this setting, in light of the Section 2.2.3 definition.","section":"Section 5.1.2"}],"recommendation":"major_revision","confidential_remarks":"The central derivation appears sound, and the refereed stress-test concern about Proposition 1 does not reproduce on numerical checking. The main issue is a mismatch between the stated factorial design definition and the closed-form example and simulation designs; this is fixable but needs to be addressed before the factorial-design claims are reliable. I would also encourage the editor to require the supplementary proofs to be included with the revision, as the current version references an unavailable file. The paper should be of interest to the statistical methods audience for stepped-wedge trials."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me skip the throat-clearing: the genuinely new material is Theorem 1, the GLS expectation identity for multi-intervention stepped-wedge designs, plus the factorial T=3 closed form and the random-effect working model. The simulations are extensive, the PONDER example is a reasonable illustration, and the paper states its scope conditions (additivity, no interactions, balanced cluster-period size) clearly.\n\nBut Proposition 1, the headline closed-form for concurrent designs, is wrong. I checked T=3, b=0.2, m=1: plugging into Proposition 1 gives E(theta)=2.75 delta_1 + 1.00 delta_2. The paper's own single-intervention formula immediately below gives 1.25 delta_1 - 0.25 delta_2, and direct GLS computation matches the latter. For m=2, T=3, b=0.2, Proposition 1 gives E(theta_1)=1.875 delta_11 + 0.775 delta_12 + 0.875 delta_21 + 0.225 delta_22. Under constant effects this equals 2.65 delta_1 + 1.10 delta_2, contradicting Corollary 1's unbiasedness claim. So Equation (8) and Figure 6, which are built on Proposition 1, have unreliable quantitative predictions for concurrent designs. The qualitative phenomenon—non-uniform weights and cross-intervention contamination—probably survives, but the stated formula is not a valid special case of Theorem 1.\n\nThe other soft spots are smaller. All technical proofs are deferred to a supplement that is not in the preprint, so I could not verify Theorem 1 or Proposition 2 algebraically. The PONDER analysis is not reproducible from the preprint alone. Both are fixable by posting the supplement and code/data.\n\nMy recommendation: send to peer review, but only as a major revision. The topic matters, Theorem 1 and the factorial results are likely salvageable, and the simulations give useful guidance. But the concurrent closed form has to be corrected and verified before anyone should use Equation (8) or Figure 6. This is exactly the kind of paper where referees should be asked to check the algebra, not just the story.","headline":"Extends exposure-time bias results to multiple interventions, but the concurrent-design closed form (Proposition 1) is internally inconsistent and must be fixed before the paper can be used.","tokens_in":27157,"tokens_out":13710,"would_cite":false,"duration_ms":109326,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In multi-intervention stepped-wedge trials, the constant-effect estimator is generally a biased weighted average of exposure-time effects and can even have the wrong sign.","keywords":["stepped-wedge cluster-randomized trials","multiple interventions","time-varying treatment effects","exposure time","linear mixed effects models","model misspecification","concurrent design","factorial design"],"falsifier":"Simulate data from a factorial stepped-wedge design in which the two interventions interact, fit the constant-effect model, and compare the Monte Carlo expectation of the estimator with the $\\mathbf{H}\\boldsymbol{\\delta}$ formula computed under the additive no-interaction model; a systematic divergence would show the no-interaction premise is load-bearing. A simpler check is to simulate under model (4) with unequal cluster-period sizes and see whether the closed-form H from Section 4 still matches the empirical expectation, since the derivation fixes $n_{ij}=n$.","tokens_in":26118,"feed_emoji":"📊","tokens_out":10668,"duration_ms":84588,"temperature":0.7,"pith_summary":"Stepped-wedge cluster-randomized trials are often analyzed under a constant treatment effect, but an intervention's effect may grow, lag, or taper with cumulative exposure. This paper shows that when multiple interventions are rolled out in concurrent or factorial stepped-wedge designs and the true effects depend on exposure time, the standard constant-effect estimator does not recover the exposure-time-averaged effect. Instead, it converges to a weighted average of exposure-time-specific effects, with weights that depend on the design, the variance components, and the number of periods, plus contamination from co-rolled interventions. The bias can be large enough to reverse the apparent direction of an effect, and coverage of confidence intervals can fall far below nominal. The paper then offers two time-varying models, a fixed exposure-time model and a random-effects working model, and evaluates them by simulation and on the PONDER-ICU trial.","feed_headline":"Stepped-wedge estimates can reverse sign when effects vary over time","feed_subtitle":"Standard constant-effect analyses can report the wrong sign for an intervention when effects change with exposure time.","key_machinery":"The load-bearing object is the weight matrix $\\mathbf{H}$ in Theorem 1, defined as $\\mathbf{H} = [\\sum_i \\mathbf{W}(\\mathbf{X}_i) - \\mathbf{S}\\bar{\\mathbf{X}}]^{-1}[\\sum_i \\mathbf{W}(\\mathbf{X}_i,\\mathbf{Z}_i) - \\mathbf{S}\\bar{\\mathbf{Z}}]$, where $\\mathbf{W}(\\mathbf{X}_i) = \\mathbf{X}_i'\\boldsymbol{\\Sigma}_i^{-1}\\mathbf{X}_i$ and $\\mathbf{W}(\\mathbf{X}_i,\\mathbf{Z}_i) = \\mathbf{X}_i'\\boldsymbol{\\Sigma}_i^{-1}\\mathbf{Z}_i$. It converts the true exposure-time-specific effect vector $\\boldsymbol{\\delta}$ into the expectation of the misspecified constant-effect estimator. For concurrent designs the paper derives $\\mathbf{H} = \\frac{1}{c}(\\mathbf{I}_m + \\frac{d}{g}\\mathbf{J}_m) \\otimes \\mathbf{r}' - \\frac{1}{g}\\mathbf{J}_m \\otimes \\mathbf{v}'$, where $\\mathbf{r}$ and $\\mathbf{v}$ are vectors of period-dependent weights depending on a variance-ratio parameter $b$; for factorial designs $\\mathbf{H}$ has the block form $\\begin{pmatrix} \\mathbf{h}_1' & \\mathbf{h}_2' \\\\ \\mathbf{h}_2' & \\mathbf{h}_1' \\end{pmatrix}$. The matrix encodes both non-uniform weighting across exposure periods and cross-intervention contamination.","core_discovery":"The central claim is Theorem 1: if the data are generated by the exposure-time-dependent model with a cluster random intercept, then the generalized least squares estimator from the constant-effect model has expectation $E(\\hat{\\boldsymbol{\\theta}}) = \\mathbf{H}\\boldsymbol{\\delta}$, where $\\boldsymbol{\\delta}$ stacks the exposure-time-specific effects of all interventions. In a concurrent design, $\\mathbf{H}$ has an explicit Kronecker form in which each intervention's estimate is a weighted sum of its own exposure-time effects plus a remainder that aggregates the other interventions' effects, and the weights are not uniform across exposure periods. In a factorial design, $\\mathbf{H}$ has a symmetric two-block structure, so each intervention's estimate mixes its own exposure-time trajectory with the other intervention's. Consequently the constant-effect estimator is generally biased for $\\Delta_k$, the arithmetic average over exposure periods, and simulations show that with lagged effects the bias can flip the estimated sign. Only when every intervention's effect is truly constant does the estimator become unbiased.","pith_inferences":["The paper does not propose a bias-corrected estimator, but the closed-form H suggests a plug-in correction: estimate the variance-ratio parameter b and subtract the contaminating term from the constant-effect estimate to recover an approximately unbiased estimator of the exposure-time-averaged effect.","The factorial closed forms are derived only for two interventions; a natural extension is that H for factorial designs with more than two interventions retains a symmetric block pattern, but the exact structure is not given in the paper.","If treatment effects vary by calendar time as well as exposure time, the bias formulas would need a second dimension; the paper lists this as future work, and the contamination between interventions is likely to be larger in that setting.","In the PONDER-ICU analysis the time-varying fixed model gives opposite-signed estimates from the other two models; the authors attribute this to instability, but an equally plausible reading is that the constant and random models share a misspecification bias that the fixed model avoids at the cost of variance."],"forward_implications":["For a concurrent design with m interventions, the constant-effect estimate for intervention k is a weighted average of its own exposure-time effects plus a term aggregating the other interventions' effects, so the estimate for one arm can be distorted by the temporal trajectory of a different arm.","Unbiasedness of the constant-effect estimator for intervention k fails even if intervention k has a constant effect, as long as the other interventions' effects vary with exposure time; full unbiasedness requires all interventions to have constant effects.","Ignoring exposure-time heterogeneity can produce severe undercoverage of nominal 95% confidence intervals; in the simulations coverage drops to 0% for a half-period-lagged effect with T=11 and the large effect-size scenario.","A time-varying fixed treatment effect model is robust across effect shapes in the simulations, while a random treatment effect model is efficient but can be badly biased when the trajectory is lagged or curved, so the paper recommends sensitivity analyses under both.","Concurrent designs show modestly higher power than factorial designs under time-varying fixed effects in the studied scenarios, with differences typically within 5 to 10 percentage points."],"supporting_citations":[{"why":"Introduces the standard constant-effect linear mixed model for stepped-wedge trials that the paper treats as the misspecified baseline.","marker":"[2]"},{"why":"Establishes the single-intervention result that the constant-effect estimator is a weighted sum of exposure-time effects, which the paper extends to multiple interventions.","marker":"[5]"},{"why":"Provides the categorical exposure-time treatment effect specification that underlies the paper's time-varying fixed treatment effect model.","marker":"[7]"},{"why":"Supplies the random treatment effect working model and bootstrap inference approach that the paper adapts to multi-intervention designs.","marker":"[9]"},{"why":"Defines the concurrent, supplementation, and factorial multiple-intervention stepped-wedge design variations used throughout the paper.","marker":"[12]"},{"why":"Supplies the structured simulation evaluation approach used to design the paper's simulation studies.","marker":"[17]"},{"why":"Supports the claim that inference remains valid under random-effects misspecification when cluster-robust sandwich variance estimators are used.","marker":"[18]"}],"fun_headline_variants":["Constant-effect estimates can flip sign in stepped-wedge trials","Time-varying effects bias stepped-wedge estimates, even flip sign","When effects change over time, stepped-wedge estimates reverse","Stepped-wedge analysis misleads when treatment effect varies","Sign reversal risk in stepped-wedge trials with time-varying effects"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The derivation assumes the true effect of each intervention depends only on how long the cluster has been exposed, that interventions do not interact, that the cluster-level and residual variances are known, and that every cluster-period has the same number of people; if any of these fail, the stated bias formulas need not hold.","fun_headline_variants_meta":{"raw":{"variants":["Constant-effect estimates can flip sign in stepped-wedge trials","Time-varying effects bias stepped-wedge estimates, even flip sign","When effects change over time, stepped-wedge estimates reverse","Stepped-wedge analysis misleads when treatment effect varies","Sign reversal risk in stepped-wedge trials with time-varying effects"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":3017,"prompt_tokens":1037,"completion_tokens":1980,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":1896}},"tokens_in":653,"tokens_out":1980,"duration_ms":11835,"temperature":1.0,"reasoning_tokens":1896,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:56:35.184215+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data from a factorial stepped-wedge design in which the two interventions interact, fit the constant-effect model, and compare the Monte Carlo expectation of the estimator with the $\\mathbf{H}\\boldsymbol{\\delta}$ formula computed under the additive no-interaction model; a systematic divergence would show the no-interaction premise is load-bearing. A simpler check is to simulate under model (4) with unequal cluster-period sizes and see whether the closed-form H from Section 4 still matches the empirical expectation, since the derivation fixes $n_{ij}=n$.","supporting_citations":[{"cited_title":"Design and analysis of stepped wedge cluster randomized trials.Contemporary Clinical Trials 2007; 28(2): 182-191","cited_arxiv_id":null,"evidence_quote":"Introduces the standard constant-effect linear mixed model for stepped-wedge trials that the paper treats as the misspecified baseline."},{"cited_title":"Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect.Statistics in Medicine 2022; 41(22): 4311-4339","cited_arxiv_id":null,"evidence_quote":"Establishes the single-intervention result that the constant-effect estimator is a weighted sum of exposure-time effects, which the paper extends to multiple interventions."},{"cited_title":"Current issues in the design and analysis of stepped wedge trials.Contemporary Clinical Trials2015; 45","cited_arxiv_id":null,"evidence_quote":"Provides the categorical exposure-time treatment effect specification that underlies the paper's time-varying fixed treatment effect model."},{"cited_title":"Assessing exposure-time treatment effect heterogeneity in stepped-wedge cluster randomized trials.Biometrics2023; 79(3): 2551-2564","cited_arxiv_id":null,"evidence_quote":"Supplies the random treatment effect working model and bootstrap inference approach that the paper adapts to multi-intervention designs."},{"cited_title":"Proposed variations of the stepped-wedge design can be used to accommodate multiple interventions.Journal of Clinical Epidemiology 2017; 86: 160-167","cited_arxiv_id":null,"evidence_quote":"Defines the concurrent, supplementation, and factorial multiple-intervention stepped-wedge design variations used throughout the paper."},{"cited_title":"Using simulation studies to evaluate statistical methods.Statistics in medicine 2019; 38(11): 2074–2102","cited_arxiv_id":null,"evidence_quote":"Supplies the structured simulation evaluation approach used to design the paper's simulation studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that inference remains valid under random-effects misspecification when cluster-robust sandwich variance estimators are used."}],"review_version":1}