REVIEW 2 major objections 5 minor 20 references
Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions
T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read In multi-intervention stepped-wedge trials, the constant-effect estimator is generally a biased weighted average of exposure-time effects and can even have the wrong sign.
desk verdict Extends exposure-time bias results to multiple interventions, but the concurrent-design closed form (Proposition 1) is internally inconsistent and must be fixed before the paper can be used. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the weight matrix $\mathbf{H}$ in Theorem 1, defined as $\mathbf{H} = [\sum_i \mathbf{W}(\mathbf{X}_i) - \mathbf{S}\bar{\mathbf{X}}]^{-1}[\sum_i \mathbf{W}(\mathbf{X}_i,\mathbf{Z}_i) - \mathbf{S}\bar{\mathbf{Z}}]$, where $\mathbf{W}(\mathbf{X}_i) = \mathbf{X}_i'\boldsymbol{\Sigma}_i^{-1}\mathbf{X}_i$ and $\mathbf{W}(\mathbf{X}_i,\mathbf{Z}_i) = \mathbf{X}_i'\boldsymbol{\Sigma}_i^{-1}\mathbf{Z}_i$. It converts the true exposure-time-specific effect vector $\boldsymbol{\delta}$ into the expectation of the misspecified constant-effect estimator. For concurrent designs the paper derives $\mathbf{H} = \frac{1}{c}(\mathbf{I}_m + \frac{d}{g}\mathbf{J}_m) \otimes \mathbf{r}' - \frac{1}{g}\mathbf{J}_m \otimes \mathbf{v}'$, where $\mathbf{r}$ and $\mathbf{v}$ are vectors of period-dependent weights depending on a variance-ratio parameter $b$; for factorial designs $\mathbf{H}$ has the block form $\begin{pmatrix} \mathbf{h}_1' & \mathbf{h}_2' \\ \mathbf{h}_2' & \mathbf{h}_1' \end{pmatrix}$. The matrix encodes both non-uniform weighting across exposure periods and cross-intervention contamination.
What would settle it
Simulate data from a factorial stepped-wedge design in which the two interventions interact, fit the constant-effect model, and compare the Monte Carlo expectation of the estimator with the $\mathbf{H}\boldsymbol{\delta}$ formula computed under the additive no-interaction model; a systematic divergence would show the no-interaction premise is load-bearing. A simpler check is to simulate under model (4) with unequal cluster-period sizes and see whether the closed-form H from Section 4 still matches the empirical expectation, since the derivation fixes $n_{ij}=n$.
Extended reading notes
Core claim
The central claim is Theorem 1: if the data are generated by the exposure-time-dependent model with a cluster random intercept, then the generalized least squares estimator from the constant-effect model has expectation $E(\hat{\boldsymbol{\theta}}) = \mathbf{H}\boldsymbol{\delta}$, where $\boldsymbol{\delta}$ stacks the exposure-time-specific effects of all interventions. In a concurrent design, $\mathbf{H}$ has an explicit Kronecker form in which each intervention's estimate is a weighted sum of its own exposure-time effects plus a remainder that aggregates the other interventions' effects, and the weights are not uniform across exposure periods. In a factorial design, $\mathbf{H}$ has a symmetric two-block structure, so each intervention's estimate mixes its own exposure-time trajectory with the other intervention's. Consequently the constant-effect estimator is generally biased for $\Delta_k$, the arithmetic average over exposure periods, and simulations show that with lagged effects the bias can flip the estimated sign. Only when every intervention's effect is truly constant does the estimator become unbiased.
Load-bearing premise
The derivation assumes the true effect of each intervention depends only on how long the cluster has been exposed, that interventions do not interact, that the cluster-level and residual variances are known, and that every cluster-period has the same number of people; if any of these fail, the stated bias formulas need not hold.
Editorial extensions
If this is right
- For a concurrent design with m interventions, the constant-effect estimate for intervention k is a weighted average of its own exposure-time effects plus a term aggregating the other interventions' effects, so the estimate for one arm can be distorted by the temporal trajectory of a different arm.
- Unbiasedness of the constant-effect estimator for intervention k fails even if intervention k has a constant effect, as long as the other interventions' effects vary with exposure time; full unbiasedness requires all interventions to have constant effects.
- Ignoring exposure-time heterogeneity can produce severe undercoverage of nominal 95% confidence intervals; in the simulations coverage drops to 0% for a half-period-lagged effect with T=11 and the large effect-size scenario.
- A time-varying fixed treatment effect model is robust across effect shapes in the simulations, while a random treatment effect model is efficient but can be badly biased when the trajectory is lagged or curved, so the paper recommends sensitivity analyses under both.
- Concurrent designs show modestly higher power than factorial designs under time-varying fixed effects in the studied scenarios, with differences typically within 5 to 10 percentage points.
Reading between the lines
- The paper does not propose a bias-corrected estimator, but the closed-form H suggests a plug-in correction: estimate the variance-ratio parameter b and subtract the contaminating term from the constant-effect estimate to recover an approximately unbiased estimator of the exposure-time-averaged effect.
- The factorial closed forms are derived only for two interventions; a natural extension is that H for factorial designs with more than two interventions retains a symmetric block pattern, but the exact structure is not given in the paper.
- If treatment effects vary by calendar time as well as exposure time, the bias formulas would need a second dimension; the paper lists this as future work, and the contamination between interventions is likely to be larger in that setting.
- In the PONDER-ICU analysis the time-varying fixed model gives opposite-signed estimates from the other two models; the authors attribute this to instability, but an equally plausible reading is that the constant and random models share a misspecification bias that the fixed model avoids at the cost of variance.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops methodology for stepped-wedge cluster-randomized trials with multiple interventions when treatment effects vary by exposure time. The authors state a linear mixed model with exposure-time-specific fixed effects (model (4)) and a working random-effects variant (model (6)), and derive the expectation of the constant-effect GLS estimator under model (4). Theorem 1 expresses E(theta-hat) = H delta. Proposition 1 gives a closed-form H for concurrent designs, Proposition 2 gives a block form for factorial designs with a fully explicit T=3 example, and Corollary 1 gives unbiasedness conditions. Two simulation studies compare constant, time-varying fixed, and random-effects estimators and compare power of concurrent and factorial designs. The method is illustrated on the PONDER-ICU trial. The central qualitative claim is that the constant-effect estimator is a weighted sum of exposure-time-specific effects with possibly non-uniform or negative weights and cross-intervention contamination, so it is generally biased for the exposure-time-averaged effect.
Significance. Assuming the algebra is correct, this is a useful extension of Kenny et al. to multi-intervention stepped-wedge designs. Theorem 1 is a straightforward GLS expectation identity, but the closed-form weight matrices for concurrent designs are nontrivial and are the main technical contribution. I spot-checked Proposition 1 numerically: for T=3, b=0.2, m=1 it yields (1.25, -0.25), matching the paper's own single-intervention formula, and for m=2 it yields own-effect row sums of 1 and cross-effect row sums of 0, consistent with Corollary 1. The alleged internal inconsistency in Proposition 1 therefore does not reproduce, and appears to be based on an algebraic slip in the stress-test note. The simulation studies are extensive, with 500 replicates, bootstrap inference, and publicly available R code, and the coverage results are a useful caution about random-effect working models. The main weakness is that the factorial-design definition and the example and simulation designs do not cohere, which needs repair before the factorial-design claims can be taken at face value.
major comments (2)
- [Section 4.2] The T=3 closed-form example uses four design matrices X1-X4, but the factorial design defined in Section 2.2.3 has I = m(T-2) = 2 clusters for T=3. The sequences X2 and X3, in which only one intervention is ever introduced and the second intervention never starts, are not factorial sequences under that definition. Please reconcile the definition with the example, or explicitly state that the example uses a different, augmented factorial design.
- [Section 5.1.2] Section 5.1.2 states that both concurrent and factorial simulations use 2(T-1) clusters, which for T=5 gives 8 clusters, whereas the factorial design of Section 2.2.3 has 6 clusters. Section 5.2.2 explicitly describes augmenting the factorial design with two extra clusters, but Section 5.1.2 does not mention this augmentation. This leaves the design behind Tables 3 and 4 and Figure 8 unspecified; please clarify the cluster configuration used in each simulation and state whether the same augmentation underlies the Section 4.2 example.
minor comments (5)
- [Section 5.1.2] There are typos: 'Intervetion' appears twice in the description of the large effect size scenarios, and in Section 4.2 the text says 'Delta1 = 5/2' where 'Delta2 = 5/2' is clearly intended.
- [Section 4.1, Eq. (8)] The phrase 'weighted average' is used for E(theta-hat_k), but the weights can be negative and need not lie in [0,1]; consider saying 'weighted sum' or defining 'average' to allow negative weights.
- [Section 4] The heading 'Large-sample behavior' is a misnomer: Theorem 1 and Proposition 1 give exact finite-sample expectations for GLS with known variance components summarized by b. Please state this assumption explicitly and note how estimation of b might affect the closed-form weights in practice.
- [Supporting Information] The proofs of Theorem 1 and Propositions 1-2 are said to be in the Supplemental Materials, but the supplementary file is not included in the arXiv posting. The revision should ensure the proofs are freely available, since the closed-form weights are the main technical content.
- [Section 5.1.2] For the factorial design with T=11, the second intervention is introduced at the fourth exposure period of the first intervention; please confirm this timing rule and state explicitly how many clusters are used for the factorial design in this setting, in light of the Section 2.2.3 definition.
Circularity Check
No circularity: the main derivation is a direct GLS expectation identity, and the simulations and case study are independent checks rather than fitted predictions.
full rationale
Theorem 1 derives E(theta-hat)=H delta by substituting the true model E(y_i)=beta+Z_i delta from equation (4) into the closed-form GLS estimator in equation (7). This is a direct expectation identity: H is defined explicitly from the design and covariance matrices, not estimated from the target estimand, and no parameter is fitted to the quantity being predicted. The subsequent Proposition 1, Proposition 2, and Corollary 1 are algebraic specializations of Theorem 1. The random treatment effect model is explicitly labeled a working model in Section 3.3, and Section 7 openly lists limitations and settings where the formulas do not apply, so there is no disguised importation of the conclusion. The simulation studies generate data from specified true exposure-time-specific effects and then compare three fitting models; they do not fit a parameter to the estimand and call the result a prediction. The PONDER application is an external empirical illustration. Citations to prior work, including Kenny et al. and Maleyeff et al., are used for context and as sources of model ideas, but the central bias decomposition is proved from the estimator and model in the paper itself. Even if Proposition 1's closed-form expression were internally inconsistent or numerically wrong, that would be a correctness or accuracy concern, not a circularity concern; the argument does not reduce to its own inputs by definition.
Assumptions & free parameters
free parameters (1)
- b (variance ratio sigma^2_alpha / (T sigma^2_alpha + sigma^2_epsilon / n)) =
not fixed; estimated from data in practice
assumptions (3)
- domain assumption Data follow a linear mixed model with cluster random intercept, independent errors, and correctly specified mean when the fitted model is correct.
- domain assumption Exposure-time effects are additive and interaction-free across interventions.
- domain assumption Constant cluster-period size n and known covariance are used for the closed-form weights.
Cite this review
Pith. "Pith review of Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions." pith.science (2026). https://pith.science/paper/Q6PPEAHX
@misc{pith2026250414109,
author = {Pith},
title = {Pith review of: Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q6PPEAHX}},
note = {Machine review of arXiv:2504.14109}
}
read the original abstract
The traditional model specification of stepped-wedge cluster-randomized trials assumes a homogeneous treatment effect across time while adjusting for fixed-time effects. However, when treatment effects vary over time, the constant effect estimator may be biased. In the general setting of stepped-wedge cluster-randomized trials with multiple interventions, we derive the expected value of the constant effect estimator when the true treatment effects depend on exposure time periods. Applying this result to concurrent and factorial stepped wedge designs, we show that the estimator represents a weighted average of exposure-time-specific treatment effects, with weights that are not necessarily uniform across exposure periods. Extensive simulation studies reveal that ignoring time heterogeneity can result in biased estimates and poor coverage of the average treatment effect. In this study, we examine two models designed to accommodate multiple interventions with time-varying treatment effects: (1) a time-varying fixed treatment effect model, which allows treatment effects to vary by exposure time but remain fixed for each time point, and (2) a random treatment effect model, where the time-varying treatment effects are modeled as random deviations from an overall mean. In the simulations considered in this study, concurrent designs generally achieve higher power than factorial designs under a time-varying fixed treatment effect model, though the differences are modest. Finally, we apply the constant effect model and both time-varying treatment effect models to data from the Prognosticating Outcomes and Nudging Decisions in the Electronic Health Record (PONDER) trial. All three models indicate a lack of treatment effect for either intervention, though they differ in the precision of their estimates, likely due to variations in modeling assumptions.
Reference graph
Works this paper leans on
-
[1]
The Gambia Hepatitis Intervention Study.Cancer Research1987; 47(21): 5782-5787
Group TGHS. The Gambia Hepatitis Intervention Study.Cancer Research1987; 47(21): 5782-5787
-
[2]
Hussey MA, Hughes JP. Design and analysis of stepped wedge cluster randomized trials.Contemporary Clinical Trials 2007; 28(2): 182-191
work page 2007
-
[3]
Li F, Hughes JP, Hemming K, Taljaard M, Melnick ER, Heagerty PJ. Mixed-effects models for the design and analysis of stepped wedge cluster randomized trials: An overview.Statistical Methods in Medical Research 2021; 30(2): 612-639
work page 2021
-
[4]
Zhang Y, Preisser JS, Turner EL, Rathouz PJ, Toles M, Li F. A general method for calculating power for GEE analysis of complete and incomplete stepped wedge cluster randomized trials.Statistical methods in medical research 2023; 32(1): 71–87
work page 2023
-
[5]
Kenny A, Voldal E, Xia F, Heagerty PJ, Hughes JP. Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect.Statistics in Medicine 2022; 41(22): 4311-4339
work page 2022
-
[6]
How to achieve model-robust inference in stepped wedge trials with model-based methods?
Wang B, Wang X, Li F. How to achieve model-robust inference in stepped wedge trials with model-based methods?. Biometrics2024; 80(4): ujae123
-
[7]
Hughes JP, Granston TS, Heagerty PJ. Current issues in the design and analysis of stepped wedge trials.Contemporary Clinical Trials2015; 45
-
[8]
NicklessA,VoyseyM,GeddesJ,YuLM,FanshaweTR.Mixedeffectsapproachtotheanalysisofthesteppedwedgecluster randomised trial-Investigating the confounding effect of time through simulation.PLoS One 2018; 13(12)
work page 2018
Show all 20 references
-
[9]
Assessing exposure-time treatment effect heterogeneity in stepped-wedge cluster randomized trials.Biometrics2023; 79(3): 2551-2564
Maleyeff L, Li F, Haneuse S, Wang R. Assessing exposure-time treatment effect heterogeneity in stepped-wedge cluster randomized trials.Biometrics2023; 79(3): 2551-2564
-
[10]
Statistics in Medicine 2024; 43(1): 49-60
LeeKM,CheungYB.Clusterrandomizedtrialdesignsformodelingtime-varyinginterventioneffects. Statistics in Medicine 2024; 43(1): 49-60
2024
-
[11]
Sample size calculations for stepped wedge designs with treatment effects that may change with the duration of time under intervention.Prevention Science2024; 25(Suppl 3): 348-355
Hughes JP, Lee WY, Troxel AB, Heagerty PJ. Sample size calculations for stepped wedge designs with treatment effects that may change with the duration of time under intervention.Prevention Science2024; 25(Suppl 3): 348-355
-
[12]
Proposed variations of the stepped-wedge design can be used to accommodate multiple interventions.Journal of Clinical Epidemiology 2017; 86: 160-167
Lyons VH, Li L, Hughes JP, Rowhani-Rahbar A. Proposed variations of the stepped-wedge design can be used to accommodate multiple interventions.Journal of Clinical Epidemiology 2017; 86: 160-167
2017
-
[13]
Statistics in Medicine 2017; 36(24): 3772-3790
MatthewsJNS,ForbesAB.Steppedwedgedesigns:insightsfromadesignofexperimentsperspective. Statistics in Medicine 2017; 36(24): 3772-3790
2017
-
[14]
Admissible multiarm stepped-wedge cluster randomized trial designs.Statistics in Medicine 2019; 38(7): 1103-1119
Grayling MJ, Mander AP, Wason JMS. Admissible multiarm stepped-wedge cluster randomized trial designs.Statistics in Medicine 2019; 38(7): 1103-1119
2019
-
[15]
Variance formulae for multiphase stepped wedge cluster randomized trial
Zhang P, Shoben A, Jackson R, Fernandez S. Variance formulae for multiphase stepped wedge cluster randomized trial. Statistics in Medicine 2020; 39(28): 4147-4168. 22 Z ChenET AL
2020
-
[16]
Power analysis for stepped wedge trials with multiple interventions.Statistics in Medicine 2022; 41(8)
Sundin P, Crespi CM. Power analysis for stepped wedge trials with multiple interventions.Statistics in Medicine 2022; 41(8)
2022
-
[17]
Using simulation studies to evaluate statistical methods.Statistics in medicine 2019; 38(11): 2074–2102
Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods.Statistics in medicine 2019; 38(11): 2074–2102
2019
-
[18]
Ouyang Y, Taljaard M, Forbes AB, Li F. Maintaining the validity of inference from linear mixed models in stepped-wedge cluster randomized trials under misspecified random-effects structures.Statistical Methods in Medical Research 2024; 33(9): 1497–1516
2024
-
[19]
Analysis of Stepped-Wedge Cluster Randomized Trials when treatment effect varies by exposure time or calendar time.arXiv preprint arXiv:2409.14706 2024
Lee KM, Turner EL, Kenny A. Analysis of Stepped-Wedge Cluster Randomized Trials when treatment effect varies by exposure time or calendar time.arXiv preprint arXiv:2409.14706 2024
2024 arXiv
-
[20]
Model-assisted analysis of covariance estimators for stepped wedge cluster randomized experiments
Chen X, Li F. Model-assisted analysis of covariance estimators for stepped wedge cluster randomized experiments. Scandinavian Journal of Statistics 2025; 52(1): 416–446
2025
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.