Pith. sign in

REVIEW 2 major objections 5 minor 20 references

Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read In multi-intervention stepped-wedge trials, the constant-effect estimator is generally a biased weighted average of exposure-time effects and can even have the wrong sign.

desk verdict Extends exposure-time bias results to multiple interventions, but the concurrent-design closed form (Proposition 1) is internally inconsistent and must be fixed before the paper can be used. read the letter →

arxiv 2504.14109 v1 pith:Q6PPEAHX submitted 2025-04-18 stat.ME

classification stat.ME
keywords stepped-wedgecluster-randomizedtrialsmultipleinterventionstime-varyingtreatmenteffectsexposuretimelinearmixedmodelsmodelmisspecificationconcurrentdesignfactorial
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Stepped-wedge cluster-randomized trials are often analyzed under a constant treatment effect, but an intervention's effect may grow, lag, or taper with cumulative exposure. This paper shows that when multiple interventions are rolled out in concurrent or factorial stepped-wedge designs and the true effects depend on exposure time, the standard constant-effect estimator does not recover the exposure-time-averaged effect. Instead, it converges to a weighted average of exposure-time-specific effects, with weights that depend on the design, the variance components, and the number of periods, plus contamination from co-rolled interventions. The bias can be large enough to reverse the apparent direction of an effect, and coverage of confidence intervals can fall far below nominal. The paper then offers two time-varying models, a fixed exposure-time model and a random-effects working model, and evaluates them by simulation and on the PONDER-ICU trial.

What carries the argument

The load-bearing object is the weight matrix $\mathbf{H}$ in Theorem 1, defined as $\mathbf{H} = [\sum_i \mathbf{W}(\mathbf{X}_i) - \mathbf{S}\bar{\mathbf{X}}]^{-1}[\sum_i \mathbf{W}(\mathbf{X}_i,\mathbf{Z}_i) - \mathbf{S}\bar{\mathbf{Z}}]$, where $\mathbf{W}(\mathbf{X}_i) = \mathbf{X}_i'\boldsymbol{\Sigma}_i^{-1}\mathbf{X}_i$ and $\mathbf{W}(\mathbf{X}_i,\mathbf{Z}_i) = \mathbf{X}_i'\boldsymbol{\Sigma}_i^{-1}\mathbf{Z}_i$. It converts the true exposure-time-specific effect vector $\boldsymbol{\delta}$ into the expectation of the misspecified constant-effect estimator. For concurrent designs the paper derives $\mathbf{H} = \frac{1}{c}(\mathbf{I}_m + \frac{d}{g}\mathbf{J}_m) \otimes \mathbf{r}' - \frac{1}{g}\mathbf{J}_m \otimes \mathbf{v}'$, where $\mathbf{r}$ and $\mathbf{v}$ are vectors of period-dependent weights depending on a variance-ratio parameter $b$; for factorial designs $\mathbf{H}$ has the block form $\begin{pmatrix} \mathbf{h}_1' & \mathbf{h}_2' \\ \mathbf{h}_2' & \mathbf{h}_1' \end{pmatrix}$. The matrix encodes both non-uniform weighting across exposure periods and cross-intervention contamination.

What would settle it

Simulate data from a factorial stepped-wedge design in which the two interventions interact, fit the constant-effect model, and compare the Monte Carlo expectation of the estimator with the $\mathbf{H}\boldsymbol{\delta}$ formula computed under the additive no-interaction model; a systematic divergence would show the no-interaction premise is load-bearing. A simpler check is to simulate under model (4) with unequal cluster-period sizes and see whether the closed-form H from Section 4 still matches the empirical expectation, since the derivation fixes $n_{ij}=n$.

Watch

Extended reading notes

Core claim

The central claim is Theorem 1: if the data are generated by the exposure-time-dependent model with a cluster random intercept, then the generalized least squares estimator from the constant-effect model has expectation $E(\hat{\boldsymbol{\theta}}) = \mathbf{H}\boldsymbol{\delta}$, where $\boldsymbol{\delta}$ stacks the exposure-time-specific effects of all interventions. In a concurrent design, $\mathbf{H}$ has an explicit Kronecker form in which each intervention's estimate is a weighted sum of its own exposure-time effects plus a remainder that aggregates the other interventions' effects, and the weights are not uniform across exposure periods. In a factorial design, $\mathbf{H}$ has a symmetric two-block structure, so each intervention's estimate mixes its own exposure-time trajectory with the other intervention's. Consequently the constant-effect estimator is generally biased for $\Delta_k$, the arithmetic average over exposure periods, and simulations show that with lagged effects the bias can flip the estimated sign. Only when every intervention's effect is truly constant does the estimator become unbiased.

Load-bearing premise

The derivation assumes the true effect of each intervention depends only on how long the cluster has been exposed, that interventions do not interact, that the cluster-level and residual variances are known, and that every cluster-period has the same number of people; if any of these fail, the stated bias formulas need not hold.

Editorial extensions

If this is right

  • For a concurrent design with m interventions, the constant-effect estimate for intervention k is a weighted average of its own exposure-time effects plus a term aggregating the other interventions' effects, so the estimate for one arm can be distorted by the temporal trajectory of a different arm.
  • Unbiasedness of the constant-effect estimator for intervention k fails even if intervention k has a constant effect, as long as the other interventions' effects vary with exposure time; full unbiasedness requires all interventions to have constant effects.
  • Ignoring exposure-time heterogeneity can produce severe undercoverage of nominal 95% confidence intervals; in the simulations coverage drops to 0% for a half-period-lagged effect with T=11 and the large effect-size scenario.
  • A time-varying fixed treatment effect model is robust across effect shapes in the simulations, while a random treatment effect model is efficient but can be badly biased when the trajectory is lagged or curved, so the paper recommends sensitivity analyses under both.
  • Concurrent designs show modestly higher power than factorial designs under time-varying fixed effects in the studied scenarios, with differences typically within 5 to 10 percentage points.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not propose a bias-corrected estimator, but the closed-form H suggests a plug-in correction: estimate the variance-ratio parameter b and subtract the contaminating term from the constant-effect estimate to recover an approximately unbiased estimator of the exposure-time-averaged effect.
  • The factorial closed forms are derived only for two interventions; a natural extension is that H for factorial designs with more than two interventions retains a symmetric block pattern, but the exact structure is not given in the paper.
  • If treatment effects vary by calendar time as well as exposure time, the bias formulas would need a second dimension; the paper lists this as future work, and the contamination between interventions is likely to be larger in that setting.
  • In the PONDER-ICU analysis the time-varying fixed model gives opposite-signed estimates from the other two models; the authors attribute this to instability, but an equally plausible reading is that the constant and random models share a misspecification bias that the fixed model avoids at the cost of variance.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper develops methodology for stepped-wedge cluster-randomized trials with multiple interventions when treatment effects vary by exposure time. The authors state a linear mixed model with exposure-time-specific fixed effects (model (4)) and a working random-effects variant (model (6)), and derive the expectation of the constant-effect GLS estimator under model (4). Theorem 1 expresses E(theta-hat) = H delta. Proposition 1 gives a closed-form H for concurrent designs, Proposition 2 gives a block form for factorial designs with a fully explicit T=3 example, and Corollary 1 gives unbiasedness conditions. Two simulation studies compare constant, time-varying fixed, and random-effects estimators and compare power of concurrent and factorial designs. The method is illustrated on the PONDER-ICU trial. The central qualitative claim is that the constant-effect estimator is a weighted sum of exposure-time-specific effects with possibly non-uniform or negative weights and cross-intervention contamination, so it is generally biased for the exposure-time-averaged effect.

Significance. Assuming the algebra is correct, this is a useful extension of Kenny et al. to multi-intervention stepped-wedge designs. Theorem 1 is a straightforward GLS expectation identity, but the closed-form weight matrices for concurrent designs are nontrivial and are the main technical contribution. I spot-checked Proposition 1 numerically: for T=3, b=0.2, m=1 it yields (1.25, -0.25), matching the paper's own single-intervention formula, and for m=2 it yields own-effect row sums of 1 and cross-effect row sums of 0, consistent with Corollary 1. The alleged internal inconsistency in Proposition 1 therefore does not reproduce, and appears to be based on an algebraic slip in the stress-test note. The simulation studies are extensive, with 500 replicates, bootstrap inference, and publicly available R code, and the coverage results are a useful caution about random-effect working models. The main weakness is that the factorial-design definition and the example and simulation designs do not cohere, which needs repair before the factorial-design claims can be taken at face value.

major comments (2)
  1. [Section 4.2] The T=3 closed-form example uses four design matrices X1-X4, but the factorial design defined in Section 2.2.3 has I = m(T-2) = 2 clusters for T=3. The sequences X2 and X3, in which only one intervention is ever introduced and the second intervention never starts, are not factorial sequences under that definition. Please reconcile the definition with the example, or explicitly state that the example uses a different, augmented factorial design.
  2. [Section 5.1.2] Section 5.1.2 states that both concurrent and factorial simulations use 2(T-1) clusters, which for T=5 gives 8 clusters, whereas the factorial design of Section 2.2.3 has 6 clusters. Section 5.2.2 explicitly describes augmenting the factorial design with two extra clusters, but Section 5.1.2 does not mention this augmentation. This leaves the design behind Tables 3 and 4 and Figure 8 unspecified; please clarify the cluster configuration used in each simulation and state whether the same augmentation underlies the Section 4.2 example.
minor comments (5)
  1. [Section 5.1.2] There are typos: 'Intervetion' appears twice in the description of the large effect size scenarios, and in Section 4.2 the text says 'Delta1 = 5/2' where 'Delta2 = 5/2' is clearly intended.
  2. [Section 4.1, Eq. (8)] The phrase 'weighted average' is used for E(theta-hat_k), but the weights can be negative and need not lie in [0,1]; consider saying 'weighted sum' or defining 'average' to allow negative weights.
  3. [Section 4] The heading 'Large-sample behavior' is a misnomer: Theorem 1 and Proposition 1 give exact finite-sample expectations for GLS with known variance components summarized by b. Please state this assumption explicitly and note how estimation of b might affect the closed-form weights in practice.
  4. [Supporting Information] The proofs of Theorem 1 and Propositions 1-2 are said to be in the Supplemental Materials, but the supplementary file is not included in the arXiv posting. The revision should ensure the proofs are freely available, since the closed-form weights are the main technical content.
  5. [Section 5.1.2] For the factorial design with T=11, the second intervention is introduced at the fourth exposure period of the first intervention; please confirm this timing rule and state explicitly how many clusters are used for the factorial design in this setting, in light of the Section 2.2.3 definition.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the main derivation is a direct GLS expectation identity, and the simulations and case study are independent checks rather than fitted predictions.

full rationale

Theorem 1 derives E(theta-hat)=H delta by substituting the true model E(y_i)=beta+Z_i delta from equation (4) into the closed-form GLS estimator in equation (7). This is a direct expectation identity: H is defined explicitly from the design and covariance matrices, not estimated from the target estimand, and no parameter is fitted to the quantity being predicted. The subsequent Proposition 1, Proposition 2, and Corollary 1 are algebraic specializations of Theorem 1. The random treatment effect model is explicitly labeled a working model in Section 3.3, and Section 7 openly lists limitations and settings where the formulas do not apply, so there is no disguised importation of the conclusion. The simulation studies generate data from specified true exposure-time-specific effects and then compare three fitting models; they do not fit a parameter to the estimand and call the result a prediction. The PONDER application is an external empirical illustration. Citations to prior work, including Kenny et al. and Maleyeff et al., are used for context and as sources of model ideas, but the central bias decomposition is proved from the estimator and model in the paper itself. Even if Proposition 1's closed-form expression were internally inconsistent or numerically wrong, that would be a correctness or accuracy concern, not a circularity concern; the argument does not reduce to its own inputs by definition.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

No new entities are postulated. The main load-bearing items are standard mixed-model distributional assumptions, the additivity/no-interaction restriction, and the constant cluster-period-size simplification used to get closed-form weights. These are stated in the paper, but they bound the scope of the formulas.

free parameters (1)
  • b (variance ratio sigma^2_alpha / (T sigma^2_alpha + sigma^2_epsilon / n)) = not fixed; estimated from data in practice
    The closed-form weights in Proposition 1 depend on b, which is a function of unknown variance components. In simulations b is set by chosen variances; in applications it must be estimated, so the bias magnitude is not parameter-free.
assumptions (3)
  • domain assumption Data follow a linear mixed model with cluster random intercept, independent errors, and correctly specified mean when the fitted model is correct.
    Models (1) through (6) assume y_ijs = beta_j + sum x_kij delta + alpha_i + epsilon, with alpha_i and epsilon normal and independent; this underpins the GLS expectation calculation in Section 4.
  • domain assumption Exposure-time effects are additive and interaction-free across interventions.
    Identifiability in Section 2.2.3 and model (4) require no interactions; the authors assume the joint effect equals the sum of individual effects in the PONDER analysis.
  • domain assumption Constant cluster-period size n and known covariance are used for the closed-form weights.
    Section 4 sets n_ij = n and defines b in terms of variance components; unequal cluster sizes or estimated variances change the exact weights, though the qualitative result likely remains.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions." pith.science (2026). https://pith.science/paper/Q6PPEAHX

@misc{pith2026250414109,
  author       = {Pith},
  title        = {Pith review of: Time-varying treatment effect models in stepped-wedge cluster-randomized trials with multiple interventions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q6PPEAHX}},
  note         = {Machine review of arXiv:2504.14109}
}
read the original abstract

The traditional model specification of stepped-wedge cluster-randomized trials assumes a homogeneous treatment effect across time while adjusting for fixed-time effects. However, when treatment effects vary over time, the constant effect estimator may be biased. In the general setting of stepped-wedge cluster-randomized trials with multiple interventions, we derive the expected value of the constant effect estimator when the true treatment effects depend on exposure time periods. Applying this result to concurrent and factorial stepped wedge designs, we show that the estimator represents a weighted average of exposure-time-specific treatment effects, with weights that are not necessarily uniform across exposure periods. Extensive simulation studies reveal that ignoring time heterogeneity can result in biased estimates and poor coverage of the average treatment effect. In this study, we examine two models designed to accommodate multiple interventions with time-varying treatment effects: (1) a time-varying fixed treatment effect model, which allows treatment effects to vary by exposure time but remain fixed for each time point, and (2) a random treatment effect model, where the time-varying treatment effects are modeled as random deviations from an overall mean. In the simulations considered in this study, concurrent designs generally achieve higher power than factorial designs under a time-varying fixed treatment effect model, though the differences are modest. Finally, we apply the constant effect model and both time-varying treatment effect models to data from the Prognosticating Outcomes and Nudging Decisions in the Electronic Health Record (PONDER) trial. All three models indicate a lack of treatment effect for either intervention, though they differ in the precision of their estimates, likely due to variations in modeling assumptions.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

20 extracted references · 19 canonical work pages

  1. [1]

    The Gambia Hepatitis Intervention Study.Cancer Research1987; 47(21): 5782-5787

    Group TGHS. The Gambia Hepatitis Intervention Study.Cancer Research1987; 47(21): 5782-5787

  2. [2]

    Design and analysis of stepped wedge cluster randomized trials.Contemporary Clinical Trials 2007; 28(2): 182-191

    Hussey MA, Hughes JP. Design and analysis of stepped wedge cluster randomized trials.Contemporary Clinical Trials 2007; 28(2): 182-191

  3. [3]

    Mixed-effects models for the design and analysis of stepped wedge cluster randomized trials: An overview.Statistical Methods in Medical Research 2021; 30(2): 612-639

    Li F, Hughes JP, Hemming K, Taljaard M, Melnick ER, Heagerty PJ. Mixed-effects models for the design and analysis of stepped wedge cluster randomized trials: An overview.Statistical Methods in Medical Research 2021; 30(2): 612-639

  4. [4]

    Zhang Y, Preisser JS, Turner EL, Rathouz PJ, Toles M, Li F. A general method for calculating power for GEE analysis of complete and incomplete stepped wedge cluster randomized trials.Statistical methods in medical research 2023; 32(1): 71–87

  5. [5]

    Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect.Statistics in Medicine 2022; 41(22): 4311-4339

    Kenny A, Voldal E, Xia F, Heagerty PJ, Hughes JP. Analysis of stepped wedge cluster randomized trials in the presence of a time-varying treatment effect.Statistics in Medicine 2022; 41(22): 4311-4339

  6. [6]

    How to achieve model-robust inference in stepped wedge trials with model-based methods?

    Wang B, Wang X, Li F. How to achieve model-robust inference in stepped wedge trials with model-based methods?. Biometrics2024; 80(4): ujae123

  7. [7]

    Current issues in the design and analysis of stepped wedge trials.Contemporary Clinical Trials2015; 45

    Hughes JP, Granston TS, Heagerty PJ. Current issues in the design and analysis of stepped wedge trials.Contemporary Clinical Trials2015; 45

  8. [8]

    NicklessA,VoyseyM,GeddesJ,YuLM,FanshaweTR.Mixedeffectsapproachtotheanalysisofthesteppedwedgecluster randomised trial-Investigating the confounding effect of time through simulation.PLoS One 2018; 13(12)

Show all 20 references
  1. [9]

    Assessing exposure-time treatment effect heterogeneity in stepped-wedge cluster randomized trials.Biometrics2023; 79(3): 2551-2564

    Maleyeff L, Li F, Haneuse S, Wang R. Assessing exposure-time treatment effect heterogeneity in stepped-wedge cluster randomized trials.Biometrics2023; 79(3): 2551-2564

  2. [10]

    Statistics in Medicine 2024; 43(1): 49-60

    LeeKM,CheungYB.Clusterrandomizedtrialdesignsformodelingtime-varyinginterventioneffects. Statistics in Medicine 2024; 43(1): 49-60

  3. [11]

    Sample size calculations for stepped wedge designs with treatment effects that may change with the duration of time under intervention.Prevention Science2024; 25(Suppl 3): 348-355

    Hughes JP, Lee WY, Troxel AB, Heagerty PJ. Sample size calculations for stepped wedge designs with treatment effects that may change with the duration of time under intervention.Prevention Science2024; 25(Suppl 3): 348-355

  4. [12]

    Proposed variations of the stepped-wedge design can be used to accommodate multiple interventions.Journal of Clinical Epidemiology 2017; 86: 160-167

    Lyons VH, Li L, Hughes JP, Rowhani-Rahbar A. Proposed variations of the stepped-wedge design can be used to accommodate multiple interventions.Journal of Clinical Epidemiology 2017; 86: 160-167

  5. [13]

    Statistics in Medicine 2017; 36(24): 3772-3790

    MatthewsJNS,ForbesAB.Steppedwedgedesigns:insightsfromadesignofexperimentsperspective. Statistics in Medicine 2017; 36(24): 3772-3790

  6. [14]

    Admissible multiarm stepped-wedge cluster randomized trial designs.Statistics in Medicine 2019; 38(7): 1103-1119

    Grayling MJ, Mander AP, Wason JMS. Admissible multiarm stepped-wedge cluster randomized trial designs.Statistics in Medicine 2019; 38(7): 1103-1119

  7. [15]

    Variance formulae for multiphase stepped wedge cluster randomized trial

    Zhang P, Shoben A, Jackson R, Fernandez S. Variance formulae for multiphase stepped wedge cluster randomized trial. Statistics in Medicine 2020; 39(28): 4147-4168. 22 Z ChenET AL

  8. [16]

    Power analysis for stepped wedge trials with multiple interventions.Statistics in Medicine 2022; 41(8)

    Sundin P, Crespi CM. Power analysis for stepped wedge trials with multiple interventions.Statistics in Medicine 2022; 41(8)

  9. [17]

    Using simulation studies to evaluate statistical methods.Statistics in medicine 2019; 38(11): 2074–2102

    Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods.Statistics in medicine 2019; 38(11): 2074–2102

  10. [18]

    Ouyang Y, Taljaard M, Forbes AB, Li F. Maintaining the validity of inference from linear mixed models in stepped-wedge cluster randomized trials under misspecified random-effects structures.Statistical Methods in Medical Research 2024; 33(9): 1497–1516

  11. [19]

    Analysis of Stepped-Wedge Cluster Randomized Trials when treatment effect varies by exposure time or calendar time.arXiv preprint arXiv:2409.14706 2024

    Lee KM, Turner EL, Kenny A. Analysis of Stepped-Wedge Cluster Randomized Trials when treatment effect varies by exposure time or calendar time.arXiv preprint arXiv:2409.14706 2024

  12. [20]

    Model-assisted analysis of covariance estimators for stepped wedge cluster randomized experiments

    Chen X, Li F. Model-assisted analysis of covariance estimators for stepped wedge cluster randomized experiments. Scandinavian Journal of Statistics 2025; 52(1): 416–446

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.