REVIEW 2 major objections 6 minor 35 references
Minimax designs for causal effects in temporal experiments with treatment habituation
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper proves that the minimax optimal design for temporal experiments with treatment habituation is a completely randomized pulse design that is deliberately unbalanced, assigning more units to the always-treated and always-control…
desk verdict A correct and genuinely new minimax result for temporal designs, with a scope that is narrower than the abstract suggests and a load-bearing support assumption worth flagging to any referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the class of pulse designs: each unit is assigned one of $T+1$ fixed treatment histories, namely always treated ($1$), always control ($0$), or a pulse $e_t$ that treats only at time $t$, and a design is a distribution over the resulting assignment matrices. The argument's workhorse is Lemma 4, which shows that when the support of potential outcomes is bounded and permutation-invariant, an adversarial schedule can simultaneously push every column variance to the common maximum $V^*$, so the worst-case risk of the symmetrized completely randomized design becomes exactly $V^*$ times $[(T-1)/N_1 + (T-1)/N_0 + 2\sum_{t=2}^T 1/N_{e_t}]$. Minimizing that expression over counts yields the unbalanced completely randomized design, and the closed-form continuous solution follows by Lagrange multipliers. Later theorems recycle the same reduction with an augmented-control estimator, a weighted loss, and $k$-order carryovers, showing the minimax design remains completely randomized in each case.
What would settle it
Enumerate, for small $N$ and $T$, a bounded permutation-invariant outcome set that is deliberately not product-formed—for example, restrict schedules so that high variance in the always-treated column forces low variance in pulse columns—and compare the true maximized risk of the Eq. (9) allocation against a balanced allocation; if balanced randomization has lower worst-case risk, Theorem 1's characterization fails.
Extended reading notes
Core claim
The central claim is Theorem 1: for any bounded, permutation-invariant set of potential outcome schedules, the minimax-optimal pulse design is the completely randomized design whose counts solve the integer program in Eq. (9), minimizing $(T-1)/N_1' + (T-1)/N_0' + 2\sum_{t=2}^T 1/N_{e_t}'$ over allocations of the $N$ units. The continuous relaxation (Proposition 1) has the closed-form solution $\tilde N_1 = \tilde N_0 = N/(2+\sqrt{2(T-1)})$ and $\tilde N_{e_t} = \sqrt{2/(T-1)} \, N/(2+\sqrt{2(T-1)})$, which preserves symmetry between always-treated and always-control arms while giving each pulse arm a smaller share that shrinks relative to the persistent arms as $T$ grows. The same result extends to wedge designs because, under the non-anticipating-outcomes assumption, a pulse and a wedge coincide on all periods up to the treatment start. The paper also shows that using pulse units as augmented controls (Theorem 2), weighting instantaneous versus habituation losses (Theorem 3), and recycling units under $k$-order carryovers (Theorem 4) all keep the same completely-randomized structure while changing the optimal counts.
Load-bearing premise
The minimax proof assumes that the set of possible outcome schedules is product-formed from one bounded, permutation-invariant set of column vectors, so a single adversarial vector can maximize every arm's variance at once; if variance patterns differ across arms and times such that this simultaneous maximization is impossible, the Eq. (9) objective is not the true worst-case risk and the proposed split may not be minimax.
Editorial extensions
If this is right
- An experimenter with only $N$ and $T$ can construct the minimax design by plugging into the closed-form allocation (or solving Eq. (9) for integers) and using difference-in-means estimators for the instantaneous and habituation effects, with no outcome model to estimate.
- Under the theorem, the standard balanced completely randomized design is not minimax for temporal habituation experiments: the always-treated and always-control arms should receive more units than any pulse arm, with the imbalance growing in $T$.
- The augmented-control design of Theorem 2 lowers worst-case risk relative to both balanced randomization and the basic minimax design, with simulations in the paper showing up to about 20% reduction in maximum risk compared to balanced randomization.
- Because every optimal design in the paper is completely randomized, the usual randomization-based inference tools apply: the proposed estimators are unbiased, their variance formulas are known, and conservative confidence intervals can be constructed.
- The same worst-case-optimality conclusions hold for wedge designs, so the results cover stepped-wedge-style rollout schedules used in clinical and policy settings, not only pulse experiments.
Reading between the lines
- Editorial inference: the paper's worst-case guarantee is built for adversarial outcome schedules, but real habituation data have correlated unit effects; under models with stable heterogeneity, a balanced or model-based design may come closer to expected-risk optimality, and the gap between worst-case and expected-risk designs is a trade-off the paper leaves open.
- Editorial inference: because wedge and pulse assignments coincide before treatment under non-anticipation, the same unbalanced allocation logic should transfer to stepped-wedge cluster trials with carryover, a claim that could be tested by simulation on existing stepped-wedge datasets.
- Editorial inference: if outcome variance is systematically higher at early times, the permutation-invariance assumption is violated in a structured way; a natural extension would solve the worst-case allocation with time-specific variance bounds and compare the resulting counts with the paper's single-formula allocation.
- Editorial inference: practitioners could run a sensitivity check by fitting a simple outcome model with correlated potential outcomes across arms and times, then comparing the achieved risk of the proposed allocation against balanced randomization under that model; the product-formed support assumption predicts the proposed design wins only at the adversarial extreme.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the design of temporal experiments in which units can receive treatment or control at each of T periods, and potential outcomes may exhibit habituation: the treatment effect attenuates with repeated exposure. Adopting the randomization-based potential-outcomes framework, the authors define two estimands at each period t, the habituation effect λ_t and the instantaneous effect δ_t, based on contrasts among always-treated, always-control, and pulse assignments. They then consider pulse designs (and, by extension, wedge designs), propose plug-in estimators, and seek the design minimizing the worst-case squared-error risk over a support of potential outcome schedules. Under a support assumption that is permutation-invariant and product-formed, Theorem 1 characterizes the minimax design as a completely randomized design with arm counts solving the integer program in Eq. (9); Proposition 1 gives the continuous-relaxation allocation, which is unbalanced. Subsequent results extend the framework to augmented controls, user-specified weights ρ on the two loss components, and k-order carryover. The proof is built from symmetry reduction lemmas and an exact worst-case risk computation, with simulations illustrating the risk reduction relative to a balanced completely randomized design.
Significance. If the result is taken at face value, the paper contributes a nontrivial extension of classical minimax design results (Wu 1981; Li et al. 1983) to longitudinal experiments, with an explicit and surprising recommendation to unbalance the allocation in favor of always-treated and always-control arms. The derivation is self-contained, does not rely on parametric outcome models, and produces easily computable allocation formulas; the proof strategy of symmetrizing arbitrary designs (Lemma 2) and then computing the exact worst-case risk on the enlarged support (Lemma 4) is elegant and appears internally consistent. The extensions in Sections 3.4 and 4 address practically relevant estimands and estimation strategies, and the simulation study supports the claimed risk improvements in the minimax sense. The main value is the formal minimization framework and the explicit designs; the main caveat is the restrictiveness of the support and design class relative to the motivating applications.
major comments (2)
- [§3.2–3.3, Definition 3, Lemma 4, Eq. (9)] The support assumption in Definition 3 is not merely permutation-invariant; it is product-formed, in that each column of every assignment-specific matrix Y(z) is an arbitrary element of a common set Y. Lemma 4 then constructs the adversarial schedule Yopt whose columns are all identical to a single maximum-variance vector y*, so that V_h^{(t)} = V* for every arm h and period t while the covariance terms V_{0,et}^{(t)} and V_{1,et}^{(t)} vanish. This schedule requires the potential outcomes under always-treated, always-control, and pulse assignments to be simultaneously equal to the same vector at every period. That is inconsistent with the habituation structure motivating the paper, where treatment histories impose ordering or attenuation constraints linking Y(1), Y(0), and Y(e_t) at the same time t. Consequently, for any constrained support reflecting habituation, the exact worst-case risk in Eq. (9) is not the true maximized risk, and the completely randomized design of Theorem 1 is not necessarily minimax. The authors should either justify the product-form support as an appropriate model for the habituation setting, or reframe Theorem 1 as a finite-population minimax result for the enlarged support and clearly state its limited applicability to the motivating applications.
- [Abstract; §2.1 Definition 1; Theorem 1] The abstract claims optimality 'in a large class of practical designs,' but the formal result is restricted to pulse designs (and, via Remark 4, wedge designs): each unit's assignment vector must lie in E = {1, 0, e_1,...,e_T}. This excludes many practical temporal designs, including arbitrary treatment sequences, switchback designs, and designs with carryover-adaptive assignments. The class H is narrow, and the paper does not provide evidence that pulse/wedge designs dominate or approximate broader classes in realistic settings. This scope should be qualified in the abstract and discussion; as written, the claim is an overstatement.
minor comments (6)
- [§2.2.2, Eq. (4) and (5)] The sums run 't = 2,...,N'; given N units and T periods, these should be 't = 2,...,T'.
- [§3.4, Theorem 2, Eq. (13), and Proposition 2] In the displayed objective, the last sum is written as '∑_{t=2}^T 1/N'_et' but should be '∑_{t=2}^T 1/N'_t'; the statement of Proposition 2 repeats this inconsistency and also appears to duplicate the term '∑ 1/N'_et'.
- [§4.1, Proposition 3] The formula for N_0 contains a stray comma: '𝓁(1+√(ρ(T−1)))c2,' should read '𝓁(1+√(ρ(T−1)))c2 +'.
- [Figure 1 caption] The legend shows only two of the three designs; the blue BCRD line is described in the caption but is not labeled in the legend.
- [§3.2, Definition 3] Calling the class 'permutation-invariant' is misleading because the defining property is the product structure Y = Y(Y); a term such as 'product-formed and permutation-invariant support' would be clearer.
- [§5.2, Figure 2] The distinction between the left and right panels ('only augmented minimax design has augmented units' versus 'both designs have augmented units') should be stated in the main text or in a more informative caption.
Circularity Check
No circularity: Theorem 1 is derived from an uncalibrated minimax game; the worst-case schedule is an adversarial construction feasible under the stated product-formed permutation-invariance assumption, and the allocation formulas follow from convex optimization, not from fitted inputs or self-citation.
full rationale
The paper's central derivation is self-contained. In Lemma 4, the worst-case risk is reduced to V* times the objective in Eq. (9) by constructing a schedule Yopt whose columns are all equal to a single vector y that maximizes the column variance V*. That construction is feasible only because Definition 3 defines the schedule support as product-formed: each column of each potential-outcome matrix can be chosen independently from the same bounded, permutation-invariant set Y. This is a stated support assumption, not a quantity fitted to data, and it does not smuggle in the target design. The constant V* cancels out of the minimization, so the minimax allocation is obtained by minimizing the coefficient sum in Eq. (9), which is a convex optimization problem solved in Proposition 1 by the Lagrangian method. No parameter is fitted to the risk or to the estimands, and no 'prediction' is a renamed fit: the allocation formulas, the closed forms for N1, N0, and Net, and the augmented-control solutions in Propositions 2 and 3 all follow from first-order conditions of explicitly stated objectives. The extensions use user-chosen weights rho and carryover orders k, but these are modelling choices, not calibrated constants. The paper's reliance on prior minimax results (Wu 1981; Li et al. 1983) is external and independent, and it is used as motivation rather than as a load-bearing derivation step; the temporal extension requires the new lemmas in the appendix. The one self-citation with author overlap, Toulis and Parkes (2016), appears in Remark 2 only to support an interpretive suggestion about extrapolating habituation effects and is not part of the minimax proof. The main limitations, which the paper itself flags or implies, are scope limitations rather than circularity: Theorem 1 optimizes only within pulse designs (with wedge designs by Remark 4), and the worst-case support under Definition 3 is product-formed, which may be stronger than realistic habituation constraints. These affect applicability of the optimality claim, but they do not make the derivation equivalent to its inputs.
Assumptions & free parameters
free parameters (2)
- rho (weight on habituation loss) =
user-specified, not fitted
- k (carryover window) =
user-specified, not fitted
assumptions (6)
- domain assumption Assumption 1: no interference (potential outcome of unit i depends only on its own treatment sequence)
- domain assumption Assumption 2: non-anticipating outcomes (outcome at time t depends only on treatments up to t)
- domain assumption Assumption 3: k-order carryover (effects of a pulse vanish after k periods)
- domain assumption Permutation-invariance and product-form of the set of potential outcome schedules
- domain assumption Boundedness of potential outcomes
- standard math Randomization-based inference framework: potential outcomes fixed, assignment random
Cite this review
Pith. "Pith review of Minimax designs for causal effects in temporal experiments with treatment habituation." pith.science (2026). https://pith.science/paper/4QOVHAOF
@misc{pith2026190803531,
author = {Pith},
title = {Pith review of: Minimax designs for causal effects in temporal experiments with treatment habituation},
year = {2026},
howpublished = {\url{https://pith.science/paper/4QOVHAOF}},
note = {Machine review of arXiv:1908.03531}
}
read the original abstract
Randomized experiments are the gold standard for estimating the causal effects of an intervention. In the simplest setting, each experimental unit is randomly assigned to receive treatment or control, and then the outcomes in each treatment arm are compared. In many settings, however, randomized experiments need to be executed over several time periods such that treatment assignment happens at each time period. In such temporal experiments, it has been observed that the effects of an intervention on a given unit may be large when the unit is first exposed to it, but then it often attenuates, or even vanishes, after repeated exposures. This phenomenon is typically due to units' habituation to the intervention, or some other general form of learning, such as when users gradually start to ignore repeated mails sent by a promotional campaign. This paper proposes randomized designs for estimating causal effects in temporal experiments when habituation is present. We show that our designs are minimax optimal in a large class of practical designs. Our analysis is based on the randomization framework of causal inference, and imposes no parametric modeling assumptions on the outcomes.
Figures
Reference graph
Works this paper leans on
-
[1]
Abrahamse, W., L. Steg, C. Vlek, and T. Rothengatter (2005). A review of intervention studies aimed at household energy conservation. Journal of environmental psychology\/ 25\/ (3), 273--291
work page 2005
-
[2]
Allcott, H. and S. Mullainathan (2010). Behavior and energy policy. Science\/ 327\/ (5970), 1204--1205
work page 2010
-
[3]
Allcott, H. and T. Rogers (2014). The short-run and long-run effects of behavioral interventions: Experimental evidence from energy conservation. American Economic Review\/ 104\/ (10), 3003--37
work page 2014
- [4]
-
[5]
Bojinov, I. and N. Shephard (2019). Time series experiments and causal estimands: exact randomization tests and trading. Journal of the American Statistical Association\/ , 1--36
work page 2019
-
[6]
Brown, C. A. and R. J. Lilford (2006). The stepped wedge trial design: a systematic review. BMC medical research methodology\/ 6\/ (1), 54
work page 2006
-
[7]
Brown Jr, B. W. (1980). The crossover experiment for clinical trials. Biometrics\/ , 69--79
1980
-
[8]
Chatterjee, P., D. L. Hoffman, and T. P. Novak (2003). Modeling the clickstream: Implications for web-based advertising efforts. Marketing Science\/ 22\/ (4), 520--541
work page 2003
Show all 35 references
-
[9]
Copas, A. J., J. J. Lewis, J. A. Thompson, C. Davey, G. Baio, and J. R. Hargreaves (2015). Designing a stepped wedge trial: three main designs, carry-over effects and randomisation approaches. Trials\/ 16\/ (1), 352
2015
-
[10]
Cox, D. R. (1958). Planning of experiments
1958
-
[11]
Hahn, R. and R. Metcalfe (2016). The impact of behavioral science experiments on energy policy. Economics of Energy & Environmental Policy\/ 5\/ (2), 27--44
2016
-
[12]
Hainmueller, J., D. J. Hopkins, and T. Yamamoto (2014). Causal inference in conjoint analysis: Understanding multidimensional choices via stated preference experiments. Political analysis\/ 22\/ (1), 1--30
2014
-
[13]
Hargreaves, J. R., A. J. Copas, E. Beard, D. Osrin, J. J. Lewis, C. Davey, J. A. Thompson, G. Baio, K. L. Fielding, and A. Prost (2015). Five questions to consider before conducting a stepped wedge trial. Trials\/ 16\/ (1), 350
2015
-
[14]
Heckman, J. J., J. E. Humphries, and G. Veramendi (2016). Dynamic treatment effects. Journal of econometrics\/ 191\/ (2), 276--292
2016
-
[15]
Heckman, J. J. and E. Vytlacil (2005). Structural equations, treatment effects, and econometric policy evaluation 1. Econometrica\/ 73\/ (3), 669--738
2005
-
[16]
O'Brien, and D
Hohnhold, H., D. O'Brien, and D. Tang (2015). Focusing on the long-term: It's good for users and business. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , pp.\ 1849--1858. ACM
2015
-
[17]
Tingley, and T
Imai, K., D. Tingley, and T. Yamamoto (2011). Experimental designs for identifying causal mechanisms. Journal of the Royal Statistical Society, Series B\/ , 1--27
2011
-
[18]
Imbens, G. W. and D. B. Rubin (2015). Causal inference in statistics, social, and biomedical sciences . Cambridge University Press
2015
-
[19]
Ji, X., G. Fink, P. J. Robyn, D. S. Small, et al. (2017). Randomization inference for stepped-wedge cluster-randomized trials: an application to community-based health insurance. The Annals of Applied Statistics\/ 11\/ (1), 1--20
2017
-
[20]
Kim, S. and M. S. Wogalter (2009). Habituation, dishabituation, and recovery effects in visual warnings. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting , Volume 53, pp.\ 1612--1616. Sage Publications Sage CA: Los Angeles, CA
2009
-
[21]
Longbotham, D
Kohavi, R., R. Longbotham, D. Sommerfield, and R. M. Henne (2009). Controlled experiments on the web: survey and practical guide. Data mining and knowledge discovery\/ 18\/ (1), 140--181
2009
-
[22]
Nahum-Shani, K
Lei, H., I. Nahum-Shani, K. Lynch, D. Oslin, and S. A. Murphy (2012). A" smart" design for building individualized treatment sequences. Annual review of clinical psychology\/ 8 , 21--48
2012
-
[23]
Li, K.-C. et al. (1983). Minimaxity for randomized designs: some general results. The Annals of Statistics\/ 11\/ (1), 225--239
1983
-
[24]
Liberali, G., T. S. Gruca, and W. M. Nique (2011). The effects of sensitization and habituation in durable goods markets. European journal of operational research\/ 212\/ (2), 398--410
2011
-
[25]
Neyman, J. S. (1923). On the application of probability theory to agricultural experiments. essay on principles. section 9.(tlanslated and edited by dm dabrowska and tp speed, statistical science (1990), 5, 465-480). Annals of Agricultural Sciences\/ 10 , 1--51
1923
-
[26]
Binik, I
Prost, A., A. Binik, I. Abubakar, A. Roy, M. De Allegri, C. Mouchoux, T. Dreischulte, H. Ayles, J. J. Lewis, and D. Osrin (2015). Logistic, ethical, and political dimensions of stepped wedge trials: critical review and case studies. Trials\/ 16\/ (1), 351
2015
-
[27]
Robins, J. M. (1997). Causal inference from complex longitudinal data. In Latent variable modeling and applications to causality , pp.\ 69--117. Springer
1997
-
[28]
Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology\/ 66\/ (5), 688
1974
-
[29]
o lander, A., T. Frisell, R. Kuja-Halkola, S. \
Sj \"o lander, A., T. Frisell, R. Kuja-Halkola, S. \"O berg, and J. Zetterqvist (2016). Carryover effects in sibling comparison designs. Epidemiology\/ 27\/ (6), 852--858
2016
-
[30]
Toh, S. and M. A. Hern \'a n (2008). Causal inference from longitudinal studies with baseline randomization. The international journal of biostatistics\/ 4\/ (1)
2008
-
[31]
Toulis, P. and D. C. Parkes (2016). Long-term causal effects via behavioral game theory. In Advances in Neural Information Processing Systems , pp.\ 2604--2612
2016
-
[32]
Wathieu, L. (2004). Consumer habituation. Management Science\/ 50\/ (5), 587--596
2004
-
[33]
Wellek, S. and M. Blettner (2012). On the proper use of the crossover design in clinical trials: part 18 of a series on evaluation of scientific publications. Deutsches \"A rzteblatt International\/ 109\/ (15), 276
2012
-
[34]
Wu, C.-F. (1981). On the robustness and efficiency of some randomized designs. The Annals of Statistics\/ 9\/ (6), 1168--1177
1981
-
[35]
Tiwana , S
Yan , J., B. Tiwana , S. Ghosh , H. Liu , and S. Chatterjee (2019, Jan). Measuring Long-term Impact of Ads on LinkedIn Feed . arXiv e-prints\/ , arXiv:1902.03098
2019 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.