{"id":"3a7b9880-848e-4047-9ecb-02ec2d0c78e0","arxiv_id":"1908.03531","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In temporal experiments with treatment habituation, the minimax optimal pulse design is an unbalanced completely randomized design that assigns more units to always-treated and always-control arms and fewer to each pulse arm.","lead":"This paper derives randomized experimental designs that minimize worst-case estimation error for causal effects in temporal experiments where treatments habituate. The optimal designs allocate more units to always-treated and always-control arms and fewer to each single-pulse arm, rather than spreading units equally across all arms.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 is internally valid under the product-formed permutation-invariant support; the load-bearing limitation is that optimality holds only within pulse/wedge designs and for a worst-case support that is more restrictive than the abstract suggests.","rationale":"I read the paper as a self-contained decision-theoretic result: for a fixed class of pulse designs, fixed plug-in estimators, and a product-formed permutation-invariant worst-case support, Theorem 1 derives the minimax completely randomized allocation. The proof chain is coherent: Lemma 1 establishes permutation equivariance of the loss, Lemma 2 symmetrizes any design without increasing worst-case risk, Lemma 3 represents the symmetrized design as a mixture of completely randomized designs, and Lemma 4 identifies a single adversarial schedule that simultaneously maximizes all column variances and zeroes all cross-variance terms. The resulting optimization in Eq. (9) is then minimized by a completely randomized design with the stated counts. I found no algebraic or logical error in this argument. The reader's weakest-assumption identification is accurate and is the same concern I would emphasize: the product-form support is what makes the simultaneous worst-case schedule possible, and the theorem's scope is limited to pulse and wedge designs rather than 'a large class of practical designs' as the abstract states. Both are scope conditions rather than proof failures. They justify a conditional reading but do not require changing the reader's verdict, which is already CONDITIONAL. The proposed brute-force test would empirically demonstrate whether the product-form assumption is essential by comparing the unconstrained product-form minimax design with the minimax design over a realistic constrained support.","tokens_in":24382,"tokens_out":20973,"duration_ms":218124,"concrete_test":"For a small instance such as N=5 and T=3, enumerate all pulse designs and a constrained set of potential outcome schedules that satisfies a monotone-habituation restriction, for example 0 <= Yit(1)-Yit(et) <= Yit(et)-Yit(0) <= M with outcomes bounded by a fixed constant, and solve the exact minimax game by brute force. Compare the optimal allocation with the solution of Eq. (9); if the allocations differ, the product-formed support is essential for Theorem 1's conclusion. Separately, over the product-formed support, evaluate whether a design outside Z(E) union Z(E*) achieves lower worst-case risk than the Theorem 1 design; this would directly test the scope limitation in the abstract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The mathematical argument in Lemma 4 is consistent: because the support Y(Y) is product-formed, the adversarial schedule Yopt whose columns are all equal to a common maximal-variance vector y* is feasible, so every column variance attains V* while the cross-variance terms V0,et and V1,et vanish. This makes the worst-case risk exactly V* times the objective in Eq. (9). The minimax reduction via symmetrization (Lemma 2) is also sound, so I found no internal gap in Theorem 1 under the stated assumptions. The load-bearing concern is the support assumption itself. Definition 3 calls the schedule set 'permutation-invariant', but it is actually product-formed: each of the T+1 assignment-specific matrices can have each column chosen independently from a single bounded set Y. Real potential outcome schedules are not generally freely exchangeable across arms and time periods; for example, habituation imposes ordering or magnitude constraints linking Y(1), Y(0), and Y(et) at the same time t. Under such a constrained support, the schedule Yopt need not be feasible, so the worst-case risk can be strictly smaller than the Eq. (9) expression, and the completely randomized counts of Theorem 1 need not be minimax for that support. In addition, the theorem only optimizes within pulse designs (and, by Remark 4, wedge designs), so the abstract's claim of optimality in 'a large class of practical designs' overstates the scope. These are limitations of applicability, not inconsistencies in the proof, but they are load-bearing for the practical reading of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the design of temporal experiments in which units can receive treatment or control at each of T periods, and potential outcomes may exhibit habituation: the treatment effect attenuates with repeated exposure. Adopting the randomization-based potential-outcomes framework, the authors define two estimands at each period t, the habituation effect λ_t and the instantaneous effect δ_t, based on contrasts among always-treated, always-control, and pulse assignments. They then consider pulse designs (and, by extension, wedge designs), propose plug-in estimators, and seek the design minimizing the worst-case squared-error risk over a support of potential outcome schedules. Under a support assumption that is permutation-invariant and product-formed, Theorem 1 characterizes the minimax design as a completely randomized design with arm counts solving the integer program in Eq. (9); Proposition 1 gives the continuous-relaxation allocation, which is unbalanced. Subsequent results extend the framework to augmented controls, user-specified weights ρ on the two loss components, and k-order carryover. The proof is built from symmetry reduction lemmas and an exact worst-case risk computation, with simulations illustrating the risk reduction relative to a balanced completely randomized design.","tokens_in":24668,"tokens_out":8569,"duration_ms":84351,"significance":"If the result is taken at face value, the paper contributes a nontrivial extension of classical minimax design results (Wu 1981; Li et al. 1983) to longitudinal experiments, with an explicit and surprising recommendation to unbalance the allocation in favor of always-treated and always-control arms. The derivation is self-contained, does not rely on parametric outcome models, and produces easily computable allocation formulas; the proof strategy of symmetrizing arbitrary designs (Lemma 2) and then computing the exact worst-case risk on the enlarged support (Lemma 4) is elegant and appears internally consistent. The extensions in Sections 3.4 and 4 address practically relevant estimands and estimation strategies, and the simulation study supports the claimed risk improvements in the minimax sense. The main value is the formal minimization framework and the explicit designs; the main caveat is the restrictiveness of the support and design class relative to the motivating applications.","major_comments":[{"comment":"The support assumption in Definition 3 is not merely permutation-invariant; it is product-formed, in that each column of every assignment-specific matrix Y(z) is an arbitrary element of a common set Y. Lemma 4 then constructs the adversarial schedule Yopt whose columns are all identical to a single maximum-variance vector y*, so that V_h^{(t)} = V* for every arm h and period t while the covariance terms V_{0,et}^{(t)} and V_{1,et}^{(t)} vanish. This schedule requires the potential outcomes under always-treated, always-control, and pulse assignments to be simultaneously equal to the same vector at every period. That is inconsistent with the habituation structure motivating the paper, where treatment histories impose ordering or attenuation constraints linking Y(1), Y(0), and Y(e_t) at the same time t. Consequently, for any constrained support reflecting habituation, the exact worst-case risk in Eq. (9) is not the true maximized risk, and the completely randomized design of Theorem 1 is not necessarily minimax. The authors should either justify the product-form support as an appropriate model for the habituation setting, or reframe Theorem 1 as a finite-population minimax result for the enlarged support and clearly state its limited applicability to the motivating applications.","section":"§3.2–3.3, Definition 3, Lemma 4, Eq. (9)"},{"comment":"The abstract claims optimality 'in a large class of practical designs,' but the formal result is restricted to pulse designs (and, via Remark 4, wedge designs): each unit's assignment vector must lie in E = {1, 0, e_1,...,e_T}. This excludes many practical temporal designs, including arbitrary treatment sequences, switchback designs, and designs with carryover-adaptive assignments. The class H is narrow, and the paper does not provide evidence that pulse/wedge designs dominate or approximate broader classes in realistic settings. This scope should be qualified in the abstract and discussion; as written, the claim is an overstatement.","section":"Abstract; §2.1 Definition 1; Theorem 1"}],"minor_comments":[{"comment":"The sums run 't = 2,...,N'; given N units and T periods, these should be 't = 2,...,T'.","section":"§2.2.2, Eq. (4) and (5)"},{"comment":"In the displayed objective, the last sum is written as '∑_{t=2}^T 1/N'_et' but should be '∑_{t=2}^T 1/N'_t'; the statement of Proposition 2 repeats this inconsistency and also appears to duplicate the term '∑ 1/N'_et'.","section":"§3.4, Theorem 2, Eq. (13), and Proposition 2"},{"comment":"The formula for N_0 contains a stray comma: '𝓁(1+√(ρ(T−1)))c2,' should read '𝓁(1+√(ρ(T−1)))c2 +'.","section":"§4.1, Proposition 3"},{"comment":"The legend shows only two of the three designs; the blue BCRD line is described in the caption but is not labeled in the legend.","section":"Figure 1 caption"},{"comment":"Calling the class 'permutation-invariant' is misleading because the defining property is the product structure Y = Y(Y); a term such as 'product-formed and permutation-invariant support' would be clearer.","section":"§3.2, Definition 3"},{"comment":"The distinction between the left and right panels ('only augmented minimax design has augmented units' versus 'both designs have augmented units') should be stated in the main text or in a more informative caption.","section":"§5.2, Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The mathematical core of the paper is sound under the stated assumptions, and the proof of Theorem 1 is convincing. However, the gap between the motivating habituation setting and the product-formed support assumption, together with the overstatement of the design class in the abstract, is substantial. A careful revision that repositions the contribution and discusses the support assumption's implications would be needed. I do not see a fundamental flaw that would require rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look. It extends the old Wu (1981) / Li et al. (1983) minimax results to temporal experiments with habituation, and the extension is real: the balanced design stops being optimal, and the paper gives explicit allocation formulas (Proposition 1) that are easy to use. The augmented-control estimator in Section 3.4 is also a genuine contribution, and the recursive solution in Proposition 2 is nontrivial. I checked the proof structure for Theorem 1. Lemma 2's symmetrization argument is sound, and Lemma 4 correctly identifies the worst-case schedule when the support is product-formed: because each column can be chosen independently from the same set Y, an adversarial schedule can make every column attain the maximal variance while killing the cross-variance terms. So the theorem is internally valid under the stated assumptions. That deserves credit.\n\nThe soft spots are about scope and assumptions, not about gaps in the math. First, the abstract says optimality holds in “a large class of practical designs,” but the optimality is only over pulse designs (and, by Remark 4, wedge designs). That is a meaningful restriction, not a cosmetic one. Second, the permutation-invariance condition in Definition 3 is really a product-form support condition: the columns of the potential outcome matrices vary independently. Real habituation processes often impose cross-column constraints — say, ordering between Yit(1) and Yit(et), or magnitude restrictions on carryover. Under such constrained supports the adversarial schedule in Lemma 4 may not be feasible, so the true worst-case risk could be smaller and the completely randomized allocation from Theorem 1 need not be minimax for that support. This is a limitation of applicability, not a proof error, but it is load-bearing for anyone who wants to use the design in practice.\n\nThere are also small presentational issues: a few typos like “instaneous” and some index ranges that should be T rather than N. The simulations are illustrative rather than exhaustive, and no code is provided, but the theory stands on its own.\n\nWho should read this: anyone designing temporal experiments in digital platforms or behavioral interventions, and anyone working on randomization-based design theory. It deserves peer review. I would send it out with a referee request to clarify the scope language and to discuss explicitly how restrictive the product-form support is, ideally with an example where habituation constraints break it. My own verdict is a qualified yes: the core result is correct, but the practical reach is narrower than the packaging suggests.","headline":"A correct and genuinely new minimax result for temporal designs, with a scope that is narrower than the abstract suggests and a load-bearing support assumption worth flagging to any referee.","tokens_in":25209,"tokens_out":1266,"would_cite":true,"duration_ms":15845,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62K05","62C20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that the minimax optimal design for temporal experiments with treatment habituation is a completely randomized pulse design that is deliberately unbalanced, assigning more units to the always-treated and always-control…","keywords":["temporal experiments","causal inference","habituation effects","minimax design","completely randomized design","potential outcomes","pulse and wedge designs","treatment carryover"],"falsifier":"Enumerate, for small $N$ and $T$, a bounded permutation-invariant outcome set that is deliberately not product-formed—for example, restrict schedules so that high variance in the always-treated column forces low variance in pulse columns—and compare the true maximized risk of the Eq. (9) allocation against a balanced allocation; if balanced randomization has lower worst-case risk, Theorem 1's characterization fails.","tokens_in":24169,"feed_emoji":"📊","tokens_out":11857,"duration_ms":101442,"temperature":0.7,"pith_summary":"This paper asks how to randomize units in a temporal experiment when an intervention's effect is strongest at first exposure and then decays or vanishes with repetition, a phenomenon called habituation. Within the randomization-based potential-outcomes framework, with no parametric model for outcomes, it proves that the minimax pulse design—the design minimizing worst-case squared-error risk for jointly estimating instantaneous and habituation effects—is completely randomized but deliberately unbalanced. In the continuous relaxation the optimal counts are $\\tilde N_1 = \\tilde N_0 = N/(2+\\sqrt{2(T-1)})$ and $\\tilde N_{e_t} = \\sqrt{2/(T-1)} \\, N/(2+\\sqrt{2(T-1)})$, putting more units in the two persistent arms and fewer in each single-time pulse arm than a balanced design would. This gives experimenters a free, closed-form allocation rule with a worst-case guarantee, for problems such as repeated email campaigns, energy reports, and online ads, where the usual balanced randomization is not worst-case optimal.","feed_headline":"For habituation experiments, the minimax design is unbalanced","feed_subtitle":"A closed-form allocation between persistent and pulse arms minimizes worst-case error—no outcome model required.","key_machinery":"The central object is the class of pulse designs: each unit is assigned one of $T+1$ fixed treatment histories, namely always treated ($1$), always control ($0$), or a pulse $e_t$ that treats only at time $t$, and a design is a distribution over the resulting assignment matrices. The argument's workhorse is Lemma 4, which shows that when the support of potential outcomes is bounded and permutation-invariant, an adversarial schedule can simultaneously push every column variance to the common maximum $V^*$, so the worst-case risk of the symmetrized completely randomized design becomes exactly $V^*$ times $[(T-1)/N_1 + (T-1)/N_0 + 2\\sum_{t=2}^T 1/N_{e_t}]$. Minimizing that expression over counts yields the unbalanced completely randomized design, and the closed-form continuous solution follows by Lagrange multipliers. Later theorems recycle the same reduction with an augmented-control estimator, a weighted loss, and $k$-order carryovers, showing the minimax design remains completely randomized in each case.","core_discovery":"The central claim is Theorem 1: for any bounded, permutation-invariant set of potential outcome schedules, the minimax-optimal pulse design is the completely randomized design whose counts solve the integer program in Eq. (9), minimizing $(T-1)/N_1' + (T-1)/N_0' + 2\\sum_{t=2}^T 1/N_{e_t}'$ over allocations of the $N$ units. The continuous relaxation (Proposition 1) has the closed-form solution $\\tilde N_1 = \\tilde N_0 = N/(2+\\sqrt{2(T-1)})$ and $\\tilde N_{e_t} = \\sqrt{2/(T-1)} \\, N/(2+\\sqrt{2(T-1)})$, which preserves symmetry between always-treated and always-control arms while giving each pulse arm a smaller share that shrinks relative to the persistent arms as $T$ grows. The same result extends to wedge designs because, under the non-anticipating-outcomes assumption, a pulse and a wedge coincide on all periods up to the treatment start. The paper also shows that using pulse units as augmented controls (Theorem 2), weighting instantaneous versus habituation losses (Theorem 3), and recycling units under $k$-order carryovers (Theorem 4) all keep the same completely-randomized structure while changing the optimal counts.","pith_inferences":["Editorial inference: the paper's worst-case guarantee is built for adversarial outcome schedules, but real habituation data have correlated unit effects; under models with stable heterogeneity, a balanced or model-based design may come closer to expected-risk optimality, and the gap between worst-case and expected-risk designs is a trade-off the paper leaves open.","Editorial inference: because wedge and pulse assignments coincide before treatment under non-anticipation, the same unbalanced allocation logic should transfer to stepped-wedge cluster trials with carryover, a claim that could be tested by simulation on existing stepped-wedge datasets.","Editorial inference: if outcome variance is systematically higher at early times, the permutation-invariance assumption is violated in a structured way; a natural extension would solve the worst-case allocation with time-specific variance bounds and compare the resulting counts with the paper's single-formula allocation.","Editorial inference: practitioners could run a sensitivity check by fitting a simple outcome model with correlated potential outcomes across arms and times, then comparing the achieved risk of the proposed allocation against balanced randomization under that model; the product-formed support assumption predicts the proposed design wins only at the adversarial extreme."],"forward_implications":["An experimenter with only $N$ and $T$ can construct the minimax design by plugging into the closed-form allocation (or solving Eq. (9) for integers) and using difference-in-means estimators for the instantaneous and habituation effects, with no outcome model to estimate.","Under the theorem, the standard balanced completely randomized design is not minimax for temporal habituation experiments: the always-treated and always-control arms should receive more units than any pulse arm, with the imbalance growing in $T$.","The augmented-control design of Theorem 2 lowers worst-case risk relative to both balanced randomization and the basic minimax design, with simulations in the paper showing up to about 20% reduction in maximum risk compared to balanced randomization.","Because every optimal design in the paper is completely randomized, the usual randomization-based inference tools apply: the proposed estimators are unbiased, their variance formulas are known, and conservative confidence intervals can be constructed.","The same worst-case-optimality conclusions hold for wedge designs, so the results cover stepped-wedge-style rollout schedules used in clinical and policy settings, not only pulse experiments."],"supporting_citations":[{"why":"supplies the cross-sectional minimax design framework and permutation-invariance robustness interpretation that the paper extends to temporal experiments.","marker":"Wu (1981)"},{"why":"provides the general minimaxity results for randomized designs that the paper extends from the static balanced setting to the temporal unbalanced setting.","marker":"Li et al. (1983)"},{"why":"supplies the fixed-potential-outcomes randomization framework in which the design problem is set.","marker":"Neyman (1923)"},{"why":"defines causal effects as contrasts of potential outcomes, the template for the habituation and instantaneous estimands.","marker":"Rubin (1974)"},{"why":"gives the difference-in-means variance formulas and conservative inference used in the proof of Lemma 4 and in the randomization-inference appendix.","marker":"Imbens and Rubin, 2015"},{"why":"provides the dynamic potential-outcomes and non-anticipating-outcomes setup that the temporal estimands and augmented controls rely on.","marker":"Bojinov and Shephard (2019)"},{"why":"supplies empirical evidence of habituation in repeated energy-report treatments that motivates the estimands.","marker":"Allcott and Rogers (2014)"},{"why":"motivates the problem through ad blindness and introduces the pulse- and wedge-style designs the paper formalizes.","marker":"Hohnhold et al. (2015)"}],"fun_headline_variants":["Minimax temporal experiments: split units unevenly","Optimal habituation trials: fewer units for pulse arms","Closed-form minimax design for treatment habituation","Unbalanced allocation is minimax-optimal for habituation","Pulse arms need fewer units in minimax temporal designs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The minimax proof assumes that the set of possible outcome schedules is product-formed from one bounded, permutation-invariant set of column vectors, so a single adversarial vector can maximize every arm's variance at once; if variance patterns differ across arms and times such that this simultaneous maximization is impossible, the Eq. (9) objective is not the true worst-case risk and the proposed split may not be minimax.","fun_headline_variants_meta":{"raw":{"variants":["Minimax temporal experiments: split units unevenly","Optimal habituation trials: fewer units for pulse arms","Closed-form minimax design for treatment habituation","Unbalanced allocation is minimax-optimal for habituation","Pulse arms need fewer units in minimax temporal designs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1766,"prompt_tokens":1004,"completion_tokens":762,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":620,"completion_tokens_details":{"reasoning_tokens":684}},"tokens_in":620,"tokens_out":762,"duration_ms":7884,"temperature":1.0,"reasoning_tokens":684,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:10:48.593162+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Enumerate, for small $N$ and $T$, a bounded permutation-invariant outcome set that is deliberately not product-formed—for example, restrict schedules so that high variance in the always-treated column forces low variance in pulse columns—and compare the true maximized risk of the Eq. (9) allocation against a balanced allocation; if balanced randomization has lower worst-case risk, Theorem 1's characterization fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the fixed-potential-outcomes randomization framework in which the design problem is set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the dynamic potential-outcomes and non-anticipating-outcomes setup that the temporal estimands and augmented controls rely on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies empirical evidence of habituation in repeated energy-report treatments that motivates the estimands."},{"cited_title":"O'Brien, and D","cited_arxiv_id":null,"evidence_quote":"motivates the problem through ad blindness and introduces the pulse- and wedge-style designs the paper formalizes."}],"review_version":1}