{"id":"ff1ec7ef-89fa-4b1d-b335-ac42210d1c62","arxiv_id":"2607.27376","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Training the CUPAC covariate and its regression coefficient with a between-cell-weighted loss improves switchback estimator power, with gains concentrated in within-noise-dominated regimes.","lead":"This paper derives a new training loss and analysis-time coefficient for CUPAC covariate adjustment in switchback experiments, weighting between-cell prediction error by a factor read from the cell-size distribution to shrink estimator variance. If valid, it gives experimentation teams a simple design-dependent formula for improving statistical power in cluster-by-time randomized tests without changing the experiment design.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq (2) is imported from an unpublished companion and appears internally inconsistent; the central misalignment factor inherits this unresolved premise.","rationale":"The reader correctly identifies Eq (2) as the pivot; I agree. My pass sharpens this into a testable internal inconsistency. The equal-cell-size/σ²_res=0 case yields a direct contradiction, so the missing derivation is not merely an exposition gap—the formula as written cannot be exactly correct for the stated model. However, the correction (dropping 1/n̄ from the macro coefficient) leaves the qualitative alignment claim intact for the large-¯n regime targeted by the paper, so a conditional verdict is appropriate rather than rejection. The simulation provides empirical evidence that the aligned estimator helps, but it does not validate the exact variance identity. A revised paper should either derive Eq (2) in full or replace the macro weight with the correct design weight; the headline 'power-optimal' should be qualified until then. The reader's CONDITIONAL verdict therefore stands unchanged.","tokens_in":14644,"tokens_out":13850,"duration_ms":142082,"concrete_test":"Derive Var(τ̂) from DGP (1) for the individual-level OLS difference-in-means with cell sizes n_b (e.g., delta method or exact finite-B computation) and compare the coefficient of σ²_macro with (1/n̄+1+cv²). Equivalently, evaluate Eq (2) at n̄=1, cv²=0, σ²_res=0: if it returns 8σ²_macro/(JH) instead of the direct 4σ²_macro/(JH), the formula is falsified. Then re-run the simulation with the corrected macro weight (1+cv²) and check whether the aligned estimator's advantage persists.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is Eq (2), cited to Pankratev [2026b] and never derived. This is not just an unproven import: a direct calculation for the stated DGP and estimator contradicts it. Take n_b=1 for all cells and σ²_res=0, so every cell is a single observation and all outcome variance is macro. Then τ̂ is the difference in means of B iid cell-level outcomes, with B=JH, and Var(τ̂) ≈ 4σ²_macro/B. Eq (2) with cv²=0, n̄=1, S_macro=1 gives 8σ²_macro/B — a factor of 2. More generally, for the individual-level OLS difference-in-means with unequal cell sizes, the macro coefficient should be (1+cv²), not (1/n̄+1+cv²); the extra 1/n̄ term appears to be an error. Since Proposition 1 and Definition 1 simply substitute into Eq (2), the claimed variance identity (13) and the 'power-optimal' loss L_power inherit this error. The misalignment factor 1+n̄(1+cv²) would become n̄(1+cv²) under the corrected coefficient; the practical conclusion is nearly unchanged for large n̄, but the central theorem as stated is not correct. The simulation gives empirical support for a qualitative benefit but cannot validate the exact weighting.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a control-using-prediction-as-covariate (CUPAC) methodology for switchback experiments with unequal cluster sizes. Its central claim is that the variance of the CUPAC-adjusted treatment-effect estimator is an affine functional of within-cell and between-cell prediction losses, and that training the control variate and estimating the residualization coefficient with the power-weighted loss L_power maximizes statistical power. The paper derives a decomposition of estimator variance, introduces L_power and a per-level adjustment coefficient, proves an upper-bound surrogate theorem, and reports a Monte Carlo study suggesting that the aligned estimator remains effective as the macro variance share decreases. The key quantitative premise is Eq. (2), a variance formula imported from an unpublished companion paper, which is not derived in the manuscript and appears to be incorrect in an important limiting case.","tokens_in":15000,"tokens_out":16128,"duration_ms":181427,"significance":"If the central variance identity were correct, the paper would identify a practically important and largely overlooked issue: in high-density, low-intraclass-correlation switchbacks, unit-level MSE training can systematically underweight between-cell prediction error, and a simple design-dependent reweighting could yield meaningful power gains. The Monte Carlo design is thoughtful: the factorial comparison isolates the training-loss and analysis-coefficient corrections, the balanced regime serves as a control, and the reported between-cell correlations provide a mechanistic check. The qualitative direction of the results is plausible and the simulation supports a real benefit. However, the quantitative claims—including the exact misalignment factor and the optimality of L_power—rest on Eq. (2), which is neither derived nor publicly verifiable and appears to be wrong as stated. The practical prescription may survive a corrected derivation, but the paper as written does not establish its main theorem.","major_comments":[{"comment":"Equation (2) is the load-bearing premise of the paper, but it is neither derived nor accompanied by an error bound; the only citation is to an unpublished working paper (Pankratev 2026b). Worse, it is internally contradicted by the paper's own setup. Set n_b=1 for every cell and σ²_res=0. Then each cell contains one observation equal to the cell-level random effect, treatment is at the cell level, and τ̂ is the difference in means of B iid cell outcomes, so Var(τ̂)=4σ²_macro/B. Equation (2) with n̄=1, cv²=0, S_macro=1 gives 8σ²_macro/B, off by a factor of 2. More generally, for the usual unit-weighted difference of cell means the macro coefficient is 1+cv² rather than 1/n̄+1+cv². Since Proposition 1, Definition 1, Proposition 2, and Corollary 1 all substitute into Eq. (2), the claimed misalignment factor 1+n̄(1+cv²) inherits this error; under the corrected coefficient it would be n̄(1+cv","section":"§3.1, Eq. (2)"},{"comment":"Proposition 1 substitutes MSE_macro(g)=E[\\bar{e}_b²] into the macro slot of Eq. (2). But \\bar{e}_b is not a pure between-cell residual: Eq. (10) shows it contains \\bar{ε}_b, the cell average of the within-cell noise. In the estimator variance, this cell-mean sampling noise should receive the within-cell coefficient (approximately 1/n̄), not the macro weight b of Eq. (14). Thus even if Eq. (2) were correct for raw outcomes, Eq. (13) would not be the variance of the adjusted estimator. A concrete check: if g predicts the macro component perfectly, so e=ε, the true remaining variance is approximately 4σ²_res/(JH n̄). Equation (13) instead gives approximately (4σ²_res/(JH n̄))(2+cv²), a factor of roughly 2+cv² too large. The residual process does not have the same variance-component structure as the original model, so the formal substitution used in the proof of Proposition 1 is invalid.","section":"§3.2–3.3, Eqs. (10)–(13)"},{"comment":"The label 'power-optimal' is stronger than what is actually proved. Definition 1 identifies L_power with the calibrated variance, so minimizing it is exact only under Assumption 1, where θ=1 and g is conditionally unbiased. For the deployed estimator with per-level θ estimated from data, Theorem 1 gives only the upper bound JH/4 Var ≤ L_power, with equality only under level-calibration. An upper-bound surrogate can have a different minimizer than the true objective; the paper does not show that minimizing L_power maximizes power for the estimator actually used unless calibration is enforced. The formal claims should be restated as a calibration result plus an upper-bound approximation, and the title/abstract should match that level of support.","section":"§3.3–3.4, Definition 1 and Theorem 1"}],"minor_comments":[{"comment":"No Monte Carlo standard errors are reported for the SE ratios or power estimates. With 1,000 replications, power estimates near 0.5 have Monte Carlo standard errors of about 1.5 percentage points, and the SE ratios need error bars. In addition, only relative standard errors are reported; without the absolute variance of the unadjusted estimator, the simulation cannot validate Eq. (2) or Eq. (13).","section":"§5, Table 1"},{"comment":"The conclusion states that the paper 'gave gradient–Hessian and closed-form Ridge formulations,' but no such formulations appear in Sections 1–6. Either add them or remove the claim.","section":"§7"},{"comment":"The text repeatedly calls the variance decomposition 'exact' (e.g., the opening of Section 3), while Eq. (2) carries an approximation sign. Please reconcile this language and specify the approximation error.","section":"§3 opening and Eq. (2)"},{"comment":"The central identity is cited to an unpublished working paper with a 2026 arXiv identifier. The paper should either make that companion publicly available or reproduce the derivation in an appendix.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands. The counterexample to Eq. (2) is decisive: in the n_b=1, σ²_res=0 limiting case the formula is off by a factor of two, and the same error propagates into the central misalignment factor. The core idea—that unit-level MSE can underweight between-cell error and that a design-weighted loss helps—is plausible and likely salvageable, and the simulation supports a qualitative benefit. But the exact 'power-optimal' claims cannot stand until Eq. (2) is correctly derived and all downstream formulas are updated. I would advise the editor that the manuscript needs a substantive revision, not just editorial changes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Readable and genuinely useful in its practical message, but the exact weighting rests on a variance identity the paper does not derive and, on a simple check, gets wrong. The idea survives; the formula doesn't.\n\nThe new thing here is real: for switchback experiments, both stages of the CUPAC workflow should be aligned to the design — train the control variate on a composite loss that up-weights between-cell error, and estimate the adjustment coefficient under the same loss. The paper formalizes this cleanly, and the per-level coefficient result (Corollary 2) is a good sharpening. The simulation is also built to test the mechanism rather than just demonstrate a win, and the authors are candid about overfitting risk and the small-cell bias of the macro target. If I were building a production system, I would adopt the qualitative principle: put more training weight on cell means, and fit per-level coefficients.\n\nThe soft spot is load-bearing. Eq (2) is imported from an unpublished companion and never derived. A direct calculation for the stated DGP and estimator shows it is not correct: with one observation per cell and zero within-cell variance, the difference-in-means has variance 4σ²/B, but Eq (2) gives 8σ²/B. The source is an extra 1/n̄ term multiplying S_macro; the macro weight should be (1 + cv²), not (1/n̄ + 1 + cv²). Proposition 1, Definition 1, and the claimed misalignment factor 1 + n̄(1+cv²) inherit this. The corrected factor would be n̄(1+cv²), so for the large n̄ the paper targets, the practical conclusion is nearly unchanged — but the central theorem as stated is wrong. Also, 'power-optimal' is an overstatement; Theorem 1 gives an upper-bound surrogate, not an exact optimum. No code or production validation.\n\nSo the paper is a good idea in need of a fix — not a finished result. A serious referee should see it, mainly to insist on a proper derivation of the variance identity and a corrected simulation. I'd be glad to read the revision, but I wouldn't cite it in this form.","headline":"Useful practical intuition, but the load-bearing variance identity is imported from an unpublished companion and, on a simple calculation, wrong; the reweighting idea survives, the exact formula does not.","tokens_in":15456,"tokens_out":4641,"would_cite":false,"duration_ms":45736,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Switchback experiments lose power when control variates are trained on unit-level error; the paper shows that reweighting the training loss and analysis coefficient by a design-derived factor restores power.","keywords":["switchback experiments","CUPAC","control variates","variance reduction","cluster-randomized designs","power-optimal loss","between-cell variance","cell-size imbalance"],"falsifier":"Simulate the unadjusted switchback estimator under a known multilevel DGP with varying cell-size distributions and compare the empirical estimator variance to Eq (2). If the true variance deviates materially from the affine form in Eq (2), or if the substitution of residual MSE components is not accurate for finite numbers of cells, then the claimed weighting factor 1 + n-bar(1 + cv^2) is not the correct correction.","tokens_in":14547,"feed_emoji":"📊","tokens_out":2840,"duration_ms":33869,"temperature":0.7,"pith_summary":"The paper claims that in switchback experiments, standard control-variate adjustment is systematically misaligned with what drives statistical power. It derives an affine identity linking the adjusted estimator's variance to within-cell and between-cell prediction losses, and shows that the between-cell loss is under-weighted by a factor of 1 + n-bar(1 + cv^2) in both covariate training and coefficient estimation. The proposed fix is a composite loss L_power and a matched analysis coefficient, each re-weighted by the same design factor. If correct, this means conventional unit-level MSE training leaves substantial variance reduction on the table whenever within-cell noise dominates, and that a simple, measurable correction can recover it.","feed_headline":"Train control variates on cell means to lift switchback power","feed_subtitle":"A design-derived reweighting keeps standard error flat as within-cell noise grows, raising power from 0.41 to 0.77.","key_machinery":"The central object is the variance identity of Proposition 1, which expresses the CUPAC-adjusted estimator variance as a fixed affine combination of within-cell and between-cell prediction losses, with weights a = 1/n-bar and b = 1/n-bar + 1 + cv^2. The ratio b/a = 1 + n-bar(1 + cv^2) is the single design quantity that quantifies how much ordinary unit-level MSE under-weights between-cell error. This same ratio drives both the power-optimal training loss and the matched analysis-time coefficient, and the paper uses it to show that training and analysis are coupled: L_power is an upper bound on the variance that survives after re-estimating the coefficient.","core_discovery":"The paper establishes that, under a multilevel model of switchback outcomes, the variance of the CUPAC-adjusted treatment-effect estimator is approximately (4/(JH))[ MSE_within/n-bar + MSE_macro(1/n-bar + 1 + cv^2) ]. Because ordinary unit-level MSE splits into MSE_within + MSE_macro with equal weights, it under-weights the between-cell component by exactly the factor 1 + n-bar(1 + cv^2). The paper therefore defines the power-optimal loss L_power = (1/n-bar)MSE_within + (1/n-bar + 1 + cv^2)MSE_macro, and shows that the analysis-time residualization coefficient should be estimated under the same composite loss, or per level. When both stages are aligned, Monte Carlo simulations show the estim","pith_inferences":["A testable extension is to apply the same 1 + n-bar(1 + cv^2) weighting to control variates in cluster-randomized trials outside switchbacks, since the variance functional is a special case of classical cluster imbalance.","The paper's per-level coefficient result implies that, once coefficients are fit per level, the only thing a training loss can improve is the between-cell correlation; model families should therefore be compared on that single quantity in switchback settings.","Because L_power is an upper bound on deployed variance, recalibrating the covariate to be level-calibrated could make the bound tight; this suggests an explicit calibration step as a practical complement to the proposed loss.","The load-bearing variance formula Eq (2) is imported from an unpublished companion paper, so a natural next step is to derive explicit finite-B error bounds; until then, the practical guarantees rest on an approximation."],"forward_implications":["If the central claim is correct, standard CUPAC practice in switchbacks is not merely suboptimal but quantifiably so, leaving a variance-reduction gap that grows with within-cell noise and cell-size imbalance.","The correction factor is read directly from the cell-size distribution, so no tuning parameter is needed; the paper's formula turns a design constant into a training and analysis weight.","The training-loss and coefficient corrections are complementary: the paper's simulations show that either alone is insufficient, while the aligned pair keeps relative standard error flat across regimes.","The method does not inflate false positive rates in simulations, so aligning the objectives improves power without sacrificing test validity.","The largest gains are predicted for metrics with low intraclass correlation and high mean cell density, which are common in fine-grained event data."],"fun_headline_variants":["Reweight CUPAC loss to sharpen switchback power","Balance within and between noise for switchback tests","Optimal CUPAC weight lifts power in switchbacks","Split noise by cluster to boost switchback power"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire argument rests on the unadjusted estimator variance formula Eq (2), which the paper cites to an unpublished companion paper and does not derive or error-bound; if that formula is wrong or only a poor approximation for realistic cell-size distributions, the claimed misalignment factor and the L_power correction collapse.","fun_headline_variants_meta":{"raw":{"variants":["Reweight CUPAC loss to sharpen switchback power","Balance within and between noise for switchback tests","Optimal CUPAC weight lifts power in switchbacks","Split noise by cluster to boost switchback power"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1152,"prompt_tokens":755,"completion_tokens":397,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":334}},"tokens_in":499,"tokens_out":397,"duration_ms":4132,"temperature":1.0,"reasoning_tokens":334,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T08:29:55.864242+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the unadjusted switchback estimator under a known multilevel DGP with varying cell-size distributions and compare the empirical estimator variance to Eq (2). If the true variance deviates materially from the affine form in Eq (2), or if the substitution of residual MSE components is not accurate for finite numbers of cells, then the claimed weighting factor 1 + n-bar(1 + cv^2) is not the correct correction.","supporting_citations":[],"review_version":1}