{"id":"1dd17a5a-5cfd-4eb5-b35d-f9d7faa0bf09","arxiv_id":"2504.21595","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The authors construct anytime-valid rank tests that monitor treatment effects in real time for difference-in-differences and synthetic control estimators, preserving Type I error control at data-dependent stopping times.","lead":"This paper develops statistical tests that let researchers check, in real time, whether a treatment has an effect as new data arrive, without pre-committing to a sample size. It does this by turning treatment-effect estimates into ranks and using 'anytime-valid' p-values that control error at every moment.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 2 omits assumptions on the training period, so SCM estimates need not be exchangeable under the stated null; the SCM anytime-valid application is not justified as written.","rationale":"The core anytime-valid rank constructions (Theorems 1 and 2) are sound conditional e-value arguments; my read does not question them. The vulnerability is in the application layer: the paper's headline promise is real-time inference for program evaluation via exchangeability of treatment estimates. For SCM, Proposition 2 is the only formal justification that btau_t is exchangeable under H0. As stated, the proposition omits any condition on the training period used to estimate SCM weights. The attack shows the conditions can hold while btau_t is non-exchangeable, because the weights inherit dependence from training that is not mediated by the evaluation-period exchangeable block. This differs from the acknowledged serial-dependence limitation in Table 2 (where the evaluation block itself is non-exchangeable); it is a proof gap that causes size distortion even in a setting the paper claims is valid. The fix is straightforward: add an assumption that the training period is independent of the evaluation period, or that the full error/factor sequence (including training) is exchangeable, matching Abadie and Zhao (2021). This is a revision-level issue, so I concur with the reader's CONDITIONAL verdict. The reader's weakest_assumption points at exchangeability generally; I am flagging a more specific insufficiency in how exchangeability is verified for SCM, which the reader did not explicitly identify.","tokens_in":29105,"tokens_out":28605,"duration_ms":288123,"concrete_test":"Simulate the IFE model (21) with r=0, no covariates, N=1 control, and sample sizes like Table 1 (T0=50, TB=25, post length 30). Draw eps_it i.i.d. N(0,1) except force eps_{1,T0+1}=eps_{1,T0} and eps_{2,T0+1}=eps_{2,T0}, so the evaluation block B∪{T0+1,...,T} is i.i.d. (hence exchangeable) but the training period is correlated with the first post-period. Estimate the SCM weight from E, compute btau_t for t in B∪post under tau=0, and apply the paper's anytime-valid reduced-rank Gaussian test at fixed time T0+12. If the rejection rate at alpha=0.05 over 2000 repetitions substantially exceeds 0.05 (say >0.12), Proposition 2's conditions are insufficient; if it stays near 0.05, the counterexample does not land. A permutation test of exchangeability of btau_t can serve as an independent check.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The bridge from 'no treatment effect' to 'btau_t exchangeable' for SCM is Proposition 2 (Section 5.2.2). It assumes only that {(lambda_t,theta_t)} and {eps_t} over t in B∪{T0+1,...,T} are exchangeable and independent, but btau_t is computed with SCM weights estimated from the training period E = {1,...,T0}\\B. Nothing constrains the training period or its dependence on the evaluation period, so the weights can be correlated with evaluation-period errors. Concretely, with r=0, no covariates, one control unit, btau_t = eps_1t - w eps_2t, where w is a function of training data. If eps_{1,T0+1} is perfectly correlated with eps_{1,T0} (and similarly for unit 2) while the evaluation-block eps are i.i.d., the proposition's conditions hold, yet w is correlated with eps_{T0+1} but not eps_{T0+2}; hence (btau_{T0+1}, btau_{T0+2}) is not exchangeable. The anytime-valid tests then fail to control Type I error for the no-treatment null in a setting the proposition claims is valid. Proposition 1 (DiD) avoids this by assuming exchangeability over all t including training; Proposition 2 needs an analogous full-exchangeability or explicit independence condition for E.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops anytime-valid tests for program evaluation by testing the null hypothesis that treatment-effect estimators are exchangeable over pre- and post-treatment periods. The main construction converts the estimators into sequential ranks or reduced sequential ranks and builds sequential e-values from arbitrary non-negative test statistics (Theorems 1 and 2), yielding p-processes with finite-sample type-I error control at all data-dependent stopping times via Ville's inequality. Theorem 3 shows that, under a post-exchangeable alternative, the reduced-rank plug-in statistic dominates the sequential-rank plug-in statistic in expected log-growth. The methodology is illustrated for difference-in-differences and synthetic control in an interactive fixed-effects model, with simulations comparing anytime-valid tests to fixed-T permutation tests under exact and violated exchangeability, including a block-based remedy for serial dependence.","tokens_in":29406,"tokens_out":12926,"duration_ms":142754,"significance":"If the results hold, the paper makes a useful contribution: it supplies finite-sample anytime-valid inference for a class of program-evaluation settings where the number of post-treatment observations need not be pre-specified, and it exploits the pre-treatment batch through reduced sequential ranks. The core e-value proofs are clean and correct: Theorems 1 and 2 are valid conditional e-value constructions, and the Jensen argument in Theorem 3 is sound. The simulations are informative and honestly document size distortions under serial dependence. The main caveat is that the synthetic-control bridge from 'no treatment effect' to 'exchangeable estimators' is not justified as stated, and the abstract's 'optimal' claim is stronger than what is proved.","major_comments":[{"comment":"The proposition imposes assumptions only on t in B union {T0+1,...,T}, but the SCM weights are estimated on the training set E = {1,...,T0}\\B. Nothing in the stated conditions prevents the training period from being dependent with the evaluation period, so the weight vector can be correlated with some evaluation errors but not others. Concretely, take r=0, no covariates, one control unit, so btau_t = eps_1t - w eps_2t with w a function of training data; let E contain t=T0, set eps_1,T0 = eps_1,T0+1 and eps_2,T0 = eps_2,T0+1, and let the remaining evaluation errors be i.i.d. N(0,1). The proposition's evaluation-block conditions then hold, yet w is correlated with eps_T0+1 and not with eps_T0+2, so (btau_T0+1, btau_T0+2) is not exchangeable. Proposition 2 therefore does not establish the claimed exchangeability bridge for SCM. An additional condition analogous to Proposition 1's full exchangeability over t=1,...,T, or an explicit independence restriction on E relative to B union post-treatment periods, is needed.","section":"Section 5.2.2, Proposition 2"},{"comment":"The abstract's phrase 'optimal finite-sample valid sequential tests' and the text's claim of 'a p-value with a type of optimal shrinkage rate' overstate what Theorems 1 and 2 prove. Log-optimality is established only for a fully specified simple alternative with known conditional density (St proportional to gt). For the composite alternatives used in Section 4.1 and the adaptive mixtures of Section 4.2 and Appendix B, no optimality theorem is proved; Theorem 3 gives only a dominance relation between two plug-in statistics. The claims should be qualified, for example as 'log-optimal under a correctly specified simple alternative'.","section":"Abstract and Section 1.1"}],"minor_comments":[{"comment":"The proof of Theorem 3 relies on the fact that the normalization constant in the plug-in e-value can be set to 1, but this is stated only in a footnote. Move this point into the main text or add a cross-reference in the proof of Theorem 3 to avoid confusion.","section":"Section 4.1, footnote 3 and Appendix A.3"},{"comment":"The appendix says the reduced-rank Gaussian statistic conditions on the pre-treatment outcomes, while the main text emphasizes that inference may only depend on the coarsened rank filtration. The reason for marginalizing over the pre-treatment outcomes (to maintain measurability with respect to the rank filtration) should be explained more prominently in the main text.","section":"Section 4.2 and Appendix C"},{"comment":"The discounted utility E[U_S] uses P[H0 rejected by test S at t' <= t] without specifying the probability measure; it should be stated that this probability is evaluated under the alternative.","section":"Section 6.1.3, Eq. (27)"},{"comment":"The abstract's claim that the methods 'control size even under mild exchangeability violations' should be calibrated against the B=1 results in Table 2, where rejection rates reach 0.20-0.26 under serial dependence. The text should clarify that this robustness applies to mild violations or relies on the block structure.","section":"Abstract and Table 2"},{"comment":"There are small typographical issues in the proofs, including 'demoninator' for 'denominator', and the convention 0/0=1 should be stated before the first use in Theorem 1 rather than after it.","section":"Appendix A.1 and A.2"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the central e-value machinery is sound. The main obstacle is Proposition 2, which is a formal statement that currently overclaims: the training period is unconstrained, so the SCM application is not justified as written. This is fixable by adding an appropriate assumption and should be verified against the cited Theorem 2 of Abadie and Zhao (2021), since the paper explicitly describes Proposition 2 as a restatement of that result. The abstract's 'optimal' wording should also be corrected. I do not see a need for new simulations if the proposition is repaired, but the SCM-related claims in Sections 5 and 6 must be aligned with the corrected assumptions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nick,\n\nShort take: the core machinery is real. Theorems 1 and 2 give clean conditional e-value constructions, and Theorem 3's Jensen argument is correct. The reduced sequential rank construction tailored to a pre-treatment batch is a genuine, useful twist on the existing rank-based anytime-valid literature, and the post-exchangeable alternative is well motivated. For difference-in-differences, Proposition 1 is fine because it assumes exchangeability over all time points, including training.\n\nThe soft spot is Proposition 2 for synthetic control. It only assumes exchangeability of (lambda_t,theta_t) and epsilon_t over B union post-treatment. But the SCM weights are estimated from the training period E, and nothing in the proposition constrains how E relates to the evaluation period. The stress-test counterexample is valid: with one control unit, r=0, and a weight w estimated from training data, you can have the first post-treatment error perfectly correlated with the last training error while later post errors are i.i.d. The proposition's stated conditions hold, yet the resulting btau_t sequence is not exchangeable. As written, the SCM anytime-valid application is not justified. The fix is straightforward: assume the full error sequence, training included, is exchangeable, or assume the training errors are independent of the evaluation errors. The simulations effectively do this, so the empirical part survives, but the proposition needs correcting.\n\nThe abstract's \"optimal\" claim is also a bit strong; it is log-optimality under a simple alternative, not a global optimality. And a reader will notice the lack of code and the absence of a real-data application. These are minor relative to the Proposition 2 issue.\n\nWho gets value from this paper? People working on sequential monitoring of treatment effects and the anytime-valid inference community. The reduced-rank trick is likely to be reused. It deserves a serious referee: the core contribution is meaningful, the proofs are careful, and the main flaw is repairable. I would send it out, with a referee asked to verify the SCM assumptions and push for a softened optimality claim.","headline":"The anytime-valid rank-test core is solid and worth engaging, but the SCM bridge in Proposition 2 is missing an assumption on the training period and needs a fix before the application is justified.","tokens_in":29933,"tokens_out":4358,"would_cite":true,"duration_ms":45881,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L10","62G10","62F03","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Rank tests let you monitor treatment effects with exact error control","keywords":["anytime-valid inference","sequential ranks","reduced sequential ranks","exchangeability","e-values","test martingales","difference-in-differences","synthetic control"],"falsifier":"Run the Section 6.1 difference-in-differences simulation under the null with independent normal errors and the stopping rule “reject at the first time $p_t\\le\\alpha$”; the fraction of simulated paths that ever reject must not exceed $\\alpha$ at any horizon. A second run with AR(1) errors should show the same procedure rejecting on up to 20–26% of paths with block size one, which would locate the failure in the exchangeability premise rather than in the rank construction.","tokens_in":28913,"feed_emoji":"📈","tokens_out":10140,"duration_ms":96458,"temperature":0.7,"pith_summary":"This paper tries to establish that program evaluation can be done in real time without pre-specifying how many post-treatment observations will be collected. The trick is to interpret “no treatment effect” as exchangeability of the treatment-effect estimators over the blank and post-treatment periods, and to convert those estimators into sequential ranks. From the ranks the paper constructs e-values whose running product is a test martingale, and the reciprocal of that martingale is an anytime-valid p-value: under the null, the probability that the p-value ever drops below $\\alpha$ is at most $\\alpha$, at every data-dependent stopping time. If this is right, difference-in-differences and synthetic control analyses can be monitored continuously, rejecting early when evidence is strong or continuing past a conventional end date without losing the Type-I error guarantee.","feed_headline":"Rank tests let you monitor treatment effects with exact error control","feed_subtitle":"Anytime-valid p-values from ranks allow early rejection or continued monitoring under exact Type I control.","key_machinery":"The central objects are the sequential rank $R_t$, the position of the newest treatment estimate among all previous estimates, and the reduced sequential rank $\\tilde{R}_t$, its position among the $T_0$ pre-treatment estimates only. Their role is to reduce the highly composite exchangeability null to a simple probabilistic statement: under the null, $R_t$ is uniform on $\\{1,\\ldots,t\\}$, while $\\tilde{R}_t$ has the categorical distribution with cell probabilities $q_i^t=(1+\\#\\{s<t:\\tilde{R}_s=i\\})/t$. This reduction makes it possible to write down closed-form sequential e-values, whose running product is a test martingale; Ville's inequality converts that martingale into an anytime-valid p-value. The theorem structure allows any non-negative test statistic that is conditionally independent of the current rank, and for a simple alternative the log-optimal statistic is proportional to the conditional density of the rank, which is what the Gaussian and plug-in versions implement.","core_discovery":"The paper's central claim is that anytime-valid p-values for the null hypothesis that the treatment-effect estimators $\\hat{\\tau}_t$ are exchangeable can be built from reduced information. Under the exchangeability null, the sequential rank $R_t$ of the newest treatment estimate among all previous estimates is uniform on its support, while the reduced sequential rank $\\tilde{R}_t$, which ranks the newest post-treatment estimate among the $T_0$ pre-treatment estimates only, follows a categorical distribution with probabilities $q_i^t=(1+\\#\\{s<t:\\tilde{R}_s=i\\})/t$. Theorems 1 and 2 show that for any non-negative test statistic that is independent of the current rank given the past, the normalized ratio is a sequential e-value; choosing the statistic proportional to the conditional density under a simple alternative gives the log-optimal test martingale. The running product $W_t$ is a test martingale, and $p_t=1/W_t$ satisfies the anytime-validity bound in Eq. (3), so rejecting at the first time $p_t\\le\\alpha$, or at any later data-dependent time, controls the Type-I error exactly in finite samples. The paper also claims that for post-exchangeable alternatives the reduced-rank construction dominates the sequential-rank plug-in construction in expected log-growth (Theorem 3), and that in the interactive fixed-effects model both difference-in-differences and synthetic control estimators are exchangeable under the null under the conditions of Propositions 1 and 2.","pith_inferences":["A natural extension is to use the same rank reduction on other exchangeable test statistics from program evaluation, such as event-study coefficients or conformal prediction residuals, whenever a batch of pre-treatment values is available at the start of monitoring.","The paper's size distortions under serial dependence suggest the anytime-valid p-value itself can serve as a continuous diagnostic for non-exchangeability: a run of low p-values under a supposedly null treatment is evidence that the estimator's blank-period distribution is not representative of the post-treatment period.","Because the reduced-rank construction intentionally discards the ordering of post-treatment observations, it points to a design trade-off: choosing which statistic to monitor is also choosing which alternatives the test can detect, so monitoring several coarsened statistics simultaneously could protect against both mean shifts and dynamic effects at the cost of a small regret bound."],"forward_implications":["Under the exact exchangeability conditions of Propositions 1 and 2, decision makers can stop at the first time the anytime-valid p-value falls below $\\alpha$, or continue past any pre-planned $T$, and still keep the probability of ever rejecting under the null at most $\\alpha$.","Repeatedly applying a fixed-$T$ permutation test is shown to inflate size to about 20% in the DiD simulation, while the anytime-valid tests remain below $\\alpha$ for up to 1000 post-treatment observations.","For the discounted-utility criterion in Eq. (27), the anytime-valid Gaussian reduced-rank test is preferred over every fixed-$T$ test for discount factors $\\delta\\ge 0.8$ in the stylized DiD setting.","When the post-treatment data are exchangeable among themselves, the reduced-rank construction has weakly higher expected log-growth than the sequential-rank plug-in construction (Theorem 3), and the adaptive mixture over effect sizes recovers the growth rate of the best candidate up to a $\\log k/t$ regret bound.","In the interactive fixed-effects model, both DiD and SCM estimators are exchangeable under the null under the stated assumptions, so the same anytime-valid machinery applies to both estimators."],"supporting_citations":[{"why":"Establishes that counterfactual and synthetic control estimators are exchangeable under the null and supplies the fixed-T permutation inference this paper extends to anytime-valid settings.","marker":"Chernozhukov et al. (2021)"},{"why":"Provides the fixed-T test and the interactive fixed-effects conditions under which DiD and SCM treatment estimates are exchangeable.","marker":"Abadie and Zhao (2021)"},{"why":"Introduces online testing of exchangeability via sequential ranks and the uniformity of ranks under the null.","marker":"Vovk et al. (2003)"},{"why":"Shows that admissible anytime-valid sequential tests must rely on nonnegative martingales, justifying the test-martingale construction.","marker":"Ramdas et al. (2020)"},{"why":"Supplies the log-optimality and safe-testing results that turn likelihood ratios into growth-optimal e-values.","marker":"Grünwald et al. (2024)"},{"why":"Basis for the plug-in method that learns the alternative density from the past ranks.","marker":"Fedorova et al. (2012)"},{"why":"Provides Ville's inequality, which converts a test martingale into an anytime-valid p-value.","marker":"Ville (1939)"},{"why":"Defines the interactive fixed-effects model used to derive exchangeability conditions for DiD and SCM estimators.","marker":"Bai (2009)"}],"fun_headline_variants":["Anytime-valid rank tests enable real-time program evaluation","Monitor treatment effects in real time with anytime-valid ranks","Real-time inference for program evaluation via rank tests","Early rejection or continued monitoring: anytime-valid rank tests","Rank tests give exact error control for sequential evaluation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole construction rests on the treatment-effect estimates being exchangeable over the blank and post-treatment periods when the treatment has no effect; if serial dependence or imperfect factor-loading reconstruction breaks that, the anytime-valid size guarantee no longer holds, and the paper's own simulations show rejection rates can climb to 0.20–0.26 with block size one.","fun_headline_variants_meta":{"raw":{"variants":["Anytime-valid rank tests enable real-time program evaluation","Monitor treatment effects in real time with anytime-valid ranks","Real-time inference for program evaluation via rank tests","Early rejection or continued monitoring: anytime-valid rank tests","Rank tests give exact error control for sequential evaluation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000263,"raw_usage":{"total_tokens":1636,"prompt_tokens":1014,"completion_tokens":622,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":544}},"tokens_in":630,"tokens_out":622,"duration_ms":5695,"temperature":1.0,"reasoning_tokens":544,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:58:19.302416+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Section 6.1 difference-in-differences simulation under the null with independent normal errors and the stopping rule “reject at the first time $p_t\\le\\alpha$”; the fraction of simulated paths that ever reject must not exceed $\\alpha$ at any horizon. A second run with AR(1) errors should show the same procedure rejecting on up to 20–26% of paths with block size one, which would locate the failure in the exchangeability premise rather than in the rank construction.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that counterfactual and synthetic control estimators are exchangeable under the null and supplies the fixed-T permutation inference this paper extends to anytime-valid settings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Basis for the plug-in method that learns the alternative density from the past ranks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Ville's inequality, which converts a test martingale into an anytime-valid p-value."}],"review_version":1}