{"id":"6cc7f9bb-b8c8-421c-9f52-0bd56ac75d69","arxiv_id":"2607.20156","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Forecast-based counterfactual methods can identify aggregate treatment and spillover effects without clean controls or a known exposure mapping, whereas control-based methods cannot under widespread interference.","lead":"This paper compares two ways of building counterfactuals — using untreated control groups versus forecasting what would have happened — for measuring treatment and spillover effects when treatments leak across units. It argues forecast-based methods can still identify overall policy effects even when every unit is contaminated by spillovers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FBCM2/FBCM3 (structural stability and sufficient predictability) are the load-bearing premises; the simulations average out common post-treatment shocks across replications, so the central claim's practical scope is untested.","rationale":"The paper's internal logic is sound under FBCM1-FBCM3, and the identification proofs are straightforward. The weakest point is indeed FBCM3, as the reader states. My concern sharpens this: the simulation protocol averages out the most important violation of FBCM2/FBCM3, namely a common post-treatment shock, so the simulation evidence does not speak to the central claim's robustness in the realistic setting where a citywide or macro shock coincides with the policy. The proposed test would quantify this. This does not change the verdict from CONDITIONAL: the theoretical identification claim remains valid under the assumptions, but the empirical credibility and the scope of the 'more credible' claim are conditional on no such shock. No reason to reject.","tokens_in":36352,"tokens_out":7029,"duration_ms":70036,"concrete_test":"Modify the Section 4.2 simulation to add a common post-treatment shock (e.g., set lambda_6 = 2 for all counties in every replication) and compare MLCM/FAT bias for ATOT, ASEU, and PAPE against CS-DR using true clean controls. If FBCM biases rise by roughly the shock magnitude while CS-DR stays near zero, the FBCM advantage disappears exactly when FBCM2 is violated; repeat over a grid of shock magnitudes and report bias/RMSE. Also report the single-replication distribution of FBCM PAPE bias to show how often a shock would flip the substantive conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that FBCMs identify ATOT/PAPE under pervasive spillovers depends entirely on FBCM2 (structural stability) and FBCM3 (sufficient predictability). If a common post-treatment shock hits all units, the forecast error Y_hat - Y(0,0) is common across units and does not vanish by averaging; every FBCM estimand is then biased by that shock. CBCMs with clean controls difference out common shocks, but FBCMs cannot. The paper's simulation DGP draws an independent common time shock lambda_t each period and averages Monte Carlo bias over replications, so the T0 shock cancels across replications and the near-zero MLCM/FAT biases do not reveal the method's vulnerability in a single realized sample. The empirical application (Buenos Aires car theft) is precisely a setting where citywide crime shocks are plausible; the placebo table (F1) only uses the last pre-treatment month and has low power. Thus the practical claim that FBCMs 'can more credibly identify' aggregate effects in pervasive-spillover settings is not established for the realistic case of a common shock.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a unified potential-outcomes framework for treatment and spillover effects, comparing control-based counterfactual methods (CBCMs), such as difference-in-differences, with forecast-based counterfactual methods (FBCMs), such as MLCM and FAT. It defines aggregate estimands (ATOT, ASEU, PAPE), states identifying assumptions (CBCM1–3, FBCM1–3), and proves which estimands are identified under four interference regimes ranging from no spillovers to pervasive spillovers with unobserved exposure. The theoretical results are complemented by Monte Carlo simulations based on U.S. county geography under true and misspecified exposure mappings, and by an empirical application to police protection and car theft in Buenos Aires. The central claim is that FBCMs can identify aggregate policy effects even when all units are exposed to spillovers and no clean control group exists, a setting where CBCMs generally fail.","tokens_in":36668,"tokens_out":4271,"duration_ms":46611,"significance":"If the claims hold, the paper makes a useful contribution by clarifying the identifying content of forecast-based counterfactuals in interference settings. The distinction between the existence and observability of clean controls (Cases ii.a vs. ii.b) is valuable, and the explicit treatment of estimands defined relative to the no-treatment/no-exposure baseline is careful. The simulation design is thoughtful: using true and misspecified exposure mappings and reporting target-group composition effects is informative. The paper also honestly notes several limitations, including the dependence of FBCMs on structural stability and predictability. The formal propositions are algebraically correct under the stated assumptions. The main open question is whether the practical conditions under which FBCMs work are sufficiently credible in realistic settings with common shocks and short pre-treatment histories.","major_comments":[{"comment":"These propositions claim identification of ASEU when all units are exposed but exposure intensity is unobserved. In the distance-decay DGP, exposure is continuous and every untreated unit has strictly positive exposure, so 'exposed untreated' coincides with all untreated units. This is a special case of Case (iii). If exposure is unobserved but some untreated units have zero exposure, ASEU is not identified by FBCMs. The propositions should be stated with the conditioning set explicitly defined: for universal exposure, the average over untreated units is the ASEU only because the exposed set is the entire untreated population. Otherwise, the latent-selective-interference case (Proposition FBCM3) already covers the unobserved-strata scenario. This is a clarity issue rather than a mathematical error, but it affects the interpretation of Table 4.","section":"Section 3.2.2, Propositions FBCM4–FBCM5"}],"minor_comments":[{"comment":"The sentence 'Third, we formally demonstrates that FBCMs...' contains a subject-verb agreement error ('demonstrates' should be 'demonstrate').","section":"Introduction"},{"comment":"The text says 'standard deviation 1' for common time shocks and later 'common time shocks are independently drawn from a normal distribution with standard deviation 1'; this is fine, but the AR(1) coefficient is given as 0.60 in the text while Appendix E uses 0.6. Please standardize notation.","section":"Section 4.1"},{"comment":"The notes for Table 5 refer to 'ATOT under clean-control assumption' in the row label, but the row itself is labeled 'ATOT under clean-control assumption: treated blocks vs controls more than two blocks away.' Consider making the label shorter and moving the assumption descriptor to the notes for readability.","section":"Section 5, Table 5"},{"comment":"The placebo table is an important validity check, but it covers only the first 17 days of July, which is a very short window. Please state explicitly that this is a low-power test, and if possible, report a placebo for earlier months (e.g., June) using the same recursive forecasting scheme.","section":"Appendix F, Table F1"},{"comment":"The reference list includes 'Botosaru, I., Giacomini, R., & Weidner, M. (2026)' which is an arXiv preprint; if a published version exists, please update. Also, 'Callaway, B., & Sant’Anna, P. H. C. (2021)' is cited in text but not listed in the reference list; it appears to be missing from the references.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically sound in its formal identification propositions, but the practical scope is overstated. The authors are strongly encouraged to address the common-shock vulnerability explicitly, either by extending the method or by qualifying the claims. The reliance on the authors' own MLCM as the leading FBCM example is understandable, but the novelty relative to their prior work should be clarified. The self-citation pattern is not problematic per se, but the editor may want to ensure the contribution is framed as a generalization/comparison rather than a new estimator."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper gives applied researchers a clear map of what control-based (CBCM) and forecast-based (FBCM) counterfactuals can and cannot identify under different spillover regimes. The novelty is the FBCM identification results for latent selective and pervasive interference: aggregate estimands like ATOT and PAPE can be recovered even when no clean control exists, provided the baseline process is stable and forecastable. The propositions are simple algebra, and they are correct under the stated assumptions. The simulation design is genuinely thoughtful, comparing a true queen-contiguity mapping with misspecified distance bands, and the authors properly flag when forecast-based invariance across mappings is mechanical.\n\nThe soft spots are in the gap between the identification logic and the practical claim. The load-bearing assumptions are FBCM2 (structural stability) and FBCM3 (sufficient predictability). With five pre-treatment periods, FBCM3 cannot be credibly verified. More specific problem: the simulation DGP draws a fresh common time shock each Monte Carlo replication and averages bias over replications. A realized post-treatment common shock therefore cancels out in the reported near-zero MLCM/FAT biases. In any single sample, a citywide shock is indistinguishable from the policy effect and biases every FBCM estimand, because there are no controls to difference it out. The stress-test note about this is right, and the paper does not address it. The Buenos Aires application is exactly a setting where citywide crime shocks are plausible, and the placebo check uses only the last pre-treatment month, low power.\n\nAlso: Appendix F has two incompatible descriptions of the MLCM standard errors (bootstrap vs cross-sectional dispersion), and no replication code or data are provided.\n\nNone of this sinks the paper. The identification framework is coherent and novel, and the authors are honest about the core fragility of FBCMs. But the abstract's claim that FBCMs 'can more credibly identify' aggregate effects under pervasive spillovers is not established for the common-shock case. That needs to be reframed.\n\nWho benefits: applied researchers choosing between DiD and forecast-based designs, and methodologists working on interference. It deserves a serious referee. I would send it to review, with the practical claims trimmed and a common-shock analysis added.","headline":"A genuinely useful identification framework for spillovers, but the practical claim that FBCMs can handle pervasive interference is too strong because the load-bearing stability assumption is untestable, and the simulations average away the common-shock case where the method collapses.","tokens_in":37106,"tokens_out":3877,"would_cite":true,"duration_ms":36687,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Causal effects can be identified without clean control units by replacing cross-sectional comparisons with forecast-based counterfactuals, even when spillovers reach every unit.","keywords":["potential outcomes","spillovers","interference","forecast-based counterfactuals","difference-in-differences","machine learning control method","treatment effects","policy evaluation"],"falsifier":"Apply a forecast-based method to a panel with a known structural break at a date unrelated to any policy; a systematically nonzero forecast-based gap for a placebo treatment at that date shows the stability/predictability assumption is violated. In the Buenos Aires data, forecast car theft for blocks farthest from protected institutions and check whether the positive gap observed after August 1994 persists in placebo pre-treatment periods; if it does, the gap reflects a baseline shift rather than displacement.","tokens_in":36261,"feed_emoji":"📈","tokens_out":5747,"duration_ms":53487,"temperature":0.7,"pith_summary":"This paper argues that when a policy's effects spill over to every unit—so that no untouched control group exists—causal effects can still be identified by forecasting each unit's no-policy outcome from its pre-treatment history and comparing observed outcomes to those forecasts. The authors show that control-based methods such as difference-in-differences require clean controls that are both present and observable, and collapse when spillovers are widespread or exposure is unobserved. Forecast-based methods, including interrupted time series and machine-learning control methods, instead identify aggregate policy effects such as the Average Total Effect on the Treated and the Population Average Policy Effect, without modelling how spillovers propagate. This comes at the price of assuming that the no-policy outcome process is structurally stable and sufficiently predictable, an assumption that may be harder to sustain over long horizons. Simulations and a police-protection natural experiment illustrate the trade-off.","feed_headline":"Spillovers everywhere? Forecast the counterfactual instead","feed_subtitle":"Control-based estimators fail without clean controls; forecast-based methods recover aggregate policy effects from pre-treatment dynamics.","key_machinery":"The central object is the forecast-based counterfactual: an estimate of the no-treatment/no-exposure potential outcome Y_it(0,0) obtained by extrapolating each unit's pre-treatment outcome and covariate dynamics into the post-treatment period (the Machine Learning Control Method is the leading example). The three assumptions—no anticipation, structural stability of the baseline outcome process, and sufficient predictability of that process—are what make the forecast a valid counterfactual. The forecast-based gap Y_it − Ŷ_it(0,0) then carries the causal content, identifying total effects on treated units and spillover effects on untreated units, and aggregating to the population average poli","core_discovery":"The paper's central claim is that forecast-based counterfactual methods (FBCMs) can point-identify the aggregate causal estimands ATOT and PAPE even under pervasive spillovers where every unit is exposed and no clean control exists, a setting in which control-based counterfactual methods (CBCMs) cannot identify these estimands without additional strong assumptions. The identifying device is the forecast-based gap: for each unit, forecast the no-treatment/no-exposure potential outcome Y_it(0,0) from pre-treatment dynamics, then compare the observed outcome to that forecast. Under no anticipation, structural stability of the baseline process, and sufficient predictability, the average of these","pith_inferences":["The paper's logic suggests a testable extension: in a randomized experiment with a known spillover structure, compare forecast-based estimates with the true aggregate effect under different pre-treatment lengths to map how the convergence assumption degrades with short panels.","The framework implies that when exposure is latent, direct and indirect effects are not separately recoverable, but total and population effects are—so policy evaluation can proceed without solving the often-impossible exposure-mapping problem, at least for aggregate questions.","Because forecast-based identification relies on time-series stability, long-run policy evaluation would require external validation of the baseline model (e.g., placebo forecasts in pre-treatment periods), not just cross-sectional balance checks."],"forward_implications":["Policies that operate through general-equilibrium or network channels can be evaluated even when all units are indirectly treated, provided credible forecasts of the no-policy path can be constructed.","Exposure-specific spillover estimates still require observing exposure; forecast-based methods do not remove that need.","The credibility of the design shifts from assumptions about the spillover structure to assumptions about the stability and predictability of baseline outcomes, which are more plausible for short-term horizons.","In the Buenos Aires application, forecast-based estimates find a positive population-wide effect (0.029 and 0.022 car thefts per block) even though protected blocks themselves show fewer thefts, because spillover effects on more distant blocks outweigh local deterrence.","In the simulations, misspecifying the exposure mapping biases control-based estimates through contaminated controls, while forecast-based ATOT and PAPE remain unchanged because they do not depend on the assumed mapping."],"fun_headline_variants":["No clean controls? Forecast the counterfactual","Pervasive spillovers? Forecast beats matching","Forecast-based gaps identify aggregate spillover effects","When everyone's exposed, forecast the baseline","Forecast counterfactuals for pervasive spillover settings"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire edifice rests on the assumption that the no-policy outcome process is stable and sufficiently predictable after the treatment begins: if a structural break or non-forecastable shock hits the baseline outcome right after T0, the forecast-based gap is biased, and with only five pre-treatment periods in the empirical application this convergence premise cannot be credibly verified.","fun_headline_variants_meta":{"raw":{"variants":["No clean controls? Forecast the counterfactual","Pervasive spillovers? Forecast beats matching","Forecast-based gaps identify aggregate spillover effects","When everyone's exposed, forecast the baseline","Forecast counterfactuals for pervasive spillover settings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1112,"prompt_tokens":736,"completion_tokens":376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":480,"tokens_out":376,"duration_ms":4647,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:36:34.595412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply a forecast-based method to a panel with a known structural break at a date unrelated to any policy; a systematically nonzero forecast-based gap for a placebo treatment at that date shows the stability/predictability assumption is violated. In the Buenos Aires data, forecast car theft for blocks farthest from protected institutions and check whether the positive gap observed after August 1994 persists in placebo pre-treatment periods; if it does, the gap reflects a baseline shift rather than displacement.","supporting_citations":[],"review_version":1}