{"id":"d30d09bd-c067-4b33-8ed2-57cbd45d7716","arxiv_id":"2608.13466","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Under unknown interference, the sharp identified set for a panel treatment effect is characterized exactly by feasibility of a finite linear system, enabling uniform candidatewise bootstrap inference.","lead":"Panel treatment-effect estimates are contaminated when the policy also changes outcomes in comparison units. This paper derives sharp bounds on the treatment effect without knowing the spillover pattern, reduces the problem to a finite linear system, and shows in the Arizona immigration case that the sign of the effect cannot be determined.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniform coverage theorem depends on unverified Assumption 4.2; its Hoffman-bound and positive-variance conditions can fail for model-generated systems, and the reported inversion sets are not robust without verification.","rationale":"The identification core is carefully built and the finite representation proof (Lemma 3.1, Theorem 3.1) checks out: the minimax exchange over the simplex and the box is valid, and the construction of completions in Proposition 3.1 is sound. The Monte Carlo design is thoughtful and the application is honest about the sign being unresolved. My concern is narrower: the uniform candidatewise coverage result in Theorem 4.1 is the paper's main inferential claim, and it rests on Assumption 4.2(ii)-(iii), which are high-level and not derived from primitives. The Hoffman-bound assumption (ii) can fail when pre-treatment gap vectors are near-collinear, and the fixed-row box constraints in the finite system generate zero-variance dual recession directions that can violate part (iii). The paper itself flags these as regularity conditions and offers a conservative calibration in Remark D.2 that avoids (iii), but the reported sets do not use it. This does not make the paper wrong, but it means the empirical 95% inversion sets carry an external, unverified condition. The concrete test — recomputing with the conservative calibration — would show whether the sign-unresolved conclusion is robust. Therefore I keep the reader's CONDITIONAL verdict.","tokens_in":40193,"tokens_out":14659,"duration_ms":148072,"concrete_test":"Recompute the LAW A inversion sets and the Monte Carlo endpoints using the conservative calibration of Remark D.2: reject only when √n Q̂_n(τ) exceeds ĉ_n(τ,1−α) by a positive ε, for example ε = 0.05 or 0.1, holding all other implementation choices fixed. If any reported inversion set changes by more than the grid tolerance or ceases to contain zero, the paper's empirical sign-unresolved conclusion depends on the unverified part (iii). If the sets are essentially unchanged, the qualitative conclusion survives and the concern is about theoretical generality rather than the application.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference claim (Theorem 4.1) requires Assumption 4.2(ii)-(iii), which are stated as high-level conditions with no primitive sufficient conditions. Part (ii) assumes a uniform feasibility error bound (a Hoffman bound) over all compatible systems. The coefficient matrix A(Π) contains the donor pre-treatment gaps C_t(e_k); as any such gap shrinks or the 46 donor gap vectors become nearly collinear — exactly the near-exact geometry the paper's own Monte Carlo studies — the Hoffman constant can diverge, so the uniform bound need not hold. Part (iii) requires every null candidate's nonzero dual maximizer to have certificate variance at least c_Γ. The model's fixed v± box rows create recession directions of the dual feasible set with zero variance; if the only certificates that cancel the estimated rows pass through these fixed directions, (iii) fails. The paper's Remark D.2 shows a conservative calibration can avoid (iii), but the main theorem and the reported 95% inversion sets do not use that calibration. A user cannot tell from the stated primitives when the claimed uniform coverage applies; without (ii)-(iii), the empirical inversion sets are not guaranteed to control false exclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies partial identification and inference for a scalar treatment effect in a panel with one treated aggregate unit and K-1 comparison units when comparison units may also respond to the treatment through unknown interference. It imposes Assumption 3.1, a fit-scaled comparison-validity bound that must hold uniformly over all convex donor weights, and combines it with prespecified restrictions on the treatment effect and spillover vector (Assumptions 3.2-3.3). Proposition 3.1 gives the sharp identified set as a projection; Lemma 3.1 and Theorem 3.1 show that candidate compatibility is exactly equivalent to feasibility of a finite linear system, despite the continuum of weights. Section 4 builds candidatewise compatibility tests following Goff and Mbakop (2026) and states a uniform candidatewise coverage theorem under Assumptions 4.1-4.2. The LAWA application reports 95% inversion sets that contain zero and the original Bohn et al. decline across all reported L and rho specifications. Monte Carlo designs illustrate that interior weights can sharply narrow the population identified set when pre-treatment gaps cancel, while sampling uncertainty attenuates the gain.","tokens_in":40316,"tokens_out":3849,"duration_ms":39172,"significance":"The exact finite representation is a genuine conceptual advance: it reduces a continuum of donor-weight restrictions to finitely many rows plus box constraints, and the identified set is exactly the scalar projection of a polyhedron (Corollary 3.1). The proof of Lemma 3.1 via Sion's minimax theorem is clean and the equivalence is exact rather than approximate. The paper also shows good empirical discipline: the placebo calibration is explicitly descriptive, the inversion sets are reported as candidatewise, and the Monte Carlo isolates the geometric mechanism behind interior-weight restrictions. If Theorem 4.1 were accompanied by verifiable primitive conditions, the uniform candidatewise coverage guarantee would be a meaningful addition to inference for linear feasibility problems. The application's negative result is an honest illustration of how little can be learned under unknown interference.","major_comments":[{"comment":"Assumption 4.2(ii) is a uniform feasibility error bound over all compatible systems, but no primitive sufficient conditions are given. The coefficient matrix A(Pi0) contains the donor pre-treatment gaps C_t(e_k); when those gaps are nearly collinear or some gap becomes very small, the Hoffman constant can diverge, so the uniform bound need not hold. This is not a remote pathology: the near-exact design in Section 6.1 is constructed so that the minimum average absolute pre-treatment discrepancy is 0.000201 and the full-simplex projection contracts by 86 percent, exactly the geometry in which A(Pi0) is ill-conditioned. The theorem therefore does not currently establish uniform coverage for the model family that the paper's own Monte Carlo treats as the interesting case. The authors should either provide verifiable primitive conditions (for example, a uniform lower bound on the relevant singular values, or an explicit finite Hoffman bound for the structured bridge system) or restrict the DGP class P_U accordingly.","section":"Section 4.3, Assumption 4.2(ii)"},{"comment":"Assumption 4.2(iii) requires every nonzero population dual maximizer to have certificate variance at least c_Gamma, with a nearby extreme-point clause. The fixed zero-scale rows in the bridge representation (3.11) create recession directions of the dual feasible set; if the only multipliers that cancel the estimated rows pass through such fixed directions, the condition fails. The paper's Remark D.2 shows that a shifted rejection rule preserves size without part (iii), but the main theorem and the reported 95% inversion sets do not use that calibration. As written, a user cannot determine from the stated primitives whether (iii) holds; without it, the false-exclusion guarantee in Theorem 4.1 is not operational for the reported specifications.","section":"Section 4.3, Assumption 4.2(iii) and Remark D.2"},{"comment":"The inversion set in (4.3) is defined over a continuum of candidates, but the implementation uses an adaptive grid and bisection, and Appendix C states that coverage applies only to candidates actually tested while interpolation is a numerical display. The reported endpoints and widths in Figure 3 are therefore not themselves covered by Theorem 4.1 unless every displayed candidate was actually tested. The paper should state precisely which grid points were tested, how interpolation between them is reconciled with the candidatewise guarantee, and whether the uniform coverage claim in (4.4) is for the tested set or for the displayed interval.","section":"Section 4.2 and Appendix C"}],"minor_comments":[{"comment":"The use of T for a scalar post-treatment target while T0 denotes the last pre-treatment period is easy to misread; consider a notation such as t_post to avoid confusion.","section":"Section 2"},{"comment":"The placebo benchmarks do not test Assumption 3.1; the text says this, but the figure placing L=1 below all three benchmarks could be misread as a falsification test. An explicit sentence that the zero floor and uniformity across the simplex are maintained, not calibrated, would help.","section":"Section 5.3 and Figure 3"},{"comment":"The table notes label 'Lower' and 'Upper' as rejection frequencies at compatible endpoints; consider defining 'false exclusion' explicitly in the notes, and state that the sign-aligned design is a worst-case geometry for interior-weight information in which the triangle inequality of Proposition D.1 binds.","section":"Table 2 and Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The identification and finite-representation results are solid; the main risk is the gap between the high-level Assumption 4.2 and the empirical implementation. I would recommend asking for primitive sufficient conditions for Assumption 4.2(ii)-(iii), or otherwise presenting the main empirical results under the conservative calibration of Remark D.2 so that the reported 95% sets carry a verifiable coverage guarantee."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing worth knowing about this paper is the identification result. Reducing the continuum of donor weights to finitely many linear inequalities via Sion's minimax (Lemma 3.1, Theorem 3.1) is genuinely new and cleanly proved. The sharp projection onto the treatment effect is a real contribution, and the paper deserves credit for not overclaiming: the LAW A inversion sets contain zero and the original estimate, and the sign stays unresolved. That is honest work.\n\nWhat I'd push on is the inference layer. Theorem 4.1 is a structured specialization of Goff and Mbakop, and the paper says so. But Assumption 4.2(ii)-(iii) are high-level: a uniform Hoffman bound over model-generated systems, and a certificate-variance lower bound. The stress-test note is right that both can fail in the near-exact geometries the paper itself emphasizes. A shrinking donor gap or near-collinear gaps can blow up the Hoffman constant; the fixed box rows can create zero-variance recession directions. The paper's own Remark D.2 shows a shifted rejection rule avoids part (iii), but the main theorem and the reported 95% sets do not use that calibration. So a reader cannot tell from the stated assumptions when the uniform coverage claim actually applies. The Monte Carlo is reassuring—no excessive rejection in the cells—but it is not a substitute for primitives.\n\nAssumption 3.1 is also strong: uniformity over the entire simplex plus a zero floor at exact fit. The factor-span condition in Section 3.4 gives one sufficient route, but it's not a primitive check either. I would not call this fatal; the envelope L is a sensitivity parameter and the placebo calibration is honest descriptive benchmarking. But the uniformity across all convex weights is a substantive commitment, and the application does not test it.\n\nCitation pattern looks fine—Liu (2025) and Mealli-Viviens are the right anchors, and the paper is explicit that inference borrows from Goff-Mbakop. No code is shipped, which matters less for a theory paper but would strengthen reproducibility for the application.\n\nBottom line: this deserves a serious referee. The identification half is solid and novel; the inference half needs either primitive sufficient conditions for Assumption 4.2 or a clear statement that the reported sets are conditional on those high-level conditions plus the Remark D.2 calibration. I'd take it to reading group, and I'd probably cite the identification result, but I'd be cautious about citing the uniform coverage claim without more groundwork.","headline":"The identification core is new and clean—full-simplex validity reduced to a finite LP via minimax—but the uniform coverage theorem rests on high-level Assumption 4.2 that the paper never grounds in primitives, so the empirical inversion sets should be read as conditional on unverified regularity.","tokens_in":40947,"tokens_out":1134,"would_cite":true,"duration_ms":13563,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When comparison units are also affected by the policy, the treatment effect is still partially identified.","keywords":["treatment effects","unknown interference","partial identification","panel data","synthetic control","linear programming feasibility","bootstrap inference","spillovers"],"falsifier":"A placebo test would settle the assumption: hold out each pre-treatment period in turn and compute the smallest envelope $L$ that covers that held-out movement uniformly over the full donor simplex. If the maximum of these placebo envelopes substantially exceeds the chosen $L$, or if an interior weight with near-zero pre-treatment fit still shows a large held-out gap, Assumption 3.1's uniformity and zero floor are contradicted by the pre-treatment data on which the calibration is based.","tokens_in":39876,"feed_emoji":"🎯","tokens_out":7592,"duration_ms":71617,"temperature":0.7,"pith_summary":"This paper is trying to establish that panel comparisons remain informative about a treatment effect even when the comparison units themselves respond to the policy in unknown ways. Its route is to impose, for every convex donor weight, a fit-scaled validity bound that keeps each weight's latent no-policy gap after treatment within a prespecified multiple of that same weight's average absolute gap before treatment, and to intersect those bounds with application-specific restrictions on the spillover vector. The paper claims this intersection is sharp, that compatibility of a candidate effect $\\tau^c$ is exactly the feasibility of a finite linear system, and that inverting bootstrap-calibrated feasibility tests gives uniform candidatewise coverage. A reader should care because common policies, such as Arizona's employer verification law, plausibly move comparison states too, and this framework says what can still be learned without having to name which donors are contaminated.","feed_headline":"Spillovers unknown? Treatment effect still partially identified","feed_subtitle":"A fit-scaled envelope over every donor weight plus spillover restrictions yields a sharp, finitely checkable identified set.","key_machinery":"The central object is the full-simplex fit-scaled validity envelope: for every convex weight $w$, $|C_T^{\\mathrm{obs}}(w)-w^\\top x|\\le \\frac{L}{T_0-1}\\sum_{t=2}^{T_0}|C_t(w)|$, where $x_k=\\tau-s_k$ are relative effects. The paper converts this continuum of restrictions into a finite linear system by writing each weight's allowance as a support function, applying Sion's minimax theorem to interchange the simplex and the box of auxiliary coefficients, and using the fact that a linear function on the simplex attains its maximum at a donor vertex. This yields two certificate vectors $v^-$ and $v^+$; stacked with the rows of the prespecified admissible rule, they form a fixed-dimension system whose coefficient matrix does not depend on the candidate $\\tau^c$. That candidate-invariant finite representation carries the argument from sharp identification to uniform test inversion.","core_discovery":"The central claim is Theorem 3.1: under full-simplex fit-scaled comparison validity and a finite linear representation of additional restrictions, a candidate treatment effect $\\tau^c$ is compatible if and only if the finite linear system $F(\\Pi_0,\\tau^c)=\\{\\eta: b(\\Pi_0,\\tau^c)-A(\\Pi_0)\\eta\\le 0\\}$ is nonempty. The continuum of donor weights collapses to finitely many inequalities indexed by donor vertices together with bounded auxiliary vectors $v^-$ and $v^+$, so interior weights add identifying information exactly when convex aggregation cancels pre-treatment discrepancies. Consequently the identified set for $\\tau$ is the scalar projection of a polyhedron, a closed interval, half-line, or the whole real line, and its boundedness is decided by whether the recession cone of the admissible set blocks the common-shift direction. For inference, the paper proves that the normalized Farkas score, calibrated with candidate-specific recentered bootstrap draws, uniformly controls the probability of falsely excluding each compatible candidate in large samples.","pith_inferences":["The zero floor at exact fit does more work than the multiplier $L$: an interior weight with zero average pre-treatment gap is forced to have zero latent post-treatment gap. The paper's additive-floor relaxation is the natural robustness check, and it preserves the finite-system construction.","A placebo diagnostic follows directly from the framework: hold out each pre-treatment period and test the envelope uniformly over the simplex. If several held-out periods require envelopes far above the chosen $L$, Assumption 3.1's uniformity is implausible even when every donor vertex looks fine.","Because the coefficient matrix is candidate-invariant, the same bootstrap draws could support simultaneous inference over a grid of $(L,\\rho)$ specifications, although the paper only claims candidatewise coverage within a fixed specification.","If the admissible set only restricts relative effects, the common-shift direction remains completely unanchored, so reported interval endpoints would then be artifacts of the reporting domain rather than identified bounds."],"forward_implications":["Interior donor weights are not redundant: in the paper's near-exact geometry, full-simplex validity shrinks the population identified set by 86 percent relative to donor-vertex validity.","The identified set for the treatment effect is always a polyhedron in $\\mathbb{R}$, so an application either gets an interval, a singleton, a half-line, or the whole real line.","Because checking a candidate reduces to linear programming feasibility, the same solver and bootstrap perturbations can be reused across candidates.","In the LAW A application, the 95 percent inversion sets contain both zero and the original $-1.50$ percentage point estimate at every reported specification, so the sign of the effect remains unresolved.","The uniform candidatewise coverage guarantee protects each compatible value against false exclusion without any multiplicity adjustment over candidates."],"supporting_citations":[{"why":"Supplies the normalized solvability statistic and bootstrap calibration that the paper specializes to its model-generated family of candidate systems.","marker":"Goff and Mbakop (2026)"},{"why":"The closest predecessor studying the identifying content of multiple comparison weights, which the full-simplex intersection builds on.","marker":"Liu (2025)"},{"why":"Provides the two-group difference-in-differences accounting relation under unknown interference that the many-donor framework extends.","marker":"Mealli and Viviens (2026)"},{"why":"The relative-magnitude sensitivity template for bounding post-treatment departures using pre-treatment behavior.","marker":"Rambachan and Roth (2023)"},{"why":"Calibration principle used to express the sensitivity parameter on an observed placebo scale.","marker":"Hsu and Small (2013)"},{"why":"Minimax theorem used to interchange the simplex and coefficient-box optimization in the exact finite representation.","marker":"Sion (1958)"},{"why":"Supplies the LAW A data, outcome definition, donor exclusions, and the benchmark estimate that the application re-examines.","marker":"Bohn et al. (2014)"}],"fun_headline_variants":["Sharp treatment effect set under unknown spillovers","Unknown interference? Sharp bounds on treatment effect","Finite linear check yields sharp treatment effect set","Panel data partial identification without exposure mapping","Bootstrap inversion sets treatment effect sign unresolved"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that one fixed multiplier $L$ uniformly bounds every convex donor weight's unobserved no-policy gap after treatment by that same weight's average absolute pre-treatment gap, with no floor at exact fit; if this uniformity fails for some interior weight, the sharp set and finite system characterize a different model.","fun_headline_variants_meta":{"raw":{"variants":["Sharp treatment effect set under unknown spillovers","Unknown interference? Sharp bounds on treatment effect","Finite linear check yields sharp treatment effect set","Panel data partial identification without exposure mapping","Bootstrap inversion sets treatment effect sign unresolved"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000536,"raw_usage":{"total_tokens":2577,"prompt_tokens":951,"completion_tokens":1626,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":1560}},"tokens_in":567,"tokens_out":1626,"duration_ms":12421,"temperature":1.0,"reasoning_tokens":1560,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:21:41.639982+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A placebo test would settle the assumption: hold out each pre-treatment period in turn and compute the smallest envelope $L$ that covers that held-out movement uniformly over the full donor simplex. If the maximum of these placebo envelopes substantially exceeds the chosen $L$, or if an interior weight with near-zero pre-treatment fit still shows a large held-out gap, Assumption 3.1's uniformity and zero floor are contradicted by the pre-treatment data on which the calibration is based.","supporting_citations":[{"cited_title":"Testing the Solvability of Systems of Linear Inequalities","cited_arxiv_id":"2506.06776","evidence_quote":"Supplies the normalized solvability statistic and bootstrap calibration that the paper specializes to its model-generated family of candidate systems."},{"cited_title":"Difference-in-Differences in the Presence of Unknown Interference","cited_arxiv_id":"2512.21176","evidence_quote":"Provides the two-group difference-in-differences accounting relation under unknown interference that the many-donor framework extends."},{"cited_title":"A more credible approach to parallel trends","cited_arxiv_id":null,"evidence_quote":"The relative-magnitude sensitivity template for bounding post-treatment departures using pre-treatment behavior."}],"review_version":1}