REVIEW 3 major objections 5 minor
The experiment's randomization can estimate and test the spillover map
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-10 03:53 UTC pith:CZOZXLX7
load-bearing objection Real methodological contribution; the headline empirical result needs more scrutiny the 3 major comments →
A Design-Based Approach to Testing and Inference in (Quasi-)Experiments with Spillovers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A correctly specified exposure map implies exposure sufficiency: conditional on the exposure value, the remaining variation in the treatment assignment carries no outcome-relevant information. This property generates orthogonality conditions — between outcomes and design-side residuals of assignment functions — that are testable and estimable using only the known randomization design, without any outcome model. These conditions enable both GMM estimation of the exposure parameter and overidentification testing of the exposure specification itself.
What carries the argument
The design-side residual R_{i,θ}^ψ(W) = ψ(W) − E[ψ(W) | g(W; X_i, θ)] strips from any assignment function the component explained by the candidate exposure. Its orthogonality to any outcome function — E[ϕ(Y_i) R_{i,θ₀}^ψ(W)] = 0 — is the testable consequence of exposure sufficiency. The conditional expectation is computed by repeatedly simulating from the known assignment design, so the residual is a pure design object requiring no outcome modeling.
Load-bearing premise
The framework assumes that a single parametric exposure map captures all the ways the treatment assignment affects each unit's outcome — that once you know the exposure value, no other feature of who was treated matters. If spillovers flow through multiple channels that the map does not jointly summarize, the assumption fails, and the test can detect some violations but cannot certify that the map is fully correct.
What would settle it
If one took a setting with known multi-channel spillovers (e.g., both price effects and direct network effects operating at different spatial scales) and fit a single ring exposure map, the J-test should reject — and the rejection should weaken as the map is enriched to capture both channels.
If this is right
- Researchers studying spillovers can let the data choose the radius, decay rate, or network hop count rather than reporting results under several ad hoc values, with uncertainty about that choice properly propagated into downstream estimates.
- The overidentification J-test provides a formal specification check: if a ring specification is rejected, the researcher knows spillovers extend beyond the assumed radius and can revise accordingly.
- The framework extends to quasi-experimental settings where the design is a placebo or permutation scheme, so the logic applies beyond literal randomized trials.
- An efficiency bound within the class of unit-level moments identifies which outcome transformations and assignment functions are most informative, guiding the construction of the test.
- Downstream policy estimates — treatment effects, multipliers, counterfactual predictions — can incorporate uncertainty about the exposure specification itself, not just sampling uncertainty conditional on a fixed choice.
Where Pith is reading between the lines
- If spillovers operate through multiple channels that no single parametric map jointly summarizes (e.g., both price-mediated and network-mediated effects), the exposure sufficiency hypothesis fails. The J-test may detect this but cannot diagnose which channel is missing, so rejection leaves the researcher without a constructive direction beyond widening or reshaping the map.
- The framework could be extended to formally compare competing exposure map families — ring versus gravity versus network — on the same data using their respective J-statistics, providing a principled basis for the informal specification comparisons that applied papers currently make by reporting multiple specifications side by side.
- Exposure maps for which the assignment design leaves little residual variation after conditioning (near-injective maps) are fundamentally untestable, which may explain why some specifications in applied work are empirically indistinguishable regardless of sample size.
- The contrast between the two applications — one validated, one rejected — raises the question of whether the framework's practical value depends on ex ante proximity of the original specification to the truth, and whether pre-analysis plans for spillover maps could benefit from design-based pre-registration.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops a design-based framework for testing and estimating exposure mappings in settings with spillovers. The core idea is that a correctly specified exposure map g(W; X_i, θ₀) implies an exposure sufficiency property: conditional on the exposure, the remaining variation in the assignment vector is orthogonal to outcomes. This generates testable moment conditions E[ϕ(Y_i) R_{i,θ₀}^ψ(W)] = 0, where R is a design-side residual constructed by orthogonalizing a design function ψ(W) with respect to the exposure using the known randomization distribution. These moments enable GMM estimation of the tuning parameter θ₀ and, when overidentified, a J-test of the exposure specification. The author establishes consistency and asymptotic normality under spatial and network dependence via the affinity-set framework of Chandrasekhar et al. (2023), characterizes an efficiency bound within the unit-level moment class using Riesz representation, and propagates Stage-1 uncertainty into downstream policy estimands. Two applications to large-scale anti-poverty programs illustrate the framework: the design-based test supports the 20 km radius in Muralidharan et al. (2023) but rejects the 2 km radius in Egger et al. (2022), yielding a revised local fiscal multiplier of 1.57 (vs. the original 2.5).
Significance. The paper addresses a genuine and important gap in the spillovers literature. While exposure mappings are widely used in applied work, researchers typically lack principled tools for choosing the functional form or calibrating tuning parameters, and standard errors rarely propagate this specification uncertainty. The design-based orthogonality conditions are a clean and novel contribution, transforming the exposure map from a fixed researcher choice into an object of inference. The theoretical development is careful: identification (Theorem 2.6, Corollary 2.7) is clean, the asymptotic theory (Theorems 3.3, 3.5) correctly adapts Z-estimator arguments to the affinity-set CLT, and the efficiency bound (Theorem 4.1, Appendix C) is a coherent Riesz representation argument. The applications are well-chosen and the contrasting conclusions are informative. The framework is likely to be influential in applied microeconomics.
major comments (3)
- The most load-bearing concern is at the implementation layer. Corollary 3.6 establishes the J-test under the assumption that the population moments are exactly zero at θ₀. In practice (Remark 2.8), the conditional expectation E[ψ_m(W)|g(W;X_i,θ)] is approximated by regressing ψ_m(W^(b)) on a quadratic in the simulated exposure across B ≈ 200 placebo draws. Two error sources arise: (1) Monte Carlo error of order O_p(1/√B) that is correlated across units i (since the same B draws are reused) and thus does not average out at the √N_n rate, and (2) smoothing bias from the quadratic approximation if the true conditional expectation is nonlinear. Assumption C.15(S6) requires the residualization error to be o_p(N_n^{-1/2}), but this is imposed rather than verified. If violated, the J-test may reject even correctly specified maps, which directly threatens the credibility of the Egger et al. (202
- The scope of the specification test requires more careful framing relative to Gao et al. (2026). Remark 3.7 acknowledges that no test can have power against unrestricted richer exposure-mapping alternatives, and states that the J-test has power against alternatives that keep the moment criterion bounded away from zero. However, the paper does not characterize the alternatives against which the test has nontrivial power in any formal sense. Since the J-test is a central selling point of the framework, a more explicit statement of the local alternatives the test can detect—or at least a clearer statement that non-rejection is consistent with a broad class of misspecified maps—would strengthen the contribution and manage expectations for applied users.
- The efficiency bound in Theorem 4.1 is explicitly a within-class bound for the unit-level product residualized moment space H (Remark C.12). Exposure sufficiency also implies cross-unit restrictions (Remark 2.10) that lie outside this class. The paper acknowledges this but does not pursue it. This is defensible for a first paper, but the abstract and introduction could be read as claiming a more general efficiency result. The framing should be tightened so that the efficiency claim is clearly scoped to the unit-level moment class.
minor comments (5)
- Table 2: The J-statistic degrees of freedom are stated as 6 in the notes, but the text mentions M=7 moments (3 km bands up to 20 km). If dim(θ)=1, the df should be M−1=6, which is consistent, but this should be stated explicitly in the table notes for clarity.
- Section 6.2, Stage 1: The notation for the linear spillover exposure index g_v^(m)(W; θ_m) introduces θ_m as a vector of annulus weights, but the relationship between this vector-valued parameter and the scalar θ in the general framework (Section 2) is not immediately clear. A brief clarifying sentence would help.
- Figure 2: The y-axis scales vary substantially across the three panels, making visual comparison of effect magnitudes difficult. Consider using a common scale or at least labeling the units more prominently.
- Remark 2.8 states that the conditional expectation is approximated by a 'low-order polynomial or kernel smoother.' The paper should specify the exact smoother used in each application (the text later mentions a quadratic for Egger et al.) and report sensitivity to this choice.
- The paper cites 'Chernozhukov et al. (2025)' for the median procedure used in the split-sample exercise. The reference list gives the Fisher-Schultz lecture; ensure the citation format is complete.
Circularity Check
No significant circularity: the orthogonality moments follow from exposure sufficiency plus the known design, and the support-selection procedure is a test-based choice, not a fitted-parameter-renamed-as-prediction.
full rationale
The paper's central derivation chain is self-contained. Theorem 2.6 derives exposure sufficiency (Y_i ⊥ W | g(W;X_i,θ₀)) from Hypothesis 2.5 (Y_i(w) = ẽY_i(g(w;X_i,θ₀))) plus Assumption 2.4 (randomized assignment). Corollary 2.7 then derives the orthogonality condition E[ϕ(Y_i) R_{i,θ₀}^ψ(W)] = 0 from this conditional independence. The proof in Appendix A verifies this by direct application of the law of iterated expectations: since Y_i is σ(G_i(θ₀))-measurable and E[R_{i,θ₀}^ψ(W) | G_i(θ₀)] = 0 by construction of the residual, the product has zero expectation. No step reduces to its own input by definition. The GMM estimator (Section 3) estimates θ₀ from these moments — it is not fitted to the outcomes and then used to 'predict' the same outcomes. The J-test (Corollary 3.6) checks overidentification, which is a genuine test with power against alternatives (Remark 3.7). The efficiency bound (Theorem 4.1) is a Riesz representation argument over the moment space, not a tautology. In the Egger application (Section 6.2), the support-selection procedure selects the smallest non-rejected support via a likelihood-ratio test — this is a data-driven testing procedure, not a parameter fit being renamed as a prediction. The only mild concern is that the support-selection step uses the same data for testing and estimation, but the paper addresses this with a sublocation-level split-sample exercise (Figure 2 diamonds), and this is a multiple-testing/post-selection concern, not circularity. The theoretical results cite Chandrasekhar et al. (2023) for the affinity-set CLT and van der Vaart and Wellner (1996) for the Z-estimator theorem — these are external results, not self-citations. The Borusyak and Hull (2023) recentering approach is cited as methodological inspiration but the present paper applies the orthogonalization logic to a different target (the exposure map itself rather than a downstream parameter), so this is not a self-citation chain. No step in the derivation chain reduces to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- θ₀ (exposure map parameter) =
varies by application (e.g., 23.7 km for total income in MNS; 4 km support for GiveDirectly)
- Moment dictionary {(ϕ_m, ψ_m)} =
not specified numerically; applications use annulus averages and identity outcome transform
- Number of placebo draws B =
200
- Annular instrument band width =
2 km (main), 1 km and 1.5 km (robustness, Table 7)
axioms (6)
- domain assumption Assumption 2.4: Assignment W is drawn from a known law D_n independent of potential outcomes.
- ad hoc to paper Hypothesis 2.5: There exists θ₀ such that Y_i(w) = ẽY_i(g(w; X_i, θ₀)) for all w.
- domain assumption Affinity-set dependence structure (Chandrasekhar et al. 2023).
- ad hoc to paper Shell regularity for ring exposures (Assumption B.5, HR3).
- domain assumption Mean differentiability of population moment map μ(θ) at θ₀ (Assumption B.11, AN-AFF2).
- domain assumption Graph-HAC consistency (Assumption 3.9).
invented entities (2)
-
Design-side residual R_{i,θ}^ψ(W) = ψ(W) − E[ψ(W)|g(W;X_i,θ)]
independent evidence
-
Exposure sufficiency (Y_i ⊥ W | g(W;X_i,θ₀))
independent evidence
read the original abstract
Economic policies rarely affect only their direct targets. To study these spillovers, researchers summarize who else was treated with a simple exposure measure, such as the share of treated neighbors within a radius. But for many settings, economic theory provides little guidance on choosing the functional form (e.g., ring) of that measure or its parameters (e.g., radius). We show that the data can inform both choices. Correctly specified exposure measures imply orthogonality conditions that can be used for both estimation and testing. We establish consistency and asymptotic normality of the resulting estimator under spatial and network dependence in a design-based framework, with all randomness arising from treatment assignment. We then characterize the efficient moment conditions. Applied to two large-scale anti-poverty programs, the framework supports some prior radius estimates but rejects others. In the latter case, the revised radius yields substantively different policy-effect estimates.
Figures
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.