Pith. sign in

REVIEW 3 major objections 5 minor

The experiment's randomization can estimate and test the spillover map

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-10 03:53 UTC pith:CZOZXLX7

load-bearing objection Real methodological contribution; the headline empirical result needs more scrutiny the 3 major comments →

arxiv 2607.08640 v2 pith:CZOZXLX7 submitted 2026-07-09 econ.EM math.STstat.MEstat.TH

A Design-Based Approach to Testing and Inference in (Quasi-)Experiments with Spillovers

classification econ.EM math.STstat.MEstat.TH
keywords spilloversexposure mappingsdesign-based inferencespecification testingspatial dependencenetwork interferenceGMM estimationrandomization inference
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When researchers study spillovers, they summarize who else was treated with an exposure measure — say, the share of treated neighbors within a radius. But they rarely know what functional form that measure should take or what its parameter should be. This paper shows that the same randomization that identifies treatment effects can discipline both choices. The key construction is the design-side residual: take any function of the full assignment vector, subtract what is predictable from the exposure measure under the known randomization, and what remains should be unrelated to outcomes if the exposure map is correctly specified. This orthogonality condition holds for every outcome transformation and every assignment function, generating a family of moment conditions. These moments let researchers estimate the exposure parameter (e.g., the radius) via GMM and, when the system is overidentified, test whether the proposed exposure specification is consistent with the experimental design. The entire framework is design-based: potential outcomes are fixed, and all randomness comes from the known assignment mechanism, so no outcome model is needed. The author establishes consistency and asymptotic normality under spatial and network dependence, characterizes the efficient moments, and propagates exposure-map uncertainty into downstream policy estimates. Applied to two large anti-poverty programs, the framework validates one paper's radius choice but rejects another's, with the revised radius producing a substantially smaller fiscal multiplier.

Core claim

A correctly specified exposure map implies exposure sufficiency: conditional on the exposure value, the remaining variation in the treatment assignment carries no outcome-relevant information. This property generates orthogonality conditions — between outcomes and design-side residuals of assignment functions — that are testable and estimable using only the known randomization design, without any outcome model. These conditions enable both GMM estimation of the exposure parameter and overidentification testing of the exposure specification itself.

What carries the argument

The design-side residual R_{i,θ}^ψ(W) = ψ(W) − E[ψ(W) | g(W; X_i, θ)] strips from any assignment function the component explained by the candidate exposure. Its orthogonality to any outcome function — E[ϕ(Y_i) R_{i,θ₀}^ψ(W)] = 0 — is the testable consequence of exposure sufficiency. The conditional expectation is computed by repeatedly simulating from the known assignment design, so the residual is a pure design object requiring no outcome modeling.

Load-bearing premise

The framework assumes that a single parametric exposure map captures all the ways the treatment assignment affects each unit's outcome — that once you know the exposure value, no other feature of who was treated matters. If spillovers flow through multiple channels that the map does not jointly summarize, the assumption fails, and the test can detect some violations but cannot certify that the map is fully correct.

What would settle it

If one took a setting with known multi-channel spillovers (e.g., both price effects and direct network effects operating at different spatial scales) and fit a single ring exposure map, the J-test should reject — and the rejection should weaken as the map is enriched to capture both channels.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Researchers studying spillovers can let the data choose the radius, decay rate, or network hop count rather than reporting results under several ad hoc values, with uncertainty about that choice properly propagated into downstream estimates.
  • The overidentification J-test provides a formal specification check: if a ring specification is rejected, the researcher knows spillovers extend beyond the assumed radius and can revise accordingly.
  • The framework extends to quasi-experimental settings where the design is a placebo or permutation scheme, so the logic applies beyond literal randomized trials.
  • An efficiency bound within the class of unit-level moments identifies which outcome transformations and assignment functions are most informative, guiding the construction of the test.
  • Downstream policy estimates — treatment effects, multipliers, counterfactual predictions — can incorporate uncertainty about the exposure specification itself, not just sampling uncertainty conditional on a fixed choice.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If spillovers operate through multiple channels that no single parametric map jointly summarizes (e.g., both price-mediated and network-mediated effects), the exposure sufficiency hypothesis fails. The J-test may detect this but cannot diagnose which channel is missing, so rejection leaves the researcher without a constructive direction beyond widening or reshaping the map.
  • The framework could be extended to formally compare competing exposure map families — ring versus gravity versus network — on the same data using their respective J-statistics, providing a principled basis for the informal specification comparisons that applied papers currently make by reporting multiple specifications side by side.
  • Exposure maps for which the assignment design leaves little residual variation after conditioning (near-injective maps) are fundamentally untestable, which may explain why some specifications in applied work are empirically indistinguishable regardless of sample size.
  • The contrast between the two applications — one validated, one rejected — raises the question of whether the framework's practical value depends on ex ante proximity of the original specification to the truth, and whether pre-analysis plans for spillover maps could benefit from design-based pre-registration.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper develops a design-based framework for testing and estimating exposure mappings in settings with spillovers. The core idea is that a correctly specified exposure map g(W; X_i, θ₀) implies an exposure sufficiency property: conditional on the exposure, the remaining variation in the assignment vector is orthogonal to outcomes. This generates testable moment conditions E[ϕ(Y_i) R_{i,θ₀}^ψ(W)] = 0, where R is a design-side residual constructed by orthogonalizing a design function ψ(W) with respect to the exposure using the known randomization distribution. These moments enable GMM estimation of the tuning parameter θ₀ and, when overidentified, a J-test of the exposure specification. The author establishes consistency and asymptotic normality under spatial and network dependence via the affinity-set framework of Chandrasekhar et al. (2023), characterizes an efficiency bound within the unit-level moment class using Riesz representation, and propagates Stage-1 uncertainty into downstream policy estimands. Two applications to large-scale anti-poverty programs illustrate the framework: the design-based test supports the 20 km radius in Muralidharan et al. (2023) but rejects the 2 km radius in Egger et al. (2022), yielding a revised local fiscal multiplier of 1.57 (vs. the original 2.5).

Significance. The paper addresses a genuine and important gap in the spillovers literature. While exposure mappings are widely used in applied work, researchers typically lack principled tools for choosing the functional form or calibrating tuning parameters, and standard errors rarely propagate this specification uncertainty. The design-based orthogonality conditions are a clean and novel contribution, transforming the exposure map from a fixed researcher choice into an object of inference. The theoretical development is careful: identification (Theorem 2.6, Corollary 2.7) is clean, the asymptotic theory (Theorems 3.3, 3.5) correctly adapts Z-estimator arguments to the affinity-set CLT, and the efficiency bound (Theorem 4.1, Appendix C) is a coherent Riesz representation argument. The applications are well-chosen and the contrasting conclusions are informative. The framework is likely to be influential in applied microeconomics.

major comments (3)
  1. The most load-bearing concern is at the implementation layer. Corollary 3.6 establishes the J-test under the assumption that the population moments are exactly zero at θ₀. In practice (Remark 2.8), the conditional expectation E[ψ_m(W)|g(W;X_i,θ)] is approximated by regressing ψ_m(W^(b)) on a quadratic in the simulated exposure across B ≈ 200 placebo draws. Two error sources arise: (1) Monte Carlo error of order O_p(1/√B) that is correlated across units i (since the same B draws are reused) and thus does not average out at the √N_n rate, and (2) smoothing bias from the quadratic approximation if the true conditional expectation is nonlinear. Assumption C.15(S6) requires the residualization error to be o_p(N_n^{-1/2}), but this is imposed rather than verified. If violated, the J-test may reject even correctly specified maps, which directly threatens the credibility of the Egger et al. (202
  2. The scope of the specification test requires more careful framing relative to Gao et al. (2026). Remark 3.7 acknowledges that no test can have power against unrestricted richer exposure-mapping alternatives, and states that the J-test has power against alternatives that keep the moment criterion bounded away from zero. However, the paper does not characterize the alternatives against which the test has nontrivial power in any formal sense. Since the J-test is a central selling point of the framework, a more explicit statement of the local alternatives the test can detect—or at least a clearer statement that non-rejection is consistent with a broad class of misspecified maps—would strengthen the contribution and manage expectations for applied users.
  3. The efficiency bound in Theorem 4.1 is explicitly a within-class bound for the unit-level product residualized moment space H (Remark C.12). Exposure sufficiency also implies cross-unit restrictions (Remark 2.10) that lie outside this class. The paper acknowledges this but does not pursue it. This is defensible for a first paper, but the abstract and introduction could be read as claiming a more general efficiency result. The framing should be tightened so that the efficiency claim is clearly scoped to the unit-level moment class.
minor comments (5)
  1. Table 2: The J-statistic degrees of freedom are stated as 6 in the notes, but the text mentions M=7 moments (3 km bands up to 20 km). If dim(θ)=1, the df should be M−1=6, which is consistent, but this should be stated explicitly in the table notes for clarity.
  2. Section 6.2, Stage 1: The notation for the linear spillover exposure index g_v^(m)(W; θ_m) introduces θ_m as a vector of annulus weights, but the relationship between this vector-valued parameter and the scalar θ in the general framework (Section 2) is not immediately clear. A brief clarifying sentence would help.
  3. Figure 2: The y-axis scales vary substantially across the three panels, making visual comparison of effect magnitudes difficult. Consider using a common scale or at least labeling the units more prominently.
  4. Remark 2.8 states that the conditional expectation is approximated by a 'low-order polynomial or kernel smoother.' The paper should specify the exact smoother used in each application (the text later mentions a quadratic for Egger et al.) and report sensitivity to this choice.
  5. The paper cites 'Chernozhukov et al. (2025)' for the median procedure used in the split-sample exercise. The reference list gives the Fisher-Schultz lecture; ensure the citation format is complete.

Circularity Check

0 steps flagged

No significant circularity: the orthogonality moments follow from exposure sufficiency plus the known design, and the support-selection procedure is a test-based choice, not a fitted-parameter-renamed-as-prediction.

full rationale

The paper's central derivation chain is self-contained. Theorem 2.6 derives exposure sufficiency (Y_i ⊥ W | g(W;X_i,θ₀)) from Hypothesis 2.5 (Y_i(w) = ẽY_i(g(w;X_i,θ₀))) plus Assumption 2.4 (randomized assignment). Corollary 2.7 then derives the orthogonality condition E[ϕ(Y_i) R_{i,θ₀}^ψ(W)] = 0 from this conditional independence. The proof in Appendix A verifies this by direct application of the law of iterated expectations: since Y_i is σ(G_i(θ₀))-measurable and E[R_{i,θ₀}^ψ(W) | G_i(θ₀)] = 0 by construction of the residual, the product has zero expectation. No step reduces to its own input by definition. The GMM estimator (Section 3) estimates θ₀ from these moments — it is not fitted to the outcomes and then used to 'predict' the same outcomes. The J-test (Corollary 3.6) checks overidentification, which is a genuine test with power against alternatives (Remark 3.7). The efficiency bound (Theorem 4.1) is a Riesz representation argument over the moment space, not a tautology. In the Egger application (Section 6.2), the support-selection procedure selects the smallest non-rejected support via a likelihood-ratio test — this is a data-driven testing procedure, not a parameter fit being renamed as a prediction. The only mild concern is that the support-selection step uses the same data for testing and estimation, but the paper addresses this with a sublocation-level split-sample exercise (Figure 2 diamonds), and this is a multiple-testing/post-selection concern, not circularity. The theoretical results cite Chandrasekhar et al. (2023) for the affinity-set CLT and van der Vaart and Wellner (1996) for the Z-estimator theorem — these are external results, not self-citations. The Borusyak and Hull (2023) recentering approach is cited as methodological inspiration but the present paper applies the orthogonalization logic to a different target (the exposure map itself rather than a downstream parameter), so this is not a self-citation chain. No step in the derivation chain reduces to its inputs by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 2 invented entities

The framework's main axiom is Hypothesis 2.5 (exposure sufficiency), which is the maintained hypothesis being tested rather than assumed without check. The design-side residual and exposure sufficiency property are derived constructs, not invented entities in the physical sense. The free parameters (θ₀, moment dictionary, simulation tuning) are standard estimation choices. The main concern is that the moment dictionary and band-width choices are researcher degrees of freedom that could affect finite-sample results, though the paper shows asymptotic efficiency is attained as the dictionary grows.

free parameters (4)
  • θ₀ (exposure map parameter) = varies by application (e.g., 23.7 km for total income in MNS; 4 km support for GiveDirectly)
    The central object of estimation; estimated from design-based GMM moments rather than imposed, so it is a fitted parameter in the sense that the data determine it.
  • Moment dictionary {(ϕ_m, ψ_m)} = not specified numerically; applications use annulus averages and identity outcome transform
    The choice of which outcome transformations and design functions to include in the GMM criterion is a researcher degree of freedom. The paper shows efficiency is attained as the dictionary grows, but finite-sample choices matter and are not fully specified.
  • Number of placebo draws B = 200
    Used to approximate conditional expectations E[ψ(W)|g(W;X_i,θ)] by simulation. Stated in Section 6.2 but the sensitivity to this choice is not reported.
  • Annular instrument band width = 2 km (main), 1 km and 1.5 km (robustness, Table 7)
    The width of distance bands used to construct design functions ψ_m affects the moments. Table 7 shows estimates are stable across 1, 1.5, 2 km bands.
axioms (6)
  • domain assumption Assumption 2.4: Assignment W is drawn from a known law D_n independent of potential outcomes.
    Standard in design-based inference; requires the randomization mechanism to be fully known and correctly specified. Section 2.1.
  • ad hoc to paper Hypothesis 2.5: There exists θ₀ such that Y_i(w) = ẽY_i(g(w; X_i, θ₀)) for all w.
    The exposure sufficiency restriction — that the exposure map captures all assignment-relevant information. This is the maintained hypothesis being tested. Section 2.3.
  • domain assumption Affinity-set dependence structure (Chandrasekhar et al. 2023).
    Cross-unit dependence is summarized by affinity sets A_i with bounded size and negligible outside-affinity covariance. Section 3.2.
  • ad hoc to paper Shell regularity for ring exposures (Assumption B.5, HR3).
    No-mass-point condition on weighted pairwise distances at candidate cutoff radii. Imposed as primitive; verified only informally. Appendix B.2.2.
  • domain assumption Mean differentiability of population moment map μ(θ) at θ₀ (Assumption B.11, AN-AFF2).
    Required for asymptotic normality; stated as a population-level condition rather than verified from primitive exposure-map properties. Appendix B.3.
  • domain assumption Graph-HAC consistency (Assumption 3.9).
    Consistency of the spatial/network HAC variance estimator is imposed at a high level, with reference to Conley (1999) and Kojevnikov et al. (2021) for sufficient conditions.
invented entities (2)
  • Design-side residual R_{i,θ}^ψ(W) = ψ(W) − E[ψ(W)|g(W;X_i,θ)] independent evidence
    purpose: Orthogonalizes a design function with respect to the candidate exposure map, creating the moment conditions used for estimation and testing.
    The residual is a constructed quantity, not a new physical entity. Its properties follow from the known design D_n and the exposure map g. The orthogonality E[ϕ(Y_i)R_{i,θ₀}^ψ(W)] = 0 is a theorem (Corollary 2.7), not a postulate.
  • Exposure sufficiency (Y_i ⊥ W | g(W;X_i,θ₀)) independent evidence
    purpose: The conditional independence property that makes the exposure map testable.
    Derived as Theorem 2.6 from Hypothesis 2.5 and Assumption 2.4. It is a consequence, not an additional postulate. Its testability is limited by the impossibility result of Gao et al. (2026), which the paper acknowledges (Remark 3.7).

pith-pipeline@v1.1.0-glm · 56920 in / 3846 out tokens · 691465 ms · 2026-07-10T03:53:23.272807+00:00 · methodology

0 comments
read the original abstract

Economic policies rarely affect only their direct targets. To study these spillovers, researchers summarize who else was treated with a simple exposure measure, such as the share of treated neighbors within a radius. But for many settings, economic theory provides little guidance on choosing the functional form (e.g., ring) of that measure or its parameters (e.g., radius). We show that the data can inform both choices. Correctly specified exposure measures imply orthogonality conditions that can be used for both estimation and testing. We establish consistency and asymptotic normality of the resulting estimator under spatial and network dependence in a design-based framework, with all randomness arising from treatment assignment. We then characterize the efficient moment conditions. Applied to two large-scale anti-poverty programs, the framework supports some prior radius estimates but rejects others. In the latter case, the revised radius yields substantively different policy-effect estimates.

Figures

Figures reproduced from arXiv: 2607.08640 by Yechan Park.

Figure 1
Figure 1. Figure 1: Stage-1 objective functions for total income (top), NREGS earnings (bottom left), and [PITH_FULL_IMAGE:figures/full_fig_p026_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Radius-path diagnostics for the Egger application [PITH_FULL_IMAGE:figures/full_fig_p030_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Radius-path diagnostics for expenditure and asset outcomes in the Egger application [PITH_FULL_IMAGE:figures/full_fig_p083_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Radius-path diagnostics for wealth, income, and transfer outcomes in the Egger applica [PITH_FULL_IMAGE:figures/full_fig_p084_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Radius-path diagnostics for taxes paid in the Egger application [PITH_FULL_IMAGE:figures/full_fig_p085_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.