{"id":"ff9e6b31-8053-432f-a7ad-86b0c56972e3","arxiv_id":"2506.11313","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of methods for network experiments, centered on exposure maps, identification and power, data collection, and design trade-offs across lab, field, and natural settings.","lead":"This handbook chapter explains how to design and analyze experiments when people's choices influence each other through social networks, covering lab, field, lab-in-the-field, and natural experiments. It is a reference for researchers who need to handle spillovers, network data quality, and statistical power in networked settings.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The chapter's core 'many copies' principle rests on an unproven distance-decay condition: Eq. (1.4) shows decay only for a linear-quadratic model, and no rate or diagnostic is given for when spillovers are local enough.","rationale":"The reader's weakest_assumption matches the concern I identify: the framework's reliability hinges on spillovers being sufficiently local that neighborhoods are effectively independent. I agree that this is the load-bearing condition. The chapter's exposition in §1.2 derives distance-decay only through a linear-quadratic example, and §1.3 asserts the 'many copies' requirement without stating a rate or a test. Non-linear and long-range interference can defeat the recommended designs, and the chapter provides no diagnostic beyond noting the extreme failure of perfect diffusion. This is a real gap in the central organizing claim. However, the paper is explicitly a review chapter with no new theses, so the appropriate verdict remains UNVERDICTED; the concern affects how much weight practitioners should place on the heuristic, not whether the chapter should be accepted as a new result. I therefore keep the reader's verdict unchanged.","tokens_in":30734,"tokens_out":8100,"duration_ms":92613,"concrete_test":"Re-derive the influence of distance-r nodes in the chapter's own linear-quadratic model: evaluate T_r(n) = max_{i,j: dist(i,j)≥r} ∑_{t≥r} β^t (G^t)_{ij} on a sequence of d-regular expander graphs as n→∞, for β below the spectral threshold. If T_r does not tend to zero as r grows (or decays too slowly relative to n), then the claim in §1.2 that outcomes 'become less correlated' with distance requires an additional non-expansion condition that is absent. Complement with a complex-contagion simulation on a small-world network: if a cluster-randomized design with min-cut separated clusters remains biased as cluster separation grows, the 'many copies' advice is not generally valid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in §1.2–1.3 is that a network experiment is identified and powered only when one can find many approximately independent copies of treatment exposures. The supporting derivation is the linear-quadratic game, Eqs. (1.1)–(1.4), which yields an influence kernel ∑_{t≥1} β^t (G^t)_{ij} and the statement that this 'vanishes as i and j become further apart.' But this is not a general theorem: the decay depends on β and on the growth of walks in G. In d-regular expander graphs or dense graphs, the weighted number of walks to distance r can be non-negligible even for large r when β is close to the spectral threshold. More importantly, for non-linear outcome models—complex contagion, threshold models, or general diffusion, which the chapter itself invokes in §1.6—influence need not decay with distance at all. The chapter acknowledges the extreme case of perfect diffusion (§1.3) but gives no condition or diagnostic to distinguish the workable 'limited spillover' regime from the unidentifiable one. Consequently, the practical guidance to cleave a single large network into separated neighborhoods, or to rely on many independent networks, is only as solid as an unstated structural assumption. This is the load-bearing premise; all design advice in §2 inherits it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This is a review chapter on the design and analysis of experiments when outcomes are subject to network interference. The authors develop a vocabulary centered on exposure maps, argue that an experiment is identified and powered only if the researcher can find 'many copies' of treatment exposures with enough statistical independence, and apply this principle to field, lab, and natural experiments. The chapter covers data collection and link elicitation, the choice between many independent networks and a single large network, cluster and adaptive designs, partial network data, measurement error, panel data, lab experiments with designed networks, and natural experiments exploiting random assignment to network positions. It emphasizes trade-offs between parametric structure and nonparametric generality, and it points to exact randomization tests for non-sharp null hypotheses as a way to learn about the extent of spillovers.","tokens_in":30926,"tokens_out":8334,"duration_ms":80693,"significance":"As a survey, the chapter succeeds in organizing a large literature around a practically relevant principle: network experiments are only as strong as the availability of approximately independent copies of the treatment exposure of interest. Its strengths include a clear exposition of exposure maps, a balanced treatment of many-versus-one networks, a careful account of partial network data and its use in optimal design (Reeves et al. 2024), and the inclusion of exact finite-sample tests for non-sharp null hypotheses. The authors are appropriately explicit that the principle requires limits on spillovers and that parametric assumptions may be necessary. The stress-test concern about distance decay is partly mitigated by these explicit caveats and by the §1.4 tests, which provide a practical diagnostic for the extent of spillovers; however, the chapter could more clearly state that the formal decay argument in §1.2 is model-specific. Overall, the framework is sound and useful for guiding design choices, even though it deliberately offers no new theorem.","major_comments":[],"minor_comments":[{"comment":"The sentence 'The practical implementation of Viviano et al. (2023)'s Causal Clustering is.' is broken; it should be completed, for example, 'The practical implementation of Viviano et al. (2023)'s Causal Clustering is described by the authors.'","section":"2.5.1"},{"comment":"The claim that ARD would have saved 80% of the budget while yielding the same conclusions is a strong quantitative statement attributed to unpublished J-PAL South Asia calculations; the chapter should cite the precise appendix of Breza et al. (2020) where these calculations appear, or attribute them more cautiously.","section":"2.3"},{"comment":"The expression 'Yi(Dj :j∈C i,G|Ci)' appears to be missing notation; it should be written as Yi((Dj)_{j∈Ci}, G|_{Ci}) to indicate dependence on treatments within the cluster and the induced subgraph on Ci.","section":"1.3"},{"comment":"The text says Figure 1 presents the case of 'two arms each varying from intensities 1-3, omitting the third arm,' but the example has three factors; please clarify that Figure 1 is a two-dimensional slice of the three-factor design.","section":"1.6"},{"comment":"There are duplicate and inconsistent reference entries: Banerjee et al. (2013a) and (2013b) are the same Science article under different author formats; Banerjee et al. (2023) and (2024c) appear to be the same Econometrica paper; and Chandrasekhar and Jackson (2024a) and (2024b) are identical entries. These should be consolidated.","section":"References"},{"comment":"In the opening paragraph, 'or in in some combination' should be 'or in some combination'; likewise, §3.2.1 has 'uncertain about others' .' where a word such as 'connections' has been omitted.","section":"2"},{"comment":"The decay statement following Eq. (1.4) is true for the linear-quadratic game but is not a general property of network interference. For clarity, the chapter should explicitly say that for non-linear processes (e.g., threshold or complex contagion) the 'many copies' approach requires the researcher to verify limited spillovers, and should refer readers to the §1.4 tests for that purpose.","section":"1.2"}],"recommendation":"minor_revision","confidential_remarks":"The chapter is a review rather than an original research contribution, and its central framework is sound. The only substantive issue is that the formal decay argument is illustrated with a linear-quadratic model, and the authors should make the scope of that assumption clearer. The references contain several duplicates that should be fixed in the final version. I recommend minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a handbook chapter, not a research paper, and it should be read that way. It offers no new estimator or theorem, but it does something valuable: it gives the literature on network experiments a coherent organizing vocabulary (exposure maps) and a practical walk-through of design choices across field, lab, and natural experiments. The authors know this material from the inside, and it shows. The sections on many independent networks versus a single large network, on link elicitation, and on partial network data (ARD, RDS, star sampling) are concrete and useful. The factorialization discussion makes a good, non-obvious point: pre-emptively collapsing treatment arms is equivalent to deciding you already know the 23 cells you aren't testing.\n\nThe soft spots are few but real. There's a mechanical typo in Section 2.5.1 (\"Causal Clustering is . Given...\"). The 80% budget saving from ARD is repeated without enough detail to check. More substantively, the stress-test note is right that the chapter's load-bearing idea—that you can identify and power a network experiment when you have many approximately independent copies of a treatment neighborhood—is illustrated with the linear-quadratic game, but no general condition or diagnostic is given for when spillovers are local enough. The chapter acknowledges the perfect-diffusion extreme, and it points to the CLT and estimation papers for the technical conditions, but a practitioner is left to guess whether their setting is in the identifiable regime. That is a real gap, though it is a gap in the field, not just this chapter.\n\nWho is this for? A grad student planning a field experiment, or anyone wanting a single map of the literature. It is not for theorists seeking new results. I'd send it to peer review as a handbook chapter; the authors should fix the typos and add a short paragraph flagging the conditions under which the 'many independent neighborhoods' strategy is justified. Verdict: solid review, worth engaging.","headline":"A solid, useful review chapter on designing network experiments; the central 'many copies' principle is asserted rather than sharply delimited, but this is a review, not a theory paper.","tokens_in":31524,"tokens_out":2811,"would_cite":true,"duration_ms":31534,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An experiment on a networked population is identifiable only to the extent that many independent copies of each treatment exposure exist.","keywords":["network experiments","interference","spillovers","exposure maps","field experiments","lab experiments","natural experiments","experimental design"],"falsifier":"Run or simulate a large connected network experiment in which the treatment deliberately shifts outcomes at every distance, so far-apart neighborhoods are not independent, then apply the paper's neighborhood-averaging estimator and increase $n$; if the estimated treatment effect keeps a bias that does not shrink, the limited-spillover premise is false.","tokens_in":30442,"feed_emoji":"🧪","tokens_out":7587,"duration_ms":80238,"temperature":0.7,"pith_summary":"This paper is a survey chapter, but it argues a single organizing thesis: an experiment on a networked population is identifiable and powered only to the extent that the researcher can obtain many statistically independent copies of how each node is treated, directly and through its neighbors. This can come from many separate networks, or from one large network in which spillovers decay quickly enough with distance. The authors use exposure maps to turn this thesis into practical guidance on design: whether to randomize individuals or clusters, how much network data to collect and at what precision, whether panels help, and which settings favor lab, field, or natural experiments. If the thesis is right, it gives applied researchers a usable test for when causal inference on networks is feasible and when it is not.","feed_headline":"A network experiment works only if spillovers leave independent pockets","feed_subtitle":"Identification and power reduce to having enough separate copies of how each node is treated, directly and through its neighbors.","key_machinery":"The central object is the exposure map, $d_i = f_i(D_{1:n}, G)$: a function that reduces the full treatment vector and network to the exposure status that determines node $i$'s potential outcome $Y_i(d_i)$. The related structural causal map plays the same role for causal parameters. In the paper's running example, the exposure structure is generated by best responses in a linear-quadratic network game, $Y_i = \\varepsilon_i + \\beta \\sum_j G_{ij}Y_j$, whose iterated solution makes the influence of far-away nodes decay as powers of $\\beta$ times walks in $G$. This object carries the argument because it tells the researcher whether exposures are effectively independent: if $d_i$ and $d_j$ are systematically highly correlated for most $i,j$, treatment is an aggregate shock and no amount of data rescues identification.","core_discovery":"The paper's central claim is that when the stable unit treatment value assumption fails because outcomes depend on neighbors' treatments and on network structure, causal inference still goes through if the dependence can be compressed by an exposure map into a limited set of exposure types and each type has many nearly independent replications. In the linear-quadratic example, the influence of node $j$ on node $i$ is $\\sum_{t\\ge 1}\\beta^t G^t_{ij}$, which vanishes with network distance when $\\beta$ is small enough; this is the concrete mechanism by which a single large network can supply the needed independence. The authors conclude that researchers should either collect many independent networks, or restrict attention to settings with limited and well-measured spillovers, or impose a parsimonious parametric model that lets the joint distribution of outcomes carry the identification.","pith_inferences":["The paper's 'many copies' principle implies an effective-sample-size diagnostic: a researcher could report the number of nearly independent exposure neighborhoods, much as cluster designs report effective cluster counts, as a standard accompaniment to power calculations.","The same logic suggests adaptive designs should oversample rare exposure types in pilot waves rather than only central nodes, since the binding constraint is replications of each exposure configuration, not average connectivity.","If spillovers are suspected to be long-range, the framework points to a robustness exercise: re-estimate effects under several assumed interference radii and check how much the conclusions move; the paper's own examples stop at finite neighborhoods, but the logic invites this stress test."],"forward_implications":["If spillovers decay with distance and the network is large enough, consistent estimation and inference are possible from a single network; otherwise many independent networks are required.","Cluster designs beat Bernoulli randomization when spillovers are nontrivial, with worst-case bias tied to the share of cross-cluster links and worst-case variance tied to cluster-size imbalance.","With partial or noisy network data, model-based design that estimates a generative graph model and optimizes treatment allocation over coarse groups can outperform full network data paired with naive randomization.","Exact finite-sample randomization tests can be built for nonsharp null hypotheses such as no spillovers or no second-order spillovers by constructing an artificial experiment with focal nodes and re-randomization.","Factorial designs followed by data-driven pooling let researchers study high-dimensional treatment combinations without needing a prohibitive number of independent networks."],"supporting_citations":[{"why":"Defines the stable unit treatment value assumption whose failure provides the motivation for exposure maps.","marker":"Rubin (1974, 1980)"},{"why":"Identifies the reflection problem that random variation in treatment and network structure is meant to overcome.","marker":"Manski (1993)"},{"why":"Provides the general-interference estimation framework that the chapter builds on for identification.","marker":"Aronow and Samii (2017)"},{"why":"Supplies exact p-values and nonsharp null hypothesis testing used in network experimental design.","marker":"Athey et al. (2018)"},{"why":"Gives central limit theorem conditions for dependent triangular arrays that underpin the power discussion.","marker":"Chandrasekhar et al. (2023b)"},{"why":"Introduces graph cluster randomization and the notion of cluster-level exposure.","marker":"Ugander et al. (2013)"},{"why":"Provides the causal clustering worst-case bias-variance trade-off that guides cluster versus Bernoulli design choices.","marker":"Viviano et al. (2023)"},{"why":"Shows how partial network data plus model-based optimal allocation can outperform full data with uniform randomization.","marker":"Reeves et al. (2024)"},{"why":"Supplies the many-independent-villages experiment used as the main empirical template for identification with spillovers.","marker":"Banerjee et al. (2013a)"}],"fun_headline_variants":["Network experiments need independent pockets for valid inference","Causal inference in networks requires limited spillover exposure","For network experiments, independence comes from exposure maps","Network trials: identify via replications or parametric models","Designing experiments with networks: three paths to identification"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The prescription relies on the assumption that spillovers fade with network distance quickly enough for far-apart neighborhoods to be nearly independent; if interference is long-range, dense, or hidden in unmeasured relationships, the recommended designs and power calculations lose their warrant.","fun_headline_variants_meta":{"raw":{"variants":["Network experiments need independent pockets for valid inference","Causal inference in networks requires limited spillover exposure","For network experiments, independence comes from exposure maps","Network trials: identify via replications or parametric models","Designing experiments with networks: three paths to identification"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000146,"raw_usage":{"total_tokens":1062,"prompt_tokens":706,"completion_tokens":356,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":322,"completion_tokens_details":{"reasoning_tokens":283}},"tokens_in":322,"tokens_out":356,"duration_ms":4620,"temperature":1.0,"reasoning_tokens":283,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:11:09.531393+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run or simulate a large connected network experiment in which the treatment deliberately shifts outcomes at every distance, so far-apart neighborhoods are not independent, then apply the paper's neighborhood-averaging estimator and increase $n$; if the estimated treatment effect keeps a bias that does not shrink, the limited-spillover premise is false.","supporting_citations":[],"review_version":1}