{"id":"3bde85b2-5a9f-4f55-9612-d902f942546a","arxiv_id":"2607.06506","paper_version":1,"verdict":"CONDITIONAL","confidence":"UNKNOWN","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":5,"one_line_summary":"Under cluster randomization with fixed exposure margins, the spillover kernel level is aliased with the intercept and direct effect at any sample size, and every estimator's bias decomposes exactly into a level term, a shape error, and a leakage term.","lead":"This paper shows that in spatially structured cluster randomized trials, the level of the spillover kernel is structurally non-identifiable under fixed-margin randomization, and derives exact bias decompositions for any anchored estimator. A smart generalist would read it to understand why spacing clusters apart can make spillover assumptions unfalsifiable and what design modifications can fix this.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The linear exposure mapping (Eq. 1) is the load-bearing assumption; the paper is honest about it, but the abstract's 'under any outcome model' overstates scope relative to what Proposition 1(iv) actually covers.","rationale":"The reader correctly identified the linear exposure mapping as the load-bearing assumption and correctly assessed the paper's honesty about its limitations. The theoretical results (Proposition 1, Theorem 3, Corollary 1, Proposition 4) are well-derived within the linear-predictor class, and the key algebra — that S[1](z) = M(z)·1 − q·z̃ lies in col{[1, z̃]) under H1–H2 — is verifiable from the main text without needing the supplementary proofs. The CONDITIONAL verdict is appropriate: the framework is a genuine contribution that unifies existing approaches and provides actionable design diagnostics, but verification of the supplementary proofs and shipped code would strengthen confidence. The concern about scope (linear exposure mapping) is real but acknowledged by the authors, and does not constitute an internal inconsistency or correctness error within the stated model class. The abstract's phrasing 'under any outcome model' is somewhat broader than what Proposition 1(iv) actually covers, but this is a presentation issue rather than a mathematical flaw. I do not think the verdict needs to change.","tokens_in":20265,"tokens_out":4241,"duration_ms":178512,"concrete_test":"Run a simulation where the true outcome model uses a nonlinear exposure mapping — e.g., Y_i(z) = α + τ z_{k(i)} + g(Σ_t ϕ(d_{i,t}) z_{ι(t)}) + ε_i with g(·) = log(1 + ·) or a Hill function — under H1–H3 with a fixed-margin cluster design. Fit the working model (Eq. 3) with the linear exposure mapping and compute the plug-in bias for level-loaded estimands (τ_cluster, ATE_trial, θ(d)). Compare the realized bias to the predicted level term −ℓ(w,c)·ē. If the bias is well-approximated by the level term alone (with shape/leakage terms secondary), the framework's practical guidance survives moderate nonlinearity. If the bias has a substantial component not captured by any of the three decomposition terms, the linear-predictor scope is more binding than the paper suggests and the design diagnostics (ℓ, I_0, Λ) may mislead practitioners.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies the additive linear exposure mapping (Eq. 1: Y_i(z) = α + τ z_{k(i)} + Σ_t ϕ(d_{i,t}) z_{ι(t)} + ε_i) as the load-bearing assumption. Proposition 1's level aliasing, Theorem 3's bias decomposition, the level leverage scalar ℓ(w,c), and Corollary 1's contamination bias formula all hold within this model class. The paper acknowledges in Section 7 that outside the linear model 'that assumption is no longer a single scalar and the closed-form prices... do not survive as stated.'\n\nThe concern is not that the linear model is unreasonable — it is a standard exposure mapping (Aronow and Samii, 2017; Sävje, 2024) — but that the abstract's phrasing 'at any sample size and under any outcome model' could be read as covering arbitrary outcome models, when Proposition 1(iv) actually covers the class where the law of y given z depends on (α, τ, ϕ) only through η(z) = α1 + τz̃ + S[ϕ](z). This is the linear-predictor class (linear model with arbitrary Σ, GLMs, mixed models with conditional law depending on η + Zb). It does not cover models where the exposure enters nonlinearly — e.g., saturating dose-response in the count of treated sources, or threshold effects — which are scientifically plausible for the paper's motivating examples (vector-borne disease transmission, gene drives).\n\nWithin the linear-predictor class, the algebra of Proposition 1 is straightforward and correct: under H1–H2, S[1](z) = M(z)·1 − q·z̃, which lies in col{[1, z̃]) at every allocation; under H3 the flat direction is common. Theorem 3's decomposition is an algebraic identity (exact conditional on z), with the leakage term vanishing asymptotically under Condition 1(b). The level term −ℓ(w,c)·ē is exact under H1–H3 regardless of the metric, link, or working weights — this is a genuine and clean result.\n\nThe practical question is whether the linear exposure mapping is a reasonable approximation for spatial spillover in the target applications. If the true model has, say, a satur","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This manuscript develops a framework for spatial experiments with spillover effects under cluster randomization. The central result (Proposition 1) is that under a fixed exposure set size (H1-H3), the level of the spillover kernel is aliased with the intercept and direct effect at every sample size and within the linear-predictor model class. The paper derives an exact bias decomposition (Theorem 3) for any anchored plug-in estimator into a level term, a shape error, and a leakage term, and introduces the level leverage scalar measuring each estimand's exposure to the unidentified level. Corollary 1 specializes the result to the conventional cluster-trial analysis. The paper provides a taxonomy of design augmentations that can identify the level and a decision criterion for when to augment versus anchor. A simulation study confirms the asymptotic predictions.","tokens_in":20837,"tokens_out":931,"duration_ms":290530,"significance":"The paper makes a substantive conceptual contribution by unifying several existing approaches (Bernoulli randomization, compact support assumptions, buffer designs) as special cases of an anchoring choice within a single framework. The level leverage scalar is a genuinely useful design-computable diagnostic, and the exact variance decomposition in Proposition 4(i) and the MSE floor result (Corollary 2) are clean, falsifiable predictions. The taxonomy in Table 4 is practically valuable for trialists. The mathematical arguments are sound: Proposition 1's level aliasing is a direct algebraic observation, Theorem 3 follows from GLS projection structure, and Corollary 1's contamination bias formula is a clean specialization. The paper is appropriately honest about the scope limitations of the linear exposure mapping in Section 7.","major_comments":[{"comment":"Abstract and Proposition 1(iv) scope mismatch. The abstract states the level is aliased 'at any sample size and under any outcome model.' Proposition 1(iv) actually covers the class where the law of y given z depends on (alpha, tau, phi) only through eta(z) = alpha*1 + tau*z_tilde + S[phi](z), which includes linear models with arbitrary Sigma, GLMs, and mixed models with conditional law depending on eta + Zb. This does not cover models where the exposure enters nonlinearly in the treated-source count (e.g., saturating dose-response or threshold effects), which are scientifically plausible for the motivating examples of vector-borne disease transmission and gene drives. The abstract phrasing should be corrected to accurately reflect the linear-predictor class scope of Proposition 1(iv). This is a load-bearing issue because the central claim's reach is being overstated in the paper's most-","section":null}],"minor_comments":[{"comment":"Section 1.1, paragraph 2: 'inclduing' should be 'including' in the abstract.","section":null},{"comment":"Section 2.1: The notation z_{k(i)} and z_{iota(t)} are both used for cluster treatment status; a brief clarifying remark that these refer to the same quantity under different indexing would help readers.","section":null},{"comment":"Table 2: The column header 'c_w' appears to be 'c' (the direct-effect indicator) and 'w' (the measure); consider labeling more explicitly as '(c, w)' or adding a footnote.","section":null},{"comment":"Section 4.3: The definition of the parametric working class K references 'b psi_p(.; rho)' but the notation 'b' as both a coefficient and a function prefix is slightly confusing; consider using a different symbol for the coefficient.","section":null},{"comment":"Section 5.2: 'preiod' should be 'period' (appears twice in the paragraph containing 'the intervention preiod of the study').","section":null},{"comment":"Figure 1: The y-axis labels and legend are small; consider enlarging for readability in print format.","section":null},{"comment":"Section 6: The simulation study would benefit from stating the number of Monte Carlo replications used.","section":null},{"comment":"Section 4.4, Condition 1: The role of the norm ||.||_R is introduced but its specific choice is not discussed; a brief remark on practical selection would be helpful.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about the abstract overstatement is valid and is the primary issue requiring revision. The linear exposure mapping limitation is acknowledged in Section 7, so the fix is straightforward: align the abstract with the actual scope of Proposition 1(iv). The paper is otherwise a well-executed methodological contribution that fits the journal's scope."},"author_rebuttal":{"model":"glm-5.2","summary":"The referee identifies a scope mismatch between the abstract's claim that level aliasing holds 'under any outcome model' and the actual scope of Proposition 1(iv), which covers the linear-predictor model class (linear models with arbitrary covariance, GLMs, and mixed models whose conditional law depends on parameters through the linear predictor). The referee correctly notes that models with nonlinear exposure in the treated-source count—such as saturating dose-response or threshold effects—are not covered and are scientifically plausible for the motivating examples. We agree this is a load-bearing issue: the abstract overstates the reach of the central claim and must be corrected.","responses":[{"response":"The referee is correct on all counts. The abstract phrase 'under any outcome model' overstates the scope of Proposition 1(iv), which covers the linear-predictor class: models where the conditional law of y given z depends on (alpha, tau, phi) only through eta(z) = alpha*1 + tau*z_tilde + S[phi](z). This includes linear models with arbitrary Sigma, GLMs with E(y_i|z) = g^{-1}{eta_i(z)}, and mixed models whose conditional law given random effects depends on parameters through eta(z) + Zb. It does not cover models where the exposure enters nonlinearly in the treated-source count—for example, saturating dose-response of the form f(S[phi](z)) for nonlinear f, or threshold effects. These are scientifically plausible: vector-borne disease transmission often exhibits saturating or threshold dynamics in exposure to infectious agents, and gene drive release dynamics are inherently nonlinear in the treated-source count. We will correct the abstract to read 'at any sample size and within the linear-predictor model class' (or equivalent precise language), replacing 'under any outcome model.' We will also add a sentence to the abstract noting that the linear-predictor class encompasses linear models, GLMs, and mixed models, and that nonlinear exposure mappings are discussed as a scope limitation in Section 7. The body text is already accurate: Proposition 1(iv) states its scope precisely, and Section 7 explicitly acknowledges that 'our exact results are computed within the linear exposure mapping' and that outside it 'the closed-form prices... do not survive as stated.' The mismatch is confined to the abstract's shorthand, which we will fix.","revision_made":"yes","referee_comment":"Abstract and Proposition 1(iv) scope mismatch. The abstract states the level is aliased 'at any sample size and under any outcome model.' Proposition 1(iv) actually covers the class where the law of y given z depends on (alpha, tau, phi) only through eta(z) = alpha*1 + tau*z_tilde + S[phi](z), which includes linear models with arbitrary Sigma, GLMs, and mixed models with conditional law depending on eta + Zb. This does not cover models where the exposure enters nonlinearly in the treated-source count (e.g., saturating dose-response or threshold effects), which are scientifically plausible for the motivating examples of vector-borne disease transmission and gene drives. The abstract phrasing should be corrected to accurately reflect the linear-predictor class scope of Proposition 1(iv). This is a load-bearing issue because the central claim's reach is being overstated in the paper's most-"}],"tokens_in":19950,"tokens_out":694,"duration_ms":97598,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper identifies a structural non-identifiability in spatial cluster trials and builds a useful framework around it. Under fixed-margin randomization with homogeneous exposure sets (H1–H3), the level of the spillover kernel is aliased with the intercept and direct effect at every sample size. That is Proposition 1, and it is a clean algebraic observation — s_0(z) = M(z)1 − qz̃ lies in col{[1, z̃]} regardless of the metric or randomization scheme. From there, Theorem 3 gives an exact bias decomposition (level + shape + leakage), the level leverage ℓ(w,c) is a single design-computable scalar that captures an estimand's exposure to the unidentified level, and Corollary 1 shows the conventional CRT analysis is a special case with closed-form contamination bias. The unification of Wang et al. (2025), Leung (2025), and Watson and Smith (2025) as different anchoring choices within one framework is genuinely useful organizational work, not just a literature review. The design taxonomy (Table 4) and the anchor-or-identify criterion (Proposition 4(iii)) are practically actionable and computable before unblinding. These are real contributions. The math within the linear-predictor class is straightforward and correct as far as I can follow it in the main text. The level term −ℓ(w,c)·ē being invariant to the link, working weights, and metric is a nice result. The bias floor (Corollary 2) — that more data of the same design cannot remove level bias — is the kind of thing trialists need to hear. Two soft spots, in proportion. First, the abstract says 'under any outcome model,' but Proposition 1(iv) covers the linear-predictor class: linear models with arbitrary Σ, GLMs, mixed models with conditional law depending on η + Zb. That is broader than 'linear model only' but it is not 'any outcome model.' Models with nonlinear exposure-response (saturating dose-response, threshold effects) are excluded, and these are scientifically plausible for the motivating examples. The paper is honest about this in Section 7, but the abstract overstates. Second, the full proofs are in supplementary material that is not available, and no code or data are shipped. The simulation study is thin and does not stress-test against nonlinear exposure mappings, which is exactly where the framework's boundaries matter. These are verification gaps, not reasons to dismiss the work. The core algebra is checkable from the main text for Proposition 1 and the structure of Theorem 3, but a referee needs the supplements. This paper is for methodologists working on spatial causal inference and trialists designing cluster trials with likely spillover. It deserves a serious referee who can verify the supplementary proofs and push on the scope claims. I would accept it for peer review.","headline":"Clean framework for spatial interference in cluster trials; core math is sound but abstract overstates scope and proofs/code are missing.","tokens_in":21253,"tokens_out":1464,"would_cite":true,"duration_ms":139444,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Spillover level can't be learned from cluster trials — only its shape","keywords":[],"falsifier":"If a cluster-randomized trial with fixed margins were augmented with a design feature generating residual variation in the constant-exposure vector (e.g., sentinel units with zero exposure), and the estimated level from that augmentation disagreed with the level implied by an anchor assumed valid under the canonical design, the anchor would be refuted.","tokens_in":20356,"feed_emoji":"🗺️","tokens_out":1254,"duration_ms":118005,"temperature":0.7,"pith_summary":"The paper proves that in cluster-randomized experiments with spatial spillover, the data can identify the shape of the spillover kernel (how effects vary with distance) but never its overall level (how much spillover there is in total). This non-identifiability is structural: it holds at every sample size, under every outcome model in the linear-predictor class, and cannot be broken by better statistical modeling of variance or random effects. Every estimator therefore silently relies on an anchoring assumption that fixes the level, and the paper derives an exact three-part bias decomposition for any such anchored estimator: a level term (proportional to how wrong the anchor is), a shape error (from misspecifying the kernel's form), and a leakage term (absorbed by the realized spatial geometry). A single scalar called the level leverage measures each causal estimand's exposure to the unidentified level, giving the exact level bias, the variance cost of estimating the level rather than assuming it, and the mean-squared-error floor that no additional data can remove. The conventional cluster-trial difference-in-means estimator is shown to be the special case of an implicit anchor whose contamination bias is given in closed form. Existing approaches — Bernoulli randomization, elicited decay bounds, assumed compact support — are all recast as different anchoring choices within this single framework, and the paper provides a taxonomy of design modifications that can purchase the missing level information, along with a criterion for when augmenting the design beats simply accepting an anchor.","feed_headline":"Cluster trials can't measure spillover magnitude — only its shape","feed_subtitle":"A structural non-identifiability result shows every spillover estimate rests on an untestable anchoring assumption, with bias computable in","key_machinery":"The level leverage ℓ(w,c) = m(w) + qc, a scalar that counts the per-unit exposure pairs through which an estimand references the unexposed counterfactual. It simultaneously gives the direction of non-identifiability, the exact first-order bias of any anchored estimator, the variance premium for estimating the level instead, and the MSE floor that persists as sample size grows.","core_discovery":"Under cluster randomization with fixed exposure set size, the level of the spillover kernel is aliased with the intercept and direct effect at every sample size and under every outcome model in the linear-predictor class. The data identify only the centered kernel (its shape), while every causal estimand whose level leverage is nonzero — including the direct effect, total cluster effect, and dose-response curve — depends on the unidentified level and therefore requires an anchoring assumption. The bias of any anchored plug-in decomposes exactly into a level term, a shape error, and a leakage term, each computable from the design before data collection. The conventional cluster-trial analysis","pith_inferences":[],"forward_implications":["Trialists can compute, before unblinding, exactly which estimands are vulnerable to the unidentified level and by how much, enabling informed design choices rather than post-hoc damage control.","The traditional practice of spacing clusters apart to prevent spillover is precisely the design that makes the no-spillover assumption unfalsifiable — it removes the variation needed to test it.","Design augmentations such as sentinel units, baseline periods, or randomised saturation can be compared on a common scale via the level information I₀, giving a principled criterion for when to augment versus anchor.","Estimated effects from spaced cluster trials may not transport to new geographic configurations, because the realized geometry is baked into the implicit anchor.","Distance contrasts of the spillover curve (effects at one distance minus effects at another) are identified by randomization alone and require no anchoring assumption, making them the robust estimands of choice when the level cannot be credibly anchored."],"fun_headline_variants":["Spillover kernel level is unidentifiable in cluster trials","Every spillover estimate needs an anchoring assumption","Cluster trials identify spillover shape but not magnitude","Spillover bias decomposes into level, shape, and leakage terms","Level leverage flags every estimand's exposure to untestable assumptions"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The entire framework assumes the spillover effect on each individual is a sum of contributions from nearby treated sources, each weighted by a function of distance — the so-called additive linear exposure mapping. If the true exposure-response relationship is nonlinear in the count of treated sources (for example, if two nearby sources interact synergistically rather than additively), the level-aliasing structure and the closed-form bias decomposition may not hold as stated.","fun_headline_variants_meta":{"raw":{"variants":["Spillover kernel level is unidentifiable in cluster trials","Every spillover estimate needs an anchoring assumption","Cluster trials identify spillover shape but not magnitude","Spillover bias decomposes into level, shape, and leakage terms","Level leverage flags every estimand's exposure to untestable assumptions"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":682,"prompt_tokens":614,"completion_tokens":68,"prompt_tokens_details":null},"tokens_in":614,"tokens_out":68,"duration_ms":57025,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T03:26:08.773059+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If a cluster-randomized trial with fixed margins were augmented with a design feature generating residual variation in the constant-exposure vector (e.g., sentinel units with zero exposure), and the estimated level from that augmentation disagreed with the level implied by an anchor assumed valid under the canonical design, the anchor would be refuted.","supporting_citations":[],"review_version":1}