{"id":"57013cec-f639-4f5f-a305-c99f5c11e292","arxiv_id":"2607.19908","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Defines treatment geometry and provides a protocol for making spatial and temporal exposure choices explicit in geospatial impact evaluations.","lead":"This paper introduces 'treatment geometry' as a way to define the spatial and temporal footprint of an intervention in causal studies using Earth observation data. It provides a checklist of geometry decisions—like buffer size, time windows, and spillover handling—to make impact evaluations more credible and reproducible.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Underdetermined treatment geometry: sensitivity analysis over a wrong prior cannot establish credibility, so the protocol's central value claim is not yet supported.","rationale":"Reader identified the weakest assumption as the researcher's ability to specify the causal pathway. I agree but sharpen the point: the protocol's own philosophy—'not to find a universally correct geometry'—concedes that the pathway is a prior, yet the protocol presents sensitivity analysis as the way to build confidence (Sections 4.4, 6). Sensitivity to alternatives is only informative if the alternatives span the space of credible mechanisms. The paper gives no method for constructing that space and no external check. The case studies are ex post illustrations; they cannot show that the framework would improve a researcher's initial geometry choice. Consequently the strongest version of the paper's central claim—that explicit geometry choices strengthen the credibility of causal inference—is not yet demonstrated. This does not invalidate the framework as a vocabulary or checklist, but it suggests the chapter should either soften the credibility claim or add guidance: pre-register geometry selection, use validation data when available, or adopt a formal model-selection/placement criterion. Hence I recommend conditional acceptance: the conceptual contribution stands, but the prescriptive protocol should be amended to address underdetermination and specification search. No issues of fraud, internal inconsistency, or unsupported empirical claims arise; this is an argument about the scope of the contribution.","tokens_in":28734,"tokens_out":7166,"duration_ms":86347,"concrete_test":"Take a setting with a known gold-standard exposure measure, e.g., dense PM2.5 monitors downwind of a point source. Blind one team to the monitor data and have them apply the Box 1 protocol: specify the pathway, choose primary geometry (e.g., wind-cone), define 3-5 alternative geometries (different cone angles/heights, circular buffer), and run sensitivity/placebo checks. Then compare exposure classifications from the protocol-selected geometry to monitor-based truth, and estimate the treatment effect on a simulated outcome with known magnitude under the true monitor-based exposure. Repeat over scenarios where the true transport differs from the researcher's prior (e.g., terrain channeling). If the protocol's sensitivity analysis fails to flag the systematically misclassified geometry or yields materially biased effects with stable confidence intervals across alternatives, the framework's","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central value proposition is that explicitly defining treatment geometry enhances the credibility of causal inference (Abstract; Sections 1, 6). The protocol (Box 1) instructs researchers to specify the causal pathway and test plausible alternatives. But the framework supplies no independent anchor for adjudicating among geometries. It states the true exposure process is 'only partially observed' (Section 6) and 'rarely a single correct geometry' yet offers no validation against external exposure data, no pre-registered decision rule, and no formal criterion (e.g., predictive loss against a gold standard). If the researcher's causal prior is mistaken, the set of 'plausible alternative geometries' can be uniformly wrong in the same direction; sensitivity analysis among them will show stability while the estimates remain biased. The six case studies are retrospective narratives that illustrate the framework but do not test whether following it would have corrected or revealed geometry-induced bias. Thus the claim that the protocol 'can strengthen credibility' remains an assertion. The concern is load-bearing because it targets the manuscript's stated contribution: without a way to discriminate between a right and a wrong geometry, 'documenting choices' is transparency, not credibility.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript introduces the concept of \"treatment geometry\" for geospatial impact evaluation (GIE): the spatial and temporal footprint through which exposure to a treatment is represented in data. The chapter decomposes treatment geometry into four components—source, treated unit, pathway, and temporal window—and discusses how choices about these components interact with data resolution, aggregation, temporal alignment, treatment intensity, boundary uncertainty, and directional transport. It connects geometry choices to identification concepts such as estimand clarity, comparison-group construction, SUTVA, misclassification, and nonclassical measurement error. The second half offers six case studies from conservation, air pollution, agricultural shocks, conflict, and survey georegistration, and closes with a protocol (Box 1) for documenting and stress-testing geometry choices before estimation. The contribution is a conceptual synthesis and practical checklist rather than a formal theorem or an empirical validation of the proposed framework.","tokens_in":28984,"tokens_out":4873,"duration_ms":55099,"significance":"If accepted at face value, the framework would give applied researchers a useful vocabulary for decisions that are often treated as technical preprocessing details—buffer distances, pixel-extraction rules, temporal windows, coordinate displacement—and would connect those decisions to identification. The manuscript draws on a broad and current literature, and the protocol is a valuable transparency device. Its stated goal is not to prescribe one correct geometry but to make geometry choices explicit, defensible, and reproducible. That is a real contribution. However, the paper's central value claim—that explicit treatment geometry can \"strengthen the credibility of causal inference\"—is asserted and illustrated rather than demonstrated. The case studies are retrospective illustrations, not tests of the protocol, and the framework provides no independent way to adjudicate between a correct and a uniformly misspecified set of geometry assumptions. With a recalibrated claim and one worked example that includes external validation or falsification, the chapter would be a solid methods contribution.","major_comments":[{"comment":"The core claim that explicitly defining treatment geometry \"can strengthen the credibility of causal inference\" is stronger than the protocol supports. Box 1 asks researchers to document source, unit, pathway, temporal window, and to report sensitivity to alternatives (Q10), but it supplies no independent criterion for adjudicating among geometries when the true exposure process is only partially observed, as Section 6 concedes. Sensitivity over a set of misspecified geometries can be stable while all estimates remain biased; for example, in Case Study 2, if the wind fields used to build the pollution geometry are wrong, varying cone angles will not reveal the error. The manuscript should either (a) add a validation or falsification step—such as comparing the preferred geometry to ground-truth monitors, administrative event records, or a pre-registered hold-out prediction target—or (b) r","section":"Abstract; Section 6, Box 1 (Protocol)"},{"comment":"The six case studies are retrospective narratives assembled after the fact; none is a prospective application of Box 1, and none shows that following the protocol would have corrected or revealed geometry-induced bias. Since the manuscript's value rests on the protocol being actionable, a worked example that implements the full protocol—with pre-specified alternatives, a decision rule, and a falsification exercise—would materially support the central claim. At minimum, the text should state explicitly that the cases are illustrative of the conceptual taxonomy and are not evidence for the protocol's effectiveness.","section":"Section 5, Table 1"}],"minor_comments":[{"comment":"\"Operationationalizing\" is a typo for \"Operationalizing.\"","section":"Section 2 heading"},{"comment":"There are two boxes both numbered Box 1: \"Collecting Ground Reference Data\" in Section 3.2.1 and \"A Protocol for Defining Treatment Geometry\" at the end of Section 6. These should be renumbered and cross-referenced consistently.","section":"Section 3.2.1 and Section 6"},{"comment":"Several placeholders remain in the text: Figure XX, Figure YY, and Table XX. These need to be resolved before publication.","section":"Section 3.1, Figure XX; Section 3.2.2, Figure YY; Further Resources, Table XX"},{"comment":"\"aligning exposure to data availability rather than to casual process\" should read \"causal process.\"","section":"Case Study 4, Key risks"},{"comment":"The in-text citation \"Lui et al., 2022\" does not match the reference \"Liu, Y., et al. (2022)\" in the bibliography. Also, \"Vincente-Serrano\" in the text should be \"Vicente-Serrano\" to match the reference list.","section":"Case Study 5 and References"},{"comment":"Robert Heilmayr's biography appears twice, with slightly different affiliations. This duplication should be removed.","section":"Contributors' Biographies"},{"comment":"Some entries are working papers or preprints (e.g., Grosset-Touba et al. 2024; Jordán and Heilmayr 2024; Pignède 2025). If the chapter is intended as a permanent reference, the authors should note whether these have since been peer reviewed or update the citations accordingly.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a useful methods chapter, and the central concept is sensible. The gap between the 'credibility' language and the protocol's transparency-only content is fixable in revision: either add an external-validation or pre-registration anchor for geometry choice, or carefully restate the contribution as procedural transparency and documentation. I do not see grounds for rejection. The placeholders and duplicated biography must also be cleaned up before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version: it's a competent, useful synthesis chapter, not a research contribution. It gives the field a useful name—treatment geometry—and a four-part decomposition (source, treated unit, pathway, temporal window) that genuinely clarifies how exposure definitions shape identification. If I were teaching geospatial impact evaluation, I'd assign it. The six case studies are well chosen and illustrate the main tradeoffs: spillovers, directional transport, sparse data, timing, intensive versus extensive margins, and georegistration. The chapter is honest about what it builds on—MAUP, uncertain geographic context, SUTVA/interference, nonclassical measurement error. No inflated claims there.\n\nWhat it doesn't do is produce new theory or new empirical evidence. The central claim—geometry choices matter for identification—is asserted and illustrated, not proven. For a framework chapter, that's fine; the burden is coherence and usefulness, and it meets those. The protocol in Box 1 is a sensible checklist; it won't tell you which geometry is right. The stress-test worry about underdetermination is real but not fatal: if a researcher's causal prior is wrong, sensitivity analysis over alternative geometries can be stable and still biased. The paper itself says the exposure process is 'only partially observed' and 'rarely a single correct geometry,' so it doesn't claim to solve that; it claims to make assumptions visible, testable, reproducible. That's a meaningful, modest contribution. The one thing a referee should push on is the line in the abstract and conclusion about 'strengthening credibility'—that overstates what transparency can deliver. Suggest rewording to 'reduces the risk of arbitrary choices' or 'improves interpretability.'\n\nOn the mechanics: the manuscript is rougher than it should be. There are 'Figure XX' / 'Figure YY' / 'Box XX' placeholders, typos ('Operationationalizing,' 'anAssociate,' 'Vincente-Serrano,' 'Lui'), a duplicated Heilmayr bio, and a citation mismatch (Chatterjee 2020 in text vs Chatterjee 2022 in references; Petruccelli et al. 2026 and Josephson et al. 2026 are listed as forthcoming, which is fine if confirmed). None of this is substantive, but it needs a careful edit.\n\nI'd send this to peer review. The subfield will use it as a reference, and a referee with applied spatial experience can help scope the claims and clean up the presentation. I'd accept after moderate revision rather than desk-reject.","headline":"A solid synthesis chapter that gives geospatial impact evaluation a useful vocabulary and a sensible default protocol, but it's a framework, not a new result; a referee can help clean it up and temper its 'credibility' claims.","tokens_in":29455,"tokens_out":2929,"would_cite":true,"duration_ms":30664,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The treated area you draw on a map decides what causal effect you estimate.","keywords":["treatment geometry","geospatial impact evaluation","causal identification","earth observation data","exposure assignment","spillovers","measurement error","research design"],"falsifier":"A meta-analysis of published geospatial impact evaluations that report multiple treatment geometries would falsify the core claim if it showed that estimates are rarely sensitive to changes in buffers, time windows, or aggregation units in typical settings.","tokens_in":28660,"feed_emoji":"🛰️","tokens_out":2529,"duration_ms":32048,"temperature":0.7,"pith_summary":"This chapter introduces treatment geometry—the spatial and temporal footprint through which exposure to a treatment is represented in data—as a core analytical concept for geospatial impact evaluation. It argues that choices about polygons, buffers, time windows, and assignment rules are not technical details: they define the estimand, the comparison group, and the main threats to causal identification. The paper decomposes treatment geometry into four elements—source, treated unit, pathway, and temporal window—and provides a protocol for making these choices explicit, justified, and reproducible. A sympathetic reader would take away that the credibility of causal estimates from satellite-linked studies depends less on data volume than on how exposure is encoded.","feed_headline":"How you define 'treated' changes what you find","feed_subtitle":"Satellite impact studies often pick buffers and time windows by default; this framework makes those choices visible and testable.","key_machinery":"Treatment geometry is the central object: the spatial and temporal footprint through which exposure to a policy, intervention, hazard, or environmental condition is represented. It is operationalized by decomposing it into four elements—source, treated unit, pathway, and temporal window—which together define a defensible mapping from the underlying causal process to an empirical treatment variable. The concept carries the argument by forcing transparency about assumptions that are often left implicit, such as where exposure becomes zero, how exposure travels, and which temporal window is relevant.","core_discovery":"The chapter's central claim is that any geospatial impact evaluation implicitly chooses a treatment geometry—the spatial and temporal footprint through which exposure is encoded—and that this choice is itself part of the identification strategy. Because multiple reasonable geometries can exist for the same research question, the paper proposes decomposing the geometry into the source of treatment, the treated unit, the pathway of exposure, and the temporal window. These elements jointly determine which units are treated, which serve as controls, what variation identifies the effect, and what assumptions are required for causal interpretation. Rather than prescribing a single correct geometry","pith_inferences":["A natural extension is to formalize treatment geometry as a reporting standard for geospatial impact evaluations, similar to pre-analysis plans, so that readers can see which geometry choices were made and why.","The framework could be applied to settings beyond satellite-linked socioeconomic studies, including environmental epidemiology and conservation science, where exposure surfaces are similarly constructed from buffers, plumes, or travel-time catchments.","A testable extension would be a benchmark exercise where several plausible treatment geometries are applied to the same research question to quantify how much estimates and inference vary across reasonable choices.","The four-element decomposition may also serve as a diagnostic tool for machine-learning-based exposure models, indicating where learned exposure surfaces conflict with the assumed source, pathway, or temporal window."],"forward_implications":["Researchers should define and document treatment geometry before estimation, treating buffer distances, aggregation rules, and time windows as modeling assumptions rather than defaults.","Sensitivity analysis across alternative geometries becomes a core diagnostic: if results depend on small changes in buffers, thresholds, or windows, the causal claim is correspondingly fragile.","Spillovers should be built into comparison-group construction—by excluding, modeling, or explicitly estimating indirectly treated units—rather than treated only as a post-estimation robustness check.","Spatial anonymization of survey coordinates limits which research questions are answerable; high-resolution exposure measures are more sensitive to coordinate displacement than smooth, coarse weather products.","The same underlying data can support different estimands: a binary treatment answers a different question than a continuous dose-response measure, and researchers should align the geometry with the intended causal object."],"fun_headline_variants":["Your buffer and time window are your identification strategy","In satellite impact studies, geometry is the method","Your spatial footprint is your causal assumption","How you define exposure determines your effect","The hidden choice in satellite impact studies: treatment geometry"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework assumes researchers can specify the causal pathway—source, treated unit, pathway, and temporal window—with enough confidence to choose among alternative geometries.","fun_headline_variants_meta":{"raw":{"variants":["Your buffer and time window are your identification strategy","In satellite impact studies, geometry is the method","Your spatial footprint is your causal assumption","How you define exposure determines your effect","The hidden choice in satellite impact studies: treatment geometry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001292,"raw_usage":{"total_tokens":5072,"prompt_tokens":661,"completion_tokens":4411,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":405,"completion_tokens_details":{"reasoning_tokens":4357}},"tokens_in":405,"tokens_out":4411,"duration_ms":34052,"temperature":1.0,"reasoning_tokens":4357,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T11:19:07.982197+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A meta-analysis of published geospatial impact evaluations that report multiple treatment geometries would falsify the core claim if it showed that estimates are rarely sensitive to changes in buffers, time windows, or aggregation units in typical settings.","supporting_citations":[],"review_version":1}