{"id":"73c378b5-2136-4ff5-b36c-4e1aeea4f71f","arxiv_id":"2608.10356","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Design-based prediction-powered inference for spatial data with misspecified sampling weights leaves a non-vanishing spatial remainder, so coverage can fall as labels accumulate.","lead":"This paper recasts prediction-powered inference, which combines cheap wall-to-wall maps with a small ground-truth sample, for spatial populations where labels come from survey designs rather than random draws. It derives design-matched variances and tuning rules, and shows that under a misspecified sampling model even a doubly robust estimator can suffer coverage that worsens as more labels are collected.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Coverage erosion depends on ratio-stable labelling; under alternative sample-size growth the gap may shrink and the phenomenon may vanish.","rationale":"The reader's weakest-assumption field emphasized (A7) and (A8), including the alignment condition. This stress-test focuses instead on the ratio-stability component of (A8), which is the specific hypothesis that makes the gap n-invariant and hence drives the coverage-erosion result. The concern is not an internal inconsistency: Theorem 1 is explicitly conditional on (A8), and the paper is transparent that the empirical experiment uses the ratio-stable scaling. It is a boundary concern: if ratio-stability fails under common survey-design growth mechanisms, the qualitative headline may not transfer. The proposed test is a single, reproducible experiment on the same enumerated population that directly isolates the role of the labelling-growth mechanism. Even if the concern lands, it does not invalidate the theorem's conditional statement or the paper's other contributions (design variances, power tuning, estimated-propensity theory), so the verdict remains ACCEPT/UNCHANGED. The agreement is partial because the reader also flagged the outcome-model correctness assumption and the identifiability of G_U, which are distinct concerns; the ratio-stability issue is one part of the reader's (A8) concern but is sharpened here into a testable boundary question.","tokens_in":59629,"tokens_out":16196,"duration_ms":156461,"concrete_test":"On the fully enumerated Estonian population of Section 5.3, grow the expected label count by a non-ratio-stable mechanism: fix the logistic propensity coefficients (including the intercept) and increase the expected sample size by drawing a larger Poisson sample from the same population (equivalently, raise the sampling fraction while keeping pi_i/tilde_pi_i free to vary). At several values of n (e.g., 100, 400, 1600, 6400, 8000), compute the realized G_U using the known true propensities and the pseudo-true working propensities, and record the empirical coverage of the doubly robust interval. If |G_U| shrinks with n and coverage recovers toward nominal, the paper's central phenomenon is specific to ratio-stable scaling; if |G_U| remains roughly constant and coverage still erodes, the finding is robust to the labelling-growth mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1's headline conclusion—that the conditional gap G_U is free of the label count and hence that coverage can deteriorate as labels accumulate—rests on Assumption (A8)'s 'ratio-stable labelling': the label budget is grown by multiplying both the true and working labelling intensities by a common scale t, so that h_i = pi_i / tilde_pi_i is invariant to n. This is a specific, deliberately chosen growth mechanism. In real surveys, sample sizes are often increased by other routes: raising a fixed logistic intercept while holding the covariate slopes constant, drawing a larger SRS or stratified sample, or capping inclusion probabilities below 1. Under such mechanisms the weight-ratio field v = h - hbar can change with n, and in particular can shrink toward zero as the design approaches a census, making G_U vanish and the coverage erosion disappear. The paper's own Section 5.3 experiment grows n exactly by the ratio-stable intensity scaling, so it validates the theorem under its hypothesis but does not probe how robust the phenomenon is to the labelling-growth mechanism. Section A.8 notes that the drifting logistic intercept is only approximately ratio-stable, with relative error O(n/N); the real-population experiment reaches n/N = 0.166, where this approximation may be poor. Thus the practical claim that 'coverage can deteriorate as labels grow' is demonstrated only inside a specific asymptotic regime, and the boundary of that regime is not quantified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper recasts prediction-powered inference (PPI) in a design-based finite-population framework for spatial data. The estimand is a census parameter of a fixed pixel population, with randomness coming only from the labelling mechanism. Sections 3 and 4 provide exact design variances, a threshold for when spatial balance pays, optimal allocation, design-matched power tuning, and sandwich inference for estimated propensities. The main theoretical result, Theorem 1, derives an exact conditional finite-population remainder G_U for doubly robust estimation under a misspecified propensity and a correct outcome model, with variance σ_u^2 / N_eff,v, where N_eff,v is defined by the residual correlation matrix and the weight-ratio field. Under ratio-stable labelling (A8) this remainder does not shrink with the label count, so coverage can deteriorate as n grows when spatial dependence aligns with the weight-ratio field; for i.i.d. or exchangeable residual fields it is negligible when n/N→0. The paper validates the mechanism on a fully enumerated Estonian population of 48,175 cells with only the labelling simulated, and reports Estonian LUCAS land-cover and soil-carbon applications, including an explicit reappraisal of an earlier empirical claim.","tokens_in":59942,"tokens_out":10916,"duration_ms":101674,"significance":"Theorem 1's exact variance identity (9) is a clean and non-circular calculation from the stated model, and the definition of N_eff,v as a variance ratio is substantive rather than tautological: the paper shows that spatial coherence can make N_eff,v of order one, in which case the double-robustness remainder binds at ordinary label counts. This is a genuinely new point at the intersection of PPI and survey sampling. The paper is also unusually careful about its own limitations: Remark 7 notes that the gap is not identified from labels alone under a misspecified propensity, Appendix A.6 discloses the nuisance-rate condition (19) needed when the outcome model is fitted, and Section 7 explicitly narrows the empirical lessons after the soil-carbon study. The semi-synthetic validation on a real fully enumerated population with the labelling simulated, and the scrambling control that isolates the spatial contribution, are strong confirmatory evidence. The reproducible code and the detailed provenance table for each classical result further support the paper's reliability.","major_comments":[],"minor_comments":[{"comment":"The coverage-erosion phenomenon is demonstrated only under the ratio-stable labelling growth of (A8); a brief passage in Section 7 stating that under other growth mechanisms (for example, a fixed-intercept logistic design or simple random sampling with n increasing toward a census) the weight-ratio field v and hence G_U can change with n, so that the erosion is a regime-specific warning rather than a universal law, would help calibrate the practical reading of the abstract.","section":"Section 5.3 / Assumption (A8)"},{"comment":"The double use of h as a stratum subscript and as the weight ratio in Theorem 1 is flagged in the text, but the proximity of objects such as \\bar f_{S,h} and h_i in consecutive sections is still easy to misread; adding a one-line cross-reference at the first occurrence of h_i in Section 4 would reduce the notational burden.","section":"Section 2.3"},{"comment":"In the soil-carbon section, the heading \"The corrected claim\" could be read as an erratum; renaming it \"Refined claim\" or \"What the diagnostic should read\" would better reflect that the authors are generalising a claim rather than retracting it.","section":"Section 6.3"}],"recommendation":"accept","confidential_remarks":"The only substantive stress-test concern is that the coverage-erosion phenomenon depends on the ratio-stable labelling assumption; on reading the paper, this is explicitly assumed in (A8), stated in the abstract, and the real-population experiment grows n exactly by that mechanism, so the concern does not undermine the central theorem. The paper's self-disclosures about the limits of the empirical evidence are unusually thorough and strengthen rather than weaken the contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real paper, not a translation exercise dressed up as one. The theory is mostly classical survey sampling in PPI clothing, and the author says so in Table 7. What is actually new is Theorem 1: an exact finite-population identity for the doubly robust gap under a misspecified propensity, showing the gap is controlled by N_eff,v, the effective patch count seen through the weight-ratio field, and can be O(1) under spatially coherent residuals rather than vanishing with n. The corollaries for M-estimands and the exact least-squares case are useful, and the design-matched tau correction is a real practical fix. The paper ships code and runs a semi-synthetic validation on an enumerated 48,175-cell population, which is the right way to stress-test a finite-population claim.\n\nThe empirical sections are honest. The Estonian land-cover analysis shows the i.i.d. PPI++ tuning loses on the best map while the design-matched tuning gains, and the soil-carbon study is presented as a non-verdict rather than a win. The Section 7 discussion narrows the claims carefully. Good citation practice, including pointing out where results are implicit in Kim–Haziza and Yang et al.\n\nSoft spots: the headline phenomenon depends on \"ratio-stable labelling\" — growing the label budget by scaling true and working intensities together. Under other growth routes (fixed logistic intercept, SRS, capping) the weight-ratio field v can shrink or the design approaches a census, and the gap can vanish. The paper is explicit about this and even rejects the fixed-intercept sequence, so the concern is about the boundary of the regime, not an internal contradiction. The real-population experiment uses the exact ratio-stable mechanism (log-linear intensity scaling), so the O(n/N) logistic approximation issue raised in the stress test does not invalidate the demonstration, though the paper's discussion of the logistic link at n/N = 0.166 is slightly hand-wavy. A reader should not leave thinking coverage erosion is a law; it is a mechanism that requires the designated growth path plus spatial coherence between residual field and weight-ratio field. That is what the paper says, so I don't count it as a fatal flaw.\n\nWho should read: survey-sampling people and PPI practitioners, especially remote-sensing. It deserves a serious referee; the referee's job is to check the appendix's technical conditions (A7–A9) and the claim that Section 5.3 isolates the conditional mechanism. I'd accept this for review, and I'd cite it.","headline":"Worth a serious referee: the genuinely new result is Theorem 1's exact finite-population gap identity under spatial residuals; the rest is a transparent, useful translation of survey sampling into PPI, with the ratio-stable labelling caveat properly flagged.","tokens_in":60413,"tokens_out":3365,"would_cite":true,"duration_ms":31425,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62M30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Double robustness breaks for spatial PPI when the propensity is wrong","keywords":["prediction-powered inference","design-based inference","spatial sampling","double robustness","finite population","spatially balanced sampling","inverse probability weighting"],"falsifier":"Enumerate a fixed spatial population with known outcomes and map, draw many label sets at several budgets under a deliberately misspecified propensity plus a correct outcome model, and track the conditional bias and coverage of the self-normalised AIPW estimator on the same population. Theorem 1 predicts the bias stays constant at $G_U$ and coverage falls once $\\sqrt{n}|G_U|$ grows; observing the bias shrink with $n$ or coverage recover would refute it.","tokens_in":59400,"feed_emoji":"🗺️","tokens_out":9616,"duration_ms":75662,"temperature":0.7,"pith_summary":"This paper reworks prediction-powered inference (PPI) as design-based survey inference for a fixed spatial population: the map is a phase-one census, the labels a phase-two probability sample, and the target is the census parameter the gold-standard protocol would produce. Its central result is that double robustness is asymmetric for such census estimands. If the propensity model is misspecified while the outcome model is correct, the doubly robust estimator carries a conditional finite-population gap whose size is governed by the effective number of residual patches seen by the weight-ratio field; under ratio-stable labelling that gap does not shrink with the label count, so coverage can deteriorate as labels accumulate. The paper also gives exact design variances, a threshold for when spatial balance pays, design-matched power tuning, sandwich inference for estimated propensities, and confirms the mechanism on a fully enumerated 48,175-cell population and on Estonian land-cover and soil-carbon data.","feed_headline":"A wrong propensity makes spatial PPI coverage worsen with more labels","feed_subtitle":"A fixed gap from spatially correlated map error does not shrink as labels grow; design-matched tuning restores validity.","key_machinery":"The central object is the gap identity of Theorem 1, built from the weight-ratio field $v_i = \\pi_i/\\tilde\\pi_i - \\overline{\\pi/\\tilde\\pi}$ and the residual correlation matrix $P=[\\rho_u(s_i,s_j)]$. The effective sample size $N_{\\mathrm{eff},v} = (N\\bar h)^2/(v^\\top P v)$ counts how many independent patches of residual error the misspecified weights actually see; it is $O(N)$ for i.i.d. or exchangeable residuals and can be $O(1)$ under spatially coherent dependence. This identity carries the argument because it shows that the conditional finite-population remainder is free of the label count under ratio-stable labelling, and that correcting the propensity (making $v\\equiv 0$) kills it while correcting the outcome model does not.","core_discovery":"The paper's central claim is Theorem 1: in a fixed spatial population, with the outcome model correctly specified and the propensity model misspecified, the self-normalised doubly robust PPI estimator decomposes as design error $O_p(n^{-1/2})$ plus a conditional gap $G_U = (N\\bar h)^{-1}\\sum_{i\\in U} v_i u_i$, where $h_i = \\pi_i/\\tilde\\pi_i$ is the ratio of true to pseudo-true inclusion probability, $v_i = h_i - \\bar h$, and $u_i$ is the outcome-model residual. Conditionally on the realised population, $G_U$ is a fixed number, its superpopulation variance is $\\sigma_u^2 / N_{\\mathrm{eff},v}$ with $N_{\\mathrm{eff},v} = (N\\bar h)^2/(v^\\top P v)$, and under ratio-stable labelling it does not depend on $n$. Consequently, as labels accumulate, the fixed gap is divided by a shrinking design standard error and coverage can fall; a correct propensity sets $v=0$ and removes the gap, whereas spatial correlation of $u$ can inflate $v^\\top P v$ and shrink the effective patch count.","pith_inferences":["If Theorem 1 transfers to small-area estimation, the exposure is worst where labels are few: borrowing strength across areas reduces sampling variance without touching the propensity misspecification, so the fixed gap becomes relatively larger; the paper names small-area estimation as a likely site but does not formalise it.","The paper's warning about pooled residual diagnostics suggests a general caution: any Moran-type gate for choosing a spatial variance estimator should be run on design-centred residuals, since a pooled test can flag stratum effects as spatial dependence.","A testable extension would derive the effective patch count for spatially balanced designs with $\\pi_{ij}=0$ on within-block pairs, because Theorem 1's exact identity is stated for independent Bernoulli selection and such designs may change how the gap is evaluated."],"forward_implications":["For a fixed spatial population under simple random sampling, spatial correlation of the map error never enters the design variance, so i.i.d.-style PPI intervals hold nominal coverage and the map only shortens the interval.","Under clustered labelling, using the i.i.d. variance formula can lower coverage to 58% as the error field becomes smooth; a cluster-robust variance is required.","Spatially balanced one-per-block sampling reduces the design variance exactly when between-block variation exceeds $(n-1)/(N-n)$ times the average within-block variation; below that threshold it buys nothing.","The design-matched power-tuning coefficient, which minimises the stratified design variance, differs from the pooled PPI++ coefficient whenever the design has already removed between-stratum covariance, and using the pooled value can lose precision on the best maps.","With a misspecified propensity and a correct outcome model, the doubly robust estimator's conditional bias is the fixed gap $G_U$; coverage erodes once the design standard error falls below $|G_U|$, so enlarging the label budget can make matters worse."],"supporting_citations":[{"why":"introduces prediction-powered inference and the rectifier estimator that the paper recasts in design-based form.","marker":"Angelopoulos et al. (2023a)"},{"why":"supplies the two-phase difference-estimator theory and variance formulas that Propositions 1–4 translate into PPI notation.","marker":"Särndal et al. (1992)"},{"why":"provides the nested finite-population asymptotic framework used for all statements about growing label budgets.","marker":"Isaki and Fuller (1982)"},{"why":"derives the doubly robust finite-population remainder that Theorem 1 sizes and shows is implicit in design-model double robustness.","marker":"Kim and Haziza (2014)"},{"why":"gives the design-model doubly robust algebra for combining probability and non-probability samples from which the remainder term is taken.","marker":"Yang et al. (2020)"},{"why":"supplies the data-defect correlation principle underlying the non-shrinking gap $G_U$ in the big-data paradox.","marker":"Meng (2018)"},{"why":"develops the superpopulation spatially dependent doubly robust PPI whose conditioning choice Theorem 1 contrasts with conditional finite-population inference.","marker":"Salerno et al. (2026)"}],"fun_headline_variants":["Spatial PPI: a wrong propensity gap never shrinks as labels multiply","Design-matched tuning restores spatial PPI when propensity is misspecified","More labels, worse coverage: spatial PPI's wrong-propensity trap","Spatial PPI: fixed propensity gap makes coverage degrade with labels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The key assumption is that the map-error field has zero mean under the working model and that the misspecified selection weights vary in step with that field's spatial correlation; if either fails, the non-shrinking gap disappears or changes size.","fun_headline_variants_meta":{"raw":{"variants":["Spatial PPI: a wrong propensity gap never shrinks as labels multiply","Design-matched tuning restores spatial PPI when propensity is misspecified","More labels, worse coverage: spatial PPI's wrong-propensity trap","Spatial PPI: fixed propensity gap makes coverage degrade with labels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000866,"raw_usage":{"total_tokens":3845,"prompt_tokens":1127,"completion_tokens":2718,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":743,"completion_tokens_details":{"reasoning_tokens":2638}},"tokens_in":743,"tokens_out":2718,"duration_ms":18169,"temperature":1.0,"reasoning_tokens":2638,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:22:52.036944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Enumerate a fixed spatial population with known outcomes and map, draw many label sets at several budgets under a deliberately misspecified propensity plus a correct outcome model, and track the conditional bias and coverage of the self-normalised AIPW estimator on the same population. Theorem 1 predicts the bias stays constant at $G_U$ and coverage falls once $\\sqrt{n}|G_U|$ grows; observing the bias shrink with $n$ or coverage recover would refute it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the nested finite-population asymptotic framework used for all statements about growing label budgets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"derives the doubly robust finite-population remainder that Theorem 1 sizes and shows is implicit in design-model double robustness."},{"cited_title":"K., and Song, R","cited_arxiv_id":null,"evidence_quote":"gives the design-model doubly robust algebra for combining probability and non-probability samples from which the remainder term is taken."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the data-defect correlation principle underlying the non-shrinking gap $G_U$ in the big-data paradox."}],"review_version":1}