{"id":"b3d003ff-07fc-444b-af19-df156c39778e","arxiv_id":"2608.03804","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic usability evaluation of 89 geospatial foundation models finds severe accessibility gaps, including no surveyed model offering documented uncertainty quantification.","lead":"A user-experience evaluation of 89 geospatial AI models finds most are hard for ecologists to use: over half require heavy technical setup and none provide documented uncertainty estimates. The paper introduces a seven-dimension usability rubric and diagnoses field-level gaps in trust, permanence, and multilingual support.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"D1 inter-rater reliability is under-reported and internally inconsistent (nominal κ=0.56 vs κ=0.912); the 50.6% 'inaccessible' headline is not backed by a 'highly consistent' rating process.","rationale":"The reader's weakest_assumption focused on corpus representativeness, which is a valid generalizability concern but secondary to the internal reliability of the primary measurement. The paper's own reporting contradicts itself: Section 6.6 states nominal κ=0.56 for D1, while Appendix C Table C.6 reports κ=0.912 for the same dimension without clarifying that these are different metrics (nominal vs quadratic-weighted). This is not merely a reporting inconsistency; it bears directly on whether the two raters agreed on the dimension that generates the 50.6% figure and on the 'highly consistent' wording in the abstract. Since only 20 models were double-rated, the remaining 69 ratings are single-observer and could carry systematic rater bias at the exact boundaries where disagreement was acknowledged. A third-rater sample or a full double-rating pass would settle whether the D1 distribution is reliable. This concern does not change the overall verdict (CONDITIONAL remains appropriate) but refines the condition: the authors must demonstrate D1 reliability, not just assert it. I partially agree with the reader because I also see real issues, but I prioritize measurement reliability over corpus representativeness as the more load-bearing threat to the central quantitative claim.","tokens_in":38714,"tokens_out":5887,"duration_ms":61158,"concrete_test":"Obtain the raw rating matrix for the 20-model calibration set. (1) Reproduce D1 nominal Cohen's κ and Gwet's AC1; reconcile with the 0.56 in Sec 6.6 and the 0.912 in Table C.6. (2) Report the post-discussion D1 ratings for the calibration set, and compute κ again to see if discussion fixed the disagreement. (3) Have a third, independent rater (not an author) score a random sample of 15–20 of the 69 single-rated models using the same rubric. Compare the D1 distribution from this sample to the original ratings; if the proportion at Level 0–1 differs by more than ~5 percentage points, the 50.6% headline is not robust. Also report per-rater D1 means for the full corpus to check for rater bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim is the D1 distribution: 50.6% of 89 models at Level 0–1 (Sec 5.1). The reliability evidence is contradictory: Sec 4.3 reports overall nominal κ=0.92, ordinal weighted κ=0.95; Appendix C Table C.6 reports D1 AC1=κ=0.912 (typed 'Ordinal'); but Sec 6.6 states D1 nominal κ=0.56. The table uses AC1 (a nominal statistic) for an ordinal dimension, and no table row reports the nominal κ=0.56. If nominal κ for D1 is 0.56, agreement on the key scale is only moderate. More importantly, only 20 models were double-rated; the remaining 69 were each coded by one rater after a 'discussion' that resolved the Level 2/3 disagreement. The paper never shows that the post-discussion ratings converge, nor does it report per-rater D1 distributions. Any systematic difference between raters—one more lenient, the other stricter—directly shifts the percentage of models at Levels 0/1 vs 2/3, since the boundaries are exactly where they disagreed. The 'highly consistent' claim in the abstract is not supported by the 0.56 figure. This is not a problem with the rubric's conceptual validity but with the demonstrated reliability of the primary outcome measure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a seven-dimension usability evaluation framework for geospatial foundation models (GeoFMs), motivated by a pilot expert survey (N=11) with ecology and conservation scientists. The framework covers Access & Deployment, Interaction & Customization, Trust & Transparency, Community & Support, Scientific Permanence, Multilingual Support, and Offline Usability. Two raters applied the rubric to 89 GeoFMs drawn from four existing surveys. The main empirical claims are that 50.6% of the 89 models sit at access levels 0–1 (no usable pretrained weights or weights with minimal environment guidance), no accessible model provides documented uncertainty quantification, and documentation/community support is generally thin. The paper distinguishes discriminating dimensions from field-level diagnostics and closes with recommendations for embedding-as-data, usability metadata, and uncertainty-aware design.","tokens_in":38939,"tokens_out":6263,"duration_ms":60796,"significance":"The contribution is potentially valuable: it is one of the first usability-centric, systematic evaluations of GeoFMs, with a transparent rubric and a full per-model rating table (Appendix D). The finding that uncertainty quantification is entirely absent from a large sample of models, and that most models require substantial ML expertise, is a useful challenge to the model-centric evaluation culture in this field. The authors also deserve credit for explicitly bounding their corpus, reporting detailed limitations, and making their rubric and survey instrument available in appendices. However, the central quantitative claims depend on (a) the reliability of the D1 ratings and (b) the representativeness of the corpus, and both need reinforcement before the field-level conclusions can be taken as established.","major_comments":[{"comment":"The reliability evidence for D1 is internally inconsistent. Section 4.3 reports overall nominal κ=0.92 and ordinal weighted κ=0.95; §6.6 states D1 nominal κ=0.56 (quadratic-weighted κ=0.91); Appendix C Table C.6 reports D1 AC1=κ=0.912 and labels the row 'Ordinal.' These three statements cannot all describe the same calibration sample. Since D1 is the primary outcome, the abstract's 'highly consistent' claim and the headline 50.6% figure rest on unresolved reporting. Moreover, only 20 models were double-rated; the remaining 69 were coded by one rater after a discussion that resolved the Level 2/3 boundary, but no per-rater distributions or post-discussion reliability are reported. A systematic leniency difference between raters at exactly the Level 0/1 versus 2/3 boundary would directly shift the 50.6% figure. Please reconcile the statistics, report nominal κ for D1 explicitly, and provid","section":"§4.3, §6.6, Appendix C Table C.6"},{"comment":"The corpus is assembled from four secondary surveys, and the paper explicitly notes it is bounded by them, yet the Discussion repeatedly upgrades aggregate percentages to 'field-level diagnostics' and 'field-level norms.' No sensitivity analysis is provided for corpus composition; if the four surveys under-represent industry models, non-English venues, or models released outside those surveys, the estimates (e.g., 0% with UQ, 96.8% open-core-only) could be sample-specific. At minimum, the paper should either restrict the claims to 'the 89 models in these four surveys' or provide a robustness check (e.g., leave-one-survey-out, comparison with a hand-curated list, or a characterization of the excluded margin).","section":"§4.1, §5.8, §6.5"},{"comment":"The pilot survey (N=11, single institution) is used to support the central interpretive claim that uncertainty quantification is the feature domain experts prioritize most for trust, and this drives the 'misalignment' conclusion. The survey section itself is carefully hedged as exploratory, but the Results and Discussion treat it as strong evidence ('the feature that our survey found domain experts prioritize most highly,' 'the most direct evidence of misalignment'). With 10 respondents answering the trust question and 5 selecting uncertainty visualization, this is a 50% preference in a very small, homogeneous sample. Please temper the field-level language or supplement with additional evidence.","section":"§2.2.3, §5.3, §6.2"},{"comment":"Two of the four authors are co-authors of Tessera, which is rated D1=5, receives the only full set of interaction paradigms (I1–I5) and one of the two Level-5 community scores, and is used in §6.1 as the 'existence demonstration' for the embedding-as-data-product model. The manuscript does not disclose this relationship in the rating methodology or in a competing-interests statement. For an evaluation study, the absence of disclosure makes it impossible for readers to assess potential bias in the ratings of Tessera and in the archetype comparison of Table 5. Please add a COI statement and, ideally, have an independent rater score Tessera or justify why the co-authorship does not affect the rating.","section":"§4.2, §5.9.5, §6.1"}],"minor_comments":[{"comment":"The text says 'two raters independently evaluated each model in the corpus,' but the described procedure has only a 20-model overlap double-rated; the remaining models are single-rated. Please rephrase to avoid implying full double-coding.","section":"§4.2"},{"comment":"The D1 row is typed 'Ordinal' but reports AC1 and κ; Section 4.3 says AC2 is used for ordinal dimensions. Clarify which statistics are computed for ordinal versus nominal dimensions.","section":"Appendix C Table C.6"},{"comment":"Typo: 'reproducibility can remains fragile' should be 'reproducibility can remain fragile.'","section":"§5.5"},{"comment":"The row 'DOF A (Xiong et al., 2024)' appears to be a typo for 'DOFA' (or possibly duplicates 'OFA-Net'). Please correct the entry and check the model list for duplicates.","section":"Table D.7"},{"comment":"The denominator for D6 is '61 models with available data,' whereas §5.2 reports 63 accessible models. Explain why two models lack D6 data (e.g., no documentation located) so the percentages are interpretable.","section":"§5.6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is readable and within scope. The main concern is the reliability reporting, which needs to be reconciled before the headline claims can be accepted. The authors should also be asked to add a competing-interests statement regarding Tessera. The paper is timely, and the appendices are a strength; with a careful revision of the statistics and corpus claims, it could make a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a genuinely useful first systematic usability landscape for GeoFMs. The seven-dimension rubric is concrete and mostly well-grounded in HCI theory plus a small practitioner survey, and the 89-model corpus with per-model profiles is a real contribution. Findings like 50.6% of models at Level 0–1 and zero documented uncertainty quantification are plausible and would be valuable if confirmed. The paper is also transparent about the single-institution N=11 survey and the snapshot nature of the corpus.\n\nThe soft spots are real, though not fatal. The inter-rater reliability story is inconsistent where it matters most. Section 4.3 reports overall nominal kappa = 0.92 and weighted 0.95; Appendix C gives D1 AC1 = 0.912 typed as ordinal; Section 6.6 says D1 nominal kappa = 0.56. That 0.56 is on the primary access dimension, and the abstract's \"highly consistent\" does not square with it. Only 20 models were double-rated; the remaining 69 were split after a discussion that resolved the Level 2/3 boundary. Unless post-discussion per-rater agreement is shown, the 50.6% headline could shift with rater leniency. That is load-bearing for the field-level diagnostic claim.\n\nSecond, two of the four authors are Tessera co-authors, and Tessera is rated the most feature-complete model and used as an existence proof in Sections 5.9.5 and 6.1. The paper does not flag this. It does not invalidate the framework, but it should be disclosed, and ideally an independent rater should re-score that model.\n\nThe corpus-representativeness concern is real but minor: the paper explicitly bounds the corpus by four source surveys and the October 2025–January 2026 window, and no sensitivity analysis is provided. That is a stated limitation rather than a hidden flaw.\n\nBottom line: the rubric and the dataset deserve circulation. The paper needs a careful revision on IRR reporting and COI disclosure, and the field-level claims should be softened until the reliability is demonstrated. Send it to peer review.","headline":"A genuinely useful first systematic usability rubric for GeoFMs, but the reliability evidence for its headline numbers is shakier than the abstract admits, and the authors' own Tessera is scored top without disclosure: worth refereeing after careful revision.","tokens_in":39530,"tokens_out":1723,"would_cite":true,"duration_ms":19644,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A seven-dimension usability review of 89 geospatial foundation models finds that no release documents uncertainty quantification—the trust feature domain experts rank first—and that over half the models demand machine-learning expertise jus","keywords":["geospatial foundation models","usability evaluation","human-computer interaction","Earth observation","uncertainty quantification","expert elicitation","accessibility","systematic review"],"falsifier":"Two checks would settle it. (1) Independently audit the 63 accessible code repositories for runnable, documented uncertainty-estimation code: finding even one model with per-pixel or per-prediction confidence output would overturn the 0% (T3) headline, converting it from a capability gap to a documentation gap. (2) Rebuild the corpus from a different set of primary sources—industry releases, Chinese-language publications, agency products—and re-run the rubric: if the aggregate percentages (50.6% at Level 0–1, ~95% monolingual) do not survive, the field-level conclusions are an artifact of the","tokens_in":38525,"feed_emoji":"🛰️","tokens_out":13626,"duration_ms":129424,"temperature":0.7,"pith_summary":"This paper claims that geospatial foundation models—the large AI systems trained on satellite imagery to power environmental monitoring—have been built for machine-learning specialists, not for the ecologists, conservation scientists, and planners who need them. The authors surveyed eleven domain experts, turned their stated priorities into a seven-dimension usability rubric, and applied it to 89 published models. They report that 50.6% of models require users to re-train from scratch or assemble a complex software environment, and that no model provides documented uncertainty quantification—the single feature the experts most wanted for trusting outputs. The paper reads this as a field-level misalignment: development has optimized benchmark accuracy while leaving interfaces, documentation, trust features, and archival permanence underinvested. If correct, the finding reframes geo-AI progress from 'how accurate is the model?' to 'can the intended scientist actually use it?', a prerequisite for these tools to affect climate and biodiversity decisions.","feed_headline":"Zero of 89 geospatial AI models disclose their uncertainty","feed_subtitle":"Ecologists trust confidence maps above all; a systematic review of 89 models finds none exist—and half are out of reach.","key_machinery":"The carrying mechanism is the seven-dimension framework itself, a rubric turning 'can the intended user actually run this model?' into measurable levels and checklists. Access & Deployment is a 0–5 ladder from 'source code only' to 'graphical interface'; Interaction & Customization is a base level plus a checklist of documented interaction paradigms; Trust & Transparency is a four-attribute checklist including replicable benchmarks and uncertainty quantification; community support, scientific permanence, multilingual support, and offline usability complete the instrument. The framework does two jobs: wide-variance dimensions (access, interaction, community) discriminate between models today,","core_discovery":"The paper claims the geospatial foundation model field has optimized benchmark accuracy while underinvesting in user-centered infrastructure. Grounded in an eleven-expert survey, the authors built a seven-dimension usability rubric and had two raters score 89 models (93% observed agreement). Findings: 50.6% of models require pre-training or a poorly supported environment; only 6.7% reach a hosted API or GUI; zero of 63 accessible models document uncertainty quantification—the trust feature experts ranked first; 96.8% are open-core but only 1.6% archived with a DOI; 95.2% are monolingual. Wide-variance dimensions discriminate models today; near-zero-variance dimensions are field-level diagnos","pith_inferences":["The corpus is bounded by four source surveys, so the field-level percentages are only as good as that coverage; an independent corpus built from industry releases, agency products, and non-English venues is the untested check, and until it runs, the 50.6% and 0% figures should be read as claims about the surveyed landscape, not a census.","The paper's 'communicated not latent capability' stance yields a sharp test: audit the 63 accessible codebases for runnable, documented uncertainty code. Finding any would convert the headline from 'no model provides UQ' to 'no model documents UQ'—a documentation failure, arguably still a usability failure, but a materially different claim.","The expert survey's top-rated AI capability, anomaly detection (6.4/7), points to a product niche the paper does not pursue: embedding-based change-flagging tools aimed at field ecologists rather than ML engineers, which would directly exercise the offline pre-cache-and-analyze workflow the framework defines.","The paper's access trend rests on small annual samples (n=5 to n=30) and its own reading is cautious; a sharp prediction follows from its framework—if the 2025 jump to mean level 2.2 is a genuine norm shift, 2026 releases should stay at or above it and the diagnostic dimensions should finally show variance, while a fallback would indicate a selection artifact."],"forward_implications":["Practitioners get a working shortlist: only 6 of 89 models reach the top access tiers (hosted API or GUI), so any project expecting a field scientist to run a GeoFM today must plan around that short list or budget for ML-engineering support.","No ecological conclusion drawn from a GeoFM currently rests on published confidence estimates; the paper's strict threshold suggests the fix is architectural—downstream heads should report calibrated intervals and embedding-level out-of-distribution flags by default.","Because the rubric scores documented capability, the result doubles as an incentive: stating a usability profile in a model card would turn invisible gaps into selection pressure and give developers credit for work that already exists but goes unreported.","The embedding-as-data-product pattern—precomputing embeddings once and serving them as versioned, downloadable data—is the demonstrated route to closing the access gap, and the paper argues a community norm of wrapping inference code in lightweight web demos could lift most releases several levels at negligible cost.","Temporal trends move in one direction only: 2025 releases were more accessible (54% at Level 3+, up from 27% in 2024), but trust, permanence, and multilingual dimensions showed no improvement at all, so recent progress is concentrated in access alone."],"supporting_citations":[{"why":"One of the four source surveys whose model lists define the 89-model corpus; supplies the model-centric evaluation baseline the paper contrasts against.","marker":"Xiao et al. (2025)"},{"why":"Source survey for the corpus, cataloging Earth and climate foundation models by training data and benchmark performance.","marker":"Zhu et al. (2026)"},{"why":"Source survey covering vision foundation models released 2021–2024; contributes models and the release-timeframe baseline for the temporal-trend analysis.","marker":"Lu et al. (2025)"},{"why":"Source survey whose appendix catalogs 80+ GeoFMs; also the Tessera reference used as the existence proof that an open embedding-as-data product reaches top access levels.","marker":"Feng et al. (2026)"},{"why":"Supplies the AC1 agreement statistic the paper uses to report inter-rater reliability under prevalence-skewed ratings where Cohen's kappa reads paradoxically low.","marker":"Gwet (2008)"},{"why":"Documents usability struggles of domain experts with established geospatial tools; grounds the claim that GeoFM accessibility problems extend a known pattern.","marker":"Ziegler and Chasins (2023)"},{"why":"The FAIR4RS principles operationalized as Dimension 5, defining why code-on-GitHub without permanent archiving fails scientific permanence.","marker":"Barker et al. (2022)"},{"why":"Model cards, the mechanism the paper proposes extending with structured usability metadata to create selection feedback.","marker":"Mitchell et al. (2019)"},{"why":"The 'gulf of execution' concept that defines the Access & Deployment ladder (Dimension 1), the paper's main discriminator.","marker":"Norman (1988)"}],"fun_headline_variants":["89 GeoFMs scored: zero disclose uncertainty","Only 6 of 89 geospatial AI models reach hosted APIs","Geospatial AI: 95% monolingual, zero uncertainty","Ecologists' top trust feature absent from all 89 GeoFMs"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The 89-model corpus, assembled from four published model surveys, stands in for the entire geospatial foundation model landscape; if large segments (industrial models, non-English venues, or models omitted from those four surveys) differ systematically, the field-level percentages—50.6% inaccessible, 0% with documented uncertainty—would be wrong, and the paper does not test how its numbers respond to corpus composition.","fun_headline_variants_meta":{"raw":{"variants":["89 GeoFMs scored: zero disclose uncertainty","Only 6 of 89 geospatial AI models reach hosted APIs","Geospatial AI: 95% monolingual, zero uncertainty","Ecologists' top trust feature absent from all 89 GeoFMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000529,"raw_usage":{"total_tokens":2370,"prompt_tokens":712,"completion_tokens":1658,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":1586}},"tokens_in":456,"tokens_out":1658,"duration_ms":14363,"temperature":1.0,"reasoning_tokens":1586,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:03:55.043746+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two checks would settle it. (1) Independently audit the 63 accessible code repositories for runnable, documented uncertainty-estimation code: finding even one model with per-pixel or per-prediction confidence output would overturn the 0% (T3) headline, converting it from a capability gap to a documentation gap. (2) Rebuild the corpus from a different set of primary sources—industry releases, Chinese-language publications, agency products—and re-run the rubric: if the aggregate percentages (50.6% at Level 0–1, ~95% monolingual) do not survive, the field-level conclusions are an artifact of the","supporting_citations":[],"review_version":1}