Pith. sign in

REVIEW 3 major objections 5 minor 71 references

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding

T0 review · 3 major / 5 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read Ultra-high-resolution remote sensing improves when models organize a small, topology-preserving evidence set under global context, not when they simply see more pixels.

desk verdict Solid training-free UHR systems paper with real multi-benchmark gains; the “organize better” slogan is only partly isolated by the ablations. read the letter →

arxiv 2607.10120 v1 pith:WFM45NYL submitted 2026-07-11 cs.CV

classification cs.CV
keywords Vision-LanguageModelsUltra-High-ResolutionRemoteSensingTraining-FreeUnderstandingStructuredEvidenceReasoningMinimalSupportSetGlobalContextConstraintTopology-PreservingBoard
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Ultra-high-resolution remote sensing images force vision-language models to hold both a wide scene layout and sparse, answer-critical local details under tight compute budgets. The paper argues that the usual fixes—feeding larger images or running multi-round zoom search—either drop fine detail or fragment context and waste time. WeaveEarth instead treats the problem as structured evidence construction: under a cheap global thumbnail, it selects a compact, low-redundancy, spatially complementary Minimal Support Evidence Set, then weaves those patches with explicit spatial metadata and a topology-preserving board so a frozen model can reason jointly over global layout and local evidence in one pass. Across several UHR benchmarks and multiple frozen backbones, this training-free pipeline beats strong baselines and prior UHR methods, and ablations plus budget curves show the gains come from better organization, not from exposing more visual content. A sympathetic reader cares because the same budget constraint appears in any large-image multimodal setting where the answer lives in a few regions but depends on their place in the whole scene.

What carries the argument

Minimal Support Evidence Set (MSES) under Global Context Constraint, plus Structured Evidence Reasoning via Structured Evidence Metadata (SEM) and a Topology-Preserving Evidence Board (TPEB). Together they select a small complementary patch set and re-present it so the model retains spatial grounding and relative layout for joint global-local reasoning.

What would settle it

On the same UHR benchmarks, fix the visual budget and either (a) replace the greedy MSES selection with random or purely question-only top-k patches of equal size, or (b) strip SEM and the topology-preserving board while keeping the same patches: if accuracy does not fall and the rise-then-fall budget curve disappears, the claim that organization—not access—drives the gains fails.

Watch

Extended reading notes

Core claim

The paper claims that for ultra-high-resolution remote sensing, effective understanding is a problem of constructing and organizing the right evidence under global context constraints, not of expanding visual access. Its training-free framework, WeaveEarth, first builds a compact Minimal Support Evidence Set that is relevant, low-redundancy, and spatially complementary, then feeds a frozen vision-language model a unified interface of global thumbnail, structured evidence metadata, and a topology-preserving evidence board, yielding consistent accuracy gains over passive whole-image adaptation and active multi-round search.

Load-bearing premise

A training-free greedy score of question similarity plus global-thumbnail consistency, coverage, and redundancy, with a fixed small evidence budget and a hand-arranged board layout, is enough for frozen vision-language models to treat the selected patches and metadata as truly minimal yet sufficient support for spatial answers.

Editorial extensions

If this is right

  • UHR remote-sensing VQA can be improved without fine-tuning backbone VLMs by redesigning only the inference input interface.
  • Passive resolution scaling and multi-round zoom search are not the only viable routes; a single-pass structured evidence interface can match or beat them at lower latency.
  • Spatial-relation and complex-reasoning subtasks benefit most when local patches retain explicit coordinates, roles, neighbors, and relative layout.
  • Evidence quantity has a sweet spot: too few patches miss clues, too many dilute them, so minimal-yet-sufficient sets are preferable to always adding more crops.
  • The same construction-plus-organization pattern can be dropped onto different frozen open-source VLMs with stable relative gains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If encoder similarity under a global thumbnail is the bottleneck, swapping in a stronger or task-adapted cross-modal encoder could raise the ceiling without changing the rest of the pipeline.
  • The counting limitation the authors note suggests multi-scale or instance-aware evidence units as a natural next module when targets are dense and tiny.
  • The same ‘organize better, not access more’ principle may transfer to other large-image domains (pathology slides, satellite video frames) where answers depend on sparse regions and long-range layout.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. WeaveEarth is a training-free, plug-and-play framework for ultra-high-resolution remote sensing VQA that reformulates the task as structured evidence construction and reasoning under global context constraints. Stage 1 (Global-Aware Evidence Construction) scores overlapping patches with question and global-thumbnail similarity (Eq. 1), expands neighborhoods, and greedily builds a compact Minimal Support Evidence Set via Rel+αCov−βRed (Eq. 2). Stage 2 (Structured Evidence Reasoning) attaches Structured Evidence Metadata and arranges patches into a Topology-Preserving Evidence Board, which, with the global thumbnail, is fed to a frozen VLM. On LRS-VQA, MME-RealWorld, and XLRS-Bench, and across Qwen3-VL-8B, LLaVA-v1.6-7B, and IXC-2.5-7B, the method reports consistent gains over passive whole-image and active multi-round UHR baselines, with ablations (Table 3), efficiency comparisons (Fig. 3), and budget curves peaking at |S|=6 (Fig. 4) offered as evidence that gains come from better organization rather than expanded visual access.

Significance. If the organization-over-access thesis holds, the paper offers a practical and conceptually clean alternative to both costly whole-image adaptation and high-latency multi-round search for UHR RS understanding. Strengths include: (i) training-free transfer across three frozen backbones (Table 4); (ii) multi-benchmark evaluation with sample-size-weighted averages; (iii) component ablations and an accuracy–latency comparison against representative passive and active methods; (iv) a budget-sensitivity curve that rises then falls, consistent with a “minimal yet sufficient” evidence regime; and (v) public code. These make the work useful for deployment-oriented RS-VLM pipelines even if some mechanistic claims need tighter controls.

major comments (3)
  1. [§4.3–4.5, Tables 3–4, Figs. 3–4] Central claim isolation (Tables 1–4, Figs. 3–4; §4.3–4.5): The paper’s load-bearing thesis is that gains come from organizing evidence (MSES+SEM+TPEB under GCC), not from expanded visual access. Fig. 4 only varies WeaveEarth’s own |S|; Table 3 removes modules from the full system, so every ablated row still uses encoder-selected multi-patch input and residual structure. There is no matched visual-budget control that feeds the same number of high-resolution, question-relevant crops plus the global thumbnail without SEM and without topology-preserving layout (plain top-k multi-crop). Without that condition, it remains unclear whether the frozen VLM uses metadata/topology or simply benefits from several relevant high-res patches. Please add this control on at least one backbone and two benchmarks; if the unstructured multi-crop closes most of the gap, the slogan and interpretation of Figs.
  2. [§3.2, Eqs. (1)–(2)] Eq. (2) and greedy MSES construction (§3.2): Rel(S,q), Cov(S), and Red(S) are named but not defined operationally (feature aggregation for Rel; coverage metric for Cov; overlap/redundancy measure for Red). The claim that a training-free greedy maximizer yields a “minimal yet sufficient” support set therefore cannot be audited or reproduced from the main text alone. Please specify exact formulas, the encoder used for sim(·,·), values of λ/α/β, overlap/grid settings, and either a short justification that greedy is adequate or a small comparison against a stronger combinatorial baseline on a subset.
  3. [§4.1, Fig. 4] Hyperparameter and selection sensitivity (§4.1, Fig. 4): Free parameters include λ, α, β, evidence budget (default 6), patch grid/overlap, and TPEB layout. Only |S| is swept. Given that the weakest assumption is that encoder similarity under a global thumbnail plus hand-chosen budget yields answer-critical evidence, at least a limited sensitivity study for λ and (α,β)—or a clear statement that defaults transfer without retuning across backbones/benchmarks—is needed to support the training-free, plug-and-play claim.
minor comments (5)
  1. [Table 1] Table 1: WeaveEarth is listed as 8B while some compared UHR methods use 3B/7B; a short note on parameter fairness (or reporting the same backbone for all UHR methods where possible) would help readers.
  2. [Fig. 5] Figure 1 / case study (Fig. 5): “ZoomSearth” appears to be a typo for ZoomSearch; please correct consistently.
  3. [§4.6] §4.6 Limitation: Counting remains weak because small dense objects are hard in the thumbnail; the planned multi-scale extension is reasonable—consider quantifying residual counting error rates by object size if space allows.
  4. [§3.3, Fig. 2] SEM format in Fig. 2 is informative; ensure the exact prompt template that injects SEM+TPEB into each backbone is in the appendix or code for full reproducibility.
  5. [§4.2] No error bars or multi-seed variance are reported; for a systems paper this is common, but a brief note on run-to-run stability (or deterministic decoding settings) would strengthen Tables 1–4.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical training-free systems paper whose claims rest on external benchmarks and ablations, not on a derivation that reduces to its own inputs by construction.

full rationale

WeaveEarth is a training-free engineering framework (Global-Aware Evidence Construction via encoder similarities + greedy MSES under Eq. 1–2, then SEM + TPEB) evaluated by accuracy gains on external UHR RS benchmarks (LRS-VQA, MME-RealWorld, XLRS-Bench) against frozen VLM baselines and prior UHR methods (Tables 1–4). Ablations (Table 3) remove modules and report drops; budget sensitivity (Fig. 4) is an empirical curve peaking at the chosen budget of 6 rather than a tautological identity; efficiency comparisons (Fig. 3) and multi-backbone transfer are likewise external measurements. There is no self-definitional equation equating a claimed prediction to a fitted input, no uniqueness theorem imported from overlapping authors that forces the result, no ansatz smuggled via self-citation as the sole load-bearing step, and no renaming of a known identity presented as a first-principles derivation. Design choices (scoring form, budget, layout) are ordinary hyper-parameters of an empirical method; they do not make the reported accuracy numbers true by construction. Per the analyzer rules this is the expected non-finding for a self-contained systems paper.

Assumptions & free parameters 4 free parameters · 4 assumptions · 4 invented entities

The central claim rests on standard VLM/encoder assumptions, a few hand-chosen scoring and budget parameters, and several named methodological constructs (MSES, SEM, TPEB, GCC) that organize inference rather than postulate new physics. No formal proof; empirical sufficiency of the greedy support set is the load-bearing modeling bet.

free parameters (4)
  • λ (global-context weight in s_i)
    Balances question–patch vs thumbnail–patch similarity in Eq. (1); chosen by design, not derived.
  • α, β (coverage and redundancy weights)
    Trade off Cov vs Red in the MSES objective Eq. (2); free coefficients of the greedy selector.
  • evidence budget |S| (default 6)
    Support-set size is a free budget; Fig. 4 shows performance peaks at 6 then declines—method depends on this choice.
  • patch grid / overlap / TPEB layout scale
    Partitioning, anti-fragmentation expansion, and board arrangement are implementation knobs that affect which evidence is available and how topology is preserved.
assumptions (4)
  • domain assumption Frozen general VLMs can perform joint global–local spatial reasoning when given a thumbnail, a small set of patches, and explicit spatial metadata/topology layout.
    Core premise of Stage 2 Structured Evidence Reasoning (Section 3.3); without it, SEM/TPEB packaging would not help.
  • domain assumption Cross-modal encoder similarity to question and global thumbnail is a useful proxy for answer-critical local regions in UHR RS images.
    Underpins Global Context Constraint and candidate scoring in Eq. (1) (Section 3.2).
  • domain assumption A compact, low-redundancy, spatially complementary support set is preferable to more patches under fixed VLM budgets.
    Stated design thesis and supported by budget rise-then-fall curves (Fig. 4); still an assumption about model attention/capacity.
  • ad hoc to paper Greedy selection adequately approximates the combinatorial MSES objective in Eq. (2).
    Paper adopts training-free greedy ranking without proving approximation quality (Section 3.2).
invented entities (4)
  • Minimal Support Evidence Set (MSES)
    purpose: Name the compact selected patch set that is claimed to be necessary and sufficient for answering under budget.
    Methodological construct defined via Eq. (2); independent_evidence false outside this pipeline’s empirical results.
  • Structured Evidence Metadata (SEM)
    purpose: Attach box, grid, neighbors, role, and scale to each patch so the VLM has explicit spatial grounding.
    Prompt/interface object introduced in Eq. (3); utility shown only via ablations in this paper.
  • Topology-Preserving Evidence Board (TPEB)
    purpose: Visually arrange selected patches to preserve coarse relative topology for joint perception.
    New packaging component; largest ablation drop when removed (Table 3), but no external independent validation.
  • Global Context Constraint (GCC)
    purpose: Force candidate patches to be consistent with the global thumbnail, not only the question text.
    Named constraint via λ sim(p_i,g); modest ablation effect, still a paper-specific framing.

how reviews work

0 comments
Cite this review

Pith. "Pith review of WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding." pith.science (2026). https://pith.science/paper/WFM45NYL

@misc{pith2026260710120,
  author       = {Pith},
  title        = {Pith review of: WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WFM45NYL}},
  note         = {Machine review of arXiv:2607.10120}
}
read the original abstract

Ultra-High-Resolution (UHR) remote sensing image understanding requires Vision-Language Models (VLMs) to capture both the global scene layout and sparse yet task-critical local details under limited computational budgets. Existing methods mainly follow two paradigms. One is passive perception, which relies on resolution expansion or token compression and may therefore discard fine-grained details. The other is active perception, which depends on multi-round zooming and search, but suffers from high latency, contextual fragmentation, and error accumulation. We argue that a more effective path toward UHR understanding lies not in accessing more, but in organizing better. To this end, we propose WeaveEarth, a training-free framework that reformulates UHR understanding as a problem of structured evidence construction and reasoning under global context constraints. Specifically, WeaveEarth first employs Global-Aware Evidence Construction to select a compact, low-redundancy, and spatially complementary Minimal Support Evidence Set. It then introduces Structured Evidence Reasoning, which weaves local evidence, spatial metadata, and relative topology into a unified reasoning interface, thereby enhancing the VLM's ability to perform global-local joint reasoning. Extensive experiments show that WeaveEarth consistently outperforms strong baselines and existing UHR methods across multiple UHR remote sensing benchmarks and multiple frozen VLM backbones. Code is available at https://github.com/XianZhi-Ma/WeaveEarth.

Figures

Figures reproduced from arXiv: 2607.10120 by the authors.

Figure 1
Figure 1. Paradigm comparison for UHR remote sensing understanding. Answering questions over UHR images requires [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of WeaveEarth. Stage 1 (Global-Aware Evidence Construction) uses a global thumbnail and overlapping local crops to construct a compact, globally constrained Minimal Support Evidence Set. Stage 2 (Structured Evidence Reasoning) enriches selected patches with Structured Evidence Metadata and organizes them into a Topology-Preserving Evidence Board to preserve spatial grounding and relative topology. Together … view at source ↗
Figure 3
Figure 3. Accuracy-efficiency comparison of WeaveEarth [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Performance and inference time under different [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 21 linked inside Pith

  1. [1]

    2025.Claude 3.7 Sonnet and Claude Code

    Anthropic. 2025.Claude 3.7 Sonnet and Claude Code. Retrieved February 24, 2025 from https://www.anthropic.com/news/claude-3-7-sonnet

  2. [2]

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  3. [3]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...

  4. [4]

    Yuxiang Cai, Yongheng Shang, and Jianwei Yin. 2024. MultiDAN: Unsupervised, Multistage, Multisource and Multitarget Domain Adaptation for Semantic Seg- mentation of Remote Sensing Images. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 1168–1177

  5. [5]

    Antoine Carreaud, Elias Naha, Arthur Chansel, Nina Lahellec, Jan Skaloud, and Adrien Gressin. 2026. Context-Aware Semantic Segmentation via Stage-Wise Attention. (2026). arXiv preprint arXiv:2601.11310

  6. [6]

    Hengzhi Chen, Liqian Feng, Wenhua Wu, Xiaogang Zhu, Shawn Leo, and Kun Hu. 2025. F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing Segmentation. (2025). arXiv preprint arXiv:2506.07847

  7. [7]

    Yunkai Dang, Meiyi Zhu, Donghao Wang, Yizhuo Zhang, Jiacheng Yang, Qi Fan, Yuekun Yang, Wenbin Li, Feng Miao, and Yang Gao. 2025. A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs. (2025). arXiv preprint arXiv:2512.17319

  8. [8]

    Lamei Di, Bin Zhang, Yiming Wang, and Wenxia Zhang. 2025. Frequency Meets Semantics: Text-Visual Fusion with Directional Spectral Enhancement for Salient Object Detection in Optical Remote Sensing Images. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 1987–1996

Show all 71 references
  1. [9]

    Renxiang Guan, Junhong Li, Siwei Wang, Wenxuan Tu, Miaomiao Li, En Zhu, Xinwang Liu, and Ping Chen. 2025. Multi-view Graph Clustering with Dual Relation Optimization for Remote Sensing Data. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 7346–7355

  2. [10]

    Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, Lei Zhao, Zhuoyi Yang, Xiaotao Gu, Xiaohan Zhang, Guanyu Feng, Da Yin, Zihan Wang, Ji Qi, Xixuan Song, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Yuxiao Dong, a...

  3. [11]

    Yuan Hu, Jianlong Yuan, Congcong Wen, Xiaonan Lu, and Xiang Li. 2023. RSGPT: A Remote Sensing Vision Language Model and Benchmark. (2023). arXiv preprint arXiv:2307.15266

  4. [12]

    Ling Huang, Wenqian Dong, Song Xiao, Jiahui Qu, Yuanbo Yang, and Yunsong Li

  5. [13]

    InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM)

    Language-Guided Visual Prompt Compensation for Multi-Modal Remote Sensing Image Classification with Modality Absence. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 5161–5170

  6. [14]

    Zhong Ji, Changxu Meng, Yan Zhang, Haoran Wang, Yanwei Pang, and Jungong Han. 2024. Eliminate Before Align: A Remote Sensing Image-Text Retrieval Framework with Keyword Explicit Reasoning. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 1662–1671

  7. [15]

    Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Jing Li, Jiayang Li, Yue Zhou, Hongjie He, and Jonathan Li. 2025. GRASP: Geospatial pixel Reasoning viA Structured Policy learning. (2025). arXiv preprint arXiv:2508.17102

  8. [16]

    Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan. 2024. GeoChat:Grounded Large Vision- Language Model for Remote Sensing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 27831–27840

  9. [17]

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024. LLaVA- OneVision: Easy Visual Task Transfer. (2024). arXiv preprint arXiv:2408.03326

  10. [18]

    Jianhui Li, Chao Wu, Yingchao Piao, Yuchu Qin, Xiaoping Du, Lili Zhang, and Huadong Guo. 2023. How can we support the UN Sustainable Development Goals when open data is stagnant?Science Bulletin68, 12 (2023), 1216–1218

  11. [19]

    Ke Li, Di Wang, Ting Wang, Fuyu Dong, Yiming Zhang, Luyao Zhang, Xiangyu Wang, Shaofeng Li, and Quan Wang. 2026. RSVG-ZeroOV: Exploring a Training- Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images. InProceedings of the Fortieth AAAI Confer...

  12. [20]

    Ke Li, Di Wang, Haojie Xu, Haodi Zhong, and Cong Wang. 2024. Language- Guided Progressive Attention for Visual Grounding in Remote Sensing Images. IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–13

  13. [21]

    Ke Li, Ting Wang, Di Wang, Yongshan Zhu, Yiming Zhang, Tao Lei, and Quan Wang. 2026. ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery. (2026). arXiv preprint arXiv:2604.01893

  14. [22]

    Qingyun Li, Shuran Ma, Junwei Luo, Yi Yu, Yue Zhou, Fengxiang Wang, Xudong Lu, Xiaoxing Wang, Xin He, Yushi Chen, and Xue Yang. 2026. Co-Training Vision Language Models for Remote Sensing Multi-task Learning. (2026). arXiv preprint arXiv:2511.21272

  15. [23]

    Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou. 2024. RemoteCLIP: A Vision Language Foundation Model for Remote Sensing.IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1–16

  16. [24]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 26286–26296

  17. [25]

    2024.LLaV A-NeXT: Improved reasoning, OCR, and world knowledge

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024.LLaV A-NeXT: Improved reasoning, OCR, and world knowledge. Retrieved January 30, 2024 from https://llava-vl.github.io/blog/2024-01-30-llava- next/

  18. [26]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual In- struction Tuning. InProceedings of the Advances in Neural Information Processing Systems (NeurIPS). 34892–34916

  19. [27]

    Jiaqi Liu, Lang Sun, Ronghao Fu, and Bo Yang. 2026. Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models. (2026). arXiv preprint arXiv:2509.22221

  20. [28]

    Ruixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li, Wanchen Li, Lanxuan Xue, Hui Qiao, Weizhan Zhang, Deyu Meng, and Xiangyong Cao. 2025. ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks. (2025). arXiv preprint arXiv:2511.12267

  21. [29]

    Wang Liu, Puhong Duan, Xudong Kang, and Shutao Li. 2025. Squeezing Context into Patches: Towards Memory-Efficient Ultra-High Resolution Semantic Seg- mentation. InProceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI). 1603–1611

  22. [30]

    Weiqi Liu, Yongshan Zhang, Xinxin Wang, and Lefei Zhang. 2025. Deep Multi- Level Contrastive Clustering for Multi-Modal Remote Sensing Images. InPro- ceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 1239–1247

  23. [31]

    Xu Liu and Zhouhui Lian. 2024. RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts. (2024). arXiv preprint arXiv:2412.05679

  24. [32]

    Ye Liu, Shitao Song, Miaohui Wang, Hao Gao, and Jun Liu. 2025. DE-Unet: Dual- Encoder U-Net for Ultra-High Resolution Remote Sensing Image Segmentation. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 18 (2025), 12290–12302

  25. [33]

    Junwei Luo, Yingying Zhang, Xue Yang, Kang Wu, Qi Zhu, Lei Liang, Jingdong Chen, and Yansheng Li. 2025. When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning. (2025). arXiv preprint arXiv:2503.07588

  26. [34]

    Xianzhi Ma, Jianhui Li, Changhua Pei, and Hao Liu. 2025. GeoMag: A Vision- Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing. In Proceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 5441–5450. ACM MM, 2026, Rio de Janeiro, Brazil ...

  27. [35]

    Li Mi, Manon Béchaz, Zeming Chen, Antoine Bosselut, and Devis Tuia. 2025. GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 6122–6131

  28. [36]

    2024.GPT-4o mini: advancing cost-efficient intelligence

    OpenAI. 2024.GPT-4o mini: advancing cost-efficient intelligence. Retrieved July 19, 2024 from https://openai.com/index/gpt-4o-mini-advancing-cost-efficient- intelligence/

  29. [37]

    2024.Hello GPT -4o

    OpenAI. 2024.Hello GPT -4o. Retrieved May 13, 2024 from https://openai.com/ index/hello-gpt-4o/

  30. [38]

    Ruizhe Ou, Yuan Hu, Fan Zhang, Jiaxin Chen, and Yu Liu. 2025. GeoPix: Multi- Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing. (2025). arXiv preprint arXiv:2501.06828

  31. [39]

    Yuwen Pan, Rui Sun, Yuan Wang, Tianzhu Zhang, and Yongdong Zhang. 2024. Rethinking the Implicit Optimization Paradigm with Dual Alignments for Re- ferring Remote Sensing Image Segmentation. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 2031–2040

  32. [40]

    Chao Pang, Xingxing Weng, Jiang Wu, Jiayu Li, Yi Liu, Jiaxing Sun, Weijia Li, Shuai Wang, Litong Feng, Gui-Song Xia, and Conghui He. 2025. VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis. In Proceedings of the Thirty-Ninth AAAI Conference on A...

  33. [41]

    Khan, and Salman Khan

    Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad S. Khan, and Salman Khan. 2025. GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing. (2025). arXiv preprint arXiv:2501.13925

  34. [42]

    Run Shao, Ziyu Li, Zhaoyang Zhang, Linrui Xu, Xinran He, Hongyuan Yuan, Bolei He, Yongxing Dai, Yiming Yan, Yijun Chen, Wang Guo, and Haifeng Li

  35. [43]

    Asking like Socrates: Socrates helps VLMs understand remote sensing images. (2025). arXiv preprint arXiv:2511.22396

  36. [44]

    Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Ming- wei Zhu, and Jianwei Yin. 2025. ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration. (2025). arXiv preprint arXiv:2411.16044

  37. [45]

    João Daniel Silva, João Magalhães, Devis Tuia, and Bruno Martins. 2024. Large Language Models for Captioning and Retrieving Remote Sensing Images. (2024). arXiv preprint arXiv:2402.06475

  38. [46]

    Di Wang, Shunyu Liu, Wentao Jiang, Fengxiang Wang, Yi Liu, Xiaolei Qin, Zhim- ing Luo, Chaoyang Zhou, Haonan Guo, Jing Zhang, Bo Du, Dacheng Tao, and Liangpei Zhang. 2025. GeoZero: Incentivizing Reasoning from Scratch on Geospa- tial Scenes. (2025). arXiv preprint arXiv:2511.22645

  39. [47]

    Fengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang, Haotian Wang, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang, Hongzhen Wang, Wenjing Yang, Bo Du, and Jing Zhang. 2025. GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution. (2025). ...

  40. [48]

    Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yifan Zhang, Long Lan, Xue Yang, Hongda Sun, Yulin Wang, Di Wang, Jun Song, Jing Zhang, and Bo Du. 2026. GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imager...

  41. [49]

    Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yuhao Zhou, Di Wang, Yifan Zhang, Haoyu Wang, Haiyan Zhao, Hongda Sun, Long Lan, Jun Song, Yulin Wang, Jing Zhang, Wenlong Zhang, and Bo Du. 2026. Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in ...

  42. [50]

    Fengxiang Wang, Hongzhen Wang, Zonghao Guo, Di Wang, Yulin Wang, Ming- shuo Chen, Qiang Ma, Long Lan, Wenjing Yang, Jing Zhang, Zhiyuan Liu, and Maosong Sun. 2025. XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?. ...

  43. [51]

    Kang Wu, Yingying Zhang, Lixiang Ru, Bo Dang, Jiangwei Lao, Lei Yu, Junwei Luo, Zifan Zhu, Yue Sun, Jiahao Zhang, Qi Zhu, Jian Wang, Ming Yang, Jingdong Chen, Yongjun Zhang, and Yansheng Li. 2025. A semantic-enhanced multi- modal remote sensing foundation model for Earth obser...

  44. [52]

    Kelu Yao, Nuo Xu, Rong Yang, Yingying Xu, Zhuoyan Gao, Titinunt Kitrungrot- sakul, Yi Ren, Pu Zhang, Jin Wang, Ning Wei, and Chao Li. 2025. Falcon: A Remote Sensing Vision-Language Foundation Model (Technical Report). (2025). arXiv preprint arXiv:2503.11070

  45. [53]

    Liang Yao, Fan Liu, Delong Chen, Chuanyi Zhang, Yijun Wang, Ziyun Chen, Wei Xu, Shimin Di, and Yuhui Zheng. 2025. RemoteSAM: Towards Segment Anything for Earth Observation. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 3027–3036

  46. [54]

    Liang Yao, Fan Liu, Hongbo Lu, Chuanyi Zhang, Rui Min, Shengxiang Xu, Shimin Di, and Pai Peng. 2026. RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow. InProceedings of the Fortieth AAAI Conference on Artificial Intelligence. 11883–11891

  47. [55]

    Liang Yao, Fan Liu, Shengxiang Xu, Chuanyi Zhang, Rui Min, Shimin Di, and Yuhui Zheng. 2026. RemoteZero: Geospatial Reasoning with Zero Human Anno- tations. (2026). arXiv preprint arXiv:2605.04451

  48. [56]

    Liang Yao, Shengxiang Xu, Fan Liu, Chuanyi Zhang, Bishun Yao, Rui Min, Yongjun Li, Chaoqian Ouyang, Shimin Di, and Min-Ling Zhang. 2026. RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs. (2026). arXiv preprint arXiv:2604.07765

  49. [57]

    Bo Yuan, Danpei Zhao, Zhuoran Liu, Wentao Li, and Tian Li. 2024. Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images. InProceedings of the 32nd ACM International Conference on Multimedia (ACM MM). 2117–2126

  50. [58]

    Yang Zhan, Zhitong Xiong, and Yuan Yuan. 2025. SkyEyeGPT: Unifying remote sensing vision-language tasks via instruction tuning with large language model. ISPRS Journal of Photogrammetry and Remote Sensing221 (2025), 64–77

  51. [59]

    Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, J...

  52. [60]

    Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin, Zonghao Guo, Fengxi- ang Wang, Xue Yang, Kaiwen Wei, and Lei Wang. 2025. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding. (2025). arXiv preprint arXiv:2512.02715

  53. [61]

    Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, Jun Li, and Xuerui Mao. 2025. EarthMarker: A Visual Prompting Multimodal Large Language Model for Remote Sensing.IEEE Transactions on Geoscience and Remote Sensing63 (2025), 1–19

  54. [62]

    Wei Zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, and Xuerui Mao. 2024. Earth- GPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain.IEEE Transactions on Geoscience and Remote Sensing62 (2024), 1–20

  55. [63]

    Xu Zhang, Junyao Ge, Yang Zheng, Kaitai Guo, and Jimin Liang. 2025. Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Remote Sensing. (2025). arXiv preprint arXiv:2512.19302

  56. [64]

    Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin, and Tieniu Tan. 2025. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficu...

  57. [65]

    Zilun Zhang, Zian Guan, Tiancheng Zhao, Haozhan Shen, Tianyu Li, Yuxiang Cai, Zhonggen Su, Zhaojun Liu, Jianwei Yin, and Xiang Li. 2025. Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning. (2025). arXiv preprint arXiv:2509.21976

  58. [66]

    Zilun Zhang, Haozhan Shen, Tiancheng Zhao, Zian Guan, Bin Chen, Yuhao Wang, Xu Jia, Yuxiang Cai, Yongheng Shang, and Jianwei Yin. 2025. Enhancing Ultrahigh Resolution Remote Sensing Imagery Analysis With ImageRAG: A new framework.IEEE Geoscience and Remote Sensing Magazine13, ...

  59. [67]

    Yang Zhao, Shusheng Li, and Xueshang Feng. 2025. Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit. InProceedings of the 33rd ACM International Conference on Multimedia (ACM MM). 11862–11870

  60. [68]

    Siru Zhong, Xixuan Hao, Yibo Yan, Ying Zhang, Yangqiu Song, and Yuxuan Liang. 2024. UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross- Domain Adaptation. InProceedings of the 32nd ACM International Conference on Multimedia. 6307–6315

  61. [69]

    Yue Zhou, Jue Chen, Zilun Zhang, Penghui Huang, Ran Ding, Zhentao Zou, PengFei Gao, Yuchen Wei, Ke Li, Xue Yang, Xue Jiang, Hongxin Yang, and Jonathan Li. 2026. DVGBench: Implicit-to-explicit visual grounding benchmark in UAV imagery with large vision–language models.ISPRS Jou...

  62. [70]

    Yunqi Zhou, Chengjie Jiang, Chun Yuan, and Jing Li. 2025. Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search. (2025). arXiv preprint arXiv:2511.20460

  63. [71]

    Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.