Pith. sign in

REVIEW 21 cited by

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.13925 v1 pith:JHSWTRM3 submitted 2025-01-23 cs.CV

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

classification cs.CV
keywords datageopixelgroundinglmmsconversationgroundedcapabilitycomprehension
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS). The distinct overhead viewpoint, scale variation, and presence of small objects in high-resolution RS imagery present a unique challenge in region-level comprehension. Moreover, the development of the grounding conversation capability of LMMs within RS is hindered by the lack of granular, RS domain-specific grounded data. Addressing these limitations, we propose GeoPixel - the first end-to-end high resolution RS-LMM that supports pixel-level grounding. This capability allows fine-grained visual perception by generating interleaved masks in conversation. GeoPixel supports up to 4K HD resolution in any aspect ratio, ideal for high-precision RS image analysis. To support the grounded conversation generation (GCG) in RS imagery, we curate a visually grounded dataset GeoPixelD through a semi-automated pipeline that utilizes set-of-marks prompting and spatial priors tailored for RS data to methodically control the data generation process. GeoPixel demonstrates superior performance in pixel-level comprehension, surpassing existing LMMs in both single-target and multi-target segmentation tasks. Our methodological ablation studies validate the effectiveness of each component in the overall architecture. Our code and data will be publicly released.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Evaluating Remote Sensing Image Captions Beyond Metric Biases

    cs.CV 2026-04 unverdicted novelty 7.0

    Unfine-tuned MLLMs outperform fine-tuned models on remote sensing image captioning when captions are scored by their ability to reconstruct the source image, and a training-free self-correction method achieves SOTA pe...

  2. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

    cs.CV 2026-04 unverdicted novelty 7.0

    The work introduces the UAV Reasoning Segmentation task, the DRSeg benchmark dataset, and PixDLM as a baseline dual-path multimodal language model for reasoning-based segmentation in aerial imagery.

  3. GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing

    cs.CV 2026-04 unverdicted novelty 7.0

    GeoMeld provides a large-scale aligned multimodal remote sensing dataset with verified semantic captions and a joint pretraining method that improves downstream transfer and cross-sensor robustness in foundation models.

  4. RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

    cs.CV 2026-04 unverdicted novelty 7.0

    RemoteAgent uses RL fine-tuning on VagueEO to align MLLMs for vague EO intent recognition, handling simple tasks internally and routing dense predictions to tools via Model Context Protocol.

  5. MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

    cs.CV 2025-12 conditional novelty 7.0

    MMLandmarks supplies 197k aerial and 329k ground images plus text and GPS for 18,557 landmarks to benchmark multimodal geo-spatial understanding.

  6. OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

    cs.CV 2026-07 conditional novelty 6.0

    A new 172-category, three-query-form Earth observation benchmark shows the best open-vocabulary model reaches only 38.75% mean IoU, with MLLM-based systems leading and EO-specific models trailing.

  7. OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

    cs.CV 2026-07 conditional novelty 6.0

    OVEarth-Bench, a new open-vocabulary Earth observation benchmark with broad category coverage and diverse queries, shows MLLM-based methods outperform EO-specific ones.

  8. GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0

    ChronoBench decomposes long-term remote sensing understanding into four cognitive levels, and the GeoChrono model, using per-location temporal trajectories, achieves 78.34% accuracy—over 20 points above prior MLLMs—bu...

  9. WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A training-free evidence-selection and topology-preserving packaging pipeline improves frozen VLMs on ultra-high-resolution remote sensing VQA without multi-round search.

  10. Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning

    cs.CV 2026-07 conditional novelty 6.0

    Aligning images to multi-view caption cores while suppressing orthogonal residual text and disagreement-aware temperature improves robust zero-shot recognition and LVLM transfer.

  11. InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

    cs.AI 2026-07 unverdicted novelty 6.0

    InduceKV is a retrieval-based continual adaptation method that uses bilevel selection to build a compact set of inducing KV memories for fixed-footprint updates to multimodal LLMs.

  12. GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision

    cs.CV 2026-07 unverdicted novelty 6.0

    GeoSearcher introduces anchor-centric reasoning supervised fine-tuning and process-faithful group relative policy optimization to improve MLLM-based remote sensing visual grounding.

  13. An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

    cs.CV 2026-05 conditional novelty 6.0

    Introduces MTRS task, MTRefSeg-21K benchmark of 21K image-text-mask triplets, and MTRefSeg-R1 LVLM baseline that outperforms standard models via two-stage change-aware training.

  14. B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

    cs.CV 2026-05 unverdicted novelty 6.0

    B-GRTO extends GRPO by reusing rollouts to optimize auxiliary segmentation decoder objectives, yielding substantial gains over plain GRPO on referring segmentation tasks.

  15. WOW-Seg: A Word-free Open World Segmentation Model

    cs.CV 2026-05 conditional novelty 6.0

    WOW-Seg proposes a word-free open-world segmentation model using Mask2Token and Cascade Attention Mask modules, reporting 89.7 semantic similarity and 82.4 semantic IoU on LVIS with one-eighth the parameters of prior ...

  16. PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

    cs.CV 2026-04 conditional novelty 6.0

    The paper defines UAV reasoning segmentation, releases DRSeg—10,000 aerial images with reasoning QA and masks—and shows its PixDLM baseline beats prior reasoning-segmentation models under fine-tuning.

  17. ProtoFlow: Mitigating Forgetting in Class-Incremental Remote Sensing Segmentation via Low-Curvature Prototype Flow

    cs.CV 2026-04 unverdicted novelty 6.0

    ProtoFlow stabilizes class prototypes via low-curvature temporal flow to mitigate forgetting in class- and domain-incremental remote sensing segmentation.

  18. ProtoFlow: Mitigating Forgetting in Class-Incremental Remote Sensing Segmentation via Low-Curvature Prototype Flow

    cs.CV 2026-04 unverdicted novelty 5.5

    ProtoFlow models class prototypes as low-curvature trajectories with an explicit temporal vector field, reducing forgetting in class- and domain-incremental remote sensing segmentation by about 1.5–2.0 mIoU points.

  19. More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

    cs.CV 2026-07 conditional novelty 5.0

    An unmodified general-purpose VLM trained with multi-task RL and a SAM3 tool reaches top results on most remote sensing zero-shot benchmarks, with gains the paper attributes to training-data diversity rather than arch...

  20. ProtoFlow: Mitigating Forgetting in Class-Incremental Remote Sensing Segmentation via Low-Curvature Prototype Flow

    cs.CV 2026-04 unverdicted novelty 5.0

    ProtoFlow stabilizes class prototypes as low-curvature trajectories in a temporal vector field to mitigate forgetting and improve mIoU in class- and domain-incremental remote sensing segmentation.

  21. B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation

    cs.CV 2026-05 unverdicted novelty 4.0

    B-GRTO pre-trains a segmentation tool via bootstrapped group relative optimization on GRPO rollouts, yielding substantial gains over plain GRPO on referring segmentation benchmarks.