Pith. sign in

REVIEW 5 cited by

VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.20213 v4 pith:U7EKVKUT submitted 2024-03-29 cs.CV

classification cs.CV
keywords remotesensinghonestimagecaptionslanguagemodelquestions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Unlike prevailing remote sensing image-text datasets, in which image captions focus on a few prominent objects and their relationships, VersaD captions provide detailed information about image properties, object attributes, and the overall scene. This comprehensive captioning enables VHM to thoroughly understand remote sensing images and perform diverse remote sensing tasks. Moreover, different from existing remote sensing instruction datasets that only include factual questions, HnstD contains additional deceptive questions stemming from the non-existence of objects. This feature prevents VHM from producing affirmative answers to nonsense queries, thereby ensuring its honesty. In our experiments, VHM significantly outperforms various vision language models on common tasks of scene classification, visual question answering, and visual grounding. Additionally, VHM achieves competent performance on several unexplored tasks, such as building vectorizing, multi-label classification and honest question answering. We will release the code, data and model weights at https://github.com/opendatalab/VHM .

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training on multi-tool visual reasoning trajectories (zoom, grounding, lines) with an attention-focused RL objective improves UHR remote-sensing VQA accuracy over single-tool zoom-in and larger base models.

  2. Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Under one zero-shot protocol, general-purpose MLLMs match or outperform remote-sensing-specific MLLMs on several RS benchmarks, while RS-MLLMs keep advantages in visual grounding, RS-VQA, and ultra-high-resolution und...

  3. Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images

    cs.CV 2025-08 conditional novelty 6.0 of 10

    SegEarth-OV performs annotation-free open-vocabulary segmentation of remote-sensing images by upsampling CLIP features, removing global bias, and distilling optical knowledge into a SAR encoder.

  4. GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A remote sensing vision-language model that uses task-aware resolution adjustment and attention-based cropping to perform pixel-level segmentation alongside image- and region-level tasks.

  5. VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VectorLLM, a multimodal LLM that regresses building contour vertices token by token, reports gains of 5.6 to 13.6 AP over prior polygon extraction methods on WHU, WHU-Mix, and CrowdAI.

Pith tools