Pith. sign in

REVIEW 8 cited by

RemoteCLIP: A Vision Language Foundation Model for Remote Sensing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11029 v4 pith:HQGHKH4I submitted 2023-06-19 cs.CV

RemoteCLIP: A Vision Language Foundation Model for Remote Sensing

classification cs.CV
keywords remoteclipfoundationclassificationdatamodelsremotesensingdataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these models primarily learn low-level features and require annotated data for fine-tuning. Moreover, they are inapplicable for retrieval and zero-shot applications due to the lack of language understanding. To address these limitations, we propose RemoteCLIP, the first vision-language foundation model for remote sensing that aims to learn robust visual features with rich semantics and aligned text embeddings for seamless downstream application. To address the scarcity of pre-training data, we leverage data scaling which converts heterogeneous annotations into a unified image-caption data format based on Box-to-Caption (B2C) and Mask-to-Box (M2B) conversion. By further incorporating UAV imagery, we produce a 12 $\times$ larger pretraining dataset than the combination of all available datasets. RemoteCLIP can be applied to a variety of downstream tasks, including zero-shot image classification, linear probing, $\textit{k}$-NN classification, few-shot classification, image-text retrieval, and object counting in remote sensing images. Evaluation on 16 datasets, including a newly introduced RemoteCount benchmark to test the object counting ability, shows that RemoteCLIP consistently outperforms baseline foundation models across different model scales. Impressively, RemoteCLIP beats the state-of-the-art method by 9.14% mean recall on the RSITMD dataset and 8.92% on the RSICD dataset. For zero-shot classification, our RemoteCLIP outperforms the CLIP baseline by up to 6.39% average accuracy on 12 downstream datasets. Project website: https://github.com/ChenDelong1999/RemoteCLIP

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery

    cs.MM 2026-04 unverdicted novelty 7.0

    Geo2Sound generates geographically realistic soundscapes from satellite imagery via geospatial attribute modeling, semantic hypothesis expansion, and geo-acoustic alignment, achieving SOTA FAD of 1.765 on a new 20k-pa...

  2. RoofNet: A Global Multimodal Dataset for Roof Material Identification from Earth Observation

    cs.CE 2025-05 conditional novelty 7.0

    RoofNet is a multimodal dataset pairing high-resolution Earth observation imagery with roof material annotations from diverse global locations to support vision-language models for hazard exposure mapping.

  3. MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation

    cs.CV 2026-05 unverdicted novelty 6.0

    MSD-Score introduces multi-scale distributional scoring on von Mises-Fisher mixtures to evaluate image captions without references and reports state-of-the-art correlation with human judgments.

  4. ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding

    cs.CV 2026-04 unverdicted novelty 6.0

    ChangeQuery is a new multimodal framework for semantic disaster change analysis that combines optical and SAR data with a custom dataset and annotation pipeline to support interactive damage assessment.

  5. Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

    cs.CV 2026-07 accept novelty 5.5

    A training-free subtraction-then-attention cascade matches or beats full fusion recall on LEVIR-CC at 10–15× lower query cost; Mamba is no faster than attention at L=196; TBF cuts parameters 2.3× for a 0.007 BLEU-1 cost.

  6. SemDINO: A DINOv3-Driven Network for Cross-Temporal Semantic Alignment in Change Detection

    cs.CV 2026-06 unverdicted novelty 5.0

    SemDINO proposes a dual-branch encoder with DINOv3 features, multi-scale temporal interaction, and enhancement modules for improved semantic change detection in remote sensing.

  7. CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis

    cs.CV 2026-05 unverdicted novelty 5.0

    CropVLM is a domain-adapted vision-language model that achieves 72.51% zero-shot crop classification accuracy and superior open-set detection performance on novel species without retraining.

  8. Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift

    cs.CV 2026-04 unverdicted novelty 5.0

    Supervised fine-tuning with 0.1% labeled data outperforms all 60 tested prompt variants for CLIPSeg cloud segmentation on satellite imagery under domain shift.