Pith. sign in

REVIEW 1 cited by

SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12856 v1 pith:NWFW5VF3 submitted 2023-12-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords remotesensingdatasetimagesclassificationvlmsdiverseimage-text
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Remote sensing imagery, despite its broad applications in helping achieve Sustainable Development Goals and tackle climate change, has not yet benefited from the recent advancements of versatile, task-agnostic vision language models (VLMs). A key reason is that the large-scale, semantically diverse image-text dataset required for developing VLMs is still absent for remote sensing images. Unlike natural images, remote sensing images and their associated text descriptions cannot be efficiently collected from the public Internet at scale. In this work, we bridge this gap by using geo-coordinates to automatically connect open, unlabeled remote sensing images with rich semantics covered in OpenStreetMap, and thus construct SkyScript, a comprehensive vision-language dataset for remote sensing images, comprising 2.6 million image-text pairs covering 29K distinct semantic tags. With continual pre-training on this dataset, we obtain a VLM that surpasses baseline models with a 6.2% average accuracy gain in zero-shot scene classification across seven benchmark datasets. It also demonstrates the ability of zero-shot transfer for fine-grained object attribute classification and cross-modal retrieval. We hope this dataset can support the advancement of VLMs for various multi-modal tasks in remote sensing, such as open-vocabulary classification, retrieval, captioning, and text-to-image synthesis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Using maps and metadata as extra context for GPT-4o caption generation yields a richer remote sensing dataset, fMoW-mm, with claimed lower hallucination rates and better few-shot detection than prior datasets.

Pith tools