Pith. sign in

REVIEW 8 cited by

SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17179 v3 pith:W4ONWYDW submitted 2023-11-28 cs.CV cs.AIcs.CYcs.LG

SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

classification cs.CV cs.AIcs.CYcs.LG
keywords locationsatclipgeographictasksglobalimagerysatellitecharacteristics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be challenging, often requiring expensive data fusion or distillation from massive global imagery datasets. To address this challenge, we introduce Satellite Contrastive Location-Image Pretraining (SatCLIP). This global, general-purpose geographic location encoder learns an implicit representation of locations by matching CNN and ViT inferred visual patterns of openly available satellite imagery with their geographic coordinates. The resulting SatCLIP location encoder efficiently summarizes the characteristics of any given location for convenient use in downstream tasks. In our experiments, we use SatCLIP embeddings to improve prediction performance on nine diverse location-dependent tasks including temperature prediction, animal recognition, and population density estimation. Across tasks, SatCLIP consistently outperforms alternative location encoders and improves geographic generalization by encoding visual similarities of spatially distant environments. These results demonstrate the potential of vision-location models to learn meaningful representations of our planet from the vast, varied, and largely untapped modalities of geospatial data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MotifGen: Spatiotemporal interpolation of misaligned satellite images via multi-source generative modeling, in an application to tropical cyclones

    cs.CV 2026-06 unverdicted novelty 7.0

    MotifGen is the first multi-source generative model for spatiotemporal interpolation of misaligned microwave cyclone images from heterogeneous instruments at irregular intervals, achieving lower CRPS via self-supervis...

  2. Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery

    cs.MM 2026-04 unverdicted novelty 7.0

    Geo2Sound generates geographically realistic soundscapes from satellite imagery via geospatial attribute modeling, semantic hypothesis expansion, and geo-acoustic alignment, achieving SOTA FAD of 1.765 on a new 20k-pa...

  3. GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations

    cs.CV 2025-03 unverdicted novelty 6.0

    GAIR introduces a geo-aligned implicit representation module inside a multi-encoder contrastive SSL framework that produces location-aware embeddings and outperforms prior geo-foundation models on 22 geospatial datase...

  4. General Geospatial Inference with a Population Dynamics Foundation Model

    cs.LG 2024-11 unverdicted novelty 6.0

    A GNN-based foundation model on aggregated US geospatial data produces embeddings achieving SOTA on all 27 interpolation tasks and 25/27 extrapolation/super-resolution tasks across health, socioeconomic and environmen...

  5. MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications

    cs.CV 2026-04 unverdicted novelty 5.0

    MOMO merges sensor-specific models from three Mars orbital instruments at matched validation loss stages to form a foundation model that outperforms ImageNet, Earth observation, sensor-specific, and supervised baselin...

  6. Feature Extraction in the Remote Sensing Data Value Chain: A Systematic Review of Methods and Applications

    cs.CV 2025-10 unverdicted novelty 5.0

    A systematic review that introduces a framework for feature extraction in remote sensing, traces its evolution in the data value chain, and synthesizes trends toward unified representations and foundation models.

  7. OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics

    cs.CV 2026-05 unverdicted novelty 4.0

    OmniCD proposes a multimodal semantic-guided framework for remote sensing change detection supporting binary to zero-shot tasks, plus the RSITCD dataset, with claimed SOTA performance.

  8. CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images

    cs.CV 2025-06 unverdicted novelty 4.0

    A lightweight multi-modal CLIP pipeline predicts exact-match geographical tags on a Kaggle subset of the Geograph crowdsourced image archive by fusing image, location, and title embeddings.