Pith. sign in

WildSAT: Learning Satellite Image Representations from Wildlife Observations

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on citizen science platforms. WildSAT employs a contrastive learning approach that jointly leverages satellite images, species occurrence maps, and textual habitat descriptions to train or fine-tune models. This approach significantly improves performance on diverse satellite image recognition tasks, outperforming both ImageNet-pretrained models and satellite-specific baselines. Additionally, by aligning visual and textual information, WildSAT enables zero-shot retrieval, allowing users to search geographic locations based on textual descriptions. WildSAT surpasses recent cross-modal learning methods, including approaches that align satellite images with ground imagery or wildlife photos, demonstrating the advantages of our approach. Finally, we analyze the impact of key design choices and highlight the broad applicability of WildSAT to remote sensing and biodiversity monitoring.

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Global and Local Entailment Learning for Natural World Imagery

cs.CV · 2025-06-26 · conditional · novelty 6.0

RCME enforces a transitivity constraint in radial embeddings, producing a hierarchical vision-language model that orders taxonomic labels better and improves hierarchical classification and retrieval.

citing papers explorer

Showing 1 of 1 citing paper.

  • Global and Local Entailment Learning for Natural World Imagery cs.CV · 2025-06-26 · conditional · none · ref 10 · internal anchor

    RCME enforces a transitivity constraint in radial embeddings, producing a hierarchical vision-language model that orders taxonomic labels better and improves hierarchical classification and retrieval.