REVIEW 3 major objections 5 minor 2 cited by
WildSAT: Learning Satellite Image Representations from Wildlife Observations
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read WildSAT claims that the locations of wildlife observations, paired with Wikipedia habitat text and location embeddings, are a powerful free supervision signal for learning satellite image representations, improving 108 of 115 downstream…
desk verdict A solid, well-engineered method whose full objective demonstrably improves satellite representations; the specific contribution of wildlife supervision is only partially isolated by the ablations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the WildSAT contrastive objective, a sum of three InfoNCE-style alignment losses: Limg aligns two satellite images of the same location taken at different times (with geometric augmentation of one), Ltxt aligns image embeddings with GritLM text embeddings of a randomly sampled section of the species' Wikipedia page, and Lloc aligns image embeddings with the SINR location embedding of the observation coordinates, optionally concatenated with WorldClim2 environmental covariates. The pairing that ties these modalities together is the shared species observation location, and the loss is computed symmetrically over each minibatch. For out-of-domain pretrained models, parameter-efficient fine-tuning is used to preserve the original representations.
What would settle it
A control experiment that fine-tunes the same satellite-pretrained models on WildSAT with species labels randomly permuted across locations (keeping images, text, and location modalities intact) would settle whether the ecological pairing drives the gains; if this corrupted objective matches the real WildSAT results, the improvement is not from wildlife-based supervision.
Extended reading notes
Core claim
The paper's central claim is that species distributions encode the visual character of the landscape, and that a contrastive objective forcing satellite image embeddings to be close to the text, location, and environment embeddings of co-occurring wildlife observations produces representations that transfer better than those learned from ImageNet, self-supervision, or anthropogenic labels alone. This holds across architectures — ResNet, Swin, and ViT — and across initialization regimes, from random weights to strong satellite-specific pretraining such as Prithvi, SatlasNet, and SeCo. The authors further claim that, by aligning images with text, WildSAT enables zero-shot retrieval of satellite tiles from habitat descriptions, and that it avoids the 'forgetting' observed when other cross-modal methods (TaxaBind, GRAFT, RemoteCLIP) fine-tune CLIP.
Load-bearing premise
The paper assumes that a 512 by 512 meter satellite tile centered on a citizen-science wildlife observation is actually a picture of that species' preferred habitat, even though the exact observation date is ignored, the data mostly come from the US and Europe, and for several satellite-pretrained models the reported gains are not isolated from the effect of continued image-only training.
Editorial extensions
If this is right
- Any satellite image encoder, from random initialization to satellite-specific foundation models, can be improved by fine-tuning on WildSAT's objective, with the largest relative gains on habitat-related classes like trees, water, and low plants.
- CLIP-based remote sensing models can be fine-tuned for wildlife-aligned cross-modal retrieval without sacrificing downstream linear probing accuracy, unlike TaxaBind, GRAFT, or RemoteCLIP.
- Zero-shot text-to-image retrieval over satellite tiles becomes possible with queries ranging from general landscapes ('desert', 'rainforest') to species names ('ibex', 'house finch'), returning images of the expected habitat.
- Because the supervision comes from citizen science and open text, the training set can grow continuously without manual annotation effort.
Reading between the lines
- The wildlife signal may be largely a high-quality land-cover and habitat signal; isolating Limg for each satellite-pretrained base model would reveal how much of the gain is ecological semantics versus extra contrastive training.
- The same location-anchored alignment could be run in reverse to predict species distribution maps from satellite embeddings, effectively turning WildSAT into a species distribution model as well as a representation learner.
- Incorporating observation dates and filtering to seasonal habitat use would likely reduce label noise and sharpen the habitat signal, especially for migratory species whose observed tile may not reflect year-round habitat.
- Quantitative evaluation of zero-shot retrieval on a geographically balanced test set would test whether the US/Europe training bias limits the claimed broad applicability.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces WildSAT, a contrastive learning framework for satellite image representation learning. It fine-tunes an image encoder using three alignment losses: an image self-supervision loss on temporally and geometrically augmented Sentinel-2 tiles (Limg), a text alignment loss using Wikipedia habitat descriptions encoded by GritLM (Ltxt), and a location alignment loss using SINR location embeddings with optional environmental covariates (Lloc). Training data comes from 35.5M iNaturalist observations paired with satellite images, Wikipedia text, and WorldClim covariates. The authors evaluate via linear probing on seven classification datasets and two segmentation datasets, bird encounter-rate prediction on SatBird, and qualitative zero-shot text-to-image retrieval. They report that WildSAT improves 108 of 115 classification settings, with average gains of 7.7-17.4% overall and 4.3-10.4% excluding random-initialized models, and that fine-tuning helps a range of base models including ImageNet, CLIP, Prithvi, SatCLIP, SatlasNet, SeCo, and random-initialized networks.
Significance. The core idea — using citizen-science species observations as a global, free supervision signal for remote sensing representations — is novel and potentially valuable. The scale of the dataset (35.5M observations, 47K species), the breadth of base models considered (20 model-initialization combinations), and the release of code and data are concrete strengths. If the attribution of gains to the wildlife-specific modalities is substantiated, the paper would make a solid contribution to satellite image representation learning and to cross-modal ecological supervision. The zero-shot retrieval capability, though only qualitatively demonstrated, is an interesting byproduct. The central claim is plausible but currently under-supported by the ablation and comparison protocol, as detailed below.
major comments (3)
- [Section 6.4, Table A7] The ablation isolates the contribution of the wildlife text/location terms from the image-only augmentation term (Limg) for only two base models. For ImageNet ViT-B/16, Limg alone raises the four-dataset average from 76.9 to 82.3, while adding text and location adds only 1.3 points (to 83.6). For SeCo ResNet50, Limg alone adds 1.7 points (70.1 to 71.8) and the full method adds 5.1 (to 75.2). Since no Limg-only ablation is reported for SatlasNet, Prithvi, SatCLIP, Random, or CLIP base models, the specific contribution of wildlife supervision to the headline '108 of 115 settings' claim is not established for most settings. This is load-bearing for the abstract's claim that species distributions drive the improvement. Please report Limg-only ablations for at least the satellite-specific base models and the CLIP model, or temper the causal attribution.
- [Section 6.1, Table 2] The comparison to TaxaBind, GRAFT, and RemoteCLIP is confounded. WildSAT's CLIP run includes the image self-supervision term (Limg) and uses parameter-efficient fine-tuning (DoRa), whereas the three baselines are fine-tuned with their own cross-modal objectives (which do not include the Limg term) and, based on their original papers, with full fine-tuning or different protocols. Consequently, the claim that WildSAT 'outperforms recent cross-modal learning methods' conflates the effect of the extra image-augmentation loss and the PEFT scheme with the effect of wildlife text/location supervision. Please re-run the baselines under matched protocols (e.g., adding the same PEFT and/or Limg to the baselines) or clearly state the protocol differences and restrict the comparative claim to what is actually controlled.
- [All result tables (Tables 1-4, A1-A9)] No error bars, confidence intervals, or multiple-seed runs are reported anywhere in the paper. This is a concern because several headline improvements are small in absolute terms (e.g., FMoW 43.3 vs 39.0 in Table A1; SatBird Kenya 24.40 vs 23.90 in Table 3) and some settings show decreases after WildSAT fine-tuning (e.g., ImageNet ResNet50 on UCM, AID, RESISC45, and FMoW in Table A1). The statement in Section 6.1 that WildSAT 'significantly improves performance' needs a statistical basis or at least variance estimates over multiple seeds to rule out that the smaller gains are within noise.
minor comments (5)
- [Section 6.3] The zero-shot retrieval evaluation is qualitative only. Please provide quantitative retrieval metrics (e.g., recall@k on a held-out set with known text-image pairs) or explicitly state that retrieval is demonstrated qualitatively and not yet benchmarked.
- [Section 6.1, Figure 3] The claim that 'the addition of WildSAT improves 108 of the 115 settings' would be more informative if the few settings where performance decreases (e.g., ImageNet ResNet50 on UCM, AID, RESISC45, and FMoW in Table A1) were mentioned or visualized, so readers can judge the consistency of the effect.
- [Section 3.3] The parameter-efficient fine-tuning description is brief. Please specify which parameters are adapted for each architecture (e.g., which layers DoRa modifies, which BatchNorm parameters are tuned for ResNet50) and report the number of trainable parameters introduced.
- [Appendix B] The fact that exact observation dates are not used (only the 2017-2021 time range) is an important limitation that should be discussed in the main text, since temporal mismatch between the wildlife observation and the satellite image acquisition could weaken the habitat-image association.
- [Section 5.3] There is a typo: 'Setinel-2' should be 'Sentinel-2' in the SeCo description. Also, 'topk' in Section 6.3 should be 'top-k'.
Circularity Check
No significant circularity: WildSAT's downstream gains are measured on independent benchmarks, and its training targets (SINR location embeddings, GritLM text, Sentinel-2 tiles) are external to the evaluation; the reported ablations actually expose a confound between the wildlife losses and the image-augmentation term, which is an attribution gap rather than a circular reduction.
full rationale
WildSAT's training objective (Eq. 1) is a sum of three InfoNCE losses that align satellite image embeddings with (i) temporal/geometric augmentations of the same tile, (ii) GritLM embeddings of Wikipedia habitat text, and (iii) SINR location embeddings. None of these targets are derived from the downstream evaluation labels; the seven classification datasets and two segmentation benchmarks are external, and results are reported via linear probing or decoder-only training on fixed frozen features (Section 5). No parameter fitted to a downstream subset is later reported as a prediction, and the zero-shot retrieval experiments use the same frozen encoder with no task-specific training. The main self-citation is SINR [12], whose authors overlap with WildSAT; SINR supplies the location encoder and the underlying iNaturalist observation set. This is a data/module dependency, not a reduction: the paper's own ablation (Tab. A8) shows the full method with SINR (75.2%) is only slightly better than using raw lat/lon/env (74.2-74.9%), and the central claim that wildlife-grounded supervision improves satellite representations does not require SINR's specific embeddings. The skeptic's point about Limg (Tab. A7: for ImageNet ViT-B/16, Limg alone gives 82.3% vs 83.6% full WildSAT; for SeCo, 71.8% vs 75.2%) is an attribution gap for the headline 'wildlife observations drive the gains' across all 115 settings, but it is not circularity: the paper reports the ablation, Limg is not a downstream label, and the wildlife-specific losses do add measurable (if smaller) gains in the reported rows. The iNaturalist-tile-as-habitat premise and geographic bias (Section 6.5) are validity limitations, not circular reductions. Overall no step of the derivation is equivalent to its input by construction; score 1 reflects only the minor, non-load-bearing SINR self-citation.
Assumptions & free parameters
free parameters (4)
- InfoNCE temperature tau =
not reported in main text
- embedding dimension d =
512
- batch size =
64
- learning rate and epochs =
1e-4, 25 epochs
assumptions (5)
- domain assumption A species observation point indicates that the surrounding landscape, at the scale of a 512x512 10m Sentinel-2 tile, reflects the species' preferred habitat.
- domain assumption Wikipedia species text sections describe habitat and range in a way that is semantically aligned with satellite imagery.
- domain assumption Ignoring the exact observation date and matching observations to any 2017-2021 image of the same location is sufficient.
- domain assumption Pre-trained SINR and GritLM embeddings provide useful fixed auxiliary modalities for contrastive alignment.
- standard math InfoNCE contrastive loss is a valid objective for aligning the modalities.
Cite this review
Pith. "Pith review of WildSAT: Learning Satellite Image Representations from Wildlife Observations." pith.science (2026). https://pith.science/paper/MNJ7AARN
@misc{pith2026241214428,
author = {Pith},
title = {Pith review of: WildSAT: Learning Satellite Image Representations from Wildlife Observations},
year = {2026},
howpublished = {\url{https://pith.science/paper/MNJ7AARN}},
note = {Machine review of arXiv:2412.14428}
}
read the original abstract
Species distributions encode valuable ecological and environmental information, yet their potential for guiding representation learning in remote sensing remains underexplored. We introduce WildSAT, which pairs satellite images with millions of geo-tagged wildlife observations readily-available on citizen science platforms. WildSAT employs a contrastive learning approach that jointly leverages satellite images, species occurrence maps, and textual habitat descriptions to train or fine-tune models. This approach significantly improves performance on diverse satellite image recognition tasks, outperforming both ImageNet-pretrained models and satellite-specific baselines. Additionally, by aligning visual and textual information, WildSAT enables zero-shot retrieval, allowing users to search geographic locations based on textual descriptions. WildSAT surpasses recent cross-modal learning methods, including approaches that align satellite images with ground imagery or wildlife photos, demonstrating the advantages of our approach. Finally, we analyze the impact of key design choices and highlight the broad applicability of WildSAT to remote sensing and biodiversity monitoring.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 2 Pith papers
-
Global and Local Entailment Learning for Natural World Imagery
RCME enforces a transitivity constraint in radial embeddings, producing a hierarchical vision-language model that orders taxonomic labels better and improves hierarchical classification and retrieval.
-
EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
The paper introduces EcoWikiRS, a dataset of aerial images paired with GBIF species observations and Wikipedia habitat text, and WINCEL, a weighted InfoNCE loss that improves zero-shot EUNIS ecosystem classification.
Reference graph
Works this paper leans on
-
[1]
https://www.ebird.org
eBird. https://www.ebird.org . Accessed on 2025- 03-03. 13
2025
-
[2]
https://www.inaturalist.org
iNaturalist. https://www.inaturalist.org . Ac- cessed on 2025-03-03. 2, 13
2025
-
[3]
https : / / source
Planet, Radiant Earth Foundation, Western Cape Department of Agriculture, German Aerospace Center (DLR): A fusion dataset for crop type classification in Western Cape, South Africa. https : / / source . coop / esa / fusion - competition. Accessed on 2025-03-03. 5, 6, 13, 17
2025
-
[4]
https://www.wikipedia.org
Wikipedia. https://www.wikipedia.org. Accessed on 2025-03-03. 3, 4, 5, 14
2025
-
[5]
Satlaspretrain: A large- scale dataset for remote sensing image understanding
Favyen Bastani, Piper Wolters, Ritwik Gupta, Joe Ferdi- nando, and Aniruddha Kembhavi. Satlaspretrain: A large- scale dataset for remote sensing image understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16772–16782, 2023. 2, 3, 4, 5, 6, 7, 14, 17, 19
2023
-
[6]
Bird- snap: Large-scale fine-grained visual categorization of birds
Thomas Berg, Jiongxin Liu, Seung Woo Lee, Michelle L Alexander, David W Jacobs, and Peter N Belhumeur. Bird- snap: Large-scale fine-grained visual categorization of birds. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2014. 2
2014
-
[7]
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. InPro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9640–9649, 2021. 5, 6, 14, 17
2021
-
[8]
Remote sens- ing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu. Remote sens- ing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883, 2017. 5, 6, 8, 13, 14, 15, 18, 19
2017
Show all 97 references
-
[9]
Functional map of the world
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6172–6180, 2018. 5, 6, 13, 14, 15, 19
2018
-
[10]
Geo-aware networks for fine-grained recognition
Grace Chu, Brian Potetz, Weijun Wang, Andrew Howard, Yang Song, Fernando Brucher, Thomas Leung, and Hartwig Adam. Geo-aware networks for fine-grained recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019. 2
2019
-
[11]
The harmonized landsat and sentinel-2 surface reflectance data set
Martin Claverie, Junchang Ju, Jeffrey G Masek, Jennifer L Dungan, Eric F Vermote, Jean-Claude Roger, Sergii V Skakun, and Christopher Justice. The harmonized landsat and sentinel-2 surface reflectance data set. Remote sensing of environment, 219:145–161, 2018. 5
2018
-
[12]
Spatial implicit neural representations for global-scale species mapping
Elijah Cole, Grant Van Horn, Christian Lange, Alexander Shepard, Patrick Leary, Pietro Perona, Scott Loarie, and Oisin Mac Aodha. Spatial implicit neural representations for global-scale species mapping. In International Conference on Machine Learning, pages 6320–6342. PMLR, 2...
2023
-
[13]
Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery
Yezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu, Erik Rozi, Yutong He, Marshall Burke, David Lobell, and Stefano Ermon. Satmae: Pre-training transformers for tem- poral and multi-spectral satellite imagery. Advances in Neu- ral Information Processing Systems, 35:197–211, ...
2022
-
[14]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 3, 4, 5, 6, 14, 15, 17, 18
2009
-
[15]
Geobind: Binding text, image, and audio through satellite images
Aayush Dhakal, Subash Khanal, Srikumar Sastry, Adeel Ah- mad, and Nathan Jacobs. Geobind: Binding text, image, and audio through satellite images. In International Geoscience and Remote Sensing Symposium, 2024. 1, 2, 3
2024
-
[16]
Sat-sinr: High-resolution species distribution models through satellite imagery
Johannes Dollinger, Philipp Brun, Vivien Sainte Fare Gar- not, and Jan Dirk Wegner. Sat-sinr: High-resolution species distribution models through satellite imagery. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Infor- mation Sciences, 2024. 1, 3
2024
-
[17]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. InInternational ...
2020
-
[18]
Sentinel-1-missions-sentinel online-sentinel online
ESA. Sentinel-1-missions-sentinel online-sentinel online. Eur. Sp. Agency, 2022. 2, 4, 14
2022
-
[19]
A gnn-rnn approach for harnessing geospa- tial and temporal information: application to crop yield pre- diction
Joshua Fan, Junwen Bai, Zhiyun Li, Ariel Ortiz-Bobea, and Carla P Gomes. A gnn-rnn approach for harnessing geospa- tial and temporal information: application to crop yield pre- diction. In Proceedings of the AAAI conference on artificial intelligence, pages 11873–11881, 2022. 1
2022
-
[20]
Worldclim 2: new 1-km spatial resolution climate surfaces for global land ar- eas
Stephen E Fick and Robert J Hijmans. Worldclim 2: new 1-km spatial resolution climate surfaces for global land ar- eas. International journal of climatology , 37(12):4302– 4315, 2017. 4
2017
-
[21]
Train- ing batchnorm and only batchnorm: On the expressive power of random features in cnns
Jonathan Frankle, David J Schwab, and Ari S Morcos. Train- ing batchnorm and only batchnorm: On the expressive power of random features in cnns. In International Conference on Learning Representations, 2021. 4, 19
2021
-
[22]
Mapping fishing activities and suitable fishing grounds using nighttime satellite images and maximum entropy modelling
Rollan C Geronimo, Erik C Franklin, Russell E Brainard, Christopher D Elvidge, Mudjekeewis D Santos, Roberto Venegas, and Camilo Mora. Mapping fishing activities and suitable fishing grounds using nighttime satellite images and maximum entropy modelling. Remote Sensing, 10(10):1604,
-
[23]
Skysense: A multi- modal remote sensing foundation model towards universal interpretation for earth observation imagery
Xin Guo, Jiangwei Lao, Bo Dang, Yingying Zhang, Lei Yu, Lixiang Ru, Liheng Zhong, Ziyuan Huang, Kang Wu, Dingxiang Hu, Huimei He, Jian Wang, Jingdong Chen, Ming Yang, Yongjun Zhang, and Yansheng Li. Skysense: A multi- modal remote sensing foundation model towards universal int...
2024
-
[24]
Openstreetmap: User- generated street maps
Mordechai Haklay and Patrick Weber. Openstreetmap: User- generated street maps. IEEE Pervasive computing, 7(4):12– 18, 2008. 2
2008
-
[25]
Combining observational data and language for species range estimation
Max Hamilton, Christian Lange, Elijah Cole, Alexan- der Shepard, Samuel Heinrich, Oisin Mac Aodha, Grant Van Horn, and Subhransu Maji. Combining observational data and language for species range estimation. In Advances in Neural Information Processing Systems, 2024. 3, 4 9
2024
-
[26]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3
2016
-
[27]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 16000– 16009, 2022. 2, 5
2022
-
[28]
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019. 1, 2, 5...
2019
-
[29]
Contrastive ground-level image and remote sensing pre- training improves representation learning for natural world imagery
Andy V Huynh, Lauren E Gillespie, Jael Lopez-Saucedo, Claire Tang, Rohan Sikand, and Mois ´es Exp ´osito-Alonso. Contrastive ground-level image and remote sensing pre- training improves representation learning for natural world imagery. In European Conference on Computer Vision ,
-
[30]
First experience with sentinel-2 data for crop and tree species classifications in central europe
Markus Immitzer, Francesco Vuolo, and Clement Atzberger. First experience with sentinel-2 data for crop and tree species classifications in central europe. Remote sensing, 8(3):166,
-
[31]
Aligning geo-tagged clip representations and satellite im- agery for few-shot land use classification
Pallavi Jain, Diego Marcos, Dino Ienco, Roberto Interdo- nato, Aayush Dhakal, Nathan Jacobs, and Tristan Berchoux. Aligning geo-tagged clip representations and satellite im- agery for few-shot land use classification. In IGARSS 2024- 2024 IEEE International Geoscience and Remo...
2024
-
[32]
Foundation models for generalist geospatial artificial intelligence, 2023
Johannes Jakubik, S Roy, CE Phillips, P Fraccaro, D Godwin, B Zadrozny, D Szwarcman, C Gomes, G Nyir- jesy, B Edwards, et al. Foundation models for generalist geospatial artificial intelligence, 2023. URL https://arxiv. org/abs/2310.18660, 2023. 1, 2, 3, 5, 14, 15
2023 arXiv
-
[33]
Crop monitoring by multimodal remote sensing: A review
Priyabrata Karmakar, Shyh Wei Teng, Manzur Murshed, Shaoning Pang, Yanyu Li, and Hao Lin. Crop monitoring by multimodal remote sensing: A review. Remote Sensing Applications: Society and Environment, 33:101093, 2024. 1
2024
-
[34]
Learning tri-modal embeddings for zero-shot soundscape mapping
Subash Khanal, Srikumar Sastry, Aayush Dhakal, and Nathan Jacobs. Learning tri-modal embeddings for zero-shot soundscape mapping. In BMVC, 2023. 2
2023
-
[35]
Satclip: Global, general- purpose location embeddings with satellite imagery
Konstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey, and Marc Rußwurm. Satclip: Global, general- purpose location embeddings with satellite imagery. InAAAI Conference on Artificial Intelligence, 2024. 2, 3, 5, 14, 15, 18, 19
2024
-
[36]
Operational monitoring of illegal fishing in ghana through exploitation of satellite earth obser- vation and ais data
Andrey A Kurekin, Benjamin R Loveday, Oliver Clements, Graham D Quartly, Peter I Miller, George Wiafe, and Kwame Adu Agyekum. Operational monitoring of illegal fishing in ghana through exploitation of satellite earth obser- vation and ais data. Remote Sensing, 11(3):293, 2019. 1
2019
-
[37]
Deep learning classification of land cover and crop types using remote sensing data
Nataliia Kussul, Mykola Lavreniuk, Sergii Skakun, and An- drii Shelestov. Deep learning classification of land cover and crop types using remote sensing data. IEEE Geoscience and Remote Sensing Letters, 14(5):778–782, 2017. 1
2017
-
[38]
Geo- bench: Toward foundation models for earth monitoring
Alexandre Lacoste, Nils Lehmann, Pau Rodriguez, Evan Sherwin, Hannah Kerner, Bj¨orn L¨utjens, Jeremy Irvin, David Dao, Hamed Alemohammad, Alexandre Drouin, et al. Geo- bench: Toward foundation models for earth monitoring. Ad- vances in Neural Information Processing Systems, 36...
2023
-
[39]
A high-resolution canopy height model of the earth
Nico Lang, Walter Jetz, Konrad Schindler, and Jan Dirk Wegner. A high-resolution canopy height model of the earth. Nature Ecology & Evolution, 2023. 1
2023
-
[40]
Scaling & shifting your features: A new baseline for efficient model tuning
Dongze Lian, Daquan Zhou, Jiashi Feng, and Xinchao Wang. Scaling & shifting your features: A new baseline for efficient model tuning. Advances in Neural Information Processing Systems, 35:109–123, 2022. 4, 19
2022
-
[41]
Re- moteclip: A vision language foundation model for remote sensing
Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou. Re- moteclip: A vision language foundation model for remote sensing. IEEE Transactions on Geoscience and Remote Sensing, 2024. 2, 3, 5, 6, 15, 16
2024
-
[42]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, 2021. 4
2021
-
[43]
Overview of geo- lifeclef 2022: Predicting species presence from multi-modal remote sensing, bioclimatic and pedologic data
Titouan Lorieul, Elijah Cole, Benjamin Deneu, Maximilien Servajean, Pierre Bonnet, and Alexis Joly. Overview of geo- lifeclef 2022: Predicting species presence from multi-modal remote sensing, bioclimatic and pedologic data. In CLEF (Working Notes), pages 1940–1956, 2022. 3
2022
-
[44]
Presence- only geographical priors for fine-grained image classifica- tion
Oisin Mac Aodha, Elijah Cole, and Pietro Perona. Presence- only geographical priors for fine-grained image classifica- tion. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, 2019. 2
2019
-
[45]
Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations
Gengchen Mai, Ni Lao, Yutong He, Jiaming Song, and Stefano Ermon. Csp: Self-supervised contrastive spatial pre-training for geospatial-visual representations. In Inter- national Conference on Machine Learning , pages 23498– 23515. PMLR, 2023. 3
2023
-
[46]
Change- aware sampling and contrastive learning for satellite images
Utkarsh Mall, Bharath Hariharan, and Kavita Bala. Change- aware sampling and contrastive learning for satellite images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5261–5270, 2023. 1, 2, 15
2023
-
[47]
Remote sens- ing vision-language foundation models without annotations via ground remote alignment
Utkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl V ondrick, Bharath Hariharan, and Kavita Bala. Remote sens- ing vision-language foundation models without annotations via ground remote alignment. In The Twelfth International Conference on Learning Representations, 2024....
2024
-
[48]
Seasonal contrast: Un- supervised pre-training from uncurated remote sensing data
Oscar Manas, Alexandre Lacoste, Xavier Gir ´o-i Nieto, David Vazquez, and Pau Rodriguez. Seasonal contrast: Un- supervised pre-training from uncurated remote sensing data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9414–9423, 2021. 1, 2, ...
2021
-
[49]
Dora: Enhancing parameter- efficient fine-tuning with dynamic rank distribution
Yulong Mao, Kaiyu Huang, Changhao Guan, Ganglin Bao, Fengran Mo, and Jinan Xu. Dora: Enhancing parameter- efficient fine-tuning with dynamic rank distribution. arXiv preprint arXiv:2405.17357, 2024. 4, 19 10
2024 arXiv
-
[50]
Landsat 9: Empowering open science and appli- cations through continuity
Jeffrey G Masek, Michael A Wulder, Brian Markham, Joel McCorkel, Christopher J Crawford, James Storey, and Del T Jenstrom. Landsat 9: Empowering open science and appli- cations through continuity. Remote Sensing of Environment, 248:111968, 2020. 2
2020
-
[51]
Land cover classification and feature extraction from national agriculture imagery program (naip) orthoimagery: A review
Aaron E Maxwell, Timothy A Warner, Brian C Vanderbilt, and Christopher A Ramezan. Land cover classification and feature extraction from national agriculture imagery program (naip) orthoimagery: A review. Photogrammetric Engineer- ing & Remote Sensing, 83(11):737–747, 2017. 2
2017
-
[52]
Gen- erative representational instruction tuning
Niklas Muennighoff, SU Hongjin, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. Gen- erative representational instruction tuning. In ICLR 2024 Workshop: How Far Are We From AGI, 2024. 3
2024
-
[53]
Mmearth: Ex- ploring multi-modal pretext tasks for geospatial representa- tion learning
Vishal Nedungadi, Ankit Kariryaa, Stefan Oehmcke, Serge Belongie, Christian Igel, and Nico Lang. Mmearth: Ex- ploring multi-modal pretext tasks for geospatial representa- tion learning. In European Conference on Computer Vision,
-
[54]
In-domain representation learning for remote sensing
Maxim Neumann, Andre Susano Pinto, Xiaohua Zhai, and Neil Houlsby. In-domain representation learning for remote sensing. arXiv preprint arXiv:1911.06721, 2019. 2
1911 arXiv
-
[55]
Repre- sentation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 4, 5
2018 arXiv
-
[56]
xview3-sar: Detecting dark fishing activity using synthetic aperture radar imagery
Fernando Paolo, Tsu-ting Tim Lin, Ritwik Gupta, Bryce Goodman, Nirav Patel, Daniel Kuster, David Kroodsma, and Jared Dunnmon. xview3-sar: Detecting dark fishing activity using synthetic aperture radar imagery. Advances in Neural Information Processing Systems, 2022. 1
2022
-
[57]
Dis- count: counting in large image collections with detector- based importance sampling
Gustavo Perez, Subhransu Maji, and Daniel Sheldon. Dis- count: counting in large image collections with detector- based importance sampling. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 22294–22302,
-
[58]
Geo- plant: Spatial plant species prediction dataset
Lukas Picek, Christophe Botella, Maximilien Servajean, C´esar Leblanc, R ´emi Palard, Th ´eo Larcher, Benjamin Deneu, Diego Marcos, Pierre Bonnet, and Alexis Joly. Geo- plant: Spatial plant species prediction dataset. Advances in Neural Information Processing Systems - Dataset...
2024
-
[59]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International conference on machine learning...
2021
-
[60]
Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400, 2019
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400, 2019. 5, 15, 19
2019
-
[61]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI, pages 234–241. Springer, 2015. 5, 17
2015
-
[62]
Birdsat: Cross-view contrastive masked autoencoders for bird species classification and mapping
Srikumar Sastry, Subash Khanal, Aayush Dhakal, Di Huang, and Nathan Jacobs. Birdsat: Cross-view contrastive masked autoencoders for bird species classification and mapping. In Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, pages 7136–7145, ...
2024
-
[63]
Taxabind: A unified embedding space for ecological applications
Srikumar Sastry, Subash Khanal, Aayush Dhakal, Adeel Ah- mad, and Nathan Jacobs. Taxabind: A unified embedding space for ecological applications. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2025. 1, 2, 3, 5, 6, 7, 15, 16
2025
-
[64]
ebird: A citizen-based bird observation network in the biological sci- ences
Brian L Sullivan, Christopher L Wood, Marshall J Iliff, Rick E Bonney, Daniel Fink, and Steve Kelling. ebird: A citizen-based bird observation network in the biological sci- ences. Biological conservation, 142(10):2282–2292, 2009. 2
2009
-
[65]
Bigearthnet: A large-scale benchmark archive for remote sensing image understanding
Gencer Sumbul, Marcela Charfuelan, Beg ¨um Demir, and V olker Markl. Bigearthnet: A large-scale benchmark archive for remote sensing image understanding. In IGARSS 2019- 2019 IEEE International Geoscience and Remote Sensing Symposium, pages 5901–5904. IEEE, 2019. 2, 5, 6, 13, ...
2019
-
[66]
Ap- plications of artificial intelligence for disaster management
Wenjuan Sun, Paolo Bocchini, and Brian D Davison. Ap- plications of artificial intelligence for disaster management. Natural Hazards, 103(3):2631–2689, 2020. 1
2020
-
[67]
Satbird: a dataset for bird species distribution modeling using remote sensing and citizen science data
M ´elisande Teng, Amna Elmustafa, Benjamin Akera, Yoshua Bengio, Hager Radi, Hugo Larochelle, and David Rolnick. Satbird: a dataset for bird species distribution modeling using remote sensing and citizen science data. Advances in Neural Information Processing Systems - Dataset...
2023
-
[68]
Canopy height mapping by sentinel 1 and 2 satellite images, airborne lidar data, and machine learning
Catherine Torres de Almeida, J ´essica Gerente, Jamerson Ro- drigo dos Prazeres Campos, Francisco Caruso Gomes Junior, Lucas Antonio Providelo, Guilherme Marchiori, and Xinjian Chen. Canopy height mapping by sentinel 1 and 2 satellite images, airborne lidar data, and machine l...
2022
-
[69]
Learn- ing to interpret satellite images using wikipedia
Burak Uzkent, Evan Sheehan, Chenlin Meng, Zhongyi Tang, Marshall Burke, David Lobell, and Stefano Ermon. Learn- ing to interpret satellite images using wikipedia. In Proceed- ings of the International Joint Conference on Artificial Intel- ligence, 2019. 2, 3
2019
-
[70]
Esa worldcover: Global land cover mapping at 10 m resolution for 2020 based on sentinel-1 and 2 data
Ruben Van De Kerchove, Daniele Zanaga, Wanda Keers- maecker, Niels Souverijns, Jan Wevers, Carsten Brockmann, Alex Grosu, Audrey Paccini, Oliver Cartus, Maurizio San- toro, et al. Esa worldcover: Global land cover mapping at 10 m resolution for 2020 based on sentinel-1 and 2 d...
2020
-
[71]
The inaturalist species classification and de- tection dataset
Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and de- tection dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages...
-
[72]
Geoclip: Clip-inspired alignment be- tween locations and images for effective worldwide geo- localization
Vicente Vivanco Cepeda, Gaurav Kumar Nayak, and Mubarak Shah. Geoclip: Clip-inspired alignment be- tween locations and images for effective worldwide geo- localization. Advances in Neural Information Processing Systems, 2024. 1, 2
2024
-
[73]
Satellite image anal- 11 ysis for disaster and crisis-management support.IEEE trans- actions on geoscience and remote sensing, 45(6):1520–1528,
Stefan V oigt, Thomas Kemper, Torsten Riedlinger, Ralph Kiefl, Klaas Scholte, and Harald Mehl. Satellite image anal- 11 ysis for disaster and crisis-management support.IEEE trans- actions on geoscience and remote sensing, 45(6):1520–1528,
-
[74]
isaid: A large-scale dataset for instance segmentation in aerial images
Syed Waqas Zamir, Aditya Arora, Akshita Gupta, Salman Khan, Guolei Sun, Fahad Shahbaz Khan, Fan Zhu, Ling Shao, Gui-Song Xia, and Xiang Bai. isaid: A large-scale dataset for instance segmentation in aerial images. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision ...
2019
-
[75]
Global crop monitoring: a satellite-based hierarchical approach
Bingfang Wu, Ren ´e Gommes, Miao Zhang, Hongwei Zeng, Nana Yan, Wentao Zou, Yang Zheng, Ning Zhang, Sheng Chang, Qiang Xing, et al. Global crop monitoring: a satellite-based hierarchical approach. Remote Sensing, 7(4): 3907–3933, 2015. 1
2015
-
[76]
Aid: A benchmark data set for performance evaluation of aerial scene classification
Gui-Song Xia, Jingwen Hu, Fan Hu, Baoguang Shi, Xiang Bai, Yanfei Zhong, Liangpei Zhang, and Xiaoqiang Lu. Aid: A benchmark data set for performance evaluation of aerial scene classification. IEEE Transactions on Geoscience and Remote Sensing, 55(7):3965–3981, 2017. 2, 5, 6, 8...
2017
-
[77]
Dota: A large-scale dataset for object detection in aerial images
Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Be- longie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liang- pei Zhang. Dota: A large-scale dataset for object detection in aerial images. In Proceedings of the IEEE conference on computer vision and pattern recognition, ...
-
[78]
Bag-of-visual-words and spa- tial extensions for land-use classification
Yi Yang and Shawn Newsam. Bag-of-visual-words and spa- tial extensions for land-use classification. In Proceedings of the 18th SIGSPATIAL international conference on advances in geographic information systems, pages 270–279, 2010. 5, 6, 8, 13, 14, 15, 18, 19
2010
-
[79]
Mapping smallholder cashew plantations to inform sustainable tree crop expan- sion in benin
Leikun Yin, Rahul Ghosh, Chenxi Lin, David Hale, Christoph Weigl, James Obarowski, Junxiong Zhou, Jessica Till, Xiaowei Jia, Nanshan You, et al. Mapping smallholder cashew plantations to inform sustainable tree crop expan- sion in benin. Remote Sensing of Environment, 295:113695,
-
[80]
Deep learning in environmen- tal remote sensing: Achievements and challenges
Qiangqiang Yuan, Huanfeng Shen, Tongwen Li, Zhiwei Li, Shuwen Li, Yun Jiang, Hongzhang Xu, Weiwei Tan, Qian- qian Yang, Jiwen Wang, et al. Deep learning in environmen- tal remote sensing: Achievements and challenges. Remote sensing of Environment, 241:111716, 2020. 1
2020
-
[81]
Ecowikirs: Learning ecolog- ical representation of satellite images from weak supervision with species observations and wikipedia
Valerie Zermatten, Javiera Castillo-Navarro, Pallavi Jain, Devis Tuia, and Diego Marcos. Ecowikirs: Learning ecolog- ical representation of satellite images from weak supervision with species observations and wikipedia. In Proceedings of the Computer Vision and Pattern Recogni...
2025
-
[82]
Changen2: Multi-temporal re- mote sensing generative change foundation model
Zhuo Zheng, Stefano Ermon, Dongjun Kim, Liangpei Zhang, and Yanfei Zhong. Changen2: Multi-temporal re- mote sensing generative change foundation model. IEEE Transactions on Pattern Analysis and Machine Intelligence,
-
[83]
Deep learning in remote sensing: A comprehensive review and list of resources
Xiao Xiang Zhu, Devis Tuia, Lichao Mou, Gui-Song Xia, Liangpei Zhang, Feng Xu, and Friedrich Fraundorfer. Deep learning in remote sensing: A comprehensive review and list of resources. Geoscience and remote sensing magazine , 5 (4):8–36, 2017. 1
2017
-
[84]
So2Sat20k
Xiao Xiang Zhu, Jingliang Hu, Chunping Qiu, Yilei Shi, Jian Kang, Lichao Mou, Hossein Bagheri, Matthias Haberle, Yuansheng Hua, Rong Huang, et al. So2sat lcz42: A bench- mark data set for the classification of global local climate zones [software and data sets]. IEEE Geoscienc...
2020
-
[85]
Results of linear probing different models on seven downstream datasets without (Base) and with (+WS) WildSAT fine- tuning
[76] [8] [9] [28] [84] [65] Base +WS Base +WS Base +WS Base +WS Base +WS Base +WS Base +WS ImageNet [14] 93.2 97.5 84.4 88.9 88.2 93.0 43.8 51.4 94.5 97.3 41.8 55.2 52.3 58.2 MoCov3 [7] 94.2 95.1 86.0 86.9 89.1 90.3 51.1 52.9 95.9 97.1 47.6 56.6 51.6 57.0 CLIP [59] 94.5 96.3 8...
2017
-
[86]
Additional linear probing results on satellite image classification datasets
[76] [8] [9] [28] [84] [65] Base +WS Base +WS Base +WS Base +WS Base +WS Base +WS Base +WS ResNet50 ImageNet1K V2 [60] 94.2 93.6 87.8 86.7 90.5 90.1 47.3 46.0 95.5 96.0 36.1 46.6 55.8 57.5 ResNet50 ImageNet1K V1 [14] 92.5 93.5 90.4 88.8 85.1 84.7 40.7 37.0 88.0 94.9 38.8 48.2 ...
-
[87]
Linear probing results on downstream satellite classification datasets using models with CLIP as the base
[76] [8] [9] [28] [84] [65] TaxaBind [63] 80.5 67.7 72.6 31.2 85.2 33.9 47.6 GRAFT [47] 81.1 76.1 83.3 39.3 90.9 36.6 48.0 RemoteCLIP [41] 96.1 86.1 90.9 45.7 93.3 35.5 49.4 CLIP [59] 94.5 86.3 92.1 51.5 92.2 37.6 47.1 WildSAT (Ours) 96.3 88.0 93.0 52.8 97.1 49.7 59.1 Table A3...
-
[88]
forgetting
[65] Base +WS Base +WS ViT-B/16 Prithvi-100M [32] 28.7 50.1 34.4 53.9 ViT-B/16 SatCLIP [35] 48.8 48.8 33.4 36.1 ResNet50 SatCLIP [35] 45.8 46.3 46.7 54.3 Average 41.1 48.4 38.2 48.1 Table A4. Linear probing results on models using multispectral images . Accuracy is reported fo...
-
[89]
Ablation of parameter efficient fine-tuning (PEFT) when applied with WildSAT
[76] [8] [9] [28] [84] [65] ResNet50 ImageNet1K [60] 91.3% 82.0% 85.6% 42.1% 95.0% 47.0% 56.1% ResNet50 ImageNet1K [60] ✓ 93.6% 86.7% 90.1% 46.0% 96.0% 46.6% 57.5% ViT-B/16 CLIP [59] 82.1% 71.0% 75.3% 34.9% 93.4% 50.4% 49.0% ViT-B/16 CLIP [59] ✓ 96.3% 88.0% 93.0% 53.6% 97.1% 4...
-
[90]
‘house sparrow’ is a small, common bird typically found in urban areas
-
[91]
‘albatross’ is a large bird commonly found in the sea
-
[92]
‘sandpiper’ is a small bird that dwells in the coast
-
[93]
‘horned lark’ is a bird species found in open land such as on farmland, on prairies, and in deserts. 19
-
[94]
‘cactus’ is a type of plant commonly found in the desert
-
[95]
‘rock pigeon’ is a bird commonly found in urban and residential areas
-
[96]
‘virginia rail’ is a bird found in freshwater and brackish marshes, and sometimes salt marshes in winter
-
[97]
urban coast cactus high altitude virginia rail house sparrow albatross sandpiper horned lark mountains savanna rainforest gypsum dunes rock pigeon american marten aquatic Figure A4
‘american marten’ is a North American mammal that is found in forests, and broadly distributed in North America from Alaska and Canada to New York. urban coast cactus high altitude virginia rail house sparrow albatross sandpiper horned lark mountains savanna rainforest gypsum ...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.