REVIEW 4 cited by
SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Self-supervised pre-training bears potential to generate expressive representations without human annotation. Most pre-training in Earth observation (EO) are based on ImageNet or medium-size, labeled remote sensing (RS) datasets. We share an unlabeled RS dataset SSL4EO-S12 (Self-Supervised Learning for Earth Observation - Sentinel-1/2) to assemble a large-scale, global, multimodal, and multi-seasonal corpus of satellite imagery from the ESA Sentinel-1 \& -2 satellite missions. For EO applications we demonstrate SSL4EO-S12 to succeed in self-supervised pre-training for a set of methods: MoCo-v2, DINO, MAE, and data2vec. Resulting models yield downstream performance close to, or surpassing accuracy measures of supervised learning. In addition, pre-training on SSL4EO-S12 excels compared to existing datasets. We make openly available the dataset, related source code, and pre-trained models at https://github.com/zhu-xlab/SSL4EO-S12.
Forward citations
Cited by 4 Pith papers
-
Above-ground Biomass Estimation with Geospatial Foundation Models
A benchmark of geospatial foundation models for biomass regression shows that pre-computed embedding products, especially AlphaEarth Foundations, outperform both frozen weight-distributed GFMs and a fully supervised s...
-
Probing Geospatial SSL Representations with Environmental Signals
Self-supervised satellite imagery representations encode physically meaningful environmental signals (ERA5 variables) that correlate with downstream task performance, particularly for agriculture and disaster domains.
-
The View From Space: Navigating Instrumentation Differences with EOFMs
EOFM embeddings are strongly partitioned by sensor architecture, so matching spectral bands is not enough to make cross-sensor embedding search reliable.
-
UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations
UrbanFusion trains a multimodal location encoder with stochastic fusion (contrastive alignment plus latent reconstruction under random modality masking) and edges out prior GeoAI models on most of 41 urban tasks, with...
Discussion (0). Continue with ORCID to comment.