UniverSat is a ViT-style model with a universal patch encoder enabling self-supervised training on heterogeneous multimodal Earth observation data from varying resolutions and sensors.
Lightweight, pre-trained transformers for remote sensing timeseries
9 Pith papers cite this work. Polarity classification is still indexing.
abstract
Machine learning methods for satellite data have a range of societally relevant applications, but labels used to train models can be difficult or impossible to acquire. Self-supervision is a natural solution in settings with limited labeled data, but current self-supervised models for satellite data fail to take advantage of the characteristics of that data, including the temporal dimension (which is critical for many applications, such as monitoring crop growth) and availability of data from many complementary sensors (which can significantly improve a model's predictive performance). We present Presto (the Pretrained Remote Sensing Transformer), a model pre-trained on remote sensing pixel-timeseries data. By designing Presto specifically for remote sensing data, we can create a significantly smaller but performant model. Presto excels at a wide variety of globally distributed remote sensing tasks and performs competitively with much larger models while requiring far less compute. Presto can be used for transfer learning or as a feature extractor for simple models, enabling efficient deployment at scale.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
Fusing embeddings from four Earth models (AlphaEarth, Tessera, GeoCLIP, SatCLIP) outperforms the best single model on four of six tasks, with gains depending on task and location.
CropNet, a lightweight CNN jointly convolving spectral and temporal dimensions, learns invariant crop signatures from multispectral time series and outperforms larger models under geographic domain shifts on the new CropGlobe benchmark spanning eight countries.
TESSERA learns robust label-efficient embeddings from irregular multi-modal EO time series via Barlow Twins plus global shuffling and mix-based regularizers, delivering SOTA accuracy on classification, segmentation and regression tasks while releasing planetary-scale embeddings and code.
MOMO merges sensor-specific models from three Mars orbital instruments at matched validation loss stages to form a foundation model that outperforms ImageNet, Earth observation, sensor-specific, and supervised baselines on nine Mars-Bench tasks.
A systematic review that introduces a framework for feature extraction in remote sensing, traces its evolution in the data value chain, and synthesizes trends toward unified representations and foundation models.
Ensemble of U-Net (temporal channels stacked) and fine-tuned Prithvi-EO-2.0 reaches 68.32 ±1 accuracy on viticulture potential, ranking 2nd of 7 in ImageCLEF AI4Agri 2026.
Sentinel-2-specific foundation models outperform ImageNet on multi-continent crop mapping, with 100 labels achieving high overall accuracy but 900 required to address class imbalance.
citing papers explorer
-
UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation
UniverSat is a ViT-style model with a universal patch encoder enabling self-supervised training on heterogeneous multimodal Earth observation data from varying resolutions and sensors.
-
Better Together: Evaluating the Complementarity of Earth Embedding Models
Fusing embeddings from four Earth models (AlphaEarth, Tessera, GeoCLIP, SatCLIP) outperforms the best single model on four of six tasks, with gains depending on task and location.
-
Invariant Features for Global Crop Type Classification
CropNet, a lightweight CNN jointly convolving spectral and temporal dimensions, learns invariant crop signatures from multispectral time series and outperforms larger models under geographic domain shifts on the new CropGlobe benchmark spanning eight countries.
-
TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis
TESSERA learns robust label-efficient embeddings from irregular multi-modal EO time series via Barlow Twins plus global shuffling and mix-based regularizers, delivering SOTA accuracy on classification, segmentation and regression tasks while releasing planetary-scale embeddings and code.
-
MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications
MOMO merges sensor-specific models from three Mars orbital instruments at matched validation loss stages to form a foundation model that outperforms ImageNet, Earth observation, sensor-specific, and supervised baselines on nine Mars-Bench tasks.
-
Feature Extraction in the Remote Sensing Data Value Chain: A Systematic Review of Methods and Applications
A systematic review that introduces a framework for feature extraction in remote sensing, traces its evolution in the data value chain, and synthesizes trends toward unified representations and foundation models.
-
Predicting Viticulture Potential through an Ensemble of U-Net and a Geospatial Foundation Model
Ensemble of U-Net (temporal channels stacked) and fine-tuned Prithvi-EO-2.0 reaches 68.32 ±1 accuracy on viticulture potential, ranking 2nd of 7 in ImageCLEF AI4Agri 2026.
-
On the Generalizability of Foundation Models for Crop Type Mapping
Sentinel-2-specific foundation models outperform ImageNet on multi-continent crop mapping, with 100 labels achieving high overall accuracy but 900 required to address class imbalance.
- OlmoEarth v1.2: A more efficient family of OlmoEarth models