REVIEW 11 cited by
Lightweight, Pre-trained Transformers for Remote Sensing Timeseries
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Machine learning methods for satellite data have a range of societally relevant applications, but labels used to train models can be difficult or impossible to acquire. Self-supervision is a natural solution in settings with limited labeled data, but current self-supervised models for satellite data fail to take advantage of the characteristics of that data, including the temporal dimension (which is critical for many applications, such as monitoring crop growth) and availability of data from many complementary sensors (which can significantly improve a model's predictive performance). We present Presto (the Pretrained Remote Sensing Transformer), a model pre-trained on remote sensing pixel-timeseries data. By designing Presto specifically for remote sensing data, we can create a significantly smaller but performant model. Presto excels at a wide variety of globally distributed remote sensing tasks and performs competitively with much larger models while requiring far less compute. Presto can be used for transfer learning or as a feature extractor for simple models, enabling efficient deployment at scale.
Forward citations
Cited by 11 Pith papers
-
How Usable Are Geospatial Foundation Models? A Systematic Evaluation of 89 Models
A systematic usability evaluation of 89 geospatial foundation models finds severe accessibility gaps, including no surveyed model offering documented uncertainty quantification.
-
TESSERA v2: Scaling Pixel-wise Earth Foundation Models
Downstream-driven scaling of pixel-wise Barlow Twins EO models favors large encoders and matched data over projectors, and distillation yields compact Matryoshka students that lead multi-task embedding benchmarks.
-
Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets
A locality-aware embedding anomaly detector identifies label errors in global crop reference data; conservative cleaning raises WorldCereal crop-type macro-F1 in all five tested regions.
-
Position Prediction Self-Supervised Learning for Multimodal Satellite Imagery Semantic Segmentation
Applying LOCA's relative position prediction to multimodal satellite imagery achieves higher flood-segmentation IoU than MAE-style baselines on Sen1Floods11, but the improvement is largely confounded by test-set hyper...
-
Benchmarking Deep Learning Models for Dense Event Classification of Offshore Wind Infrastructure in Sentinel-1 Time Series
A supervised BiLSTM improves dense event classification of offshore wind Sentinel-1 time series over the rule-based baseline (AUCEditSim 0.7853 to 0.8509), and the resulting labels expose regional deployment dynamics.
-
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
A structured protocol for deploying geospatial foundation models is introduced and validated in WorldCereal, where fine-tuned Presto outperforms a fully-supervised CatBoost baseline in crop mapping.
-
Farm-Level, In-Season Crop Identification for India
A Google DeepMind team built a transformer-based system that maps 12 crops across India at farm level, in-season, with state-level area agreement of 94% (winter) and 75% (monsoon) against the 2023-24 census.
-
High-Resolution Live Fuel Moisture Content (LFMC) Maps for Wildfire Risk from Multimodal Earth Observation Data
Fine-tuning the pretrained Galileo model on the Globe-LFMC dataset yields 10 m wall-to-wall live fuel moisture maps with RMSE 18.91, about 20% better than a randomly initialized model.
-
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
Frame-level captions with per-frame cross-attention and parallel multi-window denoising reduce semantic confusion in long multi-scene video generation in the authors' internal evaluation.
-
STS-NET: Spatio-Temporal Stress Network for Self-Supervised Crop Stress Detection using Satellite Image Time Series
A self-supervised 3D convolutional autoencoder trained on unlabeled satellite vegetation-index time series reports 97.98% water, 85.08% nitrogen, and 83.47% combined stress accuracy on one sugarcane farm.
-
Predicting Viticulture Potential through an Ensemble of U-Net and a Geospatial Foundation Model
Ensemble of U-Net (temporal channels stacked) and fine-tuned Prithvi-EO-2.0 reaches 68.32 ±1 accuracy on viticulture potential, ranking 2nd of 7 in ImageCLEF AI4Agri 2026.
Discussion (0). Sign in to comment.