REVIEW 29 cited by
Foundation Models for Generalist Geospatial Artificial Intelligence
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled datasets through self-supervision, and then fine-tuned for various downstream tasks with small labeled datasets. This paper introduces a first-of-a-kind framework for the efficient pre-training and fine-tuning of foundational models on extensive geospatial data. We have utilized this framework to create Prithvi, a transformer-based geospatial foundational model pre-trained on more than 1TB of multispectral satellite imagery from the Harmonized Landsat-Sentinel 2 (HLS) dataset. Our study demonstrates the efficacy of our framework in successfully fine-tuning Prithvi to a range of Earth observation tasks that have not been tackled by previous work on foundation models involving multi-temporal cloud gap imputation, flood mapping, wildfire scar segmentation, and multi-temporal crop segmentation. Our experiments show that the pre-trained model accelerates the fine-tuning process compared to leveraging randomly initialized weights. In addition, pre-trained Prithvi compares well against the state-of-the-art, e.g., outperforming a conditional GAN model in multi-temporal cloud imputation by up to 5pp (or 5.7%) in the structural similarity index. Finally, due to the limited availability of labeled data in the field of Earth observation, we gradually reduce the quantity of available labeled data for refining the model to evaluate data efficiency and demonstrate that data can be decreased significantly without affecting the model's accuracy. The pre-trained 100 million parameter model and corresponding fine-tuning workflows have been released publicly as open source contributions to the global Earth sciences community through Hugging Face.
Forward citations
Cited by 29 Pith papers
-
SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
SMARTIES, a single masked-autoencoder foundation model with spectrum-aware band projections and cross-sensor token mixup, handles multiple remote sensing sensors and transfers to unseen sensors via interpolation.
-
AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
A single JEPA-based model with scale-adaptive encoders is pre-trained on five heterogeneous Earth observation datasets and reaches state-of-the-art results across nine downstream tasks.
-
CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities
Introduces a multi-modal 100m wildfire forecasting benchmark for Canada and shows deep learning models benefit from fusing Sentinel-2 imagery with environmental predictors.
-
HeatCast: A Benchmark for Neighborhood-Scale LST Forecasting across 124 U.S. Cities
HeatCast is a new open dataset and evaluation standard for monthly 30 m land surface temperature forecasting across 124 U.S. cities, with baseline RMSEs around 7.7 K.
-
Multi-Year Geospatial Reasoning using Interannually-Consistent Historical Predictions as a Free Input Modality
Feeding a crop-type model its own interannual-fixed historical predictions, encoded as confidence-scaled categorical tokens, raises crop-only F1 by 1.6 points and rebalances precision and recall.
-
Above-ground Biomass Estimation with Geospatial Foundation Models
A benchmark of geospatial foundation models for biomass regression shows that pre-computed embedding products, especially AlphaEarth Foundations, outperform both frozen weight-distributed GFMs and a fully supervised s...
-
SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models
A fine-tuning framework that uses all available satellite bands via a residual gated adapter and allocates LoRA ranks by stage-level transferability, improving cross-sensor segmentation at lower parameter cost.
-
Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets
A locality-aware embedding anomaly detector identifies label errors in global crop reference data; conservative cleaning raises WorldCereal crop-type macro-F1 in all five tested regions.
-
Now We Know? A Systematic Comparison of TerraMind and THOR
In a controlled comparison of two geospatial foundation models, patch size and decoder type explain more of the performance difference than the choice of model itself.
-
From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring
A JEPA world model forecasts cloud-induced observation usability and recovery timing on EarthNet2021, outperforming persistence and competing with LightGBM on most splits.
-
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
A 4B model fine-tuned on tool-augmented geospatial reasoning traces outperforms larger general-purpose models on executable GIS/spectral tool-use benchmarks and matches frontier models on trajectory fidelity.
-
The View From Space: Navigating Instrumentation Differences with EOFMs
EOFM embeddings are strongly partitioned by sensor architecture, so matching spectral bands is not enough to make cross-sensor embedding search reliable.
-
An Open Benchmark Dataset for GeoAI Foundation Models for Oil Palm Mapping in Indonesia
A new open polygon dataset with 52,225 labeled land cover polygons for oil palm mapping in Riau and West Sulawesi, validated at 83% overall accuracy.
-
Trees as Gaussians: Large-Scale Individual Tree Mapping
A deep learning system detects individual large trees globally in 3 m PlanetScope imagery by regressing Gaussian heatmaps trained on 14 billion lidar-derived pseudo-labels.
-
Fine-Scale Soil Mapping in Alaska with Multimodal Machine Learning
A multimodal deep learning system produces 10 m Alaska soil and permafrost maps and finds permafrost more sensitively than random forest under spatial holdout.
-
VME: A Satellite Imagery Dataset and Benchmark for Detecting Vehicles in the Middle East and Beyond
A new Middle East vehicle-detection dataset (VME) and a combined global benchmark (CDSI) show that existing satellite-imagery detectors underperform in the Middle East, and that region-specific training data closes mo...
-
Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
Large language models, especially GPT-4 with few-shot prompts, can classify topological spatial relations between WKT-encoded geometries with roughly 0.6 to 0.66 accuracy, though errors cluster near conceptually simil...
-
How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?
Balanced, globally representative pre-training data generally outperforms region-specific sampling for two geospatial foundation models in few-shot downstream tasks, and the advantage shrinks as finetuning data grows.
-
WildSAT: Learning Satellite Image Representations from Wildlife Observations
Aligning Sentinel-2 images with geotagged wildlife observations and Wikipedia species text via contrastive learning improves satellite encoders on downstream classification and enables zero-shot text-based image retrieval.
-
Improving Satellite Imagery Masking using Multi-task and Transfer Learning
A single multi-task neural network predicts all five masks needed for global river sediment monitoring from satellite images, cutting runtime 30x and improving water and cloud shadow F1 over prior methods.
-
How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models
Remote sensing foundation models tolerate aggressive post-training width reduction (roughly 70–100% relative accuracy at 1% compute), which the authors attribute to redundant rather than sparse feature encoding.
-
UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations
UrbanFusion trains a multimodal location encoder with stochastic fusion (contrastive alignment plus latent reconstruction under random modality masking) and edges out prior GeoAI models on most of 41 urban tasks, with...
-
Finetuning AI Foundation Models to Develop Subgrid-Scale Parameterizations: A Case Study on Atmospheric Gravity Waves
Fine-tuning Prithvi WxC to predict gravity wave fluxes beats an Attention U-Net baseline on one month of ERA5 validation, including in upper stratospheric levels absent from the pre-training data.
-
High-Resolution Live Fuel Moisture Content (LFMC) Maps for Wildfire Risk from Multimodal Earth Observation Data
Fine-tuning the pretrained Galileo model on the Globe-LFMC dataset yields 10 m wall-to-wall live fuel moisture maps with RMSE 18.91, about 20% better than a randomly initialized model.
-
Geospatial Foundation Models to Enable Progress on Sustainable Development Goals
A benchmark of 16 satellite-imaging tasks mapped to the UN Sustainable Development Goals shows geospatial foundation models often beat scratch-trained networks, though not always, and that energy use should be part of...
-
Deep Self-Supervised Disturbance Mapping with the OPERA Sentinel-1 Radiometric Terrain Corrected SAR Backscatter Product
A label-free vision transformer trained on OPERA RTC-S1 radar backscatter delineates landslide, wildfire, and flood damage with F1 scores above 0.6 on three test events.
-
Defining Foundation Models for Computational Science: A Call for Clarity and Rigor
The paper defines foundation models for computational science and presents DD-FEM, a local-to-global data-driven framework inspired by finite elements, as a candidate path to meet that definition.
-
Parameter-Efficient Fine-Tuning of Multispectral Foundation Models for Hyperspectral Image Classification
KronA+ fine-tunes SpectralGPT for hyperspectral image classification using only 0.056% trainable parameters and reaches accuracy close to full fine-tuning on five public datasets.
-
MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks
A ViT-based MultiMAE pre-trained on MMEarth with split Sentinel-2 bands, elevation, and segmentation labels transfers to several EO classification and segmentation datasets.
Discussion (0). Continue with ORCID to comment.