Pith. sign in

REVIEW 29 cited by

Foundation Models for Generalist Geospatial Artificial Intelligence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.18660 v2 pith:YVGXT7QY submitted 2023-10-28 cs.CV cs.LG

classification cs.CVcs.LG
keywords modeldatamodelspre-trainedearthfine-tuningfoundationframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled datasets through self-supervision, and then fine-tuned for various downstream tasks with small labeled datasets. This paper introduces a first-of-a-kind framework for the efficient pre-training and fine-tuning of foundational models on extensive geospatial data. We have utilized this framework to create Prithvi, a transformer-based geospatial foundational model pre-trained on more than 1TB of multispectral satellite imagery from the Harmonized Landsat-Sentinel 2 (HLS) dataset. Our study demonstrates the efficacy of our framework in successfully fine-tuning Prithvi to a range of Earth observation tasks that have not been tackled by previous work on foundation models involving multi-temporal cloud gap imputation, flood mapping, wildfire scar segmentation, and multi-temporal crop segmentation. Our experiments show that the pre-trained model accelerates the fine-tuning process compared to leveraging randomly initialized weights. In addition, pre-trained Prithvi compares well against the state-of-the-art, e.g., outperforming a conditional GAN model in multi-temporal cloud imputation by up to 5pp (or 5.7%) in the structural similarity index. Finally, due to the limited availability of labeled data in the field of Earth observation, we gradually reduce the quantity of available labeled data for refining the model to evaluate data efficiency and demonstrate that data can be decreased significantly without affecting the model's accuracy. The pre-trained 100 million parameter model and corresponding fine-tuning workflows have been released publicly as open source contributions to the global Earth sciences community through Hugging Face.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SMARTIES, a single masked-autoencoder foundation model with spectrum-aware band projections and cross-sensor token mixup, handles multiple remote sensing sensors and transfers to unseen sensors via interpolation.

  2. AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A single JEPA-based model with scale-adaptive encoders is pre-trained on five heterogeneous Earth observation datasets and reaches state-of-the-art results across nine downstream tasks.

  3. CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities

    cs.CV 2025-06 conditional novelty 6.5 of 10

    Introduces a multi-modal 100m wildfire forecasting benchmark for Canada and shows deep learning models benefit from fusing Sentinel-2 imagery with environmental predictors.

  4. HeatCast: A Benchmark for Neighborhood-Scale LST Forecasting across 124 U.S. Cities

    cs.CV 2026-08 conditional novelty 6.0 of 10

    HeatCast is a new open dataset and evaluation standard for monthly 30 m land surface temperature forecasting across 124 U.S. cities, with baseline RMSEs around 7.7 K.

  5. Multi-Year Geospatial Reasoning using Interannually-Consistent Historical Predictions as a Free Input Modality

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Feeding a crop-type model its own interannual-fixed historical predictions, encoded as confidence-scaled categorical tokens, raises crop-only F1 by 1.6 points and rebalances precision and recall.

  6. Above-ground Biomass Estimation with Geospatial Foundation Models

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A benchmark of geospatial foundation models for biomass regression shows that pre-computed embedding products, especially AlphaEarth Foundations, outperform both frozen weight-distributed GFMs and a fully supervised s...

  7. SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A fine-tuning framework that uses all available satellite bands via a residual gated adapter and allocates LoRA ranks by stage-level transferability, improving cross-sensor segmentation at lower parameter cost.

  8. Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A locality-aware embedding anomaly detector identifies label errors in global crop reference data; conservative cleaning raises WorldCereal crop-type macro-F1 in all five tested regions.

  9. Now We Know? A Systematic Comparison of TerraMind and THOR

    cs.LG 2026-07 conditional novelty 6.0 of 10

    In a controlled comparison of two geospatial foundation models, patch size and decoder type explain more of the performance difference than the choice of model itself.

  10. From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring

    cs.CV 2026-07 accept novelty 6.0 of 10

    A JEPA world model forecasts cloud-induced observation usability and recovery timing on EarthNet2021, outperforming persistence and competing with LightGBM on most splits.

  11. OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 4B model fine-tuned on tool-augmented geospatial reasoning traces outperforms larger general-purpose models on executable GIS/spectral tool-use benchmarks and matches frontier models on trajectory fidelity.

  12. The View From Space: Navigating Instrumentation Differences with EOFMs

    cs.CV 2025-10 conditional novelty 6.0 of 10

    EOFM embeddings are strongly partitioned by sensor architecture, so matching spectral bands is not enough to make cross-sensor embedding search reliable.

  13. An Open Benchmark Dataset for GeoAI Foundation Models for Oil Palm Mapping in Indonesia

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new open polygon dataset with 52,225 labeled land cover polygons for oil palm mapping in Riau and West Sulawesi, validated at 83% overall accuracy.

  14. Trees as Gaussians: Large-Scale Individual Tree Mapping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A deep learning system detects individual large trees globally in 3 m PlanetScope imagery by regressing Gaussian heatmaps trained on 14 billion lidar-derived pseudo-labels.

  15. Fine-Scale Soil Mapping in Alaska with Multimodal Machine Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multimodal deep learning system produces 10 m Alaska soil and permafrost maps and finds permafrost more sensitively than random forest under spatial holdout.

  16. VME: A Satellite Imagery Dataset and Benchmark for Detecting Vehicles in the Middle East and Beyond

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new Middle East vehicle-detection dataset (VME) and a combined global benchmark (CDSI) show that existing satellite-imagery detectors underperform in the Middle East, and that region-specific training data closes mo...

  17. Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Large language models, especially GPT-4 with few-shot prompts, can classify topological spatial relations between WKT-encoded geometries with roughly 0.6 to 0.66 accuracy, though errors cluster near conceptually simil...

  18. How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Balanced, globally representative pre-training data generally outperforms region-specific sampling for two geospatial foundation models in few-shot downstream tasks, and the advantage shrinks as finetuning data grows.

  19. WildSAT: Learning Satellite Image Representations from Wildlife Observations

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Aligning Sentinel-2 images with geotagged wildlife observations and Wikipedia species text via contrastive learning improves satellite encoders on downstream classification and enables zero-shot text-based image retrieval.

  20. Improving Satellite Imagery Masking using Multi-task and Transfer Learning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single multi-task neural network predicts all five masks needed for global river sediment monitoring from satellite images, cutting runtime 30x and improving water and cloud shadow F1 over prior methods.

  21. How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Remote sensing foundation models tolerate aggressive post-training width reduction (roughly 70–100% relative accuracy at 1% compute), which the authors attribute to redundant rather than sparse feature encoding.

  22. UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations

    cs.LG 2025-10 conditional novelty 5.0 of 10

    UrbanFusion trains a multimodal location encoder with stochastic fusion (contrastive alignment plus latent reconstruction under random modality masking) and edges out prior GeoAI models on most of 41 urban tasks, with...

  23. Finetuning AI Foundation Models to Develop Subgrid-Scale Parameterizations: A Case Study on Atmospheric Gravity Waves

    physics.ao-ph 2025-09 conditional novelty 5.0 of 10

    Fine-tuning Prithvi WxC to predict gravity wave fluxes beats an Attention U-Net baseline on one month of ERA5 validation, including in upper stratospheric levels absent from the pre-training data.

  24. High-Resolution Live Fuel Moisture Content (LFMC) Maps for Wildfire Risk from Multimodal Earth Observation Data

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fine-tuning the pretrained Galileo model on the Globe-LFMC dataset yields 10 m wall-to-wall live fuel moisture maps with RMSE 18.91, about 20% better than a randomly initialized model.

  25. Geospatial Foundation Models to Enable Progress on Sustainable Development Goals

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A benchmark of 16 satellite-imaging tasks mapped to the UN Sustainable Development Goals shows geospatial foundation models often beat scratch-trained networks, though not always, and that energy use should be part of...

  26. Deep Self-Supervised Disturbance Mapping with the OPERA Sentinel-1 Radiometric Terrain Corrected SAR Backscatter Product

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A label-free vision transformer trained on OPERA RTC-S1 radar backscatter delineates landslide, wildfire, and flood damage with F1 scores above 0.6 on three test events.

  27. Defining Foundation Models for Computational Science: A Call for Clarity and Rigor

    cs.LG 2025-05 conditional novelty 4.0 of 10

    The paper defines foundation models for computational science and presents DD-FEM, a local-to-global data-driven framework inspired by finite elements, as a candidate path to meet that definition.

  28. Parameter-Efficient Fine-Tuning of Multispectral Foundation Models for Hyperspectral Image Classification

    cs.CV 2025-05 conditional novelty 4.0 of 10

    KronA+ fine-tunes SpectralGPT for hyperspectral image classification using only 0.056% trainable parameters and reaches accuracy close to full fine-tuning on five public datasets.

  29. MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A ViT-based MultiMAE pre-trained on MMEarth with split Sentinel-2 bands, elevation, and segmentation labels transfers to several EO classification and segmentation datasets.

Pith tools