Pith. sign in

REVIEW 16 cited by

Foundation Models for Generalist Geospatial Artificial Intelligence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.18660 v2 pith:YVGXT7QY submitted 2023-10-28 cs.CV cs.LG

classification cs.CVcs.LG
keywords modeldatamodelspre-trainedearthfine-tuningfoundationframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled datasets through self-supervision, and then fine-tuned for various downstream tasks with small labeled datasets. This paper introduces a first-of-a-kind framework for the efficient pre-training and fine-tuning of foundational models on extensive geospatial data. We have utilized this framework to create Prithvi, a transformer-based geospatial foundational model pre-trained on more than 1TB of multispectral satellite imagery from the Harmonized Landsat-Sentinel 2 (HLS) dataset. Our study demonstrates the efficacy of our framework in successfully fine-tuning Prithvi to a range of Earth observation tasks that have not been tackled by previous work on foundation models involving multi-temporal cloud gap imputation, flood mapping, wildfire scar segmentation, and multi-temporal crop segmentation. Our experiments show that the pre-trained model accelerates the fine-tuning process compared to leveraging randomly initialized weights. In addition, pre-trained Prithvi compares well against the state-of-the-art, e.g., outperforming a conditional GAN model in multi-temporal cloud imputation by up to 5pp (or 5.7%) in the structural similarity index. Finally, due to the limited availability of labeled data in the field of Earth observation, we gradually reduce the quantity of available labeled data for refining the model to evaluate data efficiency and demonstrate that data can be decreased significantly without affecting the model's accuracy. The pre-trained 100 million parameter model and corresponding fine-tuning workflows have been released publicly as open source contributions to the global Earth sciences community through Hugging Face.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SMARTIES, a single masked-autoencoder foundation model with spectrum-aware band projections and cross-sensor token mixup, handles multiple remote sensing sensors and transfers to unseen sensors via interpolation.

  2. CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities

    cs.CV 2025-06 conditional novelty 6.5 of 10

    Introduces a multi-modal 100m wildfire forecasting benchmark for Canada and shows deep learning models benefit from fusing Sentinel-2 imagery with environmental predictors.

  3. Above-ground Biomass Estimation with Geospatial Foundation Models

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A benchmark of geospatial foundation models for biomass regression shows that pre-computed embedding products, especially AlphaEarth Foundations, outperform both frozen weight-distributed GFMs and a fully supervised s...

  4. SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A fine-tuning framework that uses all available satellite bands via a residual gated adapter and allocates LoRA ranks by stage-level transferability, improving cross-sensor segmentation at lower parameter cost.

  5. Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A locality-aware embedding anomaly detector identifies label errors in global crop reference data; conservative cleaning raises WorldCereal crop-type macro-F1 in all five tested regions.

  6. Now We Know? A Systematic Comparison of TerraMind and THOR

    cs.LG 2026-07 conditional novelty 6.0 of 10

    In a controlled comparison of two geospatial foundation models, patch size and decoder type explain more of the performance difference than the choice of model itself.

  7. From Surface Forecasting to Observability Forecasting: A Latent World Model for Cloud-Aware EO Monitoring

    cs.CV 2026-07 accept novelty 6.0 of 10

    A JEPA world model forecasts cloud-induced observation usability and recovery timing on EarthNet2021, outperforming persistence and competing with LightGBM on most splits.

  8. OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 4B model fine-tuned on tool-augmented geospatial reasoning traces outperforms larger general-purpose models on executable GIS/spectral tool-use benchmarks and matches frontier models on trajectory fidelity.

  9. The View From Space: Navigating Instrumentation Differences with EOFMs

    cs.CV 2025-10 conditional novelty 6.0 of 10

    EOFM embeddings are strongly partitioned by sensor architecture, so matching spectral bands is not enough to make cross-sensor embedding search reliable.

  10. An Open Benchmark Dataset for GeoAI Foundation Models for Oil Palm Mapping in Indonesia

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new open polygon dataset with 52,225 labeled land cover polygons for oil palm mapping in Riau and West Sulawesi, validated at 83% overall accuracy.

  11. Trees as Gaussians: Large-Scale Individual Tree Mapping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A deep learning system detects individual large trees globally in 3 m PlanetScope imagery by regressing Gaussian heatmaps trained on 14 billion lidar-derived pseudo-labels.

  12. Fine-Scale Soil Mapping in Alaska with Multimodal Machine Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multimodal deep learning system produces 10 m Alaska soil and permafrost maps and finds permafrost more sensitively than random forest under spatial holdout.

  13. How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Remote sensing foundation models tolerate aggressive post-training width reduction (roughly 70–100% relative accuracy at 1% compute), which the authors attribute to redundant rather than sparse feature encoding.

  14. UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations

    cs.LG 2025-10 conditional novelty 5.0 of 10

    UrbanFusion trains a multimodal location encoder with stochastic fusion (contrastive alignment plus latent reconstruction under random modality masking) and edges out prior GeoAI models on most of 41 urban tasks, with...

  15. Finetuning AI Foundation Models to Develop Subgrid-Scale Parameterizations: A Case Study on Atmospheric Gravity Waves

    physics.ao-ph 2025-09 conditional novelty 5.0 of 10

    Fine-tuning Prithvi WxC to predict gravity wave fluxes beats an Attention U-Net baseline on one month of ERA5 validation, including in upper stratospheric levels absent from the pre-training data.

  16. High-Resolution Live Fuel Moisture Content (LFMC) Maps for Wildfire Risk from Multimodal Earth Observation Data

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fine-tuning the pretrained Galileo model on the Globe-LFMC dataset yields 10 m wall-to-wall live fuel moisture maps with RMSE 18.91, about 20% better than a randomly initialized model.

Pith tools