ViTCG, a channel-grouped Vision Transformer, retrieves AOD from PACE hyperspectral data with 62% lower MSE than prior foundation models while producing spatially coherent fields.
hub
org/abs/2310.18660
19 Pith papers cite this work, alongside 13 external citations. Polarity classification is still indexing.
hub tools
representative citing papers
SpectralEarth-FM is a multisensor hierarchical transformer pretrained on a 40TB co-located HSI-MSI-SAR dataset using a JEPA-style objective and reports state-of-the-art results on hyperspectral and standard EO benchmarks.
GAIR introduces a geo-aligned implicit representation module inside a multi-encoder contrastive SSL framework that produces location-aware embeddings and outperforms prior geo-foundation models on 22 geospatial datasets across 9 tasks.
Benchmark of Prithvi, SpectralGPT, and SatMAE shows sharp performance drop under regional distribution shift, with models defaulting to common crops and missing rare ones.
AlphaEarth land-cover priors improve SAR flood segmentation IoU over SAR-only and DEM baselines across CNN and ViT backbones on held-out events like Hurricane Florence.
Prithvi-2.0-UPN achieves SOTA flood mapping on RGB datasets and superior zero-shot transfer plus rapid adaptation to new events versus baselines.
Lite ViT-Adapter with LoRA on Prithvi-EO reaches mAP@50 of 0.9479 for fallow detection, improving the baseline adapter-free approach by 25.70%.
Prithvi-EO-2.0 shows environment-dependent flood detection limits, with highest accuracy in cropland (IoU 52%) and riverine events (F1 0.69) and near-zero performance in tree cover and built-up areas across 19 global events.
Affinity-propagation clustering of Arctic VHSR imagery enables MAE pretraining of a ViT-Large encoder that outperforms ImageNet and Prithvi-EO-2.0 baselines by 5-15 percentage points in mean F1 on four downstream Arctic detection and segmentation tasks.
A fleet of sensor-specialized 22M-parameter JEPA models routed by an LLM improves LLM-as-judge scores on hydrologic questions over AlphaEarth alone with Cohen's d of 1.10.
WATCH localizes archaeological site changes to the month using temporal embedding distances and self-supervised signals on PlanetScope data, with TED achieving 55% exact-month recall and 92.5% within three months on Afghan sites.
Strong generalist vision foundation models match or outperform electro-optical specific models in remote sensing retrieval with better cross-scene stability.
European pretraining data outperforms global and other regional datasets for geospatial foundation models, with only spectral diversity showing strong correlation to downstream performance.
A prompting-based adaptation technique lets RGB-trained LMMs process multi-spectral inputs and deliver strong zero-shot gains on remote-sensing benchmarks.
AlphaEarth embeddings form a rotating 13-dimensional manifold where local geometry predicts retrieval quality, and an agentic system using nine geometric tools outperforms parametric reasoning on environmental queries.
LIANet encodes multi-temporal Earth observation data into a coordinate-based neural field that supports label-only fine-tuning for downstream tasks without access to raw imagery.
SHRUG-FM fuses geophysical OOD detection, embedding-space OOD detection, and predictive uncertainty via a shallow decision tree to let foundation models abstain from unreliable outputs on burn scar, flood, and landslide tasks.
Foundation model embeddings provide no advantage over traditional spectral features for cross-country maize yield generalization in Africa, with all methods yielding negative R² under leave-one-country-out testing due to distribution shifts.
SAM achieves ~58% accuracy delineating field boundaries from SkySat imagery without training, with gains from multi-date inputs and varied sizes, establishing proof-of-concept for data-scarce agriculture mapping.
citing papers explorer
-
Foundation AI Models for Aerosol Optical Depth Estimation from PACE Satellite Data
ViTCG, a channel-grouped Vision Transformer, retrieves AOD from PACE hyperspectral data with 62% lower MSE than prior foundation models while producing spatially coherent fields.
-
SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining
SpectralEarth-FM is a multisensor hierarchical transformer pretrained on a 40TB co-located HSI-MSI-SAR dataset using a JEPA-style objective and reports state-of-the-art results on hyperspectral and standard EO benchmarks.
-
GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations
GAIR introduces a geo-aligned implicit representation module inside a multi-encoder contrastive SSL framework that produces location-aware embeddings and outperforms prior geo-foundation models on 22 geospatial datasets across 9 tasks.
-
Benchmarking Geospatial Foundation Models for Agriculture Applications
Benchmark of Prithvi, SpectralGPT, and SatMAE shows sharp performance drop under regional distribution shift, with models defaulting to common crops and missing rare ones.
-
Beyond Backscatter: AlphaEarth Land-Cover Priors for Rapid SAR Flood Segmentation Across Foundation Backbones
AlphaEarth land-cover priors improve SAR flood segmentation IoU over SAR-only and DEM baselines across CNN and ViT backbones on held-out events like Hurricane Florence.
-
Flood Mapping from RGB imagery using a Vision Foundation Model
Prithvi-2.0-UPN achieves SOTA flood mapping on RGB datasets and superior zero-shot transfer plus rapid adaptation to new events versus baselines.
-
Adapting Prithvi-EO for Fallow Detection for Food-Water Nexus: ViT-Adapter Necks and Parameter-Efficient Backbone tuning of Geospatial Foundation Model
Lite ViT-Adapter with LoRA on Prithvi-EO reaches mAP@50 of 0.9479 for fallow detection, improving the baseline adapter-free approach by 25.70%.
-
Land cover and flood type govern the detection limits of satellite-based flood mapping across diverse global flood events
Prithvi-EO-2.0 shows environment-dependent flood detection limits, with highest accuracy in cropland (IoU 52%) and riverine events (F1 0.69) and near-zero performance in tree cover and built-up areas across 19 global events.
-
Clustering Guided Domain-Specific Pretrained Foundation Model for Very High-Resolution Arctic Remote Sensing
Affinity-propagation clustering of Arctic VHSR imagery enables MAE pretraining of a ViT-Large encoder that outperforms ImageNet and Prithvi-EO-2.0 baselines by 5-15 percentage points in mean F1 on four downstream Arctic detection and segmentation tasks.
-
Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence
A fleet of sensor-specialized 22M-parameter JEPA models routed by an LLM improves LLM-as-judge scores on hydrologic questions over AlphaEarth alone with Cohen's d of 1.10.
-
WATCH: Wide-Area Archaeological Site Tracking for Change Detection
WATCH localizes archaeological site changes to the month using temporal embedding distances and self-supervised signals on PlanetScope data, with TED achieving 55% exact-month recall and 92.5% within three months on Afghan sites.
-
Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM
Strong generalist vision foundation models match or outperform electro-optical specific models in remote sensing retrieval with better cross-scene stability.
-
Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance
European pretraining data outperforms global and other regional datasets for geospatial foundation models, with only spectral diversity showing strong correlation to downstream performance.
-
Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
A prompting-based adaptation technique lets RGB-trained LMMs process multi-spectral inputs and deliver strong zero-shot gains on remote-sensing benchmarks.
-
Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning
AlphaEarth embeddings form a rotating 13-dimensional manifold where local geometry predicts retrieval quality, and an agentic system using nine geometric tools outperforms parametric reasoning on environmental queries.
-
Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data
LIANet encodes multi-temporal Earth observation data into a coordinate-based neural field that supports label-only fine-tuning for downstream tasks without access to raw imagery.
-
SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation
SHRUG-FM fuses geophysical OOD detection, embedding-space OOD detection, and predictive uncertainty via a shallow decision tree to let foundation models abstain from unreliable outputs on burn scar, flood, and landslide tasks.
-
Do Foundation Model Embeddings Improve Cross-Country Crop Yield Generalisation? A Leave-One-Country-Out Evaluation in Sub-Saharan Africa
Foundation model embeddings provide no advantage over traditional spectral features for cross-country maize yield generalization in Africa, with all methods yielding negative R² under leave-one-country-out testing due to distribution shifts.
-
Investigating the Segment Anything Foundation Model for Mapping Smallholder Agriculture Field Boundaries Without Training Labels
SAM achieves ~58% accuracy delineating field boundaries from SkySat imagery without training, with gains from multi-date inputs and varied sizes, establishing proof-of-concept for data-scarce agriculture mapping.