GeoX is a self-play RL framework in which a single multimodal policy proposes and solves spatial problems as executable programs over image primitives, using verifiable rewards to improve base VLMs by up to 5.5 points without large curated data.
Earthgpt: A universal multimodal large language model for multisensor image comprehension in remote sensing domain.IEEE Transactions on Geoscience and Remote Sensing, 62:1–20
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
A dual-branch contrastive learning framework distills street-view semantics and temporal context into satellite representations, improving monthly carbon emission prediction using only satellite imagery at inference.
An audit of 152 papers reveals that geospatial foundation models lack standardized evaluations, training controls, and weight releases, so no one knows the state of the art.
citing papers explorer
-
GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards
GeoX is a self-play RL framework in which a single multimodal policy proposes and solves spatial problems as executable programs over image primitives, using verifiable rewards to improve base VLMs by up to 5.5 points without large curated data.
-
CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training
A dual-branch contrastive learning framework distills street-view semantics and temporal context into satellite representations, improving monthly carbon emission prediction using only satellite imagery at inference.
-
No One Knows the State of the Art in Geospatial Foundation Models
An audit of 152 papers reveals that geospatial foundation models lack standardized evaluations, training controls, and weight releases, so no one knows the state of the art.