UNIGEOCLIP creates a unified embedding for aerial imagery, street views, elevation, text, and coordinates via all-to-all contrastive alignment plus a scaled lat-long encoder, outperforming single-modality and coordinate baselines on geospatial tasks.
General Geospatial Inference with a Population Dynamics Foundation Model
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
Supporting the health and well-being of dynamic populations around the world requires governmental agencies, organizations and researchers to understand and reason over complex relationships between human behavior and local contexts in order to identify high-risk groups and strategically allocate limited resources. Traditional approaches to these classes of problems often entail developing manually curated, task-specific features and models to represent human behavior and the natural and built environment, which can be challenging to adapt to new, or even, related tasks. To address this, we introduce a Population Dynamics Foundation Model (PDFM) that aims to capture the relationships between diverse data modalities and is applicable to a broad range of geospatial tasks. We first construct a geo-indexed dataset for postal codes and counties across the United States, capturing rich aggregated information on human behavior from maps, busyness, and aggregated search trends, and environmental factors such as weather and air quality. We then model this data and the complex relationships between locations using a graph neural network, producing embeddings that can be adapted to a wide range of downstream tasks using relatively simple models. We evaluate the effectiveness of our approach by benchmarking it on 27 downstream tasks spanning three distinct domains: health indicators, socioeconomic factors, and environmental measurements. The approach achieves state-of-the-art performance on all 27 geospatial interpolation tasks, and on 25 out of the 27 extrapolation and super-resolution tasks. We combined the PDFM with a state-of-the-art forecasting foundation model, TimesFM, to predict unemployment and poverty, achieving performance that surpasses fully supervised forecasting. The full set of embeddings and sample code are publicly available for researchers.
years
2026 5verdicts
UNVERDICTED 5representative citing papers
OSMGraphCLIP learns global location embeddings from OSM graphs via multi-scale graph encoding and contrastive alignment that match or exceed satellite baselines on many socioeconomic, health, and environmental tasks.
MobFusion fuses mobility networks into foundation models via three designs and reports improved performance on income, density, and crime prediction tasks using data from three U.S. metropolitan areas.
PDFM embeddings reduce unexplained variance in subnational population estimates by a median 20.1% versus geospatial covariates, with gains strongest in larger less-developed areas but weaker transfer across scales.
Context-conditioned normalizing flows refine subnational survey distributions under severe data scarcity when conditioning covariates capture local heterogeneity.
citing papers explorer
-
UNIGEOCLIP: Unified Geospatial Contrastive Learning
UNIGEOCLIP creates a unified embedding for aerial imagery, street views, elevation, text, and coordinates via all-to-all contrastive alignment plus a scaled lat-long encoder, outperforming single-modality and coordinate baselines on geospatial tasks.
-
OSMGraphCLIP: Learning Global Location Representations from OpenStreetMap Graphs
OSMGraphCLIP learns global location embeddings from OSM graphs via multi-scale graph encoding and contrastive alignment that match or exceed satellite baselines on many socioeconomic, health, and environmental tasks.
-
Enhancing the Socioeconomic Understanding of Foundation Models with Urban Mobility
MobFusion fuses mobility networks into foundation models via three designs and reports improved performance on income, density, and crime prediction tasks using data from three U.S. metropolitan areas.
-
Geospatial foundation-model embeddings improve population estimation unevenly across space and scale
PDFM embeddings reduce unexplained variance in subnational population estimates by a median 20.1% versus geospatial covariates, with gains strongest in larger less-developed areas but weaker transfer across scales.
-
Context-Conditioned Generative Models Enable Subnational Refinement of Sparse Humanitarian Surveys
Context-conditioned normalizing flows refine subnational survey distributions under severe data scarcity when conditioning covariates capture local heterogeneity.