REVIEW 5 cited by
UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Urban socioeconomic indicator prediction aims to infer various metrics related to sustainable development in diverse urban landscapes using data-driven methods. However, prevalent pretrained models, particularly those reliant on satellite imagery, face dual challenges. Firstly, concentrating solely on macro-level patterns from satellite data may introduce bias, lacking nuanced details at micro levels, such as architectural details at a place. Secondly, the text generated by the precursor work UrbanCLIP, which fully utilizes the extensive knowledge of LLMs, frequently exhibits issues such as hallucination and homogenization, resulting in a lack of reliable quality. In response to these issues, we devise a novel framework entitled UrbanVLP based on Vision-Language Pretraining. Our UrbanVLP seamlessly integrates multi-granularity information from both macro (satellite) and micro (street-view) levels, overcoming the limitations of prior pretrained models. Moreover, it introduces automatic text generation and calibration, providing a robust guarantee for producing high-quality text descriptions of urban imagery. Rigorous experiments conducted across six socioeconomic indicator prediction tasks underscore its superior performance.
Forward citations
Cited by 5 Pith papers
-
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
A two-stage tuning method that grafts street-view images onto labeled satellite maps gives small vision-language models street-level address localization accuracy well above direct fine-tuning.
-
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
A fine-tuned small multimodal LLM outperforms much larger general models on urban tasks in a new benchmark, with caveats about benchmark overlap with training data.
-
Satellites Reveal Mobility: A Commuting Origin-destination Flow Generator for Global Cities
Satellite imagery plus population is enough to generate commuting origin-destination flows that closely match models using detailed sociodemographic and point-of-interest data.
-
AirRadar: Inferring Nationwide Air Quality in China with Deep Neural Networks
A deep network with mask tokens, local and global spatial learners, and learned context weights infers PM2.5 at unmonitored locations across China with reported MAE of 6.41 to 8.11 at 25% to 75% missing stations.
-
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.
Discussion (0). Continue with ORCID to comment.