Across four benchmark datasets and 25 vision-language models, GPT-4.1 is the most accurate geolocator, reaching 61% Recall@1km on social-media-like images while all models struggle on street-level imagery.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Assessing the Geolocation Capabilities, Limitations and Societal Risks of Generative Vision-Language Models
Across four benchmark datasets and 25 vision-language models, GPT-4.1 is the most accurate geolocator, reaching 61% Recall@1km on social-media-like images while all models struggle on street-level imagery.