GeoSkill lets vision-language models improve geolocation accuracy and reasoning by maintaining an evolving Skill-Graph that grows through autonomous analysis of successful and failed rollouts on web-scale image data.
arXiv preprint arXiv:2406.18572 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Expert-guided VLMs produce accessibility ratings from street-view images that show negative correlation and distributional similarity with GPS-derived wheelchair dwell times as a mobility-friction proxy.
citing papers explorer
-
Skill-Conditioned Visual Geolocation for Vision-Language Models
GeoSkill lets vision-language models improve geolocation accuracy and reasoning by maintaining an evolving Skill-Graph that grows through autonomous analysis of successful and failed rollouts on web-scale image data.
-
Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View
Expert-guided VLMs produce accessibility ratings from street-view images that show negative correlation and distributional similarity with GPS-derived wheelchair dwell times as a mobility-friction proxy.