A two-stage tuning method that grafts street-view images onto labeled satellite maps gives small vision-language models street-level address localization accuracy well above direct fine-tuning.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
A two-stage tuning method that grafts street-view images onto labeled satellite maps gives small vision-language models street-level address localization accuracy well above direct fine-tuning.