Pith. sign in

AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In this study, we introduce a new problem raised by social media and photojournalism, named Image Address Localization (IAL), which aims to predict the readable textual address where an image was taken. Existing two-stage approaches involve predicting geographical coordinates and converting them into human-readable addresses, which can lead to ambiguity and be resource-intensive. In contrast, we propose an end-to-end framework named AddressCLIP to solve the problem with more semantics, consisting of two key ingredients: i) image-text alignment to align images with addresses and scene captions by contrastive learning, and ii) image-geography matching to constrain image features with the spatial distance in terms of manifold learning. Additionally, we have built three datasets from Pittsburgh and San Francisco on different scales specifically for the IAL problem. Experiments demonstrate that our approach achieves compelling performance on the proposed datasets and outperforms representative transfer learning methods for vision-language models. Furthermore, extensive ablations and visualizations exhibit the effectiveness of the proposed method. The datasets and source code are available at https://github.com/xsx1001/AddressCLIP.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

GPS as a Control Signal for Image Generation

cs.CV · 2025-01-21 · conditional · novelty 7.0

A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.

citing papers explorer

Showing 1 of 1 citing paper.

  • GPS as a Control Signal for Image Generation cs.CV · 2025-01-21 · conditional · none · ref 96 · internal anchor

    A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.