REVIEW 3 cited by
Where am I? Cross-View Geo-localization with Natural Language Descriptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ .
Forward citations
Cited by 3 Pith papers
-
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
An unmodified general-purpose VLM trained with multi-task RL and a SAM3 tool reaches top results on most remote sensing zero-shot benchmarks, with gains the paper attributes to training-data diversity rather than arch...
-
Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework
A concept bottleneck projecting images and GPS into a subspace of geographic concepts improves GeoCLIP from 10.8 to 13.2 percent top-1 km accuracy on Im2GPS3k and adds semantic explanations.
-
VICI: VLM-Instructed Cross-view Image-localisation
A retrieval-plus-VLM-reranking pipeline with drone-image augmentation improves street-to-satellite matching accuracy on limited-FOV queries.
Discussion (0). Sign in to comment.