MVL-Loc fuses CLIP image and text features with per-scene pose heads and reports state-of-the-art multi-scene relocalization accuracy on 7Scenes and Cambridge Landmarks.
Image-based localization using lstms for structured feature correlation,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization
MVL-Loc fuses CLIP image and text features with per-scene pose heads and reports state-of-the-art multi-scene relocalization accuracy on 7Scenes and Cambridge Landmarks.