A dual-scale cross-attention transformer trained on an HDBSCAN-repartitioned SF-XL dataset reports higher recall than several global and two-stage VPR methods on five benchmark datasets, with 512-dim descriptors and about 30% less training data.
Learned contextual feature reweighting for image geo-localization
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
DSFormer: A Dual-Scale Cross-Learning Transformer for Visual Place Recognition
A dual-scale cross-attention transformer trained on an HDBSCAN-repartitioned SF-XL dataset reports higher recall than several global and two-stage VPR methods on five benchmark datasets, with 512-dim descriptors and about 30% less training data.