REVIEW 124 references
Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A survey of visual, LiDAR, and cross-modal place recognition with a unified code library, but riddled with errors and disclaimer-ridden experimental comparisons.
desk verdict A useful taxonomy but an unreliable benchmark; the survey's central numbers are self-disclaimed and the paper needs major correction before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The authors also built a GitHub repository that merges many public place recognition implementations into one code base and ran the methods themselves to produce comparison tables. This is a useful idea: a single benchmark with consistent code would let researchers compare methods fairly.
However, the paper has serious quality problems. Several tables are mislabeled, some dataset names appear to be wrong, the formula for Recall@N is garbled, and the same figure is presented twice. Most importantly, the experimental section repeatedly says that some results were obtained by running the code with settings that may not match the original papers, so the numbers may be inconsistent. This makes the benchmark conclusions unreliable. The survey also claims to be the first to cover all three modalities, despite citing earlier surveys that cover parts of them, without showing a detailed comparison. Because a survey's value depends on accuracy and completeness, these issues undermine the paper's central purpose.
Extended reading notes
Core claim
The paper's central claim, stated in the introduction, is: "To the best of our knowledge, this survey represents the first comprehensive survey of place recognition methods, including VPR, LPR, and CMPR methods, developed over the past decade," and that "most publicly accessible place recognition methods are merged into a single code base for the first time." If true, the paper would be a one-stop reference for the field and a reproducible comparison of SOTA methods.
Load-bearing premise
The paper's experimental comparison, which is one of its two advertised contributions, assumes that the authors' runs of each public codebase, with their chosen parameter settings, produce results that are comparable to those reported in the original papers. This assumption is load-bearing because the survey uses the resulting Recall@N tables to rank methods and draw conclusions about which architectures are superior. The paper itself repeatedly undermines this assumption: Sections 5.1.1, 5.1.2, 5.1.3, 5.2.1, and 5.2.2 state that "some parameter settings may not be consistent with the original paper, inconsistent results may occur." Without a fixed protocol and commit-pinned code, the benchmark conclusions are not reproducible.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
assumptions (3)
- standard math Precision, Recall (Eq. 1) and Recall@N (Eq. 2) as defined in Section 4.2 are the standard and sufficient evaluation metrics for place recognition.
- domain assumption The benchmark datasets listed in Table 11 are correctly named and described.
- domain assumption The experimental runs reported in Section 5 faithfully represent the performance of each cited method.
Cite this review
Pith. "Pith review of Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions." pith.science (2026). https://pith.science/paper/HI5JLXY4
@misc{pith2026250514068,
author = {Pith},
title = {Pith review of: Place Recognition Meet Multiple Modalitie: A Comprehensive Review, Current Challenges and Future Directions},
year = {2026},
howpublished = {\url{https://pith.science/paper/HI5JLXY4}},
note = {Machine review of arXiv:2505.14068}
}
read the original abstract
Place recognition is a cornerstone of vehicle navigation and mapping, which is pivotal in enabling systems to determine whether a location has been previously visited. This capability is critical for tasks such as loop closure in Simultaneous Localization and Mapping (SLAM) and long-term navigation under varying environmental conditions. In this survey, we comprehensively review recent advancements in place recognition, emphasizing three representative methodological paradigms: Convolutional Neural Network (CNN)-based approaches, Transformer-based frameworks, and cross-modal strategies. We begin by elucidating the significance of place recognition within the broader context of autonomous systems. Subsequently, we trace the evolution of CNN-based methods, highlighting their contributions to robust visual descriptor learning and scalability in large-scale environments. We then examine the emerging class of Transformer-based models, which leverage self-attention mechanisms to capture global dependencies and offer improved generalization across diverse scenes. Furthermore, we discuss cross-modal approaches that integrate heterogeneous data sources such as Lidar, vision, and text description, thereby enhancing resilience to viewpoint, illumination, and seasonal variations. We also summarize standard datasets and evaluation metrics widely adopted in the literature. Finally, we identify current research challenges and outline prospective directions, including domain adaptation, real-time performance, and lifelong learning, to inspire future advancements in this domain. The unified framework of leading-edge place recognition methods, i.e., code library, and the results of their experimental evaluations are available at https://github.com/CV4RA/SOTA-Place-Recognitioner.
Figures
Figures from the paper (21 more)
Reference graph
Works this paper leans on
-
[1]
B. Fang, G. Mei, X. Yuan, L. Wang, Z. Wang, J. Wang, Visual slam for robot navigation in healthcare facility, Pattern recognition 113 (2021) 107822. 53
2021
-
[2]
K. A. Tsintotas, L. Bampis, A. Gast eratos, The revisiting problem in si- multaneous localization and mapping: A survey on visual loop closure de- tection, IEEE Transactions on Intelligent Transportation Systems 23 (11) (2022) 19929–19953
2022
-
[3]
I. A. Kazerouni, L. Fitzgerald, G. Dooly, D. Toal, A survey of state-of-the- art on visual slam, Expert Systems with Applications 205 (2022) 117734
2022
-
[4]
Lauri, D
M. Lauri, D. Hsu, J. Pajarinen, Partially observable markov decision pro- cesses in robotics: A survey, IEEE Transactions on Robotics 39 (1) (2022) 21–40
2022
-
[5]
G u o , F
H . G u o , F . W u , Y . Q i n , R . L i , K . L i , K . L i , R e c e n t t r e n d s i n t a s k a n d motion planning for robotics: A surv ey, ACM Computing Surveys 55 (13s) (2023) 1–36
2023
- [6]
- [7]
-
[8]
Z. Li, P. Xu, Cspformer: A cross-sp atial pyramid transformer for visual place recognition, Neurocomputing 580 (2024) 127472
2024
Show all 124 references
-
[9]
M. A. Khan, H. E. Sayed, S. Malik, T. Zia, J. Khan, N. Alkaabi, H. Igna- tious, Level-5 autonomous driving—are we there yet? a review of research literature, ACM Computing Surveys (CSUR) 55 (2) (2022) 1–38
2022
-
[10]
Y. D. Yasuda, L. E. G. Martins, F. A. Cappabianco, Autonomous vi- sual navigation for mobile robots: A systematic literature review, ACM Computing Surveys (CSUR) 53 (1) (2020) 1–34. 54
2020
-
[11]
Munoz-Salinas, M
R. Munoz-Salinas, M. J. Marin-Jime nez, R. Medina-Carnicer, Spm-slam: Simultaneous localization and mapping with squared planar markers, Pat- tern Recognition 86 (2019) 156–171
2019
-
[12]
Nahavandi, R
S. Nahavandi, R. Alizadehsani, D. Nahavandi, S. Mohamed, N. Mohajer, M. Rokonuzzaman, I. Hossain, A comp rehensive review on autonomous navigation, ACM Computing Surveys (2022)
2022
-
[13]
F. Gu, X. Hu, M. Ramezani, D. Acharya, K. Khoshelham, S. Valaee, J. Shang, Indoor localization improved by spatial context—a survey, ACM Computing Surveys (CSUR) 52 (3) (2019) 1–35
2019
-
[14]
Cruz-Mota, I
J. Cruz-Mota, I. Bogdanova, B. Paquier, M. Bierlaire, J.-P. Thiran, Scale invariant feature transform on the sphere: Theory and applications, In- ternational journal of computer vision 98 (2012) 217–241
2012
-
[15]
H. Bay, A. Ess, T. Tuytelaars, L. Van Gool, Speeded-up robust features (surf), Computer Vision and Image Understanding 110 (2008) 346–359
2008
-
[16]
Lowry, N
S. Lowry, N. Su¨nderhauf, P. Newman, J. J. Leonard, D. Cox, P. Corke, M. J. Milford, Visual place recognit ion: A survey, IEEE Transactions on Robotics 32 (1) (2015) 1–19
2015
-
[17]
Zhang, L
X. Zhang, L. Wang, Y. Su, Visual place recognition: A survey from deep learning perspective, Pattern Recognition 113 (2021) 107760
2021
-
[18]
Z. Li, P. Xu, T. Shang, Cwpformer: Towards high-performance visual place recognition for robot with cross-weight attention learning, IEEE Transactions on Artificial Intelligence (2025)
2025
-
[19]
Izquierdo, J
S. Izquierdo, J. Civera, Optimal tr ansport aggregation for visual place recognition, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 17658–17668
2024
-
[20]
Hausler, P
S. Hausler, P. Moghadam, Pair-vpr: Place-aware pre-training and con- trastive pair classification for visual place recognition with vision trans- formers, IEEE Robotics and Automation Letters (2025). 55
2025
-
[21]
X i a , L
Y . X i a , L . S h i , Z . D i n g , J . F . H e n r i q u e s , D . C r e m e r s , T e x t 2 l o c : 3 d point cloud localization from natural language, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 14958–14967
2024
-
[22]
Kolmet, Q
M. Kolmet, Q. Zhou, A. Oˇsep, L. Leal-Taix´e, Text2pos: Text-to-point- cloud cross-modal localization, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2022, pp. 6687–6696
2022
-
[23]
G. Wang, H. Fan, M. Kankanhalli, Te xt to point cloud localization with relation-enhanced transformer, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 2023, pp. 2501–2509
2023
-
[24]
Barros, R
T. Barros, R. Pereira, L. Garrote, C. Premebida, U. J. Nunes, Place recog- nition survey: An update on deep le arning approaches, arXiv preprint arXiv:2106.10458 (2021)
2021 arXiv
-
[25]
Z h a n g , P
Y . Z h a n g , P . S h i , J . L i , L i d a r - b ased place recognition for autonomous driving: A survey, ACM Computing Surveys 57 (4) (2024) 1–36
2024
-
[26]
K. Luo, H. Yu, X. Chen, Z. Yang, J. Wang, P. Cheng, A. Mian, 3d point cloud-based place recognition: a survey, Artificial Intelligence Re- view 57 (4) (2024) 83
2024
-
[27]
H. J. S. Feder, J. J. Leonard, C. M. Smith, Adaptive mobile robot naviga- tion and mapping, The International Journal of Robotics Research 18 (7) (1999) 650–668
1999
-
[28]
Thrun, M
S. Thrun, M. Montemerlo, The graph slam algorithm with applications to large-scale mapping of urban structures, The International Journal of Robotics Research 25 (5-6) (2006) 403–429
2006
-
[29]
Montemerlo, S
M. Montemerlo, S. Thrun, D. Koller, B. Wegbreit, et al., Fastslam: A factored solution to the simultaneous localization and mapping problem, AAAI 593598 (2002) 593–598. 56
2002
-
[30]
M. G. Dissanayake, P. Newman, S. Clark, H. F. Durrant-Whyte, M. Csorba, A solution to the simultaneous localization and map building (slam) problem, IEEE Transactions on Robotics and Automation 17 (3) (2001) 229–241
2001
-
[31]
Mur-Artal, J
R. Mur-Artal, J. M. M. Montiel, J. D. Tardos, Orb-slam: A versatile and accurate monocular slam system, IEEE Transactions on Robotics 31 (5) (2015) 1147–1163
2015
-
[32]
Engel, V
J. Engel, V. Koltun, D. Cremers, Direct sparse odometry, IEEE Transac- tions on Pattern Analysis and Machine Intelligence 40 (3) (2017) 611–625
2017
-
[33]
S. Wang, R. Clark, H. Wen, N. Trigoni, Deepvo: Towards end-to-end vi- sual odometry with deep recurrent convolutional neural networks, in: 2017 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2017, pp . 2043–2050
2017
-
[34]
Radwan, A
N. Radwan, A. Valada, W. Burgard, Vlocnet++: Deep multitask learn- ing for semantic visual localization and odometry, IEEE Robotics and Automation Letters 3 (4 ) (2018) 4407–4414
2018
-
[35]
J. Liu, G. Wang, C. Jiang, Z. Liu, H. Wang, Translo: A window-based masked point transformer framework for large-scale lidar odometry, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 2023, pp. 1683–1691
2023
-
[36]
Taketomi, H
T. Taketomi, H. Uchiyama, S. Ikeda, Visual slam algorithms: A survey from 2010 to 2016, IPSJ transactions on computer vision and applications 9 (1) (2017) 16
2017
-
[37]
Y. Wang, Y. Tian, J. Chen, K. Xu, X. Ding, A survey of visual slam in dynamic environment: The evolution from geometric to semantic approaches, IEEE Transactions on Instrumentation and Measurement (2024). 57
2024
-
[38]
C. Yan, D. Qu, D. Xu, B. Zhao, Z. Wang, D. Wang, X. Li, Gs-slam: Dense visual slam with 3d gaussian splatting, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19595–19604
2024
-
[39]
M . L i , S . L i u , H . Z h o u , G . Z h u , N . C h e n g , T . D e n g , H . W a n g , S g s - slam: Semantic gaussian splatting for neural dense slam, in: European Conference on Computer Vision, Springer, 2024, pp. 163–179
2024
-
[40]
M. J. Milford, G. F. Wyeth, Seqslam: Visual route-based navigation for sunny summer days and stormy winter nights, in: 2012 IEEE International Conference on Robotics and Automation, IEEE, 2012, pp. 1643–1649
2012
-
[41]
Talbot, S
B. Talbot, S. Garg, M. Milford, Openseqslam2. 0: An open source tool- box for visual place recognition under changing conditions, in: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2018, pp. 7758–7765
2018
-
[42]
P. Yin, R. A. Srivatsan, Y. Chen, X. Li, H. Zhang, L. Xu, L. Li, Z. Jia, J. Ji, Y. He, Mrs-vpr: a multi-resolu tion sampling base d global visual place recognition method, in: 2019 In ternational Conference on Robotics and Automation (ICRA), IEEE, 2019, pp. 7137–7142
2019
-
[43]
Cummins, P
M. Cummins, P. Newman, Fab-map: Probabilistic localization and map- ping in the space of appearance, The International Journal of Robotics Research 27 (6) (2008) 647–665
2008
-
[44]
G´alvez-L´opez, J
D. G´alvez-L´opez, J. D. Tardos, Bags of binary words for fast place recog- nition in image sequences, IEEE Transactions on Robotics 28 (5) (2012) 1188–1197
2012
-
[45]
Z. Chen, O. Lam, A. Jacobson, M. Milford, Convolutional neural network- based place recognition, arXiv preprint arXiv:1411.1509 (2014)
2014 arXiv
-
[46]
Su¨nderhauf, S
N. Su¨nderhauf, S. Shirazi, F. Dayoub, B. Upcroft, M. Milford, On the performance of convnet features for place recognition, in: 2015 IEEE/RSJ 58 International Conference on intelligent robots and Systems (IROS), IEEE, 2015, pp. 4297–4304
2015
-
[47]
L o w r y , G
S . L o w r y , G . W y e t h , M . M i l f o r d , U n s u p e r v i s e d o n l i n e l e a r n i n g o f condition-invariant images for place recognition, in: Australasian Con- ference on Robotics and Automation, Vol. 2014, 2014
2014
-
[48]
Finman, L
R. Finman, L. Paull, J. J. Leonard, Toward object-based place recognition in dense rgb-d maps, in: ICRA Workshop Visual Place Recognition in Changing Environments, Seattle, WA, Vol. 76, 2015, p. 480
2015
-
[49]
A r a n d j e l o v i c , P
R . A r a n d j e l o v i c , P . G r o n a t , A . T o r i i , T . P a j d l a , J . S i v i c , N e t v l a d : C n n architecture for weakly supervised place recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 5297–5307
2016
-
[50]
Jin Kim, E
H. Jin Kim, E. Dunn, J.-M. Frahm, Learned contextual feature reweight- ing for image geo-localization, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2136–2145
2017
-
[51]
Panphattarasap, A
P. Panphattarasap, A. Calway, Visual place recognition using landmark distribution descriptor s, in: Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part IV 13, Springer, 2017, pp. 487–502
2016
-
[52]
Radenovi´c, G
F. Radenovi´c, G. Tolias, O. Chum, Fine-tuning cnn image retrieval with no human annotation, IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (7) (2018) 1655–1668
2018
-
[53]
J . Y u , C . Z h u , J . Z h a n g , Q . H u a n g , D . T a o , S p a t i a l p y r a m i d - e n h a n c e d netvlad with weighted triplet loss for place recognition, IEEE Transactions on Neural Networks and Learning Systems 31 (2) (2019) 661–674
2019
-
[54]
Y. Ge, H. Wang, F. Zhu, R. Zhao, H. Li, Self-supervising fine-grained region similarities for large-scale image localization, in: Computer Vision– 59 ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16, Springer, 2020, pp. 369–386
2020
-
[55]
B. Cao, A. Araujo, J. Sim, Unifying deep local and global features for im- age search, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16, Springer, 2020, pp. 726–743
2020
-
[56]
Khaliq, S
A. Khaliq, S. Ehsan, Z. Chen, M. Milford, K. McDonald-Maier, A holistic visual place recognition approach using lightweight cnns for significant viewpoint and appearance changes, IEEE Transactio ns on Robotics 36 (2) (2019) 561–569
2019
-
[57]
Hausler, S
S. Hausler, S. Garg, M. Xu, M. Milford, T. Fischer, Patch-netvlad: Multi- scale fusion of locally-gl obal descriptors for place recognition, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14141–14152
2021
-
[58]
Ali-Bey, B
A. Ali-Bey, B. Chaib-Draa, P. Giguere, Mixvpr: Feature mixing for visual place recognition, in: Proceedings of the IEEE/ CVF winter Conference on Applications of Computer Vision, 2023, pp. 2998–3007
2023
-
[59]
Berton, G
G. Berton, G. Trivigno, B. Caputo, C. Masone, Eigenp laces: Training viewpoint robust models for visual place recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 11080–11090
2023
-
[60]
Berton, C
G. Berton, C. Masone, B. Caputo, Reth inking visual geo-localization for large-scale applications, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4878–4888
2022
-
[61]
Ali-bey, B
A. Ali-bey, B. Chaib-draa, P. Gigu`ere, Gsv-cities: Toward appropriate supervised visual place recognition, Neurocomputing 513 (2022) 194–203
2022
-
[62]
R. Wang, Y. Shen, W. Zuo, S. Zhou , N. Zheng, Transvpr: Transformer- based place recognition wi th multi-level attention aggregation, in: Pro- 60 ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 13648–13657
2022
-
[63]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khali- dov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al., Di- nov2: Learning robust visual features without supervision, arXiv preprint arXiv:2304.07193 (2023)
2023 arXiv
-
[64]
Sarlin, D
P.-E. Sarlin, D. DeTone, T. Malisiewicz, A. Rabinovich, Superglue: Learn- ing feature matching with graph neural networks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 4938–4947
2020
-
[65]
Y. Xu, P. Shamsolmoali, E. Granger, C. Nicodeme, L. Gardes, J. Yang, Transvlad: Multi-scale attention-based global descriptors for visual geo- localization, in: Proceedings of th e IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2840–2849
2023
-
[66]
F . L u , S . D o n g , L . Z h a n g , B . L i u , X . L a n , D . J i a n g , C . Y u a n , D e e p homography estimation for visual pl ace recognition, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 2024, pp. 10341– 10349
2024
-
[67]
Ali-Bey, B
A. Ali-Bey, B. Chaib-draa, P. Gigu`ere, Boq: A place is worth a bag of learnable queries, in: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recogn ition, 2024, pp. 17794–17803
2024
-
[68]
F . L u , X . L a n , L . Z h a n g , D . J i a n g , Y . W a n g , C . Y u a n , C r i c a v p r : C r o s s - image correlation-aware representation learning for visual place recogni- tion, in: Proceedings of the IEEE/CV F Conference on Computer Vision and Pattern Recognitio...
2024
-
[69]
F . L u , L . Z h a n g , X . L a n , S . D o n g , Y . W a n g , C . Y u a n , T o w a r d s s e a m - less adaptation of pre-trained models for visual place recognition, arXiv preprint arXiv:2402.14505 (2024). 61
2024 arXiv
-
[70]
J. Hu, C. Mao, C. Tan, H. Li, H. Liu, M. Zheng, Progeo: Gener- ating prompts through image-text cont rastive learning for visual geo- localization, in: International Conference on Artificial Neural Networks, Springer, 2024, pp. 448–462
2024
-
[71]
K. Garg, S. S. Puligilla, S. Kolathaya, M. Krishna, S. Garg, Revisit any- thing: Visual place recognition via image segment retrieval, in: European Conference on Computer Vision, Springer, 2024, pp. 326–343
2024
-
[72]
Tzachor, B
I. Tzachor, B. Lerner, M. Levy, M. Green, T. B. Shalev, G. Habib, D. Samuel, N. K. Zailer, O. Shimshi, N. Darshan, et al., Effovpr: Ef- fective foundation model utilization for visual place recognition, arXiv preprint arXiv:2405.18065 (2024)
2024
-
[73]
W. Zuo, L. Liu, Y. Li, Y. Shen, F. Xiang, J. Xin, N. Zheng, Prgs: Patch- to-region graph search for visual pl ace recognition, Pattern Recognition (2025) 111673
2025
-
[74]
F . L u , T . J i n , X . L a n , L . Z h a n g , Y . L i u , Y . W a n g , C . Y u a n , S e l a v p r + + : Towards seamless adaptation of foun dation models for efficient place recognition, arXiv preprint arXiv:2502.16601 (2025)
2025
-
[75]
Keetha, A
N. Keetha, A. Mishra, J. Karhade, K. M. Jatavallabhu la, S. Scherer, M. Krishna, S. Garg, Anyloc: Towards universal visual place recognition, IEEE Robotics and Automation Le tters 9 (2) (2023) 1286–1293
2023
-
[76]
S. Zhu, L. Yang, C. Chen, M. Shah, X. Shen, H. Wang, R2former: Unified retrieval and reranking transformer for place recognition, in: Proceedings of the IEEE/CVF Conference on Comp uter Vision and Pattern Recogni- tion, 2023, pp. 19370–19380
2023
-
[77]
Hansen, B
P. Hansen, B. Browning, Visual pl ace recognition using hmm sequence matching, in: 2014 IEEE/RSJ Internat ional Conference on Intelligent Robots and Systems, IEEE, 2014, pp. 4549–4555. 62
2014
-
[78]
Mohan, D
M. Mohan, D. G´alvez-L´opez, C. Monteleoni, G. Sibley, Environment se- lection and hierarchical place recogn ition, in: 2015 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2015, pp. 5487– 5494
2015
-
[79]
Revaud, J
J. Revaud, J. Almaz´an, R. S. Rezende, C. R. d. Souza, Learning with average precision: Training image retrieval with a listwise loss, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 5107–5116
2019
-
[80]
L. Liu, H. Li, Y. Dai, Stochastic attraction-repulsion embedding for large scale image localization, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2570–2579
2019
-
[81]
Caron, H
M. Caron, H. Touvron, I. Misra, H. J´egou, J. Mairal, P. Bojanowski, A. Joulin, Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2021, pp. 9650–9660
2021
-
[82]
Y a n , Y
L . Y a n , Y . C u i , Y . C h e n , D . L i u , H i e r a r c h i c a l a t t e n t i o n f u s i o n f o r g e o - localization, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processi ng (ICASSP), IEEE, 2021, pp. 2220– 2224
2021
-
[83]
Leyva-Vallina, N
M. Leyva-Vallina, N. Strisciuglio, N. Petkov, Generalized contrastive optimization of siamese networks fo r place recognition, arXiv preprint arXiv:2103.06638 (2021)
2021 arXiv
-
[84]
C. R. Qi, H. Su, K. Mo, L. J. Guibas, Pointnet: Deep learning on point sets for 3d classification and segm entation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 652–660
2017
-
[85]
C. R. Qi, L. Yi, H. Su, L. J. Guibas, Pointnet++: Deep hierarchical feature learning on point sets in a metric space, Advances in neural infor- mation processing systems 30 (2017). 63
2017
-
[86]
M. A. Uy, G. H. Lee, Pointnetvlad: Deep point cloud based retrieval for large-scale place recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4470–4479
2018
-
[87]
Zhang, C
W. Zhang, C. Xiao, Pcan: 3d atte ntion map learning using contex- tual information for point cloud based retrieval, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 12436–12445
2019
-
[88]
J. Du, R. Wang, D. Cremers, Dh3d: Deep hierarchical 3d descriptors for robust large-scale 6dof relocalization, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceed- ings, Part IV 16, Springer, 2020, pp. 744–762
2020
-
[89]
Cattaneo, M
D. Cattaneo, M. Vaghi, A. Valada, Lcdnet: Deep loop closure detection and point cloud registration for lidar slam, IEEE Transactions on Robotics 38 (4) (2022) 2074–2093
2022
-
[90]
Y. Zhou, Y. Wang, F. Poiesi, Q. Qin, Y. Wan, Loop closure detection using local 3d deep descriptors, I EEE Robotics and Automation Letters 7 (3) (2022) 6335–6342
2022
-
[91]
H. Kim, J. Choi, T. Sim, G. Kim, Y. Cho, Narrowing your fov with solid: Spatially organized and lightweight global descriptor for fov-constrained lidar place recognition, IEEE Robotics and Automation Letters (2024)
2024
-
[92]
A. Zeng, S. Song, M. Nießner, M. Fish er, J. Xiao, T. Funkhouser, 3dmatch: Learning local geometric descriptors from rgb-d reconstructions, in: Pro- ceedings of the IEEE conference on computer vision and pattern recogni- tion, 2017, pp. 1802–1811
2017
-
[93]
Gojcic, C
Z. Gojcic, C. Zhou, J. D. Wegner, A. Wieser, The perfect match: 3d point cloud matching with smoothed densities, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5545–5554. 64
2019
-
[94]
M. Y. Chang, S. Yeon, S. Ryu, D. Lee, Spoxelnet: Sp herical voxel-based deep place recognition for 3d point clouds of crowded indoor spaces, in: 2020 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems (IROS), IEEE, 2020, pp. 8564–8570
2020
-
[95]
S. Siva, Z. Nahman, H. Zhang, Voxe l-based representation learning for place recognition based on 3d point clouds, in: 2020 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS), IEEE, 2020, pp. 8351–8357
2020
-
[96]
Komorowski, Minkloc3d: Point cloud based large-scale place recogni- tion, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp
J. Komorowski, Minkloc3d: Point cloud based large-scale place recogni- tion, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 1790–1799
2021
-
[97]
J. Komorowski, Improving point cl oud based place recognition with ranking-based loss and large batch training, in: 2022 26th international conference on pattern recognition (ICPR), IEEE, 2022, pp. 3699–3705
2022
-
[98]
Vidanapathirana, M
K. Vidanapathirana, M. Ramezani, P. Moghadam, S. Sridharan, C. Fookes, Logg3d-net: Locally guided global descriptor learning for 3d place recognition, in: 2022 International Conference on Robotics and Au- tomation (ICRA), IEEE, 2022, pp. 2215–2221
2022
-
[99]
L u o , S
L . L u o , S . Z h e n g , Y . L i , Y . F a n , B . Y u , S . - Y . C a o , J . L i , H . - L . S h e n , Bevplace: Learning lidar-based place recognition using bird’s eye view images, in: Proceedings of the IEEE/ CVF International Conference on Computer Vision, 2023, pp. 8700–8709
2023
-
[100]
J u n g , W
M . J u n g , W . Y a n g , D . L e e , H . G i l , G . K i m , A . K i m , H e l i p r : H e t e r o g e - neous lidar dataset for inter-lidar place recognition under spatiotemporal variations, The International Journal of Robotics Research 43 (12) (2024) 1867–1883
2024
-
[101]
Xu, Y.-C
T.-X. Xu, Y.-C. Guo, Z. Li, G. Yu, Y.-K. Lai, S.-H. Zhang, Transloc3d: 65 Point cloud based large-scale place recognition using adaptive receptive fields, arXiv preprint arXiv:2105.11605 (2021)
2021 arXiv
-
[102]
Z. Zhou, C. Zhao, D. Adolfsson, S. Su, Y. Gao, T. Duckett, L. Sun, Ndt- transformer: Large-scale 3d point cloud localisation using the normal distribution transform representation, in: 2021 IEEE international con- ference on robotics and automation (ICRA), IEEE, 2021, pp. 5654–5660
2021
-
[103]
Z. Hou, Y. Yan, C. Xu, H. Kong, Hitpr: Hierarchical transformer for place recognition in point cloud, in: 2022 International Conference on Robotics and Automation (ICRA), IEEE, 2022, pp. 2612–2618
2022
-
[104]
J. Ma, J. Zhang, J. Xu, R. Ai, W. Gu, X. Chen, Overlaptransformer: An efficient and yaw-angle-invariant transformer network for lidar-based place recognition, IEEE Robotics an d Automation Letters 7 (3) (2022) 6958–6965
2022
-
[105]
Barros, L
T. Barros, L. Garrote, R. Pereira, C. Premebida, U. J. Nunes, Attdlnet: Attention-based deep network for 3d lidar place recognition, in: Iberian Robotics conference, Springer, 2022, pp. 309–320
2022
-
[106]
J. Ma, X. Chen, J. Xu, G. Xiong, Se qot: A spatial–temp oral transformer network for place recognition using sequential lidar data, IEEE Transac- tions on Industrial Electronics 70 (8) (2022) 8225–8234
2022
-
[107]
Z. Fan, Z. Song, H. Liu, Z. Lu, J. He, X. Du, Svt-net: Super light-weight sparse voxel transformer for large scale place recognition, in: Proceedings of the AAAI conference on artificial intelligence, Vol. 36, 2022, pp. 551– 560
2022
-
[108]
R. G. Goswami, N. Patel, P. Krishnamurthy, F. Khorrami, Salsa: Swift adaptive lightweight self-attention for enhanced lidar place recognition, IEEE Robotics and Automation Letters (2024). 66
2024
-
[109]
K a n g , M
S . K a n g , M . Y . L i a o , Y . X i a , O . W y s o c k i , B . J u t z i , D . C r e m e r s , O p a l : Visibility-aware lidar-to-openstreetmap place recognition via adaptive ra- dial fusion, arXiv preprint arXiv:2504.19258 (2025)
2025 arXiv
-
[110]
X. Wang, G. Tian, J. Zhao, S. Tao, Q. Gu, Q. Yu, T. Feng, Ranking- aware continual learning for lidar place recognition, arXiv preprint arXiv:2505.07198 (2025)
2025 arXiv
-
[111]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G . S a s t r y , A . A s k e l l , P . M i s h k i n , J . C l a r k , e t a l . , L e a r n i n g t r a n s f e r a b l e visual models from natural language supervision, in: International confer- ence on mach...
2021
-
[112]
Shang, Z
T. Shang, Z. Li, W. Pei, P. Xu, Z. Deng, F. Kong, Mambaplace: Text-to- point-cloud cross-modal place recognition with attention mamba mecha- nisms, arXiv preprint ar Xiv:2408.15740 (2024)
2024 arXiv
-
[113]
Shang, Z
T. Shang, Z. Li, P. Xu, Z. Deng, R. Zhang, Text-driven 3d lidar place recognition for autonomous driving, arXiv preprint arXiv:2503.18035 (2025)
2025
-
[114]
Shang, Z
T. Shang, Z. Li, P. Xu, J. Qiao, G. Chen, Z. Ruan, W. Hu, Bridging text and vision: A multi-view text-vision registration approach for cross-modal place recognition, arXiv preprint arXiv:2502.14195 (2025)
2025 arXiv
-
[115]
T or ii , J
A . T or ii , J . Si vic , T . P a j dl a , M. O k ut om i , V is u a l p la c e r eco gn it io n wi th repetitive structures, in: Proceeding s of the IEEE Conference on Com- puter Vision and Pattern Recognition, 2013, pp. 883–890
2013
-
[116]
Neuhold, T
G. Neuhold, T. Ollmann, S. Rota Bulo, P. Kontschieder, The mapillary vistas dataset for semantic understand ing of street scenes, in: Proceed- ings of the IEEE International Conference on Computer Vision, 2017, pp. 4990–4999
2017
-
[117]
A . J . G l o v e r , W . P . M a d d e r n , M . J . M i l f o r d , G . F . W y e t h , F a b - m a p + ratslam: Appearance-based slam for multiple times of day, in: 2010 IEEE 67 international conference on roboti cs and automation, IEEE, 2010, pp. 3507–3512
2010
-
[118]
T o r i i , R
A . T o r i i , R . A r a n d j e l o v i c , J . S i v i c , M . O k u t o m i , T . P a j d l a , 2 4 / 7 p l a c e recognition by view synthesis, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1808–1817
2015
-
[119]
D. Olid, J. M. F´acil, J. Civera, Single-view place recognition under sea- sonal changes, arXiv preprint arXiv:1808.06516 (2018)
2018 arXiv
-
[120]
X. Sun, Y. Xie, P. Luo, L. Wang, A dataset for benchmarking image-based localization, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 7436–7444
2017
-
[121]
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, A. Oliva, Learning deep features for scene recognition using places database, Advances in neural information processing systems 27 (2014)
2014
-
[122]
Hu a n g, C
H . Hu a n g, C . L iu , Y. Zh u, H . C h eng , T . B r a ud , S. -K . Y eu n g, 36 0 lo c: A dataset and benchmark for omnidirectional visual localization with cross- device queries, in: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2024...
2024
-
[123]
Maddern, G
W. Maddern, G. Pascoe, C. Linegar, P. Newman, 1 year, 1000 km: The oxford robotcar dataset, The International Journal of Robotics Research 36 (1) (2017) 3–15
2017
-
[124]
Huang, Z
H. Huang, Z. Nie, Z. Wang, Z. Sh ang, Cross-modal and uni-modal soft- label alignment for image-text retr ieval, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 2024, pp. 18298–18306
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.