REVIEW 3 major objections 4 minor 87 references
A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place Recognition
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Summing condition-specific place signatures lifts visual place recognition recall with no extra matching cost.
desk verdict HOPS is a simple, broadly validated trick: sum same-place reference descriptors across conditions and get better recall than the best single reference set at constant query cost. The HDC story is mostly decorative, but the empirical result is solid and worth citing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is bundling by element-wise summation of descriptor vectors, defined as $r_{\mathrm{fused},i}=\sum_{k=1}^K r_i^k$, which relies on the quasi-orthogonality of vectors in hyperdimensional spaces so that condition-specific noise components cancel while shared place features reinforce. This single formula carries the argument: it preserves single-reference matching complexity, makes new conditions stackable at any time, and is the operation varied across the experiments, whether the fused conditions are real traverses, synthetic image augmentations, or entire datasets.
What would settle it
An experiment that would settle the mechanism: for a fixed VPR descriptor, compute the difference vector $d_k=r_i^k-\bar{r}_i$ for each place and measure the cosine of the angle between $d_k$ and $\bar{r}_i$, averaged across places and conditions; if that cosine is not near zero, condition changes produce systematic non-quasi-orthogonal shifts and the noise-cancellation story is false. A second decisive test: randomly permute the correspondence of descriptors to places before summing and check whether the reported recall gains survive; if they do, the fused vector is not encoding place identity through reinforcement of shared features.
Extended reading notes
Core claim
The central claim is that hyperdimensional feature vectors of the same place under different conditions differ by a quasi-orthogonal noise component whose influence on cosine similarity is negligible, so an element-wise sum of the condition-specific descriptors lets genuine place features reinforce while transient appearance differences cancel. The resulting fused descriptor $r_{\mathrm{fused},i}=\sum_{k=1}^K r_i^k$ replaces $K$ reference sets with one vector per place, so matching stays at single-reference complexity $O(M)$ in both compute and storage. Across the Oxford RobotCar, Nordland, and SFU Mountain datasets, this fused signature improves recall@1 over the best single-condition reference set in nearly all tested cases, and it outperforms both pooling all references into one big set and averaging independent per-set distance matrices in the majority of cases, while the same summation also enables strong dimensionality reduction via Gaussian random projection and dataset-level identification of a query with accuracy above 99.7%.
Load-bearing premise
The argument depends on the assumption that, for the deep-learned descriptors being summed, differences between the same place under different conditions behave like random noise that cancels in the sum while true place features reinforce; if those condition differences are actually structured and correlated with the place signal, summing will blur instead of sharpen the descriptor.
Editorial extensions
If this is right
- A robot with several traverses of the same route can store one fused signature per place instead of $K$ reference sets, keeping query-time matching cost fixed while improving recall.
- Fusing multiple condition traverses requires no additional training and no extra query-time memory, making appearance-invariant localization cheaper than pooling or multi-reference distance averaging.
- High-dimensional descriptors such as SALAD and CricaVPR can be fused and then projected down by about 95–97% with no loss relative to the best full-size single reference, enabling substantially smaller map databases.
- Synthetic image augmentations of a single reference set can substitute for real multi-condition traverses, so the benefit is available even when only one traversal exists.
- A single fused descriptor computed over all reference images of a dataset identifies which dataset a query came from with accuracy above 99.7%, showing the sum retains dataset-level discriminative structure.
Reading between the lines
- If quasi-orthogonal noise cancellation is the true mechanism, the same sum-and-match recipe should transfer to other retrieval tasks with multiple high-dimensional views of one identity, such as person or object re-identification across cameras; the paper does not test this extension.
- The paper's own outliers at 512 dimensions suggest a practical rule of thumb: HOPS gains require descriptors large enough that cross-condition differences are quasi-orthogonal, and the data hint the threshold lies between 512 and 4096 dimensions, a boundary the paper does not derive.
- The error-density analysis indicates HOPS mostly sharpens already-close matches, pulling near-misses onto the ground truth, rather than rescuing grossly wrong retrievals; a natural unexamined application is as a free re-ranking stage inside hierarchical localization pipelines.
- The preliminary observation that descriptors from different datasets can be stacked with minimal penalty implies a compressed multi-map representation in which many environments share one vector space, with the dataset-identification result as the enabling primitive; the paper does not develop this into an actual multi-map matching method.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Hyperdimensional One Place Signatures (HOPS), a method that fuses the deep VPR descriptors of the same place captured under K different conditions by element-wise summation (Eq. 2), producing one fused descriptor per place while keeping matching complexity O(M). The method is evaluated on Oxford RobotCar, Nordland, and SFU Mountain using NetVLAD, SALAD, MixVPR, CosPlace, EigenPlaces, and CricaVPR, and is compared against the best single reference set and two multi-reference baselines (reference pooling and distance matrix averaging). The paper reports that HOPS beats the best single reference set in 87/90 cases and the multi-reference baselines in 69/90 cases, and additionally demonstrates dimensionality reduction via Gaussian random projection, fusion of synthetic image augmentations, and dataset identification. The claimed mechanism is quasi-orthogonality in hyperdimensional spaces, which would make condition-dependent differences behave as cancelable noise.
Significance. If the empirical results hold, HOPS is a useful, training-free technique: it converts K condition-specific reference traverses into one descriptor per place, improving recall while preserving single-set query-time cost. The experiments are extensive (six descriptors, three datasets, plus AnyLoc in the supplement), the evaluation protocol holds out the query set, and no parameters are fitted to test data; code is promised in the paper. The dimensionality-reduction and synthetic-augmentation results are interesting and actionable. The main weakness is that the quasi-orthogonality explanation in Section 3.2 is asserted rather than tested for learned deep descriptors, and the paper's own conclusion defers a deeper investigation of bundling effects to future work. In addition, the absence of error bars or significance tests makes the few reported failures hard to interpret. These weaknesses do not invalidate the empirical core, but they need to be addressed before the HDC-based scalability claims are fully supported.
major comments (3)
- [Section 3.2, Eq. (2)] The justification for Eq. (2) is that condition-dependent differences behave as quasi-orthogonal noise z whose influence on cosine similarity is negligible. This is an unverified assertion about learned deep descriptors. The paper does not measure the residual vectors δ_{ik} = r_{ik} - μ_i (e.g., their norms relative to μ_i or their cosine alignment with μ_i), nor does it rule out the alternative explanation that the fused descriptor is simply the empirical centroid of the K reference descriptors, which is the same as Eq. (2) up to scale. The failure cases at 512D (Tables 1 and 2: CosPlace and EigenPlaces on RobotCar Night; CosPlace on Nordland Summer) are attributed to insufficient dimensionality, but no diagnostic supports that explanation. Since the abstract and Section 3.2 make the HDC-based scalability claim load-bearing, the paper should either provide residual-geometry measurements (e.g., distributions of cos(δ_{ik}, μ_i) and ||δ_{ik}||/||μ_i||) or temper the HDC explanation and frame HOPS as condition averaging with different scope conditions.
- [Section 4.3, Tables 1-3] All recall numbers are single point estimates with no error bars, confidence intervals, or significance tests. The claim of near-unanimous improvement (28/30, 23/24, 36/36 cases) includes margins as small as a few tenths of a percent, and the three reported degradations are close to the noise level typical for retrieval metrics. Without repeated evaluations (e.g., over multiple random projection draws or bootstrapping over queries) or a paired test such as McNemar's test, the reader cannot assess whether the few failures are real or random. This is load-bearing for the central empirical claim of consistent improvement, so the authors should add uncertainty quantification for at least the key comparisons.
- [Section 3.2 / Section 4.1] The paper does not state whether the input descriptors are L2-normalized before summation. Since most VPR descriptors are L2-normalized and the unnormalized sum is dominated by the largest-norm component, the effective weighting of the K reference conditions is undefined without this detail. Please specify the preprocessing and, if normalization is applied, state this explicitly; if not, justify the choice or evaluate a normalized variant.
minor comments (4)
- [Section 3.2] The sentence 'with an additional noise vector z affecting either vector' is unclear; please define what z is, which vectors it affects, and how the quasi-orthogonality argument applies to it.
- [Abstract] The claim that HOPS 'consistently improves recall performance across all evaluated VPR methods and datasets by large margins' overstates the three outlier cases and the small margins in some cells; suggest 'in the large majority of cases' or provide a precise count.
- [Supplementary Table 8] The first 'Feat. Dim.' entry reads 8488; this appears to be a typo for 8448 (the SALAD descriptor dimension used elsewhere).
- [Section 4.3] The paragraph beginning 'There are three outlier cases' lists two RobotCar cases and one Nordland case; consider making explicit that these are the only cases across all three datasets, for clarity.
Circularity Check
No significant circularity: HOPS is a fixed, parameter-free descriptor sum evaluated on held-out query sets; the quasi-orthogonality explanation is an untested assumption, not a circular derivation.
full rationale
The central construction, rfused,i = Σ_k r^k_i (Eq. 2), is a fixed element-wise sum of VPR descriptors of the same place from non-query reference sets. No parameter is fitted to the query sets, and the reported recall improvements are measured on held-out query traverses; Section 4.3 explicitly states that only non-query sets are combined, so the outcome is an empirical finding rather than a consequence of tuning. The quoted quasi-orthogonality argument in Section 3.2 is a motivating explanation, not a derived prediction: the paper does not define the noise vector z in terms of the measured improvement, nor does it use the improvement to infer z. Citations to the authors' own prior work ([22], [49]) appear only as comparison baselines (distance-matrix averaging) or related multi-reference curation methods, and are not used to justify the central claim. The Gaussian random projection experiment is also fixed (G sampled from N(0,1/n)) and the dataset-identification experiment explicitly excludes the query set from the bundled descriptor. Consequently, no step in the paper's derivation chain reduces, by construction or by self-citation, to its own inputs; the identified weakness (that quasi-orthogonality is assumed rather than verified for deep descriptors) concerns empirical soundness, not circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Condition-specific variation in descriptors of the same place is quasi-orthogonal noise in high-dimensional space.
- domain assumption Reference traverses have direct one-to-one spatial correspondence, so the same index i is the same physical place.
- standard math Cosine distance is the correct similarity metric for VPR descriptors.
- standard math Gaussian random projection approximately preserves distances per the Johnson-Lindenstrauss lemma.
Cite this review
Pith. "Pith review of A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place Recognition." pith.science (2026). https://pith.science/paper/H5ACKRDC
@misc{pith2026241206153,
author = {Pith},
title = {Pith review of: A Hyperdimensional One Place Signature to Represent Them All: Stackable Descriptors For Visual Place Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/H5ACKRDC}},
note = {Machine review of arXiv:2412.06153}
}
read the original abstract
Visual Place Recognition (VPR) enables coarse localization by comparing query images to a reference database of geo-tagged images. Recent breakthroughs in deep learning architectures and training regimes have led to methods with improved robustness to factors like environment appearance change, but with the downside that the required training and/or matching compute scales with the number of distinct environmental conditions encountered. Here, we propose Hyperdimensional One Place Signatures (HOPS) to simultaneously improve the performance, compute and scalability of these state-of-the-art approaches by fusing the descriptors from multiple reference sets captured under different conditions. HOPS scales to any number of environmental conditions by leveraging the Hyperdimensional Computing framework. Extensive evaluations demonstrate that our approach is highly generalizable and consistently improves recall performance across all evaluated VPR methods and datasets by large margins. Arbitrarily fusing reference images without compute penalty enables numerous other useful possibilities, three of which we demonstrate here: descriptor dimensionality reduction with no performance penalty, stacking synthetic images, and coarse localization to an entire traverse or environmental section.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Database-friendly random projections: J ohnson- L indenstrauss with binary coins
Dimitris Achlioptas. Database-friendly random projections: J ohnson- L indenstrauss with binary coins. Journal of Computer and System Sciences, 66 0 (4): 0 671--687, 2003
2003
-
[3]
Gsv-cities: Toward appropriate supervised visual place recognition
Amar Ali-bey, Brahim Chaib-draa, and Philippe Gigu \`e re. Gsv-cities: Toward appropriate supervised visual place recognition. Neurocomputing, 2022
2022
-
[4]
Mixvpr: Feature mixing for visual place recognition
Amar Ali-Bey, Brahim Chaib-Draa, and Philippe Giguere. Mixvpr: Feature mixing for visual place recognition. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2998--3007, 2023
2023
-
[5]
Boq: A place is worth a bag of learnable queries
Amar Ali-bey, Brahim Chaib-draa, and Philippe Gigu \`e re. Boq: A place is worth a bag of learnable queries. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17794--17803, 2024
2024
-
[6]
All about VLAD
Relja Arandjelovic and Andrew Zisserman. All about VLAD . In IEEE Conference on Computer Vision and Pattern Recognition, pages 1578--1585, 2013
2013
-
[7]
NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. NetVLAD: CNN architecture for weakly supervised place recognition . In IEEE Conference on Computer Vision and Pattern Recognition, pages 5297--5307, 2016
2016
-
[8]
Giovanni Barbarani, Mohamad Mostafa, Hajali Bayramov, Gabriele Trivigno, Gabriele Berton, Carlo Masone, and Barbara Caputo. Are local features all you need for cross-domain visual place recognition? In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6155--6165, 2023
work page 2023
Show all 87 references
-
[9]
Speeded-up robust features (SURF)
Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool. Speeded-up robust features (SURF) . Computer Vision and Image Understanding, 110 0 (3): 0 346--359, 2008
2008
-
[10]
Rethinking visual geo-localization for large-scale applications
Gabriele Berton, Carlo Masone, and Barbara Caputo. Rethinking visual geo-localization for large-scale applications. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4878--4888, 2022
2022
-
[11]
Eigenplaces: Training viewpoint robust models for visual place recognition
Gabriele Berton, Gabriele Trivigno, Barbara Caputo, and Carlo Masone. Eigenplaces: Training viewpoint robust models for visual place recognition. In IEEE/CVF International Conference on Computer Vision, pages 11080--11090, 2023
2023
-
[12]
MeshVPR: Citywide Visual Place Recognition Using 3D Meshes
Gabriele Berton, Lorenz Junglas, Riccardo Zaccone, Thomas Pollok, Barbara Caputo, and Carlo Masone. MeshVPR: Citywide Visual Place Recognition Using 3D Meshes . In European Conference on Computer Vision, 2024
2024
-
[13]
Adaptive-attentive geolocalization from few queries: A hybrid approach
Gabriele Moreno Berton, Valerio Paolicelli, Carlo Masone, and Barbara Caputo. Adaptive-attentive geolocalization from few queries: A hybrid approach. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2918--2927, 2021
2021
-
[14]
The SFU mountain dataset: S emi-structured woodland trails under changing environmental conditions
Jake Bruce, Jens Wawerla, and Richard Vaughan. The SFU mountain dataset: S emi-structured woodland trails under changing environmental conditions. In Workshop on Visual Place Recognition in Changing Environments, IEEE International Conference on Robotics and Automation, 2015
2015
-
[15]
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age
Cesar Cadena, Luca Carlone, Henry Carrillo, Yasir Latif, Davide Scaramuzza, Jos \'e Neira, Ian Reid, and John J Leonard. Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age. IEEE Transactions on Robotics, 32 0 (6): 0 1309--1332, 2016
2016
-
[16]
Unifying deep local and global features for image search
Bingyi Cao, Andr \'e Araujo, and Jack Sim. Unifying deep local and global features for image search. In The European Conference on Computer Vision, pages 726--743, 2020
2020
-
[17]
A survey on map-based localization techniques for autonomous vehicles
Athanasios Chalvatzaras, Ioannis Pratikakis, and Angelos A Amanatiadis. A survey on map-based localization techniques for autonomous vehicles. IEEE Transactions on Intelligent Vehicles, 8 0 (2): 0 1574--1596, 2022
2022
-
[18]
Superposition of many models into one
Brian Cheung, Alexander Terekhov, Yubei Chen, Pulkit Agrawal, and Bruno Olshausen. Superposition of many models into one. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[19]
Practice makes perfect? managing and leveraging visual experiences for lifelong navigation
Winston Churchill and Paul Newman. Practice makes perfect? managing and leveraging visual experiences for lifelong navigation. In IEEE International Conference on Robotics and Automation, pages 4525--4532, 2012
2012
-
[20]
Visual categorization with bags of keypoints
Gabriella Csurka, Christopher Dance, Lixin Fan, Jutta Willamowski, and C \'e dric Bray. Visual categorization with bags of keypoints. In Workshop on Statistical Learning in Computer Vision, The European Conference on Computer Vision, pages 1--2, 2004
2004
-
[21]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[22]
Simultaneous localization and mapping: part I
Hugh Durrant-Whyte and Tim Bailey. Simultaneous localization and mapping: part I . IEEE Robotics & Automation Magazine, 13 0 (2): 0 99--110, 2006
2006
-
[23]
Event-based visual place recognition with ensembles of temporal windows
Tobias Fischer and Michael Milford. Event-based visual place recognition with ensembles of temporal windows. IEEE Robotics and Automation Letters, 5 0 (4): 0 6924--6931, 2020
2020
-
[24]
Where is your place, visual place recognition? In International Joint Conferences on Artificial Intelligence, pages 4416--4425, 2021
Sourav Garg, Tobias Fischer, and Michael Milford. Where is your place, visual place recognition? In International Joint Conferences on Artificial Intelligence, pages 4416--4425, 2021
2021
-
[25]
Classification using hyperdimensional computing: A review
Lulu Ge and Keshab K Parhi. Classification using hyperdimensional computing: A review. IEEE Circuits and Systems Magazine, 20 0 (2): 0 30--47, 2020
2020
-
[26]
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (1): 0 87--110, 2022
2022
-
[27]
Patch-NetVLAD: Multi-scale fusion of locally-global descriptors for place recognition
Stephen Hausler, Sourav Garg, Ming Xu, Michael Milford, and Tobias Fischer. Patch-NetVLAD: Multi-scale fusion of locally-global descriptors for place recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021
2021
-
[28]
Fe-fusion-vpr: Attention-based multi-scale network architecture for visual place recognition by fusing frames and events
Kuanxu Hou, Delei Kong, Junjie Jiang, Hao Zhuang, Xinjie Huang, and Zheng Fang. Fe-fusion-vpr: Attention-based multi-scale network architecture for visual place recognition by fusing frames and events. IEEE Robotics and Automation Letters, 8 0 (6): 0 3526--3533, 2023
2023
-
[29]
Dasgil: Domain adaptation for semantic and geometric-aware image-based localization
Hanjiang Hu, Zhijian Qiao, Ming Cheng, Zhe Liu, and Hesheng Wang. Dasgil: Domain adaptation for semantic and geometric-aware image-based localization. IEEE Transactions on Image Processing, 30: 0 1342--1353, 2020
2020
-
[30]
Spiking neural networks for visual place recognition via weighted neuronal assignments
Somayeh Hussaini, Michael Milford, and Tobias Fischer. Spiking neural networks for visual place recognition via weighted neuronal assignments. IEEE Robotics and Automation Letters, 7 0 (2): 0 4094--4101, 2022
2022
-
[31]
Applications of spiking neural networks in visual place recognition
Somayeh Hussaini, Michael Milford, and Tobias Fischer. Applications of spiking neural networks in visual place recognition. arXiv preprint arXiv:2311.13186, 2023
2023 arXiv
-
[32]
Optimal transport aggregation for visual place recognition
Sergio Izquierdo and Javier Civera. Optimal transport aggregation for visual place recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17658--17668, 2024
2024
-
[33]
Autonomous multisensor calibration and closed-loop fusion for slam
Adam Jacobson, Zetao Chen, and Michael Milford. Autonomous multisensor calibration and closed-loop fusion for slam. Journal of Field Robotics, 32 0 (1): 0 85--122, 2015
2015
-
[34]
Aggregating local descriptors into a compact image representation
Herv \'e J \'e gou, Matthijs Douze, Cordelia Schmid, and Patrick P \'e rez. Aggregating local descriptors into a compact image representation. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 3304--3311, 2010
2010
-
[35]
In-memory hyperdimensional computing
Geethan Karunaratne, Manuel Le Gallo, Giovanni Cherubini, Luca Benini, Abbas Rahimi, and Abu Sebastian. In-memory hyperdimensional computing. Nature Electronics, 3 0 (6): 0 327--337, 2020
2020
-
[36]
Anyloc: Towards universal visual place recognition
Nikhil Keetha, Avneesh Mishra, Jay Karhade, Krishna Murthy Jatavallabhula, Sebastian Scherer, Madhava Krishna, and Sourav Garg. Anyloc: Towards universal visual place recognition. IEEE Robotics and Automation Letters, 2023
2023
-
[37]
Classification and recall with binary hyperdimensional computing: Tradeoffs in choice of density and mapping characteristics
Denis Kleyko, Abbas Rahimi, Dmitri A Rachkovskij, Evgeny Osipov, and Jan M Rabaey. Classification and recall with binary hyperdimensional computing: Tradeoffs in choice of density and mapping characteristics. IEEE Transactions on Neural Networks and Learning Systems, 29 0 (12)...
2018
-
[38]
A survey on hyperdimensional computing aka vector symbolic architectures, part ii: Applications, cognitive models, and challenges
Denis Kleyko, Dmitri Rachkovskij, Evgeny Osipov, and Abbas Rahimi. A survey on hyperdimensional computing aka vector symbolic architectures, part ii: Applications, cognitive models, and challenges. ACM Computing Surveys, 55 0 (9): 0 1--52, 2023
2023
-
[39]
A survey on localization for autonomous vehicles
Debasis Kumar and Naveed Muhammad. A survey on localization for autonomous vehicles. IEEE Access, 2023
2023
-
[40]
Data-efficient large scale place recognition with graded similarity supervision
Mar \' a Leyva-Vallina, Nicola Strisciuglio, and Nicolai Petkov. Data-efficient large scale place recognition with graded similarity supervision. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23487--23496, 2023
2023
-
[41]
Work smart, not hard: Recalling relevant experiences for vast-scale but time-constrained localisation
Chris Linegar, Winston Churchill, and Paul Newman. Work smart, not hard: Recalling relevant experiences for vast-scale but time-constrained localisation. In IEEE International Conference on Robotics and Automation, pages 90--97, 2015
2015
-
[42]
D.G. Lowe. Object recognition from local scale-invariant features. In IEEE International Conference on Computer Vision, pages 1150--1157 vol.2, 1999
1999
-
[43]
Visual place recognition: A survey
Stephanie Lowry, Niko S \"u nderhauf, Paul Newman, John J Leonard, David Cox, Peter Corke, and Michael J Milford. Visual place recognition: A survey. IEEE Transactions on Robotics, 32 0 (1): 0 1--19, 2015
2015
-
[44]
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
Feng Lu, Xiangyuan Lan, Lijun Zhang, Dongmei Jiang, Yaowei Wang, and Chun Yuan. CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition . In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16772--16782, 2024 a
2024
-
[45]
Towards seamless adaptation of pre-trained models for visual place recognition
Feng Lu, Lijun Zhang, Xiangyuan Lan, Shuting Dong, Yaowei Wang, and Chun Yuan. Towards seamless adaptation of pre-trained models for visual place recognition. In The International Conference on Learning Representations, 2024 b
2024
-
[46]
Similarity min-max: Zero-shot day-night domain adaptation
Rundong Luo, Wenjing Wang, Wenhan Yang, and Jiaying Liu. Similarity min-max: Zero-shot day-night domain adaptation. In IEEE/CVF International Conference on Computer Vision, pages 8104--8114, 2023
2023
-
[47]
1 Year, 1000km: The Oxford RobotCar Dataset
Will Maddern, Geoff Pascoe, Chris Linegar, and Paul Newman. 1 Year, 1000km: The Oxford RobotCar Dataset . The International Journal of Robotics Research, 36 0 (1): 0 3--15, 2017
2017
-
[48]
A survey on deep visual place recognition
Carlo Masone and Barbara Caputo. A survey on deep visual place recognition. IEEE Access, 9: 0 19516--19547, 2021
2021
-
[49]
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic. Learnable pooling with context gating for video classification. arXiv preprint arXiv:1706.06905, 2017
2017 arXiv
-
[50]
Intelligent reference curation for visual place recognition via bayesian selective fusion
Timothy L Molloy, Tobias Fischer, Michael Milford, and Girish N Nair. Intelligent reference curation for visual place recognition via bayesian selective fusion. IEEE Robotics and Automation Letters, 6 0 (2): 0 588--595, 2020
2020
-
[51]
From omnidirectional images to hierarchical localization
AC Murillo, Carlos Sag \"u \'e s, Jos \'e Jes \'u s Guerrero, Toon Goedem \'e , Tinne Tuytelaars, and Luc Van Gool. From omnidirectional images to hierarchical localization. IEEE Robotics and Autonomous Systems, 55 0 (5): 0 372--382, 2007
2007
-
[52]
Hyperdimensional computing as a framework for systematic aggregation of image descriptors
Peer Neubert and Stefan Schubert. Hyperdimensional computing as a framework for systematic aggregation of image descriptors. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16938--16947, 2021
2021
-
[53]
An introduction to hyperdimensional computing for robotics
Peer Neubert, Stefan Schubert, and Peter Protzel. An introduction to hyperdimensional computing for robotics. KI-K \"u nstliche Intelligenz , 33 0 (4): 0 319--330, 2019
2019
-
[54]
Vector semantic representations as descriptors for visual place recognition
Peer Neubert, Stefan Schubert, Kenny Schlegel, and Peter Protzel. Vector semantic representations as descriptors for visual place recognition. In Robotics: Science and Systems, pages 1--11, 2021
2021
-
[55]
Large-scale image retrieval with attentive deep local features
Hyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, and Bohyung Han. Large-scale image retrieval with attentive deep local features. In IEEE/CVF International Conference on Computer Vision, pages 3456--3465, 2017
2017
-
[56]
Fisher kernels on visual vocabularies for image categorization
Florent Perronnin and Christopher Dance. Fisher kernels on visual vocabularies for image categorization. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1--8, 2007
2007
-
[57]
Large-scale image retrieval with compressed fisher vectors
Florent Perronnin, Yan Liu, Jorge S \'a nchez, and Herv \'e Poirier. Large-scale image retrieval with compressed fisher vectors. In Computer Society Conference on Computer Vision and Pattern Recognitionn, pages 3384--3391, 2010
2010
-
[58]
On model-free re-ranking for visual place recognition with deep learned local features
Tom \'a s Pivo n ka and Libor P r eu c il. On model-free re-ranking for visual place recognition with deep learned local features. IEEE Transactions on Intelligent Vehicles, 2024
2024
-
[59]
An outlook into the future of egocentric vision
Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi. An outlook into the future of egocentric vision. International Journal of Computer Vision, pages 1--57, 2024
2024
-
[60]
Cnn image retrieval learns from bow: Unsupervised fine-tuning with hard examples
Filip Radenovi \'c , Giorgos Tolias, and Ond r ej Chum. Cnn image retrieval learns from bow: Unsupervised fine-tuning with hard examples. In The European Conference on Computer Vision, pages 3--20, 2016
2016
-
[61]
Fine-tuning cnn image retrieval with no human annotation
Filip Radenovi \'c , Giorgos Tolias, and Ond r ej Chum. Fine-tuning cnn image retrieval with no human annotation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41 0 (7): 0 1655--1668, 2018
2018
-
[62]
Learning with average precision: Training image retrieval with a listwise loss
Jerome Revaud, Jon Almaz \'a n, Rafael S Rezende, and Cesar Roberto de Souza. Learning with average precision: Training image retrieval with a listwise loss. In IEEE/CVF International Conference on Computer Vision, pages 5107--5116, 2019
2019
-
[63]
Leveraging deep visual descriptors for hierarchical efficient localization
Paul-Edouard Sarlin, Fr \'e d \'e ric Debraine, Marcin Dymczyk, Roland Siegwart, and Cesar Cadena. Leveraging deep visual descriptors for hierarchical efficient localization. In Conference on Robot Learning, pages 456--465, 2018
2018
-
[64]
From coarse to fine: Robust hierarchical localization at large scale
Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12716--12725, 2019
2019
-
[65]
Superglue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4938--4947, 2020
2020
-
[66]
Lamar: Benchmarking localization and mapping for augmented reality
Paul-Edouard Sarlin, Mihai Dusmanu, Johannes L Sch \"o nberger, Pablo Speciale, Lukas Gruber, Viktor Larsson, Ondrej Miksik, and Marc Pollefeys. Lamar: Benchmarking localization and mapping for augmented reality. In European Conference on Computer Vision, pages 686--704, 2022
2022
-
[67]
A comparison of vector symbolic architectures
Kenny Schlegel, Peer Neubert, and Peter Protzel. A comparison of vector symbolic architectures. Artificial Intelligence Review, 55 0 (6): 0 4523--4555, 2022
2022
-
[68]
Visual place recognition in changing environments using additional data-inherent knowledge
M Sc Stefan Schubert. Visual place recognition in changing environments using additional data-inherent knowledge. Technische Universität Chemnitz, Chemnitz, 2023
2023
-
[69]
Visual place recognition: A tutorial
Stefan Schubert, Peer Neubert, Sourav Garg, Michael Milford, and Tobias Fischer. Visual place recognition: A tutorial. IEEE Robotics & Automation Magazine, 2023
2023
-
[70]
Omnidirectional multisensory perception fusion for long-term place recognition
Sriram Siva and Hao Zhang. Omnidirectional multisensory perception fusion for long-term place recognition. In IEEE International Conference on Robotics and Automation, pages 5175--5181, 2018
2018
-
[71]
Video google: A text retrieval approach to object matching in videos
Sivic and Zisserman. Video google: A text retrieval approach to object matching in videos. In IEEE International Conference on Computer Vision, pages 1470--1477, 2003
2003
-
[72]
Are we there yet? challenging seqslam on a 3000 km journey across all four seasons
N Sünderhauf, Peer Neubert, and Peter Protzel. Are we there yet? challenging seqslam on a 3000 km journey across all four seasons. Workshop on Long-term Autonomy, IEEE International Conference on Robotics and Automation, 2013
2013
-
[73]
Long-term visual localization revisited
Carl Toft, Will Maddern, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, et al. Long-term visual localization revisited. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44 0 (4): 0 20...
2020
-
[74]
Divide&classify: Fine-grained classification for city-wide visual geo-localization
Gabriele Trivigno, Gabriele Berton, Juan Aragon, Barbara Caputo, and Carlo Masone. Divide&classify: Fine-grained classification for city-wide visual geo-localization. In IEEE/CVF International Conference on Computer Vision, 2023
2023
-
[75]
The revisiting problem in simultaneous localization and mapping: A survey on visual loop closure detection
Konstantinos A Tsintotas, Loukas Bampis, and Antonios Gasteratos. The revisiting problem in simultaneous localization and mapping: A survey on visual loop closure detection. IEEE Transactions on Intelligent Transportation Systems, 23 0 (11): 0 19929--19953, 2022
2022
-
[76]
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
Issar Tzachor, Boaz Lerner, Matan Levy, Michael Green, Tal Berkovitz Shalev, Gavriel Habib, Dvir Samuel, Noam Korngut Zailer, Or Shimshi, Nir Darshan, et al. EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition . arXiv preprint arXiv:2405.18065, 2024
2024
-
[77]
Effective visual place recognition using multi-sequence maps
Olga Vysotska and Cyrill Stachniss. Effective visual place recognition using multi-sequence maps. IEEE Robotics and Automation Letters, 4 0 (2): 0 1730--1736, 2019
2019
-
[78]
Transvpr: Transformer-based place recognition with multi-level attention aggregation
Ruotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou, and Nanning Zheng. Transvpr: Transformer-based place recognition with multi-level attention aggregation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13648--13657, 2022
2022
-
[79]
Mapillary street-level sequences: A dataset for lifelong place recognition
Frederik Warburg, Soren Hauberg, Manuel Lopez-Antequera, Pau Gargallo, Yubin Kuang, and Javier Civera. Mapillary street-level sequences: A dataset for lifelong place recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2626--2635, 2020
2020
-
[80]
Hyperdimensional feature fusion for out-of-distribution detection
Samuel Wilson, Tobias Fischer, Niko S \"u nderhauf, and Feras Dayoub. Hyperdimensional feature fusion for out-of-distribution detection. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2644--2654, 2023
2023
-
[81]
Real-time visual place recognition based on analyzing distribution of multi-scale cnn landmarks
Zhe Xin, Xiaoguang Cui, Jixiang Zhang, Yiping Yang, and Yanqing Wang. Real-time visual place recognition based on analyzing distribution of multi-scale cnn landmarks. Journal of Intelligent & Robotic Systems, 94: 0 777--792, 2019
2019
-
[82]
Probabilistic visual place recognition for hierarchical localization
Ming Xu, Niko Snderhauf, and Michael Milford. Probabilistic visual place recognition for hierarchical localization. IEEE Robotics and Automation Letters, 6 0 (2): 0 311--318, 2020
2020
-
[83]
Dolg: Single-stage image retrieval with deep orthogonal fusion of local and global features
Min Yang, Dongliang He, Miao Fan, Baorong Shi, Xuetong Xue, Fu Li, Errui Ding, and Jizhou Huang. Dolg: Single-stage image retrieval with deep orthogonal fusion of local and global features. In IEEE/CVF International Conference on Computer Vision, pages 11772--11781, 2021
2021
-
[84]
Spatial pyramid-enhanced NetVLAD with weighted triplet loss for place recognition
Jun Yu, Chaoyang Zhu, Jian Zhang, Qingming Huang, and Dacheng Tao. Spatial pyramid-enhanced NetVLAD with weighted triplet loss for place recognition. IEEE Transactions on Neural Networks and Learning Systems, 31 0 (2): 0 661--674, 2019
2019
-
[85]
Etr: An efficient transformer for re-ranking in visual place recognition
Hao Zhang, Xin Chen, Heming Jing, Yingbin Zheng, Yuan Wu, and Cheng Jin. Etr: An efficient transformer for re-ranking in visual place recognition. In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5665--5674, 2023
2023
-
[86]
Visual place recognition: A survey from deep learning perspective
Xiwu Zhang, Lei Wang, and Yan Su. Visual place recognition: A survey from deep learning perspective. Pattern Recognition, 113: 0 107760, 2021
2021
-
[87]
R2former: Unified retrieval and reranking transformer for place recognition
Sijie Zhu, Linjie Yang, Chen Chen, Mubarak Shah, Xiaohui Shen, and Heng Wang. R2former: Unified retrieval and reranking transformer for place recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19370--19380, 2023
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.