REVIEW 4 major objections 5 minor 87 references
Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Combining appearance, geometric, and semantic signals in pose verification improves indoor localization accuracy.
desk verdict A useful, well-executed extension of DensePV combining appearance, normals, and semantics for indoor pose verification, but headline gains are likely optimistic since key variants were tuned on the test set. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is a comparison between the query photo and a synthetic view of the scene rendered from the estimated pose. The baseline DensePV computes the median inverse Euclidean distance between dense RootSIFT descriptors at corresponding pixels, $S_D(x,y,D)=\|d(Q,x,y)-d(Q_D,x,y)\|^{-1}$. The paper's modifications add three mechanisms: (1) the scan-graph, which links each database image to nearby RGB-D panoramic scans with more than 10% visual overlap and merges their 3D points to render a more complete synthetic view; (2) surface-normal consistency, where a predicted normal map $N_Q$ is compared with the rendered normal map $N_D$ by cosine similarity $S_N=N_Q^{\top}N_D$, either as the score itself (DenseNV) or as an attention weight $w=1+\max(0,S_N)/2$ multiplying the appearance similarity (DensePNV); and (3) a semantic mask from an ADE20K-trained scene parser that groups 150 classes into five superclasses and discards pixels labeled people or transient, yielding DensePV+S, DenseNV+S, and DensePNV+S. TrainPV replaces RootSIFT with a fixed fully-convolutional ResNet-18 feature extractor and a small score-regression CNN trained by cross-entropy against softmax distributions of reprojection errors.
What would settle it
Take the same pipeline on an indoor sequence whose RGB-D scans are deliberately misaligned (for example, perturb pairwise scan registrations by 5 cm, 15 cm, and 30 cm) and measure whether the reported gains of DensePV+S and DensePNV over DensePV shrink or vanish; if the multi-modal advantage persists under misregistration, the load-bearing premise is weaker than it appears, and if it collapses, the premise is confirmed.
Extended reading notes
Core claim
The paper's central claim, stated in its abstract and supported by experiments on the InLoc benchmark, is that combining appearance, geometry, and semantics considerably improves pose verification and therefore pose accuracy. Concretely, the strict criterion (0.25 m, 5 deg) rises from 38.9% with the DensePV appearance baseline to 41.3% with DensePV+S using the scan-graph, and the normal-weighted DensePNV beats DensePV by more than five percentage points at several thresholds. The paper also reports that a projective semantic-consistency measure that works outdoors performs worse than the baseline indoors, while using semantics to ignore transient objects improves accuracy. A trainable verification network trained on appearance alone surpasses the original DensePV but not the best hand-crafted multi-modal combinations, and an oracle that picks the best among several variants shows clear remaining headroom.
Load-bearing premise
The whole verification stack re-renders the scene from RGB-D scans and assumes those scans are dense, complete, and accurately registered with respect to one another; when that fails, the synthetic view is wrong, and even a correct pose can be scored low.
Editorial extensions
If this is right
- If the claim is right, the strict InLoc accuracy (0.25 m, 5 deg) improves from 38.9% to 41.3% when appearance verification is augmented with semantics and the scan-graph.
- Normal-weighted appearance (DensePNV) alone surpasses DensePV by more than five percentage points at several thresholds, so geometry helps most where appearance is ambiguous.
- The outdoor-style projective semantic consistency baseline (PSC) does not transfer indoors; semantic information helps only when used to mask unreliable transient regions.
- A trainable pose verifier using only appearance outperforms DensePV but not the hand-crafted multi-modal variants, suggesting that modality fusion, not learned scoring alone, is the active ingredient.
- An oracle that chooses the best among four variants reaches 43.5% at the strict threshold, so correct pose selection still has room to improve beyond any single proposed combination.
Reading between the lines
- Beyond the paper: the oracle experiment implies pixel-level median combination is the bottleneck; a learned or region-level fusion of normals, semantics, and appearance should close part of the gap between 41.3% and the oracle's 43.5%.
- Beyond the paper: since TrainPV's two training-data generation strategies give nearly identical results, the verification score seems insensitive to the training distribution; a zero-shot test on an unseen building would show whether the learned verifier generalizes.
- Beyond the paper: the failure of PSC indoors suggests semantic labels for indoor localization are better treated as reliability masks than as direct geometric evidence; testing semantic-region reprojection consistency would be a natural extension.
- Beyond the paper: the scan-graph's benefit should degrade smoothly with scan misregistration; measuring that degradation curve would let practitioners know which environments need more careful scan alignment before adopting the method.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses camera pose verification in indoor visual localization, specifically the stage of selecting among candidate poses. Building on the InLoc pipeline and its Dense Pose Verification (DensePV), the authors propose hand-crafted verification scores that combine appearance (dense RootSIFT descriptors), surface normals (DenseNV, DensePNV), and semantic masks (DensePV+S, DensePNV+S), as well as a scan-graph that merges multiple RGB-D scans when rendering synthetic views. They also propose TrainPV, a trainable scoring CNN. Experiments on the InLoc dataset report consistent improvements over DensePV, e.g., DensePV+S with the scan-graph reaches 41.3% at the 0.25 m/5° threshold versus 38.9% for DensePV, and DensePNV is stated to give more than 5 percentage point gains at several thresholds. The authors conclude that fusing appearance, geometry, and semantics considerably boosts pose verification and pose accuracy.
Significance. If the claims hold, the paper would be a useful advance in pose verification for indoor localization, showing that simple hand-crafted multi-modal scores can outperform a strong baseline without per-dataset training. The methods are clearly specified, the source code and training data are made public, and the oracle analysis explores complementarity of the proposed scores. These are genuine strengths. However, the experimental evidence for the central claim is weakened by design choices selected on the same 329-query InLoc test set and by the small absolute margins reported without uncertainty quantification. The core idea is defensible, but the paper currently overstates the strength of the evidence for modality fusion.
major comments (4)
- [Section 5 and Appendix C] The semantic mask variant used throughout the paper (variant C) was selected by comparing variants A, B, and C on the InLoc test set, as reported in Appendix C and Table B. This is test-set selection, not a model-agnostic comparison. At the strictest threshold the differences are negligible (variant C versus variant A: 39.8% vs. 39.8% at [0.25, 5] and 57.8% vs. 57.4% at [0.50, 5]), so the evidence that semantic masking contributes beyond appearance at this operating point is not established at the reported scale. The paper should validate the chosen mask design on a held-out subset, or at minimum report confidence intervals and test-set-selection-corrected comparisons.
- [Appendix B and Section 3.2] The no-crop modification to the normal-estimation pipeline was chosen by comparing cropped and non-cropped variants on the InLoc test set (Table A), and the 10% scan-graph overlap threshold is another design choice evaluated on the same benchmark. Thus the introduction's claim that the approaches 'do not require fine-tuning on the actual dataset' is only partially true: while no network weights are learned, discrete hyperparameters are tuned on the test queries. This creates a real risk that the reported gains are optimistic. The paper should provide a sensitivity analysis over the overlap threshold and an honest statement that discrete choices were selected on the test set, or evaluate the full pipeline on a separate scene or held-out query split.
- [Table 1] No uncertainty quantification is provided for any of the reported percentages, and all results appear to be single-run point estimates on 329 query images. The headline improvement of DensePV+S with the scan-graph over DensePV at [0.25, 5] is 2.4 percentage points, roughly 8 queries, and the [0.50, 5] improvement is 6.1 percentage points, roughly 20 queries. Without confidence intervals, repeated runs, or paired per-query analyses, it is not possible to determine whether these differences exceed chance, especially after comparing many variants and thresholds. Please report bootstrap confidence intervals or a paired significance test across queries.
- [Section 5, oracle analysis] The oracle upper bound is constructed from only four of the proposed variants (DensePV, DensePV with scan-graph, DensePV+S with scan-graph, and DensePNV with scan-graph), and the observation that DenseNV+S provides better poses than this oracle for about 9% of queries is then used to argue that the modalities are complementary. Since DenseNV+S is excluded from the oracle, this is not an inconsistency, but the oracle is not an upper bound over all proposed variants, so its support for the complementarity claim is weaker than the text implies.
minor comments (5)
- [Section 3.2] The 5x5 patch size used for surface normal estimation is presented without an ablation or reference; please clarify whether this choice is standard or was tuned.
- [Equation (10)] The relative reprojection error \tilde r_i = r_i / \min_k r_k can be undefined if the minimum error is zero; please add a small epsilon or a note about this edge case.
- [Reference list] Reference [82] contains a typo in the author list: 'Alexander Sax, , William B. Shen' has a double comma that should be removed.
- [Appendix D] The numbered list in the training-data paragraph is numbered 1), 2), 4) and skips 3); please renumber or merge the steps.
- [Abstract and Conclusion] The phrase 'significant improvements' is used in a statistical sense, but the paper reports no significance tests; please rephrase to 'consistent improvements' or add statistical support.
Circularity Check
No circularity: verification scores are defined from rendered views and compared with an external baseline; test-set variant selection is a limitation but not a definitional reduction.
full rationale
The paper's derivation chain is empirical rather than formal. The proposed verification scores (DensePV, DenseNV, DensePNV, and their semantic variants) are closed-form functions of query and re-rendered database descriptors, normals, and semantic masks; the final pose is selected by maximizing these scores. No score is defined using the InLoc ground-truth poses that are later reported in Table 1, and no scalar parameter is fitted to the 329 test queries. The trainable variant TrainPV is trained on separate video sequences with manually verified poses (Appendix D), and its feature-extraction weights are fixed during training, so its evaluation is not a restatement of its training signal. The only load-bearing references to the authors' prior work [72] are as a baseline (DensePV), a pipeline, and a benchmark; these are independently published artifacts, and the paper's improvement claim is an empirical comparison against that baseline rather than a theorem derived from it. The paper itself discloses the relevant limitations: scan density and registration assumptions in Sec. 3.2, and the selection of the semantic-mask variant (Appendix C, Table B) and no-crop normal estimation (Appendix B, Table A) using the InLoc test set. Those passages indicate a genuine risk that test-set selection inflates the reported margins, but choosing among discrete variants that generally improve over the baseline is not equivalent to defining the reported accuracy by construction. No equation reduces to a fitted input, and no load-bearing self-citation chain forces the result. Therefore there is no significant circularity.
Assumptions & free parameters
free parameters (6)
- Scan-graph overlap threshold =
10% of database image pixels
- Surface normal patch size =
5x5 pixels
- Normal weighting constants =
w = (1 + max(0, SN)) / 2
- Semantic superclass mapping =
150 ADE20K classes grouped into 5 superclasses; variant C in Appendix C
- Normal input image scale =
Longer side 256 pixels
- TrainPV training schedule =
10 epochs, Adam learning rate 1e-5
assumptions (5)
- domain assumption RGB-D scans are dense, complete, and registered accurately with respect to each other.
- domain assumption Pretrained Taskonomy normal predictor and ADE20K semantic segmentation transfer to indoor scenes.
- domain assumption Scene changes can be captured by ignoring people and transient superclasses.
- standard math Standard RANSAC, P3P, RootSIFT, and median aggregation behave as expected.
- domain assumption InLoc reference poses are accurate enough to serve as ground truth.
Cite this review
Pith. "Pith review of Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization." pith.science (2026). https://pith.science/paper/DAR2CEMY
@misc{pith2026190804598,
author = {Pith},
title = {Pith review of: Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization},
year = {2026},
howpublished = {\url{https://pith.science/paper/DAR2CEMY}},
note = {Machine review of arXiv:1908.04598}
}
read the original abstract
Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common strategy is, first, to generate multiple estimates for the camera pose from which a given query image was taken. The pose with the largest geometric consistency with the query image, e.g., in the form of an inlier count, is then selected in a second stage. While a significant amount of research has concentrated on the first stage, there is considerably less work on the second stage. In this paper, we thus focus on pose verification. We show that combining different modalities, namely appearance, geometry, and semantics, considerably boosts pose verification and consequently pose accuracy. We develop multiple hand-crafted as well as a trainable approach to join into the geometric-semantic verification and show significant improvements over state-of-the-art on a very challenging indoor dataset.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Arandjelovi ´c, Petr Gronat, Akihiko Torii, Tomas Pa- jdla, and Josef Sivic
Relja. Arandjelovi ´c, Petr Gronat, Akihiko Torii, Tomas Pa- jdla, and Josef Sivic. NetVLAD: CNN architecture for Query DensePV [72] DensePV w/ scan-graph DensePV+S DensePNV DensePNV+S 2.99m, 17.64◦ 0.42m, 1.25◦ 0.42m, 1.25◦ 0.42m, 1.25◦ 0.42m, 1.25◦ (a) 12.39m, 26.71◦ 12.39m, 26.71◦ 0.40m, 2.37◦ 0.40m, 2.37◦ 0.40m, 2.37◦ (b) 4.70m, 88.15◦ 14.29m, 9.02◦ 5...
2016
-
[2]
Three things ev- eryone should know to improve object retrieval
Relja Arandjelovi ´c and Andrew Zisserman. Three things ev- eryone should know to improve object retrieval. In Proc. CVPR, 2012
2012
-
[3]
All about VLAD
Relja Arandjelovic and Andrew Zisserman. All about VLAD. In Proc. CVPR, 2013
2013
-
[4]
Dislocation: Scalable descriptor distinctiveness for location recognition
Relja Arandjelovi ´c and Andrew Zisserman. Dislocation: Scalable descriptor distinctiveness for location recognition. In Proc. ACCV, 2014
2014
-
[5]
Visual vocabu- lary with a semantic twist
Relja Arandjelovi ´c and Andrew Zisserman. Visual vocabu- lary with a semantic twist. In Proc. ACCV, 2014
2014
-
[6]
GIS-Assisted Object Detection and Geospatial Localization
Shervin Ardeshir, Amir Roshan Zamir, Alejandro Torroella, and Mubarak Shah. GIS-Assisted Object Detection and Geospatial Localization. In Proc. ECCV, 2014
2014
-
[7]
Nikolay Atanasov, Menglong Zhu, Kostas Daniilidis, and George J. Pappas. Localization from semantic observations via the matrix permanent. Intl. J. of Robotics Research, 35(1- 3):73–99, 2016
2016
-
[8]
Mathieu Aubry, Bryan C. Russell, and Josef Sivic. Painting- to-3D Model Alignment via Discriminative Visual Elements. ACM Trans. Graph., 33(2):14:1–14:14, Apr 2014
work page 2014
Show all 87 references
-
[9]
RelocNet: Continuous Metric Learning Relocalisation using Neural Nets
Vassileios Balntas, Shuda Li, and Victor Adrian Prisacariu. RelocNet: Continuous Metric Learning Relocalisation using Neural Nets. In Proc. ECCV, 2018
2018
-
[10]
DSAC - Differentiable RANSAC for Camera Local- ization
Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. DSAC - Differentiable RANSAC for Camera Local- ization. In Proc. CVPR, 2017
2017
-
[11]
Learning less is more- 6D camera localization via 3d surface regression
Eric Brachmann and Carsten Rother. Learning less is more- 6D camera localization via 3d surface regression. In Proc. CVPR, 2018
2018
-
[12]
Geometry-Aware Learning of Maps for Cam- era Localization
Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. Geometry-Aware Learning of Maps for Cam- era Localization. In Proc. CVPR, 2018
2018
-
[13]
Graph-based discriminative learning for location recognition
Song Cao and Noah Snavely. Graph-based discriminative learning for location recognition. In Proc. CVPR, 2013
2013
-
[14]
Minimal scene descriptions from structure from motion models
Song Cao and Noah Snavely. Minimal scene descriptions from structure from motion models. In Proc. CVPR, 2014
2014
-
[15]
Robert Castle, Georg Klein, and David W. Murray. Video- rate localization in multiple maps for wearable augmented reality. In ISWC, 2008
2008
-
[16]
Lord, Julien Valentin, Luigi Di Stefano, and Philip H
Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien Valentin, Luigi Di Stefano, and Philip H. S. Torr. On- The-Fly Adaptation of Regression Forests for Online Cam- era Relocalisation. In Proc. CVPR, 2017
2017
-
[17]
Chen, Georges Baatz, Kevin K ¨oser, Sam S Tsai, Ramakrishna Vedantham, Timo Pylv ¨an¨ainen, Kimmo Roimela, Xin Chen, Jeff Bach, Marc Pollefeys, et al
David M. Chen, Georges Baatz, Kevin K ¨oser, Sam S Tsai, Ramakrishna Vedantham, Timo Pylv ¨an¨ainen, Kimmo Roimela, Xin Chen, Jeff Bach, Marc Pollefeys, et al. City- scale landmark identification on mobile devices. In Proc. CVPR, 2011
2011
-
[18]
Matching with PROSAC- progressive sample consensus
Ond ˇrej Chum and Ji ˇr´ı Matas. Matching with PROSAC- progressive sample consensus. In Proc. CVPR, 2005
2005
-
[19]
Optimal randomized RANSAC
Ond ˇrej Chum and Ji ˇr´ı Matas. Optimal randomized RANSAC. IEEE PAMI, 30(8):1472–1482, 2008
2008
-
[20]
Total recall II: Query expansion revisited
Ond ˇrej Chum, Andrej Mikulik, Michal Perdoch, and Ji ˇr´ı Matas. Total recall II: Query expansion revisited. In Proc. CVPR, 2011
2011
-
[21]
Merg- ing the Unmatchable: Stitching Visually Disconnected SfM Models
Andrea Cohen, Torsten Sattler, and Mark Pollefeys. Merg- ing the Unmatchable: Stitching Visually Disconnected SfM Models. In Proc. ICCV, 2015
2015
-
[22]
Indoor-Outdoor 3D Reconstruction Alignment
Andrea Cohen, Johannes Lutz Sch ¨onberger, Pablo Speciale, Torsten Sattler, Jan-Michael Frahm, and Marc Pollefeys. Indoor-Outdoor 3D Reconstruction Alignment. In Proc. ECCV, 2016
2016
-
[23]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. CVPR, 2009
2009
-
[24]
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In Proc. ICCV, pages 2758–2766, 2015
2015
-
[25]
Fischler and Robert C
Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Comm. ACM, 24(6):381–395, 1981
1981
-
[26]
Learning and calibrating per-location classifiers for visual place recognition
Petr Gronat, Guillaume Obozinski, Josef Sivic, and Tomas Pajdla. Learning and calibrating per-location classifiers for visual place recognition. In Proc. CVPR, 2013
2013
-
[27]
Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael N¨olle
Bert M. Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael N¨olle. Review and analysis of solutions of the three point perspective pose estimation problem.IJCV, 13(3):331– 356, 1994
1994
-
[28]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. CVPR, 2016
2016
-
[29]
From Structure-from-Motion point clouds to fast location recognition
Arnold Irschara, Christopher Zach, Jan-Michael Frahm, and Horst Bischof. From Structure-from-Motion point clouds to fast location recognition. In Proc. CVPR, 2009
2009
-
[30]
Ham- ming embedding and weak geometric consistency for large scale image search
Herve Jegou, Matthijs Douze, and Cordelia Schmid. Ham- ming embedding and weak geometric consistency for large scale image search. In Proc. ECCV, 2008
2008
-
[31]
Packing bag-of-features
Herv ´e J´egou, Matthijs Douze, and Cordelia Schmid. Packing bag-of-features. In Proc. ICCV, 2009
2009
-
[32]
Geometric loss func- tions for camera pose regression with deep learning
Alex Kendall and Roberto Cipolla. Geometric loss func- tions for camera pose regression with deep learning. InProc. CVPR, 2017
2017
-
[33]
Posenet: A convolutional network for real-time 6-dof camera relocalization
Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. In Proc. ICCV, 2015
2015
-
[34]
Learned contextual feature reweighting for image geo- localization
Hyo Jin Kim, Enrique Dunn, and Jan-Michael Frahm. Learned contextual feature reweighting for image geo- localization. In Proc. CVPR, 2017
2017
-
[35]
Kneip, D
L. Kneip, D. Scaramuzza, and R. Siegwart. A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orienta- tion. In Proc. CVPR, 2011
2011
-
[36]
Avoiding confus- ing features in place recognition
Jan Knopp, Josef Sivic, and Tomas Pajdla. Avoiding confus- ing features in place recognition. In Proc. ECCV, 2010
2010
-
[37]
Matching Features Correctly through Semantic Un- derstanding
Nikolay Kobyshev, Hayko Riemenschneider, and Luc Van Gool. Matching Features Correctly through Semantic Un- derstanding. In Proc. 3DV, 2014
2014
-
[38]
Real- Time Solution to the Absolute Pose Problem with Unknown Radial Distortion and Focal Length
Zuzana Kukelova, Martin Bujnak, and Tomas Pajdla. Real- Time Solution to the Absolute Pose Problem with Unknown Radial Distortion and Focal Length. In Proc. ICCV, 2013
2013
-
[39]
Huttenlocher
Yunpeng Li, Noah Snavely, and Daniel P. Huttenlocher. Lo- cation recognition using prioritized feature matching. In Proc. ECCV, 2010
2010
-
[40]
Huttenlocher, and Pas- cal Fua
Yunpeng Li, Noah Snavely, Daniel P. Huttenlocher, and Pas- cal Fua. Worldwide pose estimation using 3d point clouds. In Proc. ECCV, 2012
2012
-
[41]
Sinha, Michael F
Hyon Lim, Sudipta N. Sinha, Michael F. Cohen, and Matthew Uyttendaele. Real-time image-based 6-DOF local- ization in large-scale environments. In Proc. CVPR, 2012
2012
-
[42]
Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D Map
Liu Liu, Hongdong Li, and Yuchao Dai. Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D Map. In Proc. ICCV, 2017
2017
-
[43]
David G. Lowe. Distinctive image features from scale- invariant keypoints. IJCV, 60(2):91–110, 2004
2004
-
[44]
Hesch, Marc Pollefeys, and Roland Siegwart
Simon Lynen, Torsten Sattler, Michael Bosse, Joel A. Hesch, Marc Pollefeys, and Roland Siegwart. Get Out of My Lab: Large-scale, Real-Time Visual-Inertial Localization. In Proc. RSS, 2015
2015
-
[45]
Daniela Massiceti, Alexander Krull, Eric Brachmann, Carsten Rother, and Philip H.S. Torr. Random Forests versus Neural Networks - What’s Best for Camera Relocalization? In Proc. Intl. Conf. on Robotics and Automation, 2017
2017
-
[46]
Little, Julien Valentin, and Clarence W
Lili Meng, Jianhui Chen, Frederick Tung, James J. Little, Julien Valentin, and Clarence W. de Silva. Backtracking Regression Forests for Accurate Camera Relocalization. In Proc. IEEE/RSJ Conf. on Intelligent Robots and Systems , 2017
2017
-
[47]
Little, Julien Valentin, and Clarence W
Lili Meng, Frederick Tung, James J. Little, Julien Valentin, and Clarence W. de Silva. Exploiting Points and Lines in Re- gression Forests for RGB-D Camera Relocalization. InProc. IEEE/RSJ Conf. on Intelligent Robots and Systems, 2018
2018
-
[48]
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. In NIPS-W, 2017
2017
-
[49]
Object retrieval with large vocabularies and fast spatial matching
James Philbin, Ond ˇrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman. Object retrieval with large vocabularies and fast spatial matching. In Proc. CVPR, 2007
2007
-
[50]
VLocNet++: Deep multitask learning for semantic visual localization and odometry
Noha Radwan, Abhinav Valada, and Wolfram Burgard. VLocNet++: Deep multitask learning for semantic visual localization and odometry. IEEE Robotics And Automation Letters (RA-L), 3(4):4407–4414, 2018
2018
-
[51]
Convo- lutional neural network architecture for geometric matching
Ignacio Rocco, Relja Arandjelovi ´c, and Josef Sivic. Convo- lutional neural network architecture for geometric matching. In Proc. CVPR, 2017
2017
-
[52]
Salas-Moreno, Richard A
Renato F. Salas-Moreno, Richard A. Newcombe, Hauke Strasdat, Paul H. J. Kelly, and Andrew J. Davison. SLAM++: Simultaneous Localisation and Mapping at the Level of Ob- jects. In Proc. CVPR, 2013
2013
-
[53]
From coarse to fine: Robust hierarchical localization at large scale
Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Proc. CVPR, 2019
2019
-
[54]
Hyperpoints and fine vocab- ularies for large-scale location recognition
Torsten Sattler, Michal Havlena, Filip Radenovic, Konrad Schindler, and Marc Pollefeys. Hyperpoints and fine vocab- ularies for large-scale location recognition. In Proc. ICCV, 2015
2015
-
[55]
Large-scale location recognition and the geomet- ric burstiness problem
Torsten Sattler, Michal Havlena, Konrad Schindler, and Marc Pollefeys. Large-scale location recognition and the geomet- ric burstiness problem. In Proc. CVPR, 2016
2016
-
[56]
Efficient & effective prioritized matching for large-scale image-based localization
Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Efficient & effective prioritized matching for large-scale image-based localization. IEEE PAMI, 39(9):1744–1756, 2017
2017
-
[57]
Benchmarking 6DOF outdoor visual local- ization in changing conditions
Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, and Tomas Pajdla. Benchmarking 6DOF outdoor visual local- ization in changing conditions. In Proc. CVPR, 2018
2018
-
[58]
Are large-scale 3D models really necessary for accurate visual localization? In Proc
Torsten Sattler, Akihiko Torii, Josef Sivic, Marc Pollefeys, Hajime Taira, Masatoshi Okutomi, and Tomas Pajdla. Are large-scale 3D models really necessary for accurate visual localization? In Proc. CVPR, 2017
2017
-
[59]
Understanding the Limitations of CNN-based Absolute Camera Pose Regression
Torsten Sattler, Qunjie Zhou, Mark Pollefeys, and Laura Leal-Taix´e. Understanding the Limitations of CNN-based Absolute Camera Pose Regression. In Proc. CVPR, 2019
2019
-
[60]
City-Scale Location Recognition
Grant Schindler, Matthew Brown, and Richard Szeliski. City-Scale Location Recognition. In Proc. CVPR, 2007
2007
-
[61]
Structure-From-Motion Revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-From-Motion Revisited. In Proc. CVPR, 2016
2016
-
[62]
Semantic Visual Localization
Johannes Lutz Sch ¨onberger, Marc Pollefeys, Andreas Geiger, and Torsten Sattler. Semantic Visual Localization. In Proc. CVPR, 2018
2018
-
[63]
Lane- Loc: Lane marking based localization using highly accurate maps
Markus Schreiber, Carsten Kn ¨oppel, and Uwe Franke. Lane- Loc: Lane marking based localization using highly accurate maps. In Proc. IV, 2013
2013
-
[64]
Qi Shan, Changchang Wu, Brian Curless, Yasutaka Fu- rukawa, Carlos Hernandez, and Steven M. Seitz. Accurate Geo-Registration by Ground-to-Aerial Image Matching. In Proc. 3DV, 2014
2014
-
[65]
Scene co- ordinate regression forests for camera relocalization in RGB- D images
Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. Scene co- ordinate regression forests for camera relocalization in RGB- D images. In Proc. CVPR, 2013
2013
-
[66]
SIFT-realistic rendering
Dominik Sibbing, Torsten Sattler, Bastian Leibe, and Leif Kobbelt. SIFT-realistic rendering. In Proc. 3DV, 2013
2013
-
[67]
Very deep convo- lutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convo- lutional networks for large-scale image recognition. In Proc. ICLR, 2015
2015
-
[68]
Semantically Guided Geo- location and Modeling in Urban Environments
Gautam Singh and Jana Ko ˇseck´a. Semantically Guided Geo- location and Modeling in Urban Environments. In Large- Scale Visual Geo-Localization, 2016
2016
-
[69]
Video google: A text retrieval approach to object matching in videos
Josef Sivic and Andrew Zisserman. Video google: A text retrieval approach to object matching in videos. In Proc. ICCV, 2003
2003
-
[70]
Stenborg, C
E. Stenborg, C. Toft, and L. Hammarstrand. Long-term Vi- sual Localization using Semantically Segmented Images. In Proc. Intl. Conf. on Robotics and Automation, 2018
2018
-
[71]
City-Scale Localization for Cameras with Known Vertical Direction
Linus Sv ¨arm, Olof Enqvist, Fredrik Kahl, and Magnus Os- karsson. City-Scale Localization for Cameras with Known Vertical Direction. IEEE PAMI, 39(7):1455–1461, 2017
2017
-
[72]
InLoc: Indoor visual localization with dense matching and view synthesis
Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Ak- ihiko Torii. InLoc: Indoor visual localization with dense matching and view synthesis. In Proc. CVPR, 2018
2018
-
[73]
Long-term 3D Localization and Pose from Semantic Labellings
Carl Toft, Carl Olsson, and Fredrik Kahl. Long-term 3D Localization and Pose from Semantic Labellings. In Proc. ICCV Workshops, 2017
2017
-
[74]
Semantic Match Consistency for Long-Term Visual Localization
Carl Toft, Erik Stenborg, Lars Hammarstrand, Lucas Brynte, Marc Pollefeys, Torsten Sattler, and Fredrik Kahl. Semantic Match Consistency for Long-Term Visual Localization. In Proc. ECCV, 2018
2018
-
[75]
Local visual query expan- sion: Exploiting an image collection to refine local descrip- tors
Giorgos Tolias and Herv ´e J´egou. Local visual query expan- sion: Exploiting an image collection to refine local descrip- tors. Technical Report RR-8325, INRIA, 2013
2013
-
[76]
24/7 place recognition by view synthesis
Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, and Tomas Pajdla. 24/7 place recognition by view synthesis. In Proc. CVPR, 2015
2015
-
[77]
Visual place recognition with repetitive structures
Akihiko Torii, Josef Sivic, Tomas Pajdla, and Masatoshi Okutomi. Visual place recognition with repetitive structures. In Proc. CVPR, 2013
2013
-
[78]
Image- Based Localization Using LSTMs for Structured Feature Correlation
Florian Walch, Caner Hazirbas, Laura Leal-Taix ´e, Torsten Sattler, Sebastian Hilsenbeck, and Daniel Cremers. Image- Based Localization Using LSTMs for Structured Feature Correlation. In Proc. ICCV, 2017
2017
-
[79]
Exploiting 2D floor- plan for building-scale panorama RGBD alignment
Erik Wijmans and Yasutaka Furukawa. Exploiting 2D floor- plan for building-scale panorama RGBD alignment. InProc. CVPR, 2017
2017
-
[80]
Funkhouser
Fisher Yu, Jianxiong Xiao, and Thomas A. Funkhouser. Se- mantic alignment of LiDAR data at city scale. In Proc. CVPR, 2015
2015
-
[81]
X. Yu, S. Chaturvedi, C. Feng, Y . Taguchi, T.-Y . Lee, C. Fer- nandes, and S. Ramalingam. VLASE: Vehicle Localization by Aggregating Semantic Edges. In Proc. IEEE/RSJ Conf. on Intelligent Robots and Systems, 2018
2018
-
[82]
Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese
Amir Roshan Zamir, Alexander Sax, , William B. Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In Proc. CVPR, 2018
2018
-
[83]
Accurate image lo- calization based on google maps street view
Amir Roshan Zamir and Mubarak Shah. Accurate image lo- calization based on google maps street view. InProc. ECCV, 2010
2010
-
[84]
Camera Pose V oting for Large-Scale Image-Based Localization
Bernhard Zeisl, Torsten Sattler, and Marc Pollefeys. Camera Pose V oting for Large-Scale Image-Based Localization. In Proc. ICCV, 2015
2015
-
[85]
Pyramid scene parsing network
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proc. CVPR, 2017
2017
-
[86]
Scene parsing through ADE20K dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. In Proc. CVPR, 2017
2017
-
[87]
Semantic un- derstanding of scenes through the ADE20K dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fi- dler, Adela Barriuso, and Antonio Torralba. Semantic un- derstanding of scenes through the ADE20K dataset. IJCV, 127(3):302–321, 2019
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.