Pith. sign in

REVIEW 4 major objections 5 minor 87 references

Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Combining appearance, geometric, and semantic signals in pose verification improves indoor localization accuracy.

desk verdict A useful, well-executed extension of DensePV combining appearance, normals, and semantics for indoor pose verification, but headline gains are likely optimistic since key variants were tuned on the test set. read the letter →

arxiv 1908.04598 v2 pith:DAR2CEMY submitted 2019-08-13 cs.CV

classification cs.CV
keywords visuallocalizationposeverificationindoorRGB-DscanssurfacenormalssemanticsegmentationviewsynthesisInLocdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper targets the final step of indoor visual localization: after several candidate camera poses have been generated, choosing which one actually matches the query photo. It argues that the usual choice criterion, comparing the query image with a view re-rendered from RGB-D scans using appearance only, is too weak in rooms with repetitive patterns and plain walls. The central claim is that verifying a pose with appearance together with surface-normal consistency and semantic information—plus merging nearby scans to fill the rendered view—raises both pose-selection accuracy and final localization accuracy. If the claim holds, indoor localization systems should treat pose verification as a multi-modal fusion problem, not a pure appearance-matching problem.

What carries the argument

The engine is a comparison between the query photo and a synthetic view of the scene rendered from the estimated pose. The baseline DensePV computes the median inverse Euclidean distance between dense RootSIFT descriptors at corresponding pixels, $S_D(x,y,D)=\|d(Q,x,y)-d(Q_D,x,y)\|^{-1}$. The paper's modifications add three mechanisms: (1) the scan-graph, which links each database image to nearby RGB-D panoramic scans with more than 10% visual overlap and merges their 3D points to render a more complete synthetic view; (2) surface-normal consistency, where a predicted normal map $N_Q$ is compared with the rendered normal map $N_D$ by cosine similarity $S_N=N_Q^{\top}N_D$, either as the score itself (DenseNV) or as an attention weight $w=1+\max(0,S_N)/2$ multiplying the appearance similarity (DensePNV); and (3) a semantic mask from an ADE20K-trained scene parser that groups 150 classes into five superclasses and discards pixels labeled people or transient, yielding DensePV+S, DenseNV+S, and DensePNV+S. TrainPV replaces RootSIFT with a fixed fully-convolutional ResNet-18 feature extractor and a small score-regression CNN trained by cross-entropy against softmax distributions of reprojection errors.

What would settle it

Take the same pipeline on an indoor sequence whose RGB-D scans are deliberately misaligned (for example, perturb pairwise scan registrations by 5 cm, 15 cm, and 30 cm) and measure whether the reported gains of DensePV+S and DensePNV over DensePV shrink or vanish; if the multi-modal advantage persists under misregistration, the load-bearing premise is weaker than it appears, and if it collapses, the premise is confirmed.

Watch

Extended reading notes

Core claim

The paper's central claim, stated in its abstract and supported by experiments on the InLoc benchmark, is that combining appearance, geometry, and semantics considerably improves pose verification and therefore pose accuracy. Concretely, the strict criterion (0.25 m, 5 deg) rises from 38.9% with the DensePV appearance baseline to 41.3% with DensePV+S using the scan-graph, and the normal-weighted DensePNV beats DensePV by more than five percentage points at several thresholds. The paper also reports that a projective semantic-consistency measure that works outdoors performs worse than the baseline indoors, while using semantics to ignore transient objects improves accuracy. A trainable verification network trained on appearance alone surpasses the original DensePV but not the best hand-crafted multi-modal combinations, and an oracle that picks the best among several variants shows clear remaining headroom.

Load-bearing premise

The whole verification stack re-renders the scene from RGB-D scans and assumes those scans are dense, complete, and accurately registered with respect to one another; when that fails, the synthetic view is wrong, and even a correct pose can be scored low.

Editorial extensions

If this is right

  • If the claim is right, the strict InLoc accuracy (0.25 m, 5 deg) improves from 38.9% to 41.3% when appearance verification is augmented with semantics and the scan-graph.
  • Normal-weighted appearance (DensePNV) alone surpasses DensePV by more than five percentage points at several thresholds, so geometry helps most where appearance is ambiguous.
  • The outdoor-style projective semantic consistency baseline (PSC) does not transfer indoors; semantic information helps only when used to mask unreliable transient regions.
  • A trainable pose verifier using only appearance outperforms DensePV but not the hand-crafted multi-modal variants, suggesting that modality fusion, not learned scoring alone, is the active ingredient.
  • An oracle that chooses the best among four variants reaches 43.5% at the strict threshold, so correct pose selection still has room to improve beyond any single proposed combination.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the oracle experiment implies pixel-level median combination is the bottleneck; a learned or region-level fusion of normals, semantics, and appearance should close part of the gap between 41.3% and the oracle's 43.5%.
  • Beyond the paper: since TrainPV's two training-data generation strategies give nearly identical results, the verification score seems insensitive to the training distribution; a zero-shot test on an unseen building would show whether the learned verifier generalizes.
  • Beyond the paper: the failure of PSC indoors suggests semantic labels for indoor localization are better treated as reliability masks than as direct geometric evidence; testing semantic-region reprojection consistency would be a natural extension.
  • Beyond the paper: the scan-graph's benefit should degrade smoothly with scan misregistration; measuring that degradation curve would let practitioners know which environments need more careful scan alignment before adopting the method.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper addresses camera pose verification in indoor visual localization, specifically the stage of selecting among candidate poses. Building on the InLoc pipeline and its Dense Pose Verification (DensePV), the authors propose hand-crafted verification scores that combine appearance (dense RootSIFT descriptors), surface normals (DenseNV, DensePNV), and semantic masks (DensePV+S, DensePNV+S), as well as a scan-graph that merges multiple RGB-D scans when rendering synthetic views. They also propose TrainPV, a trainable scoring CNN. Experiments on the InLoc dataset report consistent improvements over DensePV, e.g., DensePV+S with the scan-graph reaches 41.3% at the 0.25 m/5° threshold versus 38.9% for DensePV, and DensePNV is stated to give more than 5 percentage point gains at several thresholds. The authors conclude that fusing appearance, geometry, and semantics considerably boosts pose verification and pose accuracy.

Significance. If the claims hold, the paper would be a useful advance in pose verification for indoor localization, showing that simple hand-crafted multi-modal scores can outperform a strong baseline without per-dataset training. The methods are clearly specified, the source code and training data are made public, and the oracle analysis explores complementarity of the proposed scores. These are genuine strengths. However, the experimental evidence for the central claim is weakened by design choices selected on the same 329-query InLoc test set and by the small absolute margins reported without uncertainty quantification. The core idea is defensible, but the paper currently overstates the strength of the evidence for modality fusion.

major comments (4)
  1. [Section 5 and Appendix C] The semantic mask variant used throughout the paper (variant C) was selected by comparing variants A, B, and C on the InLoc test set, as reported in Appendix C and Table B. This is test-set selection, not a model-agnostic comparison. At the strictest threshold the differences are negligible (variant C versus variant A: 39.8% vs. 39.8% at [0.25, 5] and 57.8% vs. 57.4% at [0.50, 5]), so the evidence that semantic masking contributes beyond appearance at this operating point is not established at the reported scale. The paper should validate the chosen mask design on a held-out subset, or at minimum report confidence intervals and test-set-selection-corrected comparisons.
  2. [Appendix B and Section 3.2] The no-crop modification to the normal-estimation pipeline was chosen by comparing cropped and non-cropped variants on the InLoc test set (Table A), and the 10% scan-graph overlap threshold is another design choice evaluated on the same benchmark. Thus the introduction's claim that the approaches 'do not require fine-tuning on the actual dataset' is only partially true: while no network weights are learned, discrete hyperparameters are tuned on the test queries. This creates a real risk that the reported gains are optimistic. The paper should provide a sensitivity analysis over the overlap threshold and an honest statement that discrete choices were selected on the test set, or evaluate the full pipeline on a separate scene or held-out query split.
  3. [Table 1] No uncertainty quantification is provided for any of the reported percentages, and all results appear to be single-run point estimates on 329 query images. The headline improvement of DensePV+S with the scan-graph over DensePV at [0.25, 5] is 2.4 percentage points, roughly 8 queries, and the [0.50, 5] improvement is 6.1 percentage points, roughly 20 queries. Without confidence intervals, repeated runs, or paired per-query analyses, it is not possible to determine whether these differences exceed chance, especially after comparing many variants and thresholds. Please report bootstrap confidence intervals or a paired significance test across queries.
  4. [Section 5, oracle analysis] The oracle upper bound is constructed from only four of the proposed variants (DensePV, DensePV with scan-graph, DensePV+S with scan-graph, and DensePNV with scan-graph), and the observation that DenseNV+S provides better poses than this oracle for about 9% of queries is then used to argue that the modalities are complementary. Since DenseNV+S is excluded from the oracle, this is not an inconsistency, but the oracle is not an upper bound over all proposed variants, so its support for the complementarity claim is weaker than the text implies.
minor comments (5)
  1. [Section 3.2] The 5x5 patch size used for surface normal estimation is presented without an ablation or reference; please clarify whether this choice is standard or was tuned.
  2. [Equation (10)] The relative reprojection error \tilde r_i = r_i / \min_k r_k can be undefined if the minimum error is zero; please add a small epsilon or a note about this edge case.
  3. [Reference list] Reference [82] contains a typo in the author list: 'Alexander Sax, , William B. Shen' has a double comma that should be removed.
  4. [Appendix D] The numbered list in the training-data paragraph is numbered 1), 2), 4) and skips 3); please renumber or merge the steps.
  5. [Abstract and Conclusion] The phrase 'significant improvements' is used in a statistical sense, but the paper reports no significance tests; please rephrase to 'consistent improvements' or add statistical support.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: verification scores are defined from rendered views and compared with an external baseline; test-set variant selection is a limitation but not a definitional reduction.

full rationale

The paper's derivation chain is empirical rather than formal. The proposed verification scores (DensePV, DenseNV, DensePNV, and their semantic variants) are closed-form functions of query and re-rendered database descriptors, normals, and semantic masks; the final pose is selected by maximizing these scores. No score is defined using the InLoc ground-truth poses that are later reported in Table 1, and no scalar parameter is fitted to the 329 test queries. The trainable variant TrainPV is trained on separate video sequences with manually verified poses (Appendix D), and its feature-extraction weights are fixed during training, so its evaluation is not a restatement of its training signal. The only load-bearing references to the authors' prior work [72] are as a baseline (DensePV), a pipeline, and a benchmark; these are independently published artifacts, and the paper's improvement claim is an empirical comparison against that baseline rather than a theorem derived from it. The paper itself discloses the relevant limitations: scan density and registration assumptions in Sec. 3.2, and the selection of the semantic-mask variant (Appendix C, Table B) and no-crop normal estimation (Appendix B, Table A) using the InLoc test set. Those passages indicate a genuine risk that test-set selection inflates the reported margins, but choosing among discrete variants that generally improve over the baseline is not equivalent to defining the reported accuracy by construction. No equation reduces to a fitted input, and no load-bearing self-citation chain forces the result. Therefore there is no significant circularity.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

No new physical entities are introduced. The central claim rests on a set of hand-selected thresholds and pretrained networks; the most fragile is the scan-graph registration assumption, which the paper itself flags.

free parameters (6)
  • Scan-graph overlap threshold = 10% of database image pixels
    Edge criterion in Sec. 3.2; chosen by hand, not cross-validated, and affects how many scans contribute to the rendered view.
  • Surface normal patch size = 5x5 pixels
    Used to fit local planes when rendering normal maps in Sec. 3.2.
  • Normal weighting constants = w = (1 + max(0, SN)) / 2
    Eq. 6; hand-designed attention weights that combine normal consistency with descriptor similarity.
  • Semantic superclass mapping = 150 ADE20K classes grouped into 5 superclasses; variant C in Appendix C
    The assignment of classes to people, transient, stable, fixed, and outdoor, and the choice to ignore people plus transient, was selected after comparing variants A, B, and C on the test set.
  • Normal input image scale = Longer side 256 pixels
    Implementation detail in Sec. 3.2; cropping is avoided based on test-set comparisons in Appendix B.
  • TrainPV training schedule = 10 epochs, Adam learning rate 1e-5
    Implementation detail for the learned verifier in Sec. 4; set by hand and likely not critical to the central claim.
assumptions (5)
  • domain assumption RGB-D scans are dense, complete, and registered accurately with respect to each other.
    Sec. 3.2: view synthesis for verification relies on projecting multiple scans into the query. The paper admits these assumptions do not always hold.
  • domain assumption Pretrained Taskonomy normal predictor and ADE20K semantic segmentation transfer to indoor scenes.
    Sec. 3.2 and 3.3: normal and semantic maps for query and database images come from networks trained on other datasets, with no fine-tuning on InLoc.
  • domain assumption Scene changes can be captured by ignoring people and transient superclasses.
    Sec. 3.3: semantics are used only to mask out pixels; if a moved object is mislabeled, informative evidence is discarded or uninformative evidence is kept.
  • standard math Standard RANSAC, P3P, RootSIFT, and median aggregation behave as expected.
    Implicit throughout Sec. 3.1; these are established tools in visual localization.
  • domain assumption InLoc reference poses are accurate enough to serve as ground truth.
    Sec. 5: all quantitative claims are evaluated against the dataset's reference poses.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization." pith.science (2026). https://pith.science/paper/DAR2CEMY

@misc{pith2026190804598,
  author       = {Pith},
  title        = {Pith review of: Is This The Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DAR2CEMY}},
  note         = {Machine review of arXiv:1908.04598}
}
read the original abstract

Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common strategy is, first, to generate multiple estimates for the camera pose from which a given query image was taken. The pose with the largest geometric consistency with the query image, e.g., in the form of an inlier count, is then selected in a second stage. While a significant amount of research has concentrated on the first stage, there is considerably less work on the second stage. In this paper, we thus focus on pose verification. We show that combining different modalities, namely appearance, geometry, and semantics, considerably boosts pose verification and consequently pose accuracy. We develop multiple hand-crafted as well as a trainable approach to join into the geometric-semantic verification and show significant improvements over state-of-the-art on a very challenging indoor dataset.

Figures

Figures reproduced from arXiv: 1908.04598 by the authors.

Figure 1
Figure 1. Using further modalities for indoor visual lo￾calization. Given a set of camera pose estimates for a query image (a, g), we seek to identify the most accurate estimate. (b, h) Due to severe occlusion and weak textures, a state￾of-the-art method [72] fails to identify the correct camera pose. To overcome those difficulties, we use several modal￾ities along with visual appearance: (top) surface normals and (bottom) se… view at source ↗
Figure 2
Figure 2. Image-scan-graph for the InLoc dataset [72]. (a) Example RGB-D panoramic scan. (b) Neighboring database image. (c) 3D points of the RGB-D panoramic scan projected onto the view of the database image. (d) Red dots show where RGB-D panoramic scans are cap￾tured. Blue lines indicate links between panoramic scans and database images, established based on visual overlap. hold in practice. Yet, our experiments show that u… view at source ↗
Figure 3
Figure 3. Network architecture for Trainable Pose Veri￾fication. Input images are passed through a feature extrac￾tion network F to obtain dense descriptors f. These are then combined by computing the descriptor similarity map SD. Finally a score regression CNN R produces the score s of the trainable pose verification model. SD. A final average pooling then aggregates the positive and negative evidence over the score map to a… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: The impact of geometric and semantic infor￾mation on the pose verification stage. We validate the performance of the proposed methods that consider addi￾tional geometric and semantic information on the InLoc dataset [72]. Each curve shows the percentage of the queries …
Figure 5
Figure 5. Figure 5: Typical failure cases of view synthesis using the scan-graph. Top: Synthetic images obtained during DensePV with the scan-graph, affected by (a) misalignment of the 3D scans to the floor plan, (b) sparsity of the 3D scans, and (c) intensity changes. Bottom: A typical f…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 79 canonical work pages

  1. [1]

    Arandjelovi ´c, Petr Gronat, Akihiko Torii, Tomas Pa- jdla, and Josef Sivic

    Relja. Arandjelovi ´c, Petr Gronat, Akihiko Torii, Tomas Pa- jdla, and Josef Sivic. NetVLAD: CNN architecture for Query DensePV [72] DensePV w/ scan-graph DensePV+S DensePNV DensePNV+S 2.99m, 17.64◦ 0.42m, 1.25◦ 0.42m, 1.25◦ 0.42m, 1.25◦ 0.42m, 1.25◦ (a) 12.39m, 26.71◦ 12.39m, 26.71◦ 0.40m, 2.37◦ 0.40m, 2.37◦ 0.40m, 2.37◦ (b) 4.70m, 88.15◦ 14.29m, 9.02◦ 5...

  2. [2]

    Three things ev- eryone should know to improve object retrieval

    Relja Arandjelovi ´c and Andrew Zisserman. Three things ev- eryone should know to improve object retrieval. In Proc. CVPR, 2012

  3. [3]

    All about VLAD

    Relja Arandjelovic and Andrew Zisserman. All about VLAD. In Proc. CVPR, 2013

  4. [4]

    Dislocation: Scalable descriptor distinctiveness for location recognition

    Relja Arandjelovi ´c and Andrew Zisserman. Dislocation: Scalable descriptor distinctiveness for location recognition. In Proc. ACCV, 2014

  5. [5]

    Visual vocabu- lary with a semantic twist

    Relja Arandjelovi ´c and Andrew Zisserman. Visual vocabu- lary with a semantic twist. In Proc. ACCV, 2014

  6. [6]

    GIS-Assisted Object Detection and Geospatial Localization

    Shervin Ardeshir, Amir Roshan Zamir, Alejandro Torroella, and Mubarak Shah. GIS-Assisted Object Detection and Geospatial Localization. In Proc. ECCV, 2014

  7. [7]

    Nikolay Atanasov, Menglong Zhu, Kostas Daniilidis, and George J. Pappas. Localization from semantic observations via the matrix permanent. Intl. J. of Robotics Research, 35(1- 3):73–99, 2016

  8. [8]

    Russell, and Josef Sivic

    Mathieu Aubry, Bryan C. Russell, and Josef Sivic. Painting- to-3D Model Alignment via Discriminative Visual Elements. ACM Trans. Graph., 33(2):14:1–14:14, Apr 2014

Show all 87 references
  1. [9]

    RelocNet: Continuous Metric Learning Relocalisation using Neural Nets

    Vassileios Balntas, Shuda Li, and Victor Adrian Prisacariu. RelocNet: Continuous Metric Learning Relocalisation using Neural Nets. In Proc. ECCV, 2018

  2. [10]

    DSAC - Differentiable RANSAC for Camera Local- ization

    Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. DSAC - Differentiable RANSAC for Camera Local- ization. In Proc. CVPR, 2017

  3. [11]

    Learning less is more- 6D camera localization via 3d surface regression

    Eric Brachmann and Carsten Rother. Learning less is more- 6D camera localization via 3d surface regression. In Proc. CVPR, 2018

  4. [12]

    Geometry-Aware Learning of Maps for Cam- era Localization

    Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. Geometry-Aware Learning of Maps for Cam- era Localization. In Proc. CVPR, 2018

  5. [13]

    Graph-based discriminative learning for location recognition

    Song Cao and Noah Snavely. Graph-based discriminative learning for location recognition. In Proc. CVPR, 2013

  6. [14]

    Minimal scene descriptions from structure from motion models

    Song Cao and Noah Snavely. Minimal scene descriptions from structure from motion models. In Proc. CVPR, 2014

  7. [15]

    Robert Castle, Georg Klein, and David W. Murray. Video- rate localization in multiple maps for wearable augmented reality. In ISWC, 2008

  8. [16]

    Lord, Julien Valentin, Luigi Di Stefano, and Philip H

    Tommaso Cavallari, Stuart Golodetz, Nicholas A. Lord, Julien Valentin, Luigi Di Stefano, and Philip H. S. Torr. On- The-Fly Adaptation of Regression Forests for Online Cam- era Relocalisation. In Proc. CVPR, 2017

  9. [17]

    Chen, Georges Baatz, Kevin K ¨oser, Sam S Tsai, Ramakrishna Vedantham, Timo Pylv ¨an¨ainen, Kimmo Roimela, Xin Chen, Jeff Bach, Marc Pollefeys, et al

    David M. Chen, Georges Baatz, Kevin K ¨oser, Sam S Tsai, Ramakrishna Vedantham, Timo Pylv ¨an¨ainen, Kimmo Roimela, Xin Chen, Jeff Bach, Marc Pollefeys, et al. City- scale landmark identification on mobile devices. In Proc. CVPR, 2011

  10. [18]

    Matching with PROSAC- progressive sample consensus

    Ond ˇrej Chum and Ji ˇr´ı Matas. Matching with PROSAC- progressive sample consensus. In Proc. CVPR, 2005

  11. [19]

    Optimal randomized RANSAC

    Ond ˇrej Chum and Ji ˇr´ı Matas. Optimal randomized RANSAC. IEEE PAMI, 30(8):1472–1482, 2008

  12. [20]

    Total recall II: Query expansion revisited

    Ond ˇrej Chum, Andrej Mikulik, Michal Perdoch, and Ji ˇr´ı Matas. Total recall II: Query expansion revisited. In Proc. CVPR, 2011

  13. [21]

    Merg- ing the Unmatchable: Stitching Visually Disconnected SfM Models

    Andrea Cohen, Torsten Sattler, and Mark Pollefeys. Merg- ing the Unmatchable: Stitching Visually Disconnected SfM Models. In Proc. ICCV, 2015

  14. [22]

    Indoor-Outdoor 3D Reconstruction Alignment

    Andrea Cohen, Johannes Lutz Sch ¨onberger, Pablo Speciale, Torsten Sattler, Jan-Michael Frahm, and Marc Pollefeys. Indoor-Outdoor 3D Reconstruction Alignment. In Proc. ECCV, 2016

  15. [23]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. CVPR, 2009

  16. [24]

    Flownet: Learning optical flow with convolutional networks

    Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In Proc. ICCV, pages 2758–2766, 2015

  17. [25]

    Fischler and Robert C

    Martin A. Fischler and Robert C. Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Comm. ACM, 24(6):381–395, 1981

  18. [26]

    Learning and calibrating per-location classifiers for visual place recognition

    Petr Gronat, Guillaume Obozinski, Josef Sivic, and Tomas Pajdla. Learning and calibrating per-location classifiers for visual place recognition. In Proc. CVPR, 2013

  19. [27]

    Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael N¨olle

    Bert M. Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael N¨olle. Review and analysis of solutions of the three point perspective pose estimation problem.IJCV, 13(3):331– 356, 1994

  20. [28]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. CVPR, 2016

  21. [29]

    From Structure-from-Motion point clouds to fast location recognition

    Arnold Irschara, Christopher Zach, Jan-Michael Frahm, and Horst Bischof. From Structure-from-Motion point clouds to fast location recognition. In Proc. CVPR, 2009

  22. [30]

    Ham- ming embedding and weak geometric consistency for large scale image search

    Herve Jegou, Matthijs Douze, and Cordelia Schmid. Ham- ming embedding and weak geometric consistency for large scale image search. In Proc. ECCV, 2008

  23. [31]

    Packing bag-of-features

    Herv ´e J´egou, Matthijs Douze, and Cordelia Schmid. Packing bag-of-features. In Proc. ICCV, 2009

  24. [32]

    Geometric loss func- tions for camera pose regression with deep learning

    Alex Kendall and Roberto Cipolla. Geometric loss func- tions for camera pose regression with deep learning. InProc. CVPR, 2017

  25. [33]

    Posenet: A convolutional network for real-time 6-dof camera relocalization

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. In Proc. ICCV, 2015

  26. [34]

    Learned contextual feature reweighting for image geo- localization

    Hyo Jin Kim, Enrique Dunn, and Jan-Michael Frahm. Learned contextual feature reweighting for image geo- localization. In Proc. CVPR, 2017

  27. [35]

    Kneip, D

    L. Kneip, D. Scaramuzza, and R. Siegwart. A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orienta- tion. In Proc. CVPR, 2011

  28. [36]

    Avoiding confus- ing features in place recognition

    Jan Knopp, Josef Sivic, and Tomas Pajdla. Avoiding confus- ing features in place recognition. In Proc. ECCV, 2010

  29. [37]

    Matching Features Correctly through Semantic Un- derstanding

    Nikolay Kobyshev, Hayko Riemenschneider, and Luc Van Gool. Matching Features Correctly through Semantic Un- derstanding. In Proc. 3DV, 2014

  30. [38]

    Real- Time Solution to the Absolute Pose Problem with Unknown Radial Distortion and Focal Length

    Zuzana Kukelova, Martin Bujnak, and Tomas Pajdla. Real- Time Solution to the Absolute Pose Problem with Unknown Radial Distortion and Focal Length. In Proc. ICCV, 2013

  31. [39]

    Huttenlocher

    Yunpeng Li, Noah Snavely, and Daniel P. Huttenlocher. Lo- cation recognition using prioritized feature matching. In Proc. ECCV, 2010

  32. [40]

    Huttenlocher, and Pas- cal Fua

    Yunpeng Li, Noah Snavely, Daniel P. Huttenlocher, and Pas- cal Fua. Worldwide pose estimation using 3d point clouds. In Proc. ECCV, 2012

  33. [41]

    Sinha, Michael F

    Hyon Lim, Sudipta N. Sinha, Michael F. Cohen, and Matthew Uyttendaele. Real-time image-based 6-DOF local- ization in large-scale environments. In Proc. CVPR, 2012

  34. [42]

    Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D Map

    Liu Liu, Hongdong Li, and Yuchao Dai. Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D Map. In Proc. ICCV, 2017

  35. [43]

    David G. Lowe. Distinctive image features from scale- invariant keypoints. IJCV, 60(2):91–110, 2004

  36. [44]

    Hesch, Marc Pollefeys, and Roland Siegwart

    Simon Lynen, Torsten Sattler, Michael Bosse, Joel A. Hesch, Marc Pollefeys, and Roland Siegwart. Get Out of My Lab: Large-scale, Real-Time Visual-Inertial Localization. In Proc. RSS, 2015

  37. [45]

    Daniela Massiceti, Alexander Krull, Eric Brachmann, Carsten Rother, and Philip H.S. Torr. Random Forests versus Neural Networks - What’s Best for Camera Relocalization? In Proc. Intl. Conf. on Robotics and Automation, 2017

  38. [46]

    Little, Julien Valentin, and Clarence W

    Lili Meng, Jianhui Chen, Frederick Tung, James J. Little, Julien Valentin, and Clarence W. de Silva. Backtracking Regression Forests for Accurate Camera Relocalization. In Proc. IEEE/RSJ Conf. on Intelligent Robots and Systems , 2017

  39. [47]

    Little, Julien Valentin, and Clarence W

    Lili Meng, Frederick Tung, James J. Little, Julien Valentin, and Clarence W. de Silva. Exploiting Points and Lines in Re- gression Forests for RGB-D Camera Relocalization. InProc. IEEE/RSJ Conf. on Intelligent Robots and Systems, 2018

  40. [48]

    Automatic differentiation in pytorch

    Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. In NIPS-W, 2017

  41. [49]

    Object retrieval with large vocabularies and fast spatial matching

    James Philbin, Ond ˇrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman. Object retrieval with large vocabularies and fast spatial matching. In Proc. CVPR, 2007

  42. [50]

    VLocNet++: Deep multitask learning for semantic visual localization and odometry

    Noha Radwan, Abhinav Valada, and Wolfram Burgard. VLocNet++: Deep multitask learning for semantic visual localization and odometry. IEEE Robotics And Automation Letters (RA-L), 3(4):4407–4414, 2018

  43. [51]

    Convo- lutional neural network architecture for geometric matching

    Ignacio Rocco, Relja Arandjelovi ´c, and Josef Sivic. Convo- lutional neural network architecture for geometric matching. In Proc. CVPR, 2017

  44. [52]

    Salas-Moreno, Richard A

    Renato F. Salas-Moreno, Richard A. Newcombe, Hauke Strasdat, Paul H. J. Kelly, and Andrew J. Davison. SLAM++: Simultaneous Localisation and Mapping at the Level of Ob- jects. In Proc. CVPR, 2013

  45. [53]

    From coarse to fine: Robust hierarchical localization at large scale

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Proc. CVPR, 2019

  46. [54]

    Hyperpoints and fine vocab- ularies for large-scale location recognition

    Torsten Sattler, Michal Havlena, Filip Radenovic, Konrad Schindler, and Marc Pollefeys. Hyperpoints and fine vocab- ularies for large-scale location recognition. In Proc. ICCV, 2015

  47. [55]

    Large-scale location recognition and the geomet- ric burstiness problem

    Torsten Sattler, Michal Havlena, Konrad Schindler, and Marc Pollefeys. Large-scale location recognition and the geomet- ric burstiness problem. In Proc. CVPR, 2016

  48. [56]

    Efficient & effective prioritized matching for large-scale image-based localization

    Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Efficient & effective prioritized matching for large-scale image-based localization. IEEE PAMI, 39(9):1744–1756, 2017

  49. [57]

    Benchmarking 6DOF outdoor visual local- ization in changing conditions

    Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, and Tomas Pajdla. Benchmarking 6DOF outdoor visual local- ization in changing conditions. In Proc. CVPR, 2018

  50. [58]

    Are large-scale 3D models really necessary for accurate visual localization? In Proc

    Torsten Sattler, Akihiko Torii, Josef Sivic, Marc Pollefeys, Hajime Taira, Masatoshi Okutomi, and Tomas Pajdla. Are large-scale 3D models really necessary for accurate visual localization? In Proc. CVPR, 2017

  51. [59]

    Understanding the Limitations of CNN-based Absolute Camera Pose Regression

    Torsten Sattler, Qunjie Zhou, Mark Pollefeys, and Laura Leal-Taix´e. Understanding the Limitations of CNN-based Absolute Camera Pose Regression. In Proc. CVPR, 2019

  52. [60]

    City-Scale Location Recognition

    Grant Schindler, Matthew Brown, and Richard Szeliski. City-Scale Location Recognition. In Proc. CVPR, 2007

  53. [61]

    Structure-From-Motion Revisited

    Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-From-Motion Revisited. In Proc. CVPR, 2016

  54. [62]

    Semantic Visual Localization

    Johannes Lutz Sch ¨onberger, Marc Pollefeys, Andreas Geiger, and Torsten Sattler. Semantic Visual Localization. In Proc. CVPR, 2018

  55. [63]

    Lane- Loc: Lane marking based localization using highly accurate maps

    Markus Schreiber, Carsten Kn ¨oppel, and Uwe Franke. Lane- Loc: Lane marking based localization using highly accurate maps. In Proc. IV, 2013

  56. [64]

    Qi Shan, Changchang Wu, Brian Curless, Yasutaka Fu- rukawa, Carlos Hernandez, and Steven M. Seitz. Accurate Geo-Registration by Ground-to-Aerial Image Matching. In Proc. 3DV, 2014

  57. [65]

    Scene co- ordinate regression forests for camera relocalization in RGB- D images

    Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. Scene co- ordinate regression forests for camera relocalization in RGB- D images. In Proc. CVPR, 2013

  58. [66]

    SIFT-realistic rendering

    Dominik Sibbing, Torsten Sattler, Bastian Leibe, and Leif Kobbelt. SIFT-realistic rendering. In Proc. 3DV, 2013

  59. [67]

    Very deep convo- lutional networks for large-scale image recognition

    Karen Simonyan and Andrew Zisserman. Very deep convo- lutional networks for large-scale image recognition. In Proc. ICLR, 2015

  60. [68]

    Semantically Guided Geo- location and Modeling in Urban Environments

    Gautam Singh and Jana Ko ˇseck´a. Semantically Guided Geo- location and Modeling in Urban Environments. In Large- Scale Visual Geo-Localization, 2016

  61. [69]

    Video google: A text retrieval approach to object matching in videos

    Josef Sivic and Andrew Zisserman. Video google: A text retrieval approach to object matching in videos. In Proc. ICCV, 2003

  62. [70]

    Stenborg, C

    E. Stenborg, C. Toft, and L. Hammarstrand. Long-term Vi- sual Localization using Semantically Segmented Images. In Proc. Intl. Conf. on Robotics and Automation, 2018

  63. [71]

    City-Scale Localization for Cameras with Known Vertical Direction

    Linus Sv ¨arm, Olof Enqvist, Fredrik Kahl, and Magnus Os- karsson. City-Scale Localization for Cameras with Known Vertical Direction. IEEE PAMI, 39(7):1455–1461, 2017

  64. [72]

    InLoc: Indoor visual localization with dense matching and view synthesis

    Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Ak- ihiko Torii. InLoc: Indoor visual localization with dense matching and view synthesis. In Proc. CVPR, 2018

  65. [73]

    Long-term 3D Localization and Pose from Semantic Labellings

    Carl Toft, Carl Olsson, and Fredrik Kahl. Long-term 3D Localization and Pose from Semantic Labellings. In Proc. ICCV Workshops, 2017

  66. [74]

    Semantic Match Consistency for Long-Term Visual Localization

    Carl Toft, Erik Stenborg, Lars Hammarstrand, Lucas Brynte, Marc Pollefeys, Torsten Sattler, and Fredrik Kahl. Semantic Match Consistency for Long-Term Visual Localization. In Proc. ECCV, 2018

  67. [75]

    Local visual query expan- sion: Exploiting an image collection to refine local descrip- tors

    Giorgos Tolias and Herv ´e J´egou. Local visual query expan- sion: Exploiting an image collection to refine local descrip- tors. Technical Report RR-8325, INRIA, 2013

  68. [76]

    24/7 place recognition by view synthesis

    Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, and Tomas Pajdla. 24/7 place recognition by view synthesis. In Proc. CVPR, 2015

  69. [77]

    Visual place recognition with repetitive structures

    Akihiko Torii, Josef Sivic, Tomas Pajdla, and Masatoshi Okutomi. Visual place recognition with repetitive structures. In Proc. CVPR, 2013

  70. [78]

    Image- Based Localization Using LSTMs for Structured Feature Correlation

    Florian Walch, Caner Hazirbas, Laura Leal-Taix ´e, Torsten Sattler, Sebastian Hilsenbeck, and Daniel Cremers. Image- Based Localization Using LSTMs for Structured Feature Correlation. In Proc. ICCV, 2017

  71. [79]

    Exploiting 2D floor- plan for building-scale panorama RGBD alignment

    Erik Wijmans and Yasutaka Furukawa. Exploiting 2D floor- plan for building-scale panorama RGBD alignment. InProc. CVPR, 2017

  72. [80]

    Funkhouser

    Fisher Yu, Jianxiong Xiao, and Thomas A. Funkhouser. Se- mantic alignment of LiDAR data at city scale. In Proc. CVPR, 2015

  73. [81]

    X. Yu, S. Chaturvedi, C. Feng, Y . Taguchi, T.-Y . Lee, C. Fer- nandes, and S. Ramalingam. VLASE: Vehicle Localization by Aggregating Semantic Edges. In Proc. IEEE/RSJ Conf. on Intelligent Robots and Systems, 2018

  74. [82]

    Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese

    Amir Roshan Zamir, Alexander Sax, , William B. Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In Proc. CVPR, 2018

  75. [83]

    Accurate image lo- calization based on google maps street view

    Amir Roshan Zamir and Mubarak Shah. Accurate image lo- calization based on google maps street view. InProc. ECCV, 2010

  76. [84]

    Camera Pose V oting for Large-Scale Image-Based Localization

    Bernhard Zeisl, Torsten Sattler, and Marc Pollefeys. Camera Pose V oting for Large-Scale Image-Based Localization. In Proc. ICCV, 2015

  77. [85]

    Pyramid scene parsing network

    Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proc. CVPR, 2017

  78. [86]

    Scene parsing through ADE20K dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. In Proc. CVPR, 2017

  79. [87]

    Semantic un- derstanding of scenes through the ADE20K dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fi- dler, Adela Barriuso, and Antonio Torralba. Semantic un- derstanding of scenes through the ADE20K dataset. IJCV, 127(3):302–321, 2019

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.