Pith. sign in

REVIEW 4 major objections 5 minor 103 references

Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that GSFFs, a 3D Gaussian scene model with a self-supervised triplane feature field, makes both feature-based and segmentation-based pose refinement state-of-the-art among rendering-based and privacy-preserving visual…

desk verdict Genuinely new 3DGS feature-field pipeline with solid localization results; the privacy-preserving claim needs stronger evidence than one qualitative inversion test. read the letter →

arxiv 2507.23569 v2 pith:PAIWVP6S submitted 2025-07-31 cs.CV

classification cs.CV
keywords visuallocalization3DGaussiansplattingfeaturefieldsposerefinementprivacy-preservingsegmentation-basedself-supervisedlearningcontrastive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a scene model built from 3D Gaussian splatting plus a learned triplane feature field can drive camera-pose refinement more accurately than existing rendering-based and privacy-preserving pipelines. The representation, called GSFFs, attaches a scale-aware feature to each 3D Gaussian by querying a triplane grid, aligns rendered features with a jointly trained 2D encoder through contrastive losses, and clusters the Gaussians to turn features into segmentation labels. Using segmentations instead of features makes localization privacy-preserving: after training, colors, the feature field, and prototypes are deleted and only geometry plus one hard cluster label per Gaussian remain. On 7Scenes, Cambridge Landmarks, Indoor6, and 12Scenes, both the feature-based and segmentation-based variants reportedly outperform prior and concurrent methods.

What carries the argument

The central object is the triplane feature field indexed by Gaussian covariance: for each Gaussian $G_i$, its center and covariance are projected onto the $xy$, $xz$, and $yz$ planes, and an RBF kernel derived from the projected 2D Gaussian samples the triplane grid; the three sampled $D$-dimensional features are averaged to give a scale-aware volumetric feature $\mathbf{g}_i$. A contrastive loss $\mathcal{L}_{\mathrm{NCE}}$ aligns rendered and 2D extracted features, a prototypical loss $\mathcal{L}_{\mathrm{PRO}}$ ties both to cluster prototypes obtained by spectral clustering of the Delaunay graph of Gaussian centers, and a cross-entropy loss $\mathcal{L}_{\mathrm{CE}}$ trains a segmentation head. Pose refinement backpropagates the alignment error through the rasterizer with pose updates on the Lie algebra $\mathfrak{se}(3)$.

What would settle it

Run the Appendix D inversion attack across all seven 7Scenes scenes and on Cambridge Landmarks, training on the same split protocol and checking whether recognisable textures or fine detail can be recovered from rendered segmentation maps or from stored geometry plus hard labels; any such recovery would refute the privacy-preserving claim.

Watch

Extended reading notes

Core claim

Gaussian Splatting Feature Fields (GSFFs) is a scene representation in which each 3D Gaussian carries a volumetric feature obtained by projecting the Gaussian onto three orthogonal triplane grids and averaging RBF-kernel-weighted queries. The feature field is trained self-supervised with a 2D encoder so that rendered feature maps align with extracted feature maps via contrastive losses, with additional prototypical and multi-view consistency losses. Spectral clustering on the Delaunay graph of Gaussian centers produces cluster prototypes; soft assignments of features to prototypes are rendered as segmentations, and a segmentation head is trained jointly. Pose is refined by backpropagating feature-metric or cross-entropy segmentation errors through the differentiable rasterizer with pose updates on the Lie algebra $\mathfrak{se}(3)$. The paper claims the feature-based pipeline (GSFFs-PRFeature) is more accurate than NeRF- and 3DGS-based rendering refinement baselines, and the segmentation pipeline (GSFFs-PRPrivacy) is more accurate than SegLoc and other privacy-preserving methods, while enabling privacy by discarding colors, features, and prototypes after training.

Load-bearing premise

The privacy-preserving variant assumes that deleting colors, the triplane feature field, and prototypes, leaving only Gaussian geometry and one hard cluster label per Gaussian, prevents an attacker from recovering texture or fine scene detail.

Editorial extensions

If this is right

  • If the central claim is correct, rendering-based localization no longer needs pre-trained feature extractors: features can be learned per scene in a self-supervised way and grounded directly in 3D geometry.
  • Privacy-preserving localization can be performed without high-dimensional descriptors: hard cluster labels stored per Gaussian and rendered densely beat sparse label reprojection (SegLoc) in accuracy.
  • The explicit finite nature of 3DGS makes the stored scene compact after label quantization, favoring cloud deployment with smaller memory footprints.
  • The coarse-to-fine hierarchy of triplane levels expands the convergence basin of pose refinement, so even coarse DenseVLAD initializations can be refined to high accuracy.
  • Because refinement backpropagates through the rasterizer, the method also benefits from better initial poses, improving HLoc and ACE estimates on Cambridge Landmarks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the privacy argument is heuristic, not formal; the paper tests only one inversion model trained on six scenes and evaluated on Chess, so an attacker with access to the rasterizer and stored geometry plus labels is not ruled out.
  • Editorial inference: the strong dependence on the number of clusters suggests that an adaptive or scene-aware prototype count could improve the trade-off between discriminative power and rendering speed.
  • Editorial inference: the sparse-refinement ablation implies that dense renderability is the active ingredient in segmentation-based refinement, which may explain why label quantization works while sparse reprojection fails.
  • Editorial inference: because GSFFs refinement improves even strong HLoc initializations, the method could serve as a generic dense refinement stage for any coarse pose estimator, not only retrieval-based pipelines.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes Gaussian Splatting Feature Fields (GSFFs), a scene representation that couples an explicit 3D Gaussian Splatting geometry model with an implicit triplane feature field, trained jointly with a 2D feature encoder in a self-supervised contrastive framework. A 3D-structure-informed spectral clustering step derives prototypes that regularize the features and convert them into segmentation labels. Two pose-refinement pipelines are built on this representation: GSFFs-PRFeature, which aligns rendered and query feature maps, and GSFFs-PRPrivacy, which aligns rendered segmentation labels and is claimed to be privacy-preserving because colors, features, and prototypes are deleted after training. The method is evaluated on 7Scenes, Cambridge Landmarks, Indoor6, and 12Scenes against rendering-, structure-, and privacy-based baselines, with ablations on clustering, multi-view consistency, scale-awareness, segmentation optimization, and sparsity.

Significance. If the claims hold, GSFFs would be a valuable dense, renderable representation for visual localization that grounds learned features in explicit 3D geometry and provides a segmentation-based privacy-preserving variant with competitive accuracy. The work has notable strengths: it explicitly backpropagates pose updates through the Gaussian rasterizer on se(3), it introduces a scale-aware feature encoding via covariance-weighted triplane queries, it includes a multi-view consistency mechanism, and it provides a reasonably extensive empirical study across several datasets and ablations, including an initialization study and timing measurements. These are concrete, reproducible contributions beyond simply applying 3DGS to localization. However, the two headline claims—state-of-the-art accuracy and privacy preservation—are not uniformly supported by the evidence as presented, which motivates the major comments below.

major comments (4)
  1. [Section 4 and Appendix D] The privacy-preserving claim is load-bearing: Section 4 states that after removing colors, the triplane feature field, and prototypes, the remaining geometry with one hard label per Gaussian 'effectively increasing the level of privacy.' The sole supporting evidence is Appendix D, which trains one inversion model on six of seven 7Scenes scenes and evaluates it only on the Chess scene, reporting qualitative images in Fig. 5 without any quantitative reconstruction metric (e.g., PSNR, SSIM, LPIPS). A single qualitative test at one operating point is insufficient to support a general privacy guarantee, especially because the paper elsewhere recommends larger label counts (e.g., 84 classes for Indoor6) and states that more classes increase discriminative power. The privacy claim should be either substantially strengthened with quantitative inversion metrics, multiple scenes, and attacks at the deployed label count, or explicitly scoped down to the tested configuration.
  2. [Table 1, Section 5.1] The abstract and introduction claim state-of-the-art performance and that 'our approaches outperform prior and concurrent work.' Table 1 shows this is not uniformly true: on Stairs, GSFFs-PRFeature has a median position error of 25.1 cm and recall of 32%, whereas GSplatLoc achieves 8.83 cm and HLoc achieves 2.9 cm. The text acknowledges this failure but does not reconcile it with the headline claim. Since the central assertion of the paper is state-of-the-art localization accuracy, the claim needs to be made dataset- and scene-specific, or the failure mode needs to be analyzed and addressed, for example by explaining why the learned feature field fails on textureless/flat scenes and whether a different initialization or geometry regularization would close the gap.
  3. [Section 4 and Appendix C.2, Table 4] The privacy-preserving variant's operating point and the privacy attack's operating point are inconsistent. The paper recommends and evaluates 84 classes for the larger Indoor6 scenes, with improved accuracy, yet Appendix D does not state how many classes were used in the inversion experiment. If the attack used the default 34-class configuration, then the privacy claim is not demonstrated at the 84-class configuration that is actually recommended for more complex scenes. Moreover, the information content of a hard label grows with K, so an attack at K=84 is a strictly harder test. The authors should disclose the label count used in the attack and provide inversion results at the label counts used in deployment.
  4. [Section 4 and Appendix D] The inversion attack only takes rendered segmentation maps as input; it does not directly test whether the retained 3D geometry (Gaussian centers, scales, rotations, opacities) plus cluster labels can be inverted, for example by optimizing an image through the differentiable rasterizer. Because the Gaussian geometry was optimized with photometric supervision, it may encode appearance-dependent structure. The paper's privacy definition excludes coarse geometry explicitly, but the claim that 'only coarse image information without any details can be recovered' needs evidence that the geometry itself does not leak fine appearance. A stronger attack that uses the full retained representation would make the privacy claim credible.
minor comments (5)
  1. [Appendix B.2 / Table 7] In Table 7, the baseline is labeled 'GoF [81]' while the text and references cite Gaussian Opacity Fields as [89]; the reference number should be corrected.
  2. [Appendix D] Please specify the exact training details of the inversion model (architecture, loss, number of classes, image resolution) and report quantitative reconstruction metrics in addition to the qualitative examples in Fig. 5.
  3. [Table 9] The column headers 'Coarse Res. Fine Steps' and 'Runtime Query (s)' are ambiguous; it would be clearer to separate coarse resolution, fine resolution, number of refinement steps, and per-query runtime into distinct columns.
  4. [Section 4] The phrase 'GSFFs-PR Privacy' in the subsection heading is missing punctuation in the rendered text; please standardize the use of subscripts or hyphens for method names throughout.
  5. [Appendix A.3] The appendix defines the normalization constant Z for the Gaussian kernel but does not specify whether it is the standard 2D Gaussian normalization; a short clarification would help reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claims rest on held-out pose refinement and external benchmarks, with self-citations used only as baselines and prior art.

full rationale

The paper's derivation chain is a learned system: a triplane feature field and a 2D encoder are trained with contrastive, prototypical, and cross-entropy losses (Eqs. 2, 3, 5), and pose refinement minimizes feature or segmentation misalignment (Eqs. 4, 6) on held-out query images. No fitted quantity is renamed as a prediction: the prototypes are learned spatial clusters, and the reported pose errors come from optimizing Eq. 4/6 against unseen queries and comparing with external baselines (HLoc, DSAC*, ACE, NeFeS, MCLoc, SSL-Nif, NeRFMatch, GS-CPR, GSplatLoc, SegLoc, DGC-GNN, GoMatch). The privacy-preserving claim rests on a definition adopted from the literature, including the authors' earlier SegLoc, and is tested by an inversion attack in Appendix D; the attack is qualitative and limited, and it tests inversion from rendered segmentations rather than directly from stored geometry plus labels, but this is an evidential weakness, not a circular reduction. Citations to the authors' prior SegLoc and SSL-Nif are used as baselines and as prior art for label-based privacy; they are externally published and independently falsifiable, so they do not make the derivation circular. No equation in the paper reduces to its own input by construction, and no uniqueness claim is imported from the authors' own work.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

The central claim rests on hand-chosen hyperparameters, most importantly the cluster count K, and on domain assumptions about pseudo-GT pose quality, local refinement basins, and the privacy definition. No new physical entities are introduced; the triplane field and prototypes are learned representations, not postulates with independent falsifiable handles. The main unstated cost is that K is selected from test-set accuracy, which can inflate reported performance.

free parameters (7)
  • Number of clusters K = 34 default; 84 on Indoor6
    Set by inspecting localization accuracy on the test scenes in Appendix C.2. Higher K helps discriminative power but hurts convergence, so the value is a hand-tuned compromise.
  • Feature dimension d = 16
    Chosen for both coarse and fine levels; determines triplane and encoder capacity and rasterizer channel count.
  • Triplane resolutions R = 256 coarse, 1024 fine
    Selected as a balance between convergence basin and fine-level accuracy in Section 5 and Appendix B.1.
  • Loss weights = 0.5 L_NCE, 0.5 L_PRO, 0.5 L_CE, 0.1 L_TVL, 0.05 L_Depth
    Fixed manually for all datasets, not optimized, but central to training and to the reported results.
  • Contrastive temperature tau = 0.05
    Used in L_NCE and L_PRO; fixed for all experiments and affects feature alignment sharpness.
  • Pose refinement settings = Learning rates 0.5/0.2 and steps 150/300 on Cambridge; 0.3/0.2 and 150/300 on Indoor6; 0.2/0.1 and 75/150 on 7Scenes
    Per-dataset tuning of the pose optimizer; directly affects final accuracy and runtime.
  • Training image resolution and preprocessing = 640x480 for 7Scenes, width 1024 for Cambridge, width 480 for Indoor6; CLAHE for Indoor6; sky masks for Cambridge
    Chosen to control memory and runtime and to handle illumination changes; these choices affect results but are not the core method.
assumptions (6)
  • standard math Alpha blending in Eq. (1) can render arbitrary per-Gaussian quantities, including features and segmentation labels, by replacing colors.
    Used throughout Sections 3 and 4 as the rendering backbone; this is the standard 3DGS rendering model.
  • ad hoc to paper A triplane grid queried with an RBF kernel parametrized by each Gaussian's projected covariance gives a scale-aware feature that supports pose refinement.
    This is the core design of GSFFs; its effectiveness is empirically validated but not derived from first principles.
  • ad hoc to paper Spectral clustering on a Delaunay graph of Gaussian centers produces meaningful spatial prototypes that improve feature discriminativeness.
    Introduced in Section 3.2; the ablation shows k-means is worse, but there is no guarantee of optimality for this clustering choice.
  • domain assumption Pseudo ground-truth poses from SfM or DSLAM are accurate enough to supervise joint training of the Gaussian model, triplane, and encoder.
    Used in all experiments; the paper cites [5], which studies the limits of such pseudo ground truth, so this assumption is load-bearing.
  • domain assumption Privacy can be equated with the inability to recover texture and color and fine-level detail, while coarse geometry and hard segmentation labels are considered non-sensitive.
    Section 4 and Appendix D follow [55,79,98]; no formal differential-privacy style guarantee is provided.
  • domain assumption The retrieved initial pose, for example from DenseVLAD, lies within the convergence basin of the coarse-to-fine pose refinement.
    Section 3.3 and Appendix C.3; refinement only optimizes locally from the initial pose, so this assumption is required for the method to work.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization." pith.science (2026). https://pith.science/paper/PAIWVP6S

@misc{pith2026250723569,
  author       = {Pith},
  title        = {Pith review of: Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PAIWVP6S}},
  note         = {Machine review of arXiv:2507.23569}
}
read the original abstract

Visual localization is the task of estimating a camera pose in a known environment. In this paper, we utilize 3D Gaussian Splatting (3DGS)-based representations for accurate and privacy-preserving visual localization. We propose Gaussian Splatting Feature Fields (GSFFs), a scene representation for visual localization that combines an explicit geometry model (3DGS) with an implicit feature field. We leverage the dense geometric information and differentiable rasterization algorithm from 3DGS to learn robust feature representations grounded in 3D. In particular, we align a 3D scale-aware feature field and a 2D feature encoder in a common embedding space through a contrastive framework. Using a 3D structure-informed clustering procedure, we further regularize the representation learning and seamlessly convert the features to segmentations, which can be used for privacy-preserving visual localization. Pose refinement, which involves aligning either feature maps or segmentations from a query image with those rendered from the GSFFs scene representation, is used to achieve localization. The resulting privacy- and non-privacy-preserving localization pipelines, evaluated on multiple real-world datasets, show state-of-the-art performances.

Figures

Figures reproduced from arXiv: 2507.23569 by the authors.

Figure 1
Figure 1. We extract 2D feature F 2D or segmentation S 2D maps from a query image with unknown pose. We estimate the pose by aligning the maps with feature F 3D or segmentation S 3D maps rendered from our Gaussian Splatting Feature Fields (GSFFs), which are learned in a self-supervised way. Using segmentations instead of features makes our localization pipeline privacy-preserving. Abstract Visual localization is the task of e… view at source ↗
Figure 2
Figure 2. Training pipeline. We extract a feature map F 2D and segmentation map S 2D from a training image with known pose (left). For each 3D Gaussian G𝑖 (right), a scale aware feature g𝑖 is extracted from a triplane representation. Spectral clustering is applied on the Delaunay graph derived from the Gaussian cloud, yielding a set of prototypes 𝑃. A label is associated with each 3D Gaussian by assigning the volumetric featu… view at source ↗
Figure 3
Figure 3. First and third lines: original image, encoder feature map [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Median pose error (cm.) (↓) and Median angle error (°) (↓) on Cambridge Landmarks for models trained with 15/34/59/84 classes. mization ultimately reaches high accuracy (especially for high resolution images) showing the robustness of GSFFs-PR. C.4. Varying resolution …
Figure 5
Figure 5. Figure 5: Left to right: original image, image inversion attack from rendering our features (middle, GSFFs-PR [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: From left to right, coarse encoder/rendered segmentation, fine encoder/rendered segmentation. [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

103 extracted references · 60 canonical work pages

  1. [1]

    Pho- tometric Bundle Adjustment for Vision-Based SLAM

    Hatem Alismail, Brett Browning, and Simon Lucey. Pho- tometric Bundle Adjustment for Vision-Based SLAM. In ACCV, 2017. 3

  2. [2]

    Cali- brated and Partially Calibrated Semi-Generalized Ho- mographies

    Snehal Bhayani, Torsten Sattler, Daniel Barath, Patrik Beliansky, Janne Heikkilä, and Zuzana Kukelova. Cali- brated and Partially Calibrated Semi-Generalized Ho- mographies. InICCV, 2021. 1, 2

  3. [3]

    6DGS: 6D Pose Esti- mationfromaSingleImageanda3DGaussianSplatting Model

    Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, and Alessio Del Bue. 6DGS: 6D Pose Esti- mationfromaSingleImageanda3DGaussianSplatting Model. InECCV, 2024. 2, 3

  4. [4]

    Visual Camera Re- Localization from RGB and RGB-D Images Using DSAC

    Eric Brachmann and Carsten Rother. Visual Camera Re- Localization from RGB and RGB-D Images Using DSAC. TPAMI, 44(9):5847–5865, 2021. 2, 7, 8, 9

  5. [5]

    On the Limits of Pseudo Ground Truth in Visual Camera Re-Localisation

    Eric Brachmann, Martin Humenberger, Carsten Rother, and Torsten Sattler. On the Limits of Pseudo Ground Truth in Visual Camera Re-Localisation. InICCV, 2021. 7

  6. [6]

    Accelerated Coordinate Encoding: Learning to Relocalize in Minutes using RGB and Poses

    Eric Brachmann, Tommaso Cavallari, and Victor Adrian Prisacariu. Accelerated Coordinate Encoding: Learning to Relocalize in Minutes using RGB and Poses. InCVPR,

  7. [7]

    Geometry-Aware Learning of Maps for Camera Localization

    Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. Geometry-Aware Learning of Maps for Camera Localization. InCVPR, 2018. 2

  8. [8]

    How Privacy-Preserving Are Line Clouds? Recovering Scene Details From 3D Lines

    Kunal Chelani, Fredrik Kahl, and Torsten Sattler. How Privacy-Preserving Are Line Clouds? Recovering Scene Details From 3D Lines. InCVPR, 2021. 3

Show all 103 references
  1. [9]

    Obfuscation Based Pri- vacy Preserving Representations are Recoverable Using Neighborhood Information

    Kunal Chelani, Assia Benbihi, Fredrik Kahl, Torsten Sattler, and Zuzana Kukelova. Obfuscation Based Pri- vacy Preserving Representations are Recoverable Using Neighborhood Information. In3DV, 2025. 2, 3

  2. [10]

    TensoRF: Tensorial Radiance Fields

    Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. TensoRF: Tensorial Radiance Fields. InECCV,

  3. [11]

    LeveragingNeuralRadianceFieldsforUncertainty- aware Visual Localization

    Le Chen, Weirong Chen, Rui Wang, and Marc Polle- feys. LeveragingNeuralRadianceFieldsforUncertainty- aware Visual Localization. InICRA, 2024. 16, 17

  4. [12]

    Direct- PoseNet: Absolute Pose Regression with Photometric Consistency

    ShuaiChen,ZiruiWang,andVictorA.Prisacariu. Direct- PoseNet: Absolute Pose Regression with Photometric Consistency. In3DV, 2021. 2

  5. [13]

    Prisacariu

    Shuai Chen, Xinghui Li, Zirui Wang, and Victor A. Prisacariu. DFNet: Enhance Absolute Pose Regression with Direct Feature Matching. InECCV, 2022. 2, 7

  6. [14]

    Prisacariu

    Shuai Chen, Yash Bhalgat, Xinghui Li, Jiawang Bian, Kejie Li, Zirui Wang, and Victor A. Prisacariu. Neural Refinement for Absolute Pose Regression with Feature Synthesis. InCVPR, 2024. 2, 3, 7, 8

  7. [15]

    Gaus- sianPro: 3D Gaussian Splatting with Progressive Propa- gation

    Kai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao, Wei Yin, Yuexin Ma, Wenping Wang, and Xuejin Chen. Gaus- sianPro: 3D Gaussian Splatting with Progressive Propa- gation. InICML, 2024. 3

  8. [16]

    Depth-Regularized Optimization for 3D Gaussian Splat- ting in Few-Shot Images

    Jaeyoung Chung, Jeongtaek Oh, and Kyoung Mu Lee. Depth-Regularized Optimization for 3D Gaussian Splat- ting in Few-Shot Images. InCVPR, 2024. 3

  9. [17]

    Sinkhorn Distances: Lightspeed Com- putation of Optimal Transport

    Marco Cuturi. Sinkhorn Distances: Lightspeed Com- putation of Optimal Transport. InNeurIPS, 2013. 5, 14

  10. [18]

    Sur la sphère vide: A la mémoire de GeorgesVoronoï

    Boris Delaunay. Sur la sphère vide: A la mémoire de GeorgesVoronoï. InProceedingsduCongrésinternational des mathématiciens, 1924. 5

  11. [19]

    Learning To Detect Scene Land- marks for Camera Localization

    Tien Do, Ondrej Miksik, Joseph DeGol, Hyun Soo Park, and Sudipta N Sinha. Learning To Detect Scene Land- marks for Camera Localization. InCVPR, 2022. 7, 9

  12. [20]

    Schönberger, Sudipta N

    Mihai Dusmanu, Johannes L. Schönberger, Sudipta N. Sinha, and Marc Pollefeys. Privacy-Preserving Image Features via Adversarial Affine Subspace Embeddings. In CVPR, 2021. 3

  13. [21]

    LSD- SLAM: Large-scale Direct Monocular SLAM

    Jakob Engel, Thomas Schöps, and Daniel Cremers. LSD- SLAM: Large-scale Direct Monocular SLAM. InECCV,

  14. [22]

    Direct Sparse Odometry.TPAMI, 40(3):611–625, 2017

    JakobEngel,VladlenKoltun,andDanielCremers. Direct Sparse Odometry.TPAMI, 40(3):611–625, 2017. 3

  15. [23]

    Privacy Preserving Partial Localization

    Marcel Geppert, Viktor Larsson, Johannes L Schön- berger, and Marc Pollefeys. Privacy Preserving Partial Localization. InCVPR, 2022. 3 10 Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization

  16. [24]

    Sparse-to-Dense Hypercolumn Matching for Long- term Visual Localization

    Hugo Germain, Guillaume Bourmaud, and Vincent Lep- etit. Sparse-to-Dense Hypercolumn Matching for Long- term Visual Localization. In3DV, 2019. 1

  17. [25]

    Fea- ture Query Networks: Neural Surface Description for Camera Pose Refinement

    Hugo Germain, Daniel DeTone, Geoffrey Pascoe, Tan- ner Schmidt, David Novotny, Richard Newcombe, Chris Sweeney, Richard Szeliski, and Vasileios Balntas. Fea- ture Query Networks: Neural Surface Description for Camera Pose Refinement. InCVPR Workshops, 2022. 3

  18. [26]

    SuGaR: Surface- Aligned Gaussian Splatting for Efficient 3D Mesh Recon- struction and High-Quality Mesh Rendering

    Antoine Guédon and Vincent Lepetit. SuGaR: Surface- Aligned Gaussian Splatting for Efficient 3D Mesh Recon- struction and High-Quality Mesh Rendering. InCVPR,

  19. [27]

    Project AutoVi- sion: Localization and 3D Scene Perception for an Au- tonomous Vehicle with a Multi-Camera System

    Lionel Heng, Benjamin Choi, Zhaopeng Cui, Marcel Geppert, Sixing Hu, Benson Kuan, Peidong Liu, Rang Nguyen, Ye Chuan Yeo, Andreas Geiger, Gim Hee Lee, Marc Pollefeys, and Torsten Sattler. Project AutoVi- sion: Localization and 3D Scene Perception for an Au- tonomous Vehicle wi...

  20. [28]

    2D Gaussian Splatting for Geomet- rically Accurate Radiance Fields

    Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian Splatting for Geomet- rically Accurate Radiance Fields. InSIGGRAPH, 2024. 3

  21. [29]

    Robust Image Retrieval-based Visual Localiza- tion using Kapture.arXiv preprint arXiv:2007.13867,

    Martin Humenberger, Yohann Cabon, Nicolas Guerin, JulienMorat, JérômeRevaud, PhilippeRerole, NoéPion, César Roberto de Souza, Vincent Leroy, and Gabriela Csurka. Robust Image Retrieval-based Visual Localiza- tion using Kapture.arXiv preprint arXiv:2007.13867,

  22. [30]

    Investigating the Role of Image Retrieval for Visual Localization.IJCV, 130(7):1811–1836, 2022

    Martin Humenberger, Yohann Cabon, Noé Pion, Philippe Weinzaepfel, Donghwan Lee, Nicolas Guérin, Torsten Sattler, and Gabriela Csurka. Investigating the Role of Image Retrieval for Visual Localization.IJCV, 130(7):1811–1836, 2022. 1, 2

  23. [31]

    Segment Any 4D Gaussians

    Shengxiang Ji, Guanjun Wu, Jiemin Fang, Jiazhong Cen, Taoran Yi, Wenyu Liu, Qi Tian, and Xinggang Wang. Segment Any 4D Gaussians. arXiv preprint arXiv:2407.04504, 2024. 3

  24. [32]

    SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM

    Nikhil Keetha, Jay Karhade, Krishna Murthy Jataval- labhula, Gengshan Yang, Sebastian Scherer, Deva Ra- manan, and Jonathon Luiten. SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM. InCVPR,

  25. [33]

    PoseNet: a Convolutional Network for Real-Time 6-DOF Camera Relocalization

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. PoseNet: a Convolutional Network for Real-Time 6-DOF Camera Relocalization. InICCV, 2015. 2, 7

  26. [34]

    3D Gaus- sian Splatting for Real-Time Radiance Field Rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 3D Gaus- sian Splatting for Real-Time Radiance Field Rendering. IEEE Transactions on Graphics, 42(4):1–14, 2023. 2

  27. [35]

    Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980,

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980,

  28. [36]

    Paired-Point Lifting for Enhanced Privacy-Preserving Visual Localization

    Chunghwan Lee, Jaihoon Kim, Chanhyuk Yun, and Je Hyeong Hong. Paired-Point Lifting for Enhanced Privacy-Preserving Visual Localization. InCVPR, 2023. 3

  29. [37]

    CLIP-GS: CLIP-Informed Gaussian Splatting for Real-time and View-consistent 3D Semantic Understanding

    Guibiao Liao, Jiankun Li, Zhenyu Bao, Xiaoqing Ye, Jingdong Wang, Qing Li, and Kanglin Liu. CLIP-GS: CLIP-Informed Gaussian Splatting for Real-time and View-consistent 3D Semantic Understanding. arXiv preprint arXiv:2404.14249, 2024. 3

  30. [38]

    Sinha, Michael F

    Hyon Lim, Sudipta N. Sinha, Michael F. Cohen, Matt Uyttendaele, and H. Jin Kim. Real-time Monocular Image-based 6-DoF Localization.International Journal of Robotics Research, 34(4–5):476–492, 2015. 1, 2

  31. [39]

    Vela, and Stan Birchfield

    Yunzhi Lin, Thomas Müller, Jonathan Tremblay, Bowen Wen, Stephen Tyree, Alex Evans, Patricio A. Vela, and Stan Birchfield. Parallel Inversion of Neural Radiance Fields for Robust Pose Estimation. InICRA, 2023. 3

  32. [40]

    Pixel-Perfect Structure-from- Motion with Featuremetric Refinement

    Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Lars- son, and Marc Pollefeys. Pixel-Perfect Structure-from- Motion with Featuremetric Refinement. InICCV, 2021. 3

  33. [41]

    GS-CPR: Efficient Camera Pose Refine- ment via 3D Gaussian Splatting

    Changkun Liu, Shuai Chen, Yash Bhalgat, Siyan Hu, Zirui Wang, Ming Cheng, Victor Adrian Prisacariu, and Tristan Braud. GS-CPR: Efficient Camera Pose Refine- ment via 3D Gaussian Splatting. InICLR, 2025. 2, 3, 7, 8, 18

  34. [42]

    A Convnet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Fe- ichtenhofer, Trevor Darrell, and Saining Xie. A Convnet for the 2020s. InCVPR, 2022. 15

  35. [43]

    Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis.arXiv preprint arXiv:2308.09713, 2023

    Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis.arXiv preprint arXiv:2308.09713, 2023. 3

  36. [44]

    Get out of my Lab: Large-scale, Real-Time Visual-Inertial Localization

    SimonLynen,TorstenSattler,MichaelBosse,JoelHesch, Marc Pollefeys, and Roland Siegwart. Get out of my Lab: Large-scale, Real-Time Visual-Inertial Localization. In RSS, 2015. 2

  37. [45]

    Loc-NeRF: Monte Carlo Lo- calization using Neural Radiance Fields

    Dominic Maggio, Marcus Abate, Jingnan Shi, Courtney Mario, and Luca Carlone. Loc-NeRF: Monte Carlo Lo- calization using Neural Radiance Fields. InICRA, 2023. 3

  38. [46]

    Kelly, and Andrew J

    Hidenobu Matsuki, Riku Murai, Paul H.J. Kelly, and Andrew J. Davison. Gaussian Splatting SLAM. InCVPR,

  39. [47]

    Efficient Privacy-Preserving Visual Localization Using 3D Ray Clouds

    Heejoon Moon, Chunghwan Lee, and Je Hyeong Hong. Efficient Privacy-Preserving Visual Localization Using 3D Ray Clouds. InCVPR, 2024. 3

  40. [48]

    LENS: Localization Enhanced by NeRF Synthesis

    Arthur Moreau, Nathan Piasco, Dzmitry Tsishkou, Bog- dan Stanciulescu, and Arnaud e La Fortelle. LENS: Localization Enhanced by NeRF Synthesis. InCoRL,

  41. [49]

    CROSSFIRE: Camera Relocalization on Self- Supervised Features from an Implicit Representation

    Arthur Moreau, Nathan Piasco, Moussab Bennehar, Dzmitry Tsishkou, Bogdan Stanciulescu, and Arnaud de La Fortelle. CROSSFIRE: Camera Relocalization on Self- Supervised Features from an Implicit Representation. In ICCV, 2023. 2, 3, 8

  42. [50]

    OoD-Pose: Camera Pose Regression From Out-of-Distribution Synthetic Views

    Tony Ng, Adrian Lopez-Rodriguez, Vassileios Balntas, and Krystian Mikolajczyk. OoD-Pose: Camera Pose Regression From Out-of-Distribution Synthetic Views. In 3DV, 2022. 2 11 Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization

  43. [51]

    Repre- sentation Learning with Contrastive Predictive Coding

    Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation Learning with Contrastive Predictive Coding. arXiv preprint arXiv:1807.03748, 2018. 5

  44. [52]

    DINOv2: Learning Robust Visual Features without Supervision

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rab...

  45. [53]

    Privacy Preserving Localization via Coordinate Permutations

    Linfei Pan, Johannes L Schönberger, Viktor Larsson, and Marc Pollefeys. Privacy Preserving Localization via Coordinate Permutations. InICCV, 2023. 3

  46. [54]

    MeshLoc: Mesh-Based Visual Localization

    Vojtech Panek, Zuzana Kukelova, and Torsten Sattler. MeshLoc: Mesh-Based Visual Localization. In ECCV,

  47. [55]

    SegLoc: Learn- ing Segmentation-Based Representations for Privacy- Preserving Visual Localization

    Maxime Pietrantoni, Martin Humenberger, Torsten Sattler, and Gabriela Csurka. SegLoc: Learn- ing Segmentation-Based Representations for Privacy- Preserving Visual Localization. InCVPR, 2023. 2, 3, 6, 7, 8, 9, 17

  48. [56]

    Self-Supervised Learning of Neural Implicit Feature Fields for Camera Pose Re- finement

    Maxime Pietrantoni, Gabriela Csurka, Martin Humen- berger, and Torsten Sattler. Self-Supervised Learning of Neural Implicit Feature Fields for Camera Pose Re- finement. In3DV, 2024. 2, 3, 7, 8

  49. [57]

    LDP-Feat: Image Features with Local Differential Privacy

    Francesco Pittaluga and Bingbing Zhuang. LDP-Feat: Image Features with Local Differential Privacy. InICCV,

  50. [58]

    Koppal, Sing Bing Kang, and Sudipta N

    Francesco Pittaluga, Sanjeev J. Koppal, Sing Bing Kang, and Sudipta N. Sinha. Revealing Scenes by Inverting Structure from Motion Reconstructions. InCVPR, 2019. 2, 3, 17

  51. [59]

    VisionTransformersforDensePrediction

    René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. VisionTransformersforDensePrediction. In ICCV,2021. 15

  52. [60]

    From Coarse to Fine: Robust Hierarchical Localization at Large Scale

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From Coarse to Fine: Robust Hierarchical Localization at Large Scale. InCVPR, 2019. 1, 2, 7, 8, 16

  53. [61]

    Back to the Feature: Learning Robust Camera Localization From Pixels To Pose

    Paul-Edouard Sarlin, Ajaykumar Unagar, Mans Larsson, HugoGermain,CarlToft,ViktorLarsson,MarcPollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, and Torsten Sattler. Back to the Feature: Learning Robust Camera Localization From Pixels To Pose. InCVPR,

  54. [62]

    Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition

    Torsten Sattler, Michal Havlena, Filip Radenović, Kon- rad Schindler, and Marc Pollefeys. Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition. In ICCV, 2015. 2

  55. [63]

    Understanding the Limitations of CNN- based Absolute Camera Pose Regression

    Torsten Sattler, Qunjie Zhou, Marc Pollefeys, and Laura Leal-Taixé. Understanding the Limitations of CNN- based Absolute Camera Pose Regression. InCVPR, 2019. 1

  56. [64]

    Schönberger, Silvano Gal- liani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger

    Thomas Schöps, Johannes L. Schönberger, Silvano Gal- liani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger. A Multi-view Stereo Benchmark with High-Resolution Images and Multi-camera Videos. In CVPR, 2017. 3

  57. [65]

    BAD SLAM: Bundle Adjusted Direct RGB-D SLAM

    Thomas Schöps, Torsten Sattler, and Marc Pollefeys. BAD SLAM: Bundle Adjusted Direct RGB-D SLAM. In CVPR, 2019. 3

  58. [66]

    Privacy Preserving Visual SLAM

    Mikiya Shibuya, Shinya Sumikura, and Ken Sakurada. Privacy Preserving Visual SLAM. InECCV, 2020. 3

  59. [67]

    Scene Coordinate Regression Forests for Camera Relocaliza- tion in RGB-D Images

    JamieShotton,BenGlocker,ChristopherZach,Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. Scene Coordinate Regression Forests for Camera Relocaliza- tion in RGB-D Images. InCVPR, 2013. 7

  60. [68]

    GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splat- ting for Improved Visual Localization.arXiv preprint arXiv:2409.16502, 2024

    Gennady Sidorov, Malik Mohrat, Ksenia Lebedeva, Ruslan Rakhimov, and Sergey Kolyubin. GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splat- ting for Improved Visual Localization.arXiv preprint arXiv:2409.16502, 2024. 2, 3, 7, 8, 18

  61. [69]

    Schönberger, Sudipta N

    Pablo Speciale, Johannes L. Schönberger, Sudipta N. Sinha, and Marc Pollefeys. Privacy Preserving Image Queries for Camera Localization. InICCV, 2019. 2, 3

  62. [70]

    Schönberger, Sing Bing Kang, Sudipta N

    Pablo Speciale, Johannes L. Schönberger, Sing Bing Kang, Sudipta N. Sinha, and Marc Pollefeys. Privacy Preserving Image-Based Localization. InCVPR, 2019. 2, 3

  63. [71]

    InLoc: Indoor Visual Localiza- tion with Dense Matching and View Synthesis.TPAMI, 43(4):1293–1307, 2021

    Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomáš Pa- jdla, and Torii Akihiko. InLoc: Indoor Visual Localiza- tion with Dense Matching and View Synthesis.TPAMI, 43(4):1293–1307, 2021. 2

  64. [72]

    Long-term Visual Localization Revisited.TPAMI, 2020

    Carl Toft, Will Maddern, Akihiko Torii, Lars Ham- marstrand, ErikStenborg, DanielSafari, MasatoshiOku- tomi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, et al. Long-term Visual Localization Revisited.TPAMI, 2020. 2

  65. [73]

    24/7 Place Recognition by View Synthesis.TPAMI, 40(2):257–271, 2018

    Akihiko Torii, Relja Arandjelović, Josef Sivic, Masatoshi Okutomi, and Tomáš Pajdla. 24/7 Place Recognition by View Synthesis.TPAMI, 40(2):257–271, 2018. 9, 16

  66. [74]

    The Unreasonable Effectiveness of Pre- Trained Features for Camera Pose Refinement

    Gabriele Trivigno, Carlo Masone, Barbara Caputo, and Torsten Sattler. The Unreasonable Effectiveness of Pre- Trained Features for Camera Pose Refinement. InCVPR,

  67. [75]

    Learning to Navigate the Energy Landscape

    Julien Valentin, Angela Dai, Matthias Nießner, Push- meet Kohli, Philip Torr, Shahram Izadi, and Cem Keskin. Learning to Navigate the Energy Landscape. In3DV. IEEE, 2016. 7, 16

  68. [76]

    GN-Net: The Gauss-Newton Loss for Multi-Weather Relocalization

    Lukas Von Stumberg, Patrick Wenzel, Qadeer Khan, and Daniel Cremers. GN-Net: The Gauss-Newton Loss for Multi-Weather Relocalization. IEEE Robotics and Automation Letters, 5(2):890–897, 2020. 2, 3

  69. [77]

    LM-Reloc: Levenberg-MarquardtBased Direct Visual Relocalization

    Lukas von Stumberg, Patrick Wenzel, Nan Yang, and DanielCremers. LM-Reloc: Levenberg-MarquardtBased Direct Visual Relocalization. In3DV, 2020. 2, 3

  70. [78]

    AtLoc: Attention Guided Camera Localization

    Bing Wang, Changhao Chen, Chris Xiaoxuan Lu, Pei- jun Zhao, Niki Trigoni, and Andrew Markham. AtLoc: Attention Guided Camera Localization. InAAAI, 2020. 2 12 Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization

  71. [79]

    DGC- GNN: Descriptor-free Geometric-Color Graph Neural Network for 2D-3D Matching

    Shuzhe Wang, Juho Kannala, and Daniel Barath. DGC- GNN: Descriptor-free Geometric-Color Graph Neural Network for 2D-3D Matching. InCVPR, 2024. 3, 6, 8, 9

  72. [80]

    Simoncelli

    ZhouWang, AlanC.Bovik, HamidR.Sheikh, andEeroP. Simoncelli. Image Quality Assessment: From Error Visibility to Structural Similarity.IEEE TIP, 13(4):600– 612, 2004. 16

  73. [81]

    Group Normalization

    Yuxin Wu and Kaiming He. Group Normalization. In ECCV, 2018. 15

  74. [82]

    SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting

    Haolin Xiong, Sairisheek Muttukuru, Rishi Upadhyay, Pradyumna Chari, and Achuta Kadambi. SparseGS: Real-Time 360° Sparse View Synthesis using Gaussian Splatting. arXiv preprint arXiv:2312.00206, 2023. 3

  75. [83]

    Deep Probabilistic Feature-metric Tracking.IEEE Robotics and Automation Letters, 6(1):223 – 230, 2021

    Binbin Xu, Andrew Davison, and Stefan Leuteneg- ger. Deep Probabilistic Feature-metric Tracking.IEEE Robotics and Automation Letters, 6(1):223 – 230, 2021. 3

  76. [84]

    Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruc- tion

    Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruc- tion. InCVPR, 2024. 3

  77. [85]

    No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images

    Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Polle- feys, Ming-Hsuan Yang, and Songyou Peng. No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images. InICLR, 2025. 3

  78. [86]

    Barron, Al- berto Rodriguez, Phillip Isola, and Tsung-Yi Lin

    Lin Yen-Chen, Pete Florence, Jonathan T. Barron, Al- berto Rodriguez, Phillip Isola, and Tsung-Yi Lin. INeRF: Inverting Neural Radiance Fields for Pose Estimation. In IROS, 2021. 3

  79. [87]

    LM- Gaussian: Boost Sparse-view 3D Gaussian Splat- ting with Large Model Priors

    Hanyang Yu, Xiaoxiao Long, and Ping Tan. LM- Gaussian: Boost Sparse-view 3D Gaussian Splat- ting with Large Model Priors. arXiv preprint arXiv:2409.03456, 2024. 3

  80. [88]

    Mip-splatting: Alias-free 3D Gaus- sian Splatting

    Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-splatting: Alias-free 3D Gaus- sian Splatting. InCVPR, 2024. 3

  81. [89]

    Gaus- sian Opacity Fields: Efficient Adaptive Surface Recon- struction in Unbounded Scenes.IEEE Transactions on Graphics, 43(6):1–15, 2024

    Zehao Yu, Torsten Sattler, and Andreas Geiger. Gaus- sian Opacity Fields: Efficient Adaptive Surface Recon- struction in Unbounded Scenes.IEEE Transactions on Graphics, 43(6):1–15, 2024. 3, 4, 14, 15, 16

  82. [90]

    SplatLoc: 3D Gaussian Splatting-based Visual Localiza- tion for Augmented Reality.IEEE Transactions on Vi- sualization and Computer Graphics, 31(5):3591–3601,

    Hongjia Zhai, Xiyu Zhang, Boming Zhao, Hai Li, Yijia He, Zhaopeng Cui, Hujun Bao, and Guofeng Zhang. SplatLoc: 3D Gaussian Splatting-based Visual Localiza- tion for Augmented Reality.IEEE Transactions on Vi- sualization and Computer Graphics, 31(5):3591–3601,

  83. [91]

    NeuraLoc: Visual Localization in Neural Implicit Map with Dual Complementary Features

    Hongjia Zhai, Boming Zhao, Hai Li, Xiaokun Pan, Yijia He, Zhaopeng Cui, Hujun Bao, and Guofeng Zhang. NeuraLoc: Visual Localization in Neural Implicit Map with Dual Complementary Features. arXiv preprint arXiv:2503.06117, 2025. 3, 7, 16, 17

  84. [92]

    Efros, Eli Shecht- man, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InCVPR, 2018. 16

  85. [93]

    Zhang and J

    W. Zhang and J. Kosecka. Image based Localization in Urban Environments. In3DPVT, 2006. 1, 2

  86. [94]

    Pixel-GS:DensityControlwithPixel- aware Gradient for 3D Gaussian Splatting

    Zheng Zhang, Wenbo Hu, Yixing Lao, Tong He, and HengshuangZhao. Pixel-GS:DensityControlwithPixel- aware Gradient for 3D Gaussian Splatting. InECCV,

  87. [95]

    PNeRFLoc: Visual localization with point-based neural radiance fields

    Boming Zhao, Luwei Yang, Mao Mao, Hujun Bao, and Zhaopeng Cui. PNeRFLoc: Visual localization with point-based neural radiance fields. InAAAI, 2024. 16, 17

  88. [96]

    Structure From Motion Using Structure-Less Resection

    Enliang Zheng and Changchang Wu. Structure From Motion Using Structure-Less Resection. InICCV, 2015. 2

  89. [97]

    ToLearnornottoLearn: VisualLocalization from Essential Matrices

    Qunjie Zhou, Torsten Sattler, Marc Pollefeys, and Laura Leal-Taixé. ToLearnornottoLearn: VisualLocalization from Essential Matrices. InICRA, 2020. 2

  90. [98]

    Is Geometry Enough for Matching in Visual Localization? In ECCV, 2022

    Qunjie Zhou, Sergio Agostinho, Aljosa Osep, and Laura Leal-Taixe. Is Geometry Enough for Matching in Visual Localization? In ECCV, 2022. 3, 6, 8, 9

  91. [99]

    The NeRFect Match: Exploring NeRF Fea- tures for Visual Localization

    Qunjie Zhou, Maxim Maximov, Or Litany, and Laura Leal-Taixé. The NeRFect Match: Exploring NeRF Fea- tures for Visual Localization. InECCV, 2024. 2, 3, 7, 8 APPENDIX In this appendix first in Appendix A, we provide addi- tionalexplanationswithregard to thecontrastivelosses, the...

  92. [100]

    A single prototype must be associated per pair of features so that the extracted/rendered features are pushed toward the same "class" in the feature space

  93. [101]

    Predictions must be as balanced as possible to avoid collapse. To solve these constraints, we resort to using optimal transport, where we frame this problem as finding a mapping𝑄∈ I R𝑁×𝐾 between pixels and prototypes that maximizes the feature similarity between the pairs of f...

  94. [102]

    with coarse/fine learning rates of 0.5/0.2 on Cam- bridge Landmarks, 0.3/0.2 on Indoor6, and 0.2/0.1 on 7Scenes respectively. The number of refinement steps for the coarse and fine level is set to 150/300 15 Gaussian Splatting Feature Fields for Privacy-Preserving Visual Local...

  95. [103]

    Additionally, the sky is masked out on Cambridge Landmarks

    for the definition of distortion) are masked out during refinement. Additionally, the sky is masked out on Cambridge Landmarks. B.4. Training time and rendering quality In Table 7 we report novel view rendering quality evalu- ated with PSNR, SSIM [80] and LPIPS [92] image met-...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.