Pith. sign in

REVIEW 3 major objections 5 minor 53 references

Geometry-Aware Camera Localization for Bronchoscopy

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read GABL localizes bronchoscope poses with 7.01 mm error at 33.6 FPS, a new state of the art on BREATH.

desk verdict A well-composed system for bronchoscopy localization with plausible benchmark gains, but the single-split, no-error-bar evaluation leaves the SOTA claim not yet proven. read the letter →

arxiv 2608.07116 v1 pith:RE3UDXTN submitted 2026-08-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords cameralocalizationbronchoscopygeometry-awaregraphneuralnetwork6-DoFposeestimationRGB-DmatchingtemporaltrackingpreoperativeCTpriors
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to establish that millimeter-level, real-time bronchoscope localization can be achieved by systematically converting preoperative CT geometry into learning targets rather than relying on generic visual features. The proposed GABL framework builds a graph of anchor points along the airway centerline, treats coarse localization as an anchor-matching classification, then refines the pose with a regressor, a temporal tracker, and an RGB–depth matching loss. On the BREATH benchmark it reports 7.01 mm mean translation error, 29.56 degrees mean rotation error, 61.04% success at 5 mm, and 33.6 FPS, cutting translation error by 8.37% and rotation error by 31.76% relative to BREATH-VL. If correct, this is a concrete demonstration that structural priors, not larger datasets, are what move bronchoscopy localization toward clinical usability.

What carries the argument

The central object is the anchor graph G=(V,E): 512 anchor points sampled along the airway centerline from preoperative CT, connected according to the topology of the bronchial tree, with each node carrying pose and depth embeddings. A three-layer graph convolutional network encodes these nodes, so that coarse localization becomes a classification over anchors instead of a continuous pose search. The fine stage regresses a residual pose, and a causal Transformer with stochastic pose dropout predicts relative pose offsets for temporal tracking. Rendered depth maps from the airway mesh provide the geometric supervision for the RGB–depth matching loss, in which sample pairs are labeled by pose distance rather than hard binary labels. The key work of this machinery is to keep every stage of estimation tied to explicit geometry, so the model does not need to reconstruct the airway from video appearance alone.

What would settle it

Compute per-case mesh-to-mesh distance between the preoperative airway surface and an intraoperative surface, from fluoroscopy, intraoperative CT, or a deformable phantom, and correlate it with per-case ATEtrans of GABL. A strong positive correlation would show the method fails when its geometric prior is violated. A simpler version is to re-render the anchor depths from a deformed mesh and measure how quickly localization error rises.

Watch

Extended reading notes

Core claim

The paper's central claim is that a unified geometry-aware framework, GABL, can estimate the 6-DoF pose of a bronchoscope by fusing preoperative structural priors with intraoperative video at three scales: structure, motion, and appearance. Structure is injected as an anchor graph built from the airway skeleton, motion as a causal Transformer that tracks relative pose offsets, and appearance as cross-modal supervision between RGB frames and rendered depth maps. The paper reports that this design achieves state-of-the-art performance on the BREATH bronchoscopy dataset, with ATEtrans of 7.01 mm and ATErot of 29.56 degrees, and that the full system runs at 33.6 FPS on an RTX 4090. The ablation results support the claim that each of the three geometric scales contributes, with the temporal tracker contributing the largest translational gain and the pose regressor the largest rotational gain.

Load-bearing premise

The load-bearing premise is that the preoperative CT-derived airway mesh and centerline remain geometrically faithful to the airway during the procedure, so that anchor poses and rendered depth maps are valid geometric priors; the paper does not validate this against intraoperative deformation or registration error.

Editorial extensions

If this is right

  • On the BREATH benchmark, GABL improves the previous best translation error from 7.65 mm to 7.01 mm and rotation error from 43.32 degrees to 29.56 degrees, raising SR-5 from 44.40% to 61.04%.
  • The 33.6 FPS inference rate means the system fits real-time clinical bronchoscopy on a consumer GPU, a 4x to 12x speedup over PANSv2, BREATH-VL, and EndoGSLAM.
  • Ablations show the temporal tracker is the largest translational contributor, removing it adds 9.47 mm ATEtrans, while the pose regressor is the largest rotational contributor, removing it raises ATErot to 107.54 degrees.
  • Random clip skipping, temporal reversal, and pose dropout at 0.75 each improve accuracy, indicating that temporal regularization is a meaningful part of the geometric design.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be to apply the same anchor-graph scheme to other tubular anatomies with known centerlines, such as colonoscopy or ureteroscopy, where the same low-texture ambiguity and real-time constraints apply.
  • Because the priors come from static preoperative CT, the most informative stress test would be a dataset with measured airway deformation; if translation error grows with deformation magnitude, the approach would need a deformable-registration extension it currently lacks.
  • The coarse-to-fine anchor classification transforms localization into a ranking problem over a small set, which suggests it could be combined with vision-language or foundation-model features as a fast pose prior in other surgical or robotic navigation settings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes GABL, a geometry-aware bronchoscope localization framework that integrates three sources of geometric information: an anchor graph built from the preoperative CT airway mesh and centerline, a causal Transformer for temporal tracking with stochastic pose dropout, and an appearance-geometry matching loss between RGB frames and rendered depth maps. The method performs coarse-to-fine localization by classifying among anchor embeddings and then regressing a refined SE(3) pose, and it fuses detector and tracker outputs during inference. On the BREATH dataset, using a fixed 56/10 case split, the authors report an ATEtrans of 7.01 mm, an ATErot of 29.56 degrees, SR-5 of 61.04%, and 33.6 FPS, claiming notable reductions over prior state-of-the-art methods including BREATH-VL and PANSv2.

Significance. If the results are robust, this would be a meaningful advance for bronchoscopy localization: GABL is a well-motivated, unified framework that demonstrates the benefit of combining structural, temporal, and cross-modal geometric supervision, and it achieves real-time inference. The ablation tables are internally consistent and show that each of the four proposed losses contributes to the final performance. However, the significance is currently limited by methodological concerns: the headline comparison rests on a single uncontrolled split with no variance estimates, and the description of some hyperparameters is ambiguous. These issues prevent the reader from assessing whether the claimed state-of-the-art margin is statistically reliable.

major comments (3)
  1. [§4.2, Table 1] The central claim of state-of-the-art performance is based on a single 56/10 case split with no random seeds, no confidence intervals, and no statistical tests. The margin over BREATH-VL in ATEtrans is only 0.64 mm and 6.5 percentage points in SR-5, which is small relative to the likely run-to-run variance on ten test procedures. Moreover, the paper does not state whether the baseline numbers in Table 1 were recomputed on the same split and with the same augmentation protocol, or whether they are copied from original publications that may use different partitions. Please provide multi-split or multi-seed results with means and standard deviations, and either rerun baselines under the identical protocol or clearly state the provenance of each baseline number.
  2. [§3.2, Eq. (4); §4.4, Table 4] The definition of the pose dropout parameter is inconsistent. Equation (4) defines p as the probability of retaining the pose embedding and states p=0.25, while Table 4 and Section 4.4 refer to a 'dropout rate' of 0.75 as the best setting and describe rate 1.0 as completely discarding pose information. If 'dropout rate' means mask probability (1-p), this must be stated explicitly and the notation should be unified; otherwise, the reported best hyperparameter is ambiguous and the regularization mechanism cannot be reproduced from the description.
  3. [§3.3, Inference Strategy] The inference-time fusion between the Pose Detector and the Pose Tracker depends on an 'adaptive threshold' whose initial value and linear growth schedule are not specified. Because this fusion rule directly determines the output trajectory, the threshold hyperparameters are part of the method and need to be reported or shown to be insensitive over a range. Without this information, the exact reported poses cannot be reproduced and the contribution of the fusion module to the final accuracy is not quantified.
minor comments (5)
  1. [§3.2 vs. §4.2] Section 3.2 states that the anchor graph is treated as undirected, while Section 4.2 says a directed graph is constructed based on the topology of the bronchial tree; please clarify which structure is actually used.
  2. [§3.2, §3.3] The Transformer input is defined as RGB embeddings plus pose embeddings, but the inference strategy states that the model takes only RGB frames as input; please explain explicitly that pose embeddings are replaced by the null vector at inference time, since no pose is observed.
  3. [Table 1 and §4.1] The SR-5 and SR-10 definitions in Section 4.1 mention only translational error thresholds; please state whether rotational error is ignored in these success-rate metrics.
  4. [Abstract and §4.3] The abstract reports a '4 times inference speedup', while Section 4.3 reports '4× to 12× speedup over prior approaches'; please make the per-baseline speedup explicit.
  5. [§4.2] The anchor count K=512 and the random roll angle sampling in anchor pose generation are not ablated or justified; a brief sensitivity analysis would strengthen the claim that the method is robust to these choices.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: GABL's pose accuracies are supervised measurements on held-out cases, not reductions of the method to its own inputs.

full rationale

GABL's pose estimates are produced by networks trained with explicit ground-truth-based losses (L_coarse, L_fine, L_track, L_match; Eqs. 8, 11, 13, 15) and evaluated with ATE/SR against held-out pose annotations. The anchor graph, anchor poses, and rendered depth maps are constructed offline from the preoperative CT mesh as inputs and supervision signals; they are not defined in terms of the predicted poses, and no equation identifies a prediction with a fitted constant or a label. The self-citations to PANS/PANSv2/BREATH-VL are from the same group and the BREATH benchmark is group-owned, but Table 1 reports concrete empirical baselines and the central SOTA claim is an experimental comparison, not a derivation from a self-cited theorem. The static-geometry assumption about the CT mesh is a generalization and correctness concern, not circularity; since the benchmark ground truth is itself CT-derived, this assumption does not make the in-benchmark numbers definitional. No load-bearing circular step was found, so the appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the fidelity of preop CT geometry to the intraoperative airway, the accuracy of BREATH ground-truth poses, and a set of hand-tuned hyperparameters. The paper provides no external validation, code, or dataset release to independently confirm these assumptions. No new physical or conceptual entities are introduced; the anchor graph and anchors are data structures built from existing CT geometry.

free parameters (4)
  • number of anchor points (K) = 512
    Set by farthest point sampling along the airway skeleton; controls the granularity of coarse localization and is a hand-chosen design value, not derived from theory.
  • pose dropout retention probability p = 0.25 (i.e., 75% dropout)
    Selected by ablation on the same benchmark (Table 4); this is a tuned hyperparameter, not a predicted constant.
  • loss weights (Lcoarse, Lfine, Ltrack, Lmatch) = 1:1:1:1
    Hand-set without ablation, yet the balance of the four objectives directly influences the final pose error.
  • random roll angle in anchor pose generation = uniform over [0, 2 pi)
    Anchor poses have an underdetermined roll that the fine regressor must learn; this makes the coarse anchor pose an incomplete geometric prior and shifts the burden to learned refinement.
assumptions (4)
  • domain assumption The airway centerline extracted from preop CT closely approximates the bronchoscope trajectory.
    Section 3.1 uses the centerline as the skeleton reference for anchor points and look-at poses; if the bronchoscope deviates from the centerline, anchor retrieval is biased.
  • domain assumption The preop airway mesh is geometrically consistent with the intraoperative airway, with no significant deformation.
    Section 3.1 renders depth maps from the preop mesh to supervise appearance-geometry matching; tissue deformation, breathing, and instrument contact would break this consistency, and no registration is performed.
  • domain assumption BREATH ground-truth 6-DoF pose annotations are accurate enough to train and evaluate.
    Section 4.1 relies on the benchmark's annotations as labels for Lcoarse, Lfine, Ltrack, and Lmatch, and as the evaluation reference.
  • standard math The inner product in Eq. 7 is a valid similarity measure in the learned shared embedding space.
    Assumes the encoders produce comparable embeddings; no normalization or metric learning beyond the cross-entropy and binary cross-entropy losses is described.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Geometry-Aware Camera Localization for Bronchoscopy." pith.science (2026). https://pith.science/paper/RE3UDXTN

@misc{pith2026260807116,
  author       = {Pith},
  title        = {Pith review of: Geometry-Aware Camera Localization for Bronchoscopy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RE3UDXTN}},
  note         = {Machine review of arXiv:2608.07116}
}
read the original abstract

Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limited training data. Compared to natural scenes, the confined anatomical structures demand millimeter-level precision, while intraoperative guidance necessitates low-latency inference. However, existing methods often fail to effectively exploit preoperative geometric priors, limiting their robustness and accuracy. To address these limitations, we propose a unified geometry-aware bronchoscope localization framework (GABL) that effectively fuses preoperative structural priors with paired intraoperative video to estimate 6-DoF camera poses. Specifically, to address visual ambiguity in complex airways, we propose a graph-guided coarse-to-fine localization scheme that effectively leverages structural priors for precise pose estimation. Furthermore, to mitigate pose jitter and bridge the visual-structural gap, we integrate a Transformer-based tracking model with a novel RGB-depth matching objective, jointly enforcing spatio-temporal and geometric consistency. Extensive experiments demonstrate that our method yields remarkable reductions of 8.37% and 31.76% in translation and rotation errors over the prior state-of-the-art, alongside 4 times inference speedup (33.6 FPS) for robust real-time bronchoscope localization. Project website: https://paulili08.github.io/GABL/.

Figures

Figures reproduced from arXiv: 2608.07116 by the authors.

Figure 1
Figure 1. Overview and performance of Geometry-Aware Bronchoscopy Localization framework (GABL). (a) GABL combines [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of GABL. The framework takes intraoperative data and anchor-based geometric priors derived from [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Visualization Results of Our Model Splatting (3DGS)-based methods, including MonoGS [26] and En￾doGSLAM [45]; depth-based representation methods, such as Endo￾FASt3r [36], VNB [5], and DARES [37]; and landmark-based meth￾ods, including PANSv2 [40] and BREATH-VL [41]. As shown in [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 38 canonical work pages

  1. [1]

    June Hong Ahn. 2020. An update on the role of bronchoscopy in the diagnosis of pulmonary disease.Yeungnam University Journal of Medicine37, 4 (2020), 253–261

  2. [2]

    Sharib Ali. 2022. Where do we stand in AI for endoscopic image analysis? Deciphering gaps and future directions.npj Digital Medicine5, 1 (2022), 184

  3. [3]

    Max Allan, Jonathan Mcleod, Congcong Wang, Jean Claude Rosenthal, Zhenglei Hu, Niklas Gard, Peter Eisert, Ke Xue Fu, Trevor Zeffiro, Wenyao Xia, et al. 2021. Stereo correspondence and reconstruction of endoscopic data challenge.arXiv preprint arXiv:2101.01133(2021)

  4. [4]

    Pablo Azagra, Carlos Sostres, Ángel Ferrández, Luis Riazuelo, Clara Tomasini, O León Barbed, Javier Morlana, David Recasens, Victor M Batlle, Juan J Gómez- Rodríguez, et al. 2023. Endomapper dataset of complete calibrated endoscopy procedures.Scientific Data10, 1 (2023), 671

  5. [5]

    Artur Banach, Franklin King, Fumitaro Masaki, Hisashi Tsukada, and Nobuhiko Hata. 2021. Visually navigated bronchoscopy using three cycle-consistent gener- ative adversarial network for depth estimation.Medical image analysis73 (2021), 102164

  6. [6]

    Yanqi Bao, Tianyu Ding, Jing Huo, Yaoli Liu, Yuxin Li, Wenbin Li, Yang Gao, and Jiebo Luo. 2025. 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities.IEEE Transactions on Circuits and Systems for Video Technology35, 7 (2025), 6832–6852. doi:10.1109/TCSVT.2025.3538684

  7. [7]

    Taylor L Bobrow, Mayank Golhar, Rohan Vijayan, Venkata S Akshintala, Juan R Garcia, and Nicholas J Durr. 2023. Colonoscopy 3D video dataset with paired depth from 2D-3D registration.Medical image analysis90 (2023), 102956

  8. [8]

    Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. 2017. DSAC - Differentiable RANSAC for Camera Localization. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

Show all 53 references
  1. [9]

    Long Chen, Wen Tang, Nigel W John, Tao Ruan Wan, and Jian Jun Zhang. 2018. SLAM-based dense surface reconstruction in monocular minimally invasive surgery and its application to augmented reality.Computer methods and programs in biomedicine158 (2018), 135–146

  2. [10]

    Joseph Cicenia, Sameer K Avasarala, and Thomas R Gildea. 2020. Navigational bronchoscopy: a guide through history, current use, and developing technology. Journal of thoracic disease12, 6 (2020), 3263

  3. [11]

    Kristoffer Mazanti Cold, Sujun Xie, Anne Orholm Nielsen, Paul Frost Clementsen, and Lars Konge. 2024. Artificial intelligence improves novices’ bronchoscopy performance: a randomized controlled trial in a simulated setting.Chest165, 2 (2024), 405–413

  4. [12]

    Jianning Deng, Peize Li, Kevin Dhaliwal, Chris Xiaoxuan Lu, and Mohsen Kha- dem. 2023. Feature-based Visual Odometry for Bronchoscopy: A Dataset and Benchmark. In2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 6557–6564. doi:10.1109/IROS55552.2...

  5. [13]

    Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, and Yanchao Yang. 2025. Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization. In Proceedings of the IEEE/CVF Conference on Computer ...

  6. [14]

    Salvatore Esposito, Matías Mattamala, Daniel Rebain, Francis Xiatian Zhang, Kevin Dhaliwal, Mohsen Khadem, and Subramanian Ramamoorthy. 2025. ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation.arXiv preprint arXiv:2509.13177(2025)

  7. [15]

    Khang Truong Giang, Soohwan Song, and Sungho Jo. 2024. Learning to Pro- duce Semi-dense Correspondences for Visual Localization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 19468–19478

  8. [16]

    Abner Guzman-Rivera, Pushmeet Kohli, Ben Glocker, Jamie Shotton, Toby Sharp, Andrew Fitzgibbon, and Shahram Izadi. 2014. Multi-Output Learning for Camera Relocalization. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  9. [17]

    Seongbo Ha, Jiung Yeon, and Hyeonwoo Yu. 2024. Rgbd gs-icp slam. InEuropean conference on computer vision. Springer, 180–197

  10. [18]

    Michel Hayoz, Christopher Hahne, Mathias Gallardo, Daniel Candinas, Thomas Kurmann, Maximilian Allan, and Raphael Sznitman. 2023. Learning how to robustly estimate camera pose in endoscopic videos.International journal of computer assisted radiology and surgery18, 7 (2023), 1185–1192

  11. [19]

    Anđela Jurić, Filip Kendeš, Ivan Marković, and Ivan Petrović. 2021. A comparison of graph optimization approaches for pose estimation in SLAM. In2021 44th international convention on information, communication and electronic technology (MIPRO). IEEE, 1113–1118

  12. [20]

    Alex Kendall and Roberto Cipolla. 2017. Geometric Loss Functions for Camera Pose Regression With Deep Learning. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  13. [21]

    Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018. Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  14. [22]

    Alex Kendall, Matthew Grimes, and Roberto Cipolla. 2015. PoseNet: A Convolu- tional Network for Real-Time 6-DOF Camera Relocalization. InProceedings of the IEEE International Conference on Computer Vision (ICCV)

  15. [23]

    Christian Kerl, Jürgen Sturm, and Daniel Cremers. 2013. Dense visual SLAM for RGB-D cameras. In2013 IEEE/RSJ International Conference on Intelligent Robots and Systems. 2100–2106. doi:10.1109/IROS.2013.6696650

  16. [24]

    Hu Lin, Chengjiang Long, Yifeng Fei, Qianchen Xia, Erwei Yin, Baocai Yin, and Xin Yang. 2024. Exploring matching rates: From keypoint selection to camera relocalization. InProceedings of the 32nd ACM International Conference on Multimedia. 506–514

  17. [25]

    Liu Liu, Hongdong Li, and Yuchao Dai. 2017. Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D Map. InProceedings of the IEEE International Conference on Computer Vision (ICCV)

  18. [26]

    Hidenobu Matsuki, Riku Murai, Paul HJ Kelly, and Andrew J Davison. 2024. Gaussian splatting slam. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 18039–18048

  19. [27]

    Carsten Moenning and Neil A. Dodgson. 2003.Fast Marching farthest point sam- pling. Technical Report UCAM-CL-TR-562. University of Cambridge, Computer Laboratory. doi:10.48456/tr-562

  20. [28]

    Morelli, F

    L. Morelli, F. Ioli, R. Beber, F. Menna, F. Remondino, and A. Vitti. 2023. COLMAP- SLAM: A FRAMEWORK FOR VISUAL ODOMETRY.The International Archives of the Photogrammetry, Remote Sensing and Spatial Information SciencesXLVIII- 1/W1-2023 (2023), 317–324. doi:10.5194/isprs-archiv...

  21. [29]

    Pierre Moulon, Pascal Monasse, Romuald Perrot, and Renaud Marlet. 2016. Open- MVG: Open multiple view geometry. InInternational Workshop on Reproducible Research in Pattern Recognition. Springer, 60–74

  22. [30]

    Riku Murai, Eric Dexheimer, and Andrew J. Davison. 2025. MASt3R-SLAM: Real- Time Dense SLAM with 3D Reconstruction Priors. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 16695–16705

  23. [31]

    Kutsev Bengisu Ozyoruk, Guliz Irem Gokceler, Taylor L Bobrow, Gulfize Coskun, Kagan Incetan, Yasin Almalioglu, Faisal Mahmood, Eva Curto, Luis Perdigoto, Marina Oliveira, et al. 2021. EndoSLAM dataset and an unsupervised monocular visual odometry and depth estimation approach ...

  24. [32]

    Laura Privitera, Irene Paraboschi, Divyansh Dixit, Owen J Arthurs, and Stefano Giuliani. 2022. Image-guided surgery and novel intraoperative devices for en- hanced visualisation in general and paediatric surgery: A review.Innovative Surgical Sciences6, 4 (2022), 161–172

  25. [33]

    Antoni Rosinol, John J Leonard, and Luca Carlone. 2023. Nerf-slam: Real-time dense monocular slam with neural radiance fields. In2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 3437–3444

  26. [34]

    Paul-Edouard Sarlin, Ajaykumar Unagar, Mans Larsson, Hugo Germain, Carl Toft, Viktor Larsson, Marc Pollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, and Torsten Sattler. 2021. Back to the Feature: Learning Robust Camera Localization From Pixels To Pose. InProceeding...

  27. [35]

    Torsten Sattler, Akihiko Torii, Josef Sivic, Marc Pollefeys, Hajime Taira, Masatoshi Okutomi, and Tomas Pajdla. 2017. Are Large-Scale 3D Models Really Neces- sary for Accurate Visual Localization?. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)

  28. [36]

    Mona Sheikh Zeinoddin, Mobarak I Hoque, Zafer Tandogdu, Greg L Shaw, Matthew J Clarkson, Evangelos B Mazomenos, and Danail Stoyanov. 2025. Endo- fast3r: endoscopic foundation model adaptation for structure from motion. In International Conference on Medical Image Computing and...

  29. [37]

    Mona Sheikh Zeinoddin, Chiara Lena, Jiongqi Qu, Luca Carlini, Mattia Magro, Seunghoi Kim, Elena De Momi, Sophia Bano, Matthew Grech-Sollars, Evangelos Mazomenos, et al. 2024. Dares: Depth anything in robotic endoscopic surgery with self-supervised vector-lora of the foundation...

  30. [38]

    Trung Duc Than, Gursel Alici, Hao Zhou, and Weihua Li. 2012. A Review of Localization Systems for Robotic Endoscopic Capsules.IEEE Transactions on Biomedical Engineering59, 9 (2012), 2387–2399. doi:10.1109/TBME.2012.2201715

  31. [39]

    Qingyao Tian, Zhen Chen, Huai Liao, Xinyan Huang, Bingyu Yang, Lujie Li, and Hongbin Liu. 2024. PANS: probabilistic airway navigation system for real-time robust bronchoscope localization. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention...

  32. [40]

    Qingyao Tian, Huai Liao, Xinyan Huang, Bingyu Yang, and Hongbin Liu. 2025. Harnessing Foundation Models for Robust and Generalizable 6-DOF Bron- choscopy Localization. InInternational Workshop on Agentic AI for Medicine. Springer, 127–135. MM ’26, November 10–14, 2026, Rio de ...

  33. [41]

    Qingyao Tian, Bingyu Yang, Huai Liao, Xinyan Huang, Junyong Li, Dong Yi, and Hongbin Liu. 2026. BREATH-VL: Vision-Language-Guided 6-DoF Bronchoscopy Localization via Semantic-Geometric Fusion. arXiv:2601.03713 [cs.CV] https: //arxiv.org/abs/2601.03713

  34. [42]

    Brostow, and Aron Monszpart

    Mehmet Ozgur Turkoglu, Eric Brachmann, Konrad Schindler, Gabriel J. Brostow, and Aron Monszpart. 2021. Visual Camera Re-Localization Using Graph Neural Networks and Relative Pose Supervision. In2021 International Conference on 3D Vision (3DV). 145–155. doi:10.1109/3DV53792.2021.00025

  35. [43]

    Guangming Wang, Lei Pan, Songyou Peng, Shaohui Liu, Chenfeng Xu, Yanzi Miao, Wei Zhan, Masayoshi Tomizuka, Marc Pollefeys, and Hesheng Wang. 2024. NeRFs in robotics: A survey.The International Journal of Robotics Research(2024), 02783649251374246

  36. [44]

    Jianyuan Wang, Christian Rupprecht, and David Novotny. 2023. PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 9773–9783

  37. [45]

    Kailing Wang, Chen Yang, Yuehao Wang, Sikuang Li, Yan Wang, Qi Dou, Xi- aokang Yang, and Wei Shen. 2024. Endogslam: Real-time dense reconstruction and tracking in endoscopic surgeries using gaussian splatting. InInternational Conference on Medical Image Computing and Computer-...

  38. [46]

    Yuntao Wang, Jinpu Zhang, Ruonan Wei, Wenbo Gao, and Yuehuan Wang. 2024. Mfrgn: Multi-scale feature representation generalization network for ground-to- aerial geo-localization. InProceedings of the 32nd ACM International Conference on Multimedia. 2574–2583

  39. [47]

    Yihong Wu, Fulin Tang, and Heping Li. 2018. Image-based camera localization: an overview.Visual Computing for Industry, Biomedicine, and Art1, 1 (2018), 8

  40. [48]

    Tao Xie, Kun Dai, Ke Wang, Ruifeng Li, Jiahe Wang, Xinyue Tang, and Lijun Zhao. 2022. A Deep Feature Aggregation Network for Accurate Indoor Camera Localization.IEEE Robotics and Automation Letters7, 2 (2022), 3687–3694. doi:10. 1109/LRA.2022.3146946

  41. [49]

    Chi Yan, Delin Qu, Dan Xu, Bin Zhao, Zhigang Wang, Dong Wang, and Xuelong Li

  42. [50]

    Luwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li, Yasutaka Furukawa, and Ping Tan. 2019. SANet: Scene Agnostic Network for Camera Localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

  43. [51]

    Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024. Depth Anything V2. InAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curr...

  44. [52]

    Huangying Zhan, Chamara Saroj Weerasekera, Jia-Wang Bian, Ravi Garg, and Ian Reid. 2021. DF-VO: What should be learnt for visual odometry?arXiv preprint arXiv:2103.00933(2021)

  45. [2024]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

    GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 19595–19604

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.