Pith. sign in

REVIEW 4 major objections 6 minor 37 references

ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

T0 review · 4 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read ROOM is a physics-based simulator that automatically transforms patient CT scans into photorealistic, multi-modal bronchoscopy training data, and fine-tuning depth models on that data improves their accuracy on external phantom data.

desk verdict A genuinely useful open-source bronchoscopy data pipeline whose core contribution stands, but whose headline transfer claim rests on thin evidence and whose tables have sloppy numeric inconsistencies. read the letter →

arxiv 2509.13177 v2 pith:TX7DQXKH submitted 2025-09-16 cs.RO

classification cs.RO
keywords bronchoscopysimulationsyntheticdatagenerationphotorealisticrenderingcontinuumrobotsmonoculardepthestimationposeCT-basedanatomicalreconstructionsim-to-realtransfer
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ROOM claims to be the first fully automated pipeline that converts patient CT scans into extensive synthetic bronchoscopy training datasets, preserving both geometric constraints and visual characteristics essential for medical navigation. The paper argues that existing pose-estimation and depth-estimation methods struggle on bronchoscopy images because of specular highlights, scarce texture, and extreme depth ranges, and shows that fine-tuning a bronchoscopy-specialized depth model on ROOM data improves its accuracy on an external phantom dataset. If correct, ROOM provides a scalable way to generate realistic medical training data without the ethical and practical constraints of collecting data from real patients, supporting progress toward autonomous bronchoscopy.

What carries the argument

The pipeline chains four components: (1) medial-axis extraction from the signed distance field of the segmented airway lumen, generating collision-free trajectories with adaptive sampling at bifurcations and high-curvature regions; (2) a constant-curvature continuum robot model with Cosserat rod kinematics, Coulomb friction, actuator noise, and soft contacts; (3) a path-tracing renderer using Principled BSDF tissue materials and a point light with exponential falloff to mimic endoscopic lighting; and (4) sensor-noise synthesis that shapes white noise to the power spectral density of real endoscopic images via bilateral filtering and Fourier-domain matching. The noise modeling is positioned a

What would settle it

Evaluate a depth model fine-tuned on ROOM data on real patient bronchoscopy videos with depth ground truth derived from CT registration; if δ1 accuracy does not improve (or degrades) relative to the un-fine-tuned model, the claim that ROOM's visual fidelity transfers to clinical data is falsified. A direct check: compare the fitted noise PSD against the actual noise statistics of a held-out set of real bronchoscopy images from a different clinical center.

Watch

Extended reading notes

Core claim

ROOM (Realistic Optical Observation in Medicine) takes a patient CT scan, segments the airway lumen, extracts a medial-axis trajectory, and simulates a continuum bronchoscope navigating that trajectory while rendering photorealistic multi-modal sensor streams: RGB with matched noise and specularities, metric depth, surface normals, optical flow, and point clouds at millimeter scales. The paper demonstrates that this synthetic data exposes a significant performance gap in current depth and pose estimators on bronchoscopy imagery, and that fine-tuning on ROOM data raises the δ1 accuracy of a bronchoscopy-specialized depth model from 65.39% to 67.70% on an external phantom dataset, with qualita

Load-bearing premise

The photorealistic rendering stack — Principled BSDF tissue materials, point-light falloff, and the noise PSD fitted from real endoscopic images — reproduces clinical bronchoscopy appearance closely enough that models fine-tuned on ROOM data transfer to real patients; quantitative validation is limited to a phantom dataset with ten image-depth pairs, while real bronchoscopy images are shown only qualitatively.

Editorial extensions

If this is right

  • ROOM can generate large-scale, patient-specific bronchoscopy training datasets without the need for clinical data collection, addressing a key data-scarcity bottleneck.
  • Fine-tuning general-purpose or domain-specialized depth models on ROOM data can improve their performance on external bronchoscopy data, suggesting a practical sim-to-real transfer strategy.
  • The multi-modal outputs (depth, normals, optical flow, point clouds) enable evaluation and training for pose estimation, 3D reconstruction, and vision-based navigation in the bronchoscopy setting.
  • The framework's modular design (swappable physics engines, rendering engines, robot models) could be extended to other endoscopic procedures where CT scans are available, such as colonoscopy and arthroscopy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether the fitted noise PSD, derived from an unspecified set of real endoscopic images, generalizes across different bronchoscope models and clinical sites; a multi-center validation would strengthen the photorealism claim.
  • Because the pipeline is fully automated from CT to rendered data, it could be used to generate densely labeled navigation trajectories for reinforcement learning, enabling closed-loop autonomy research without any manual annotation.
  • The quantitative transfer validation rests on only ten image-depth pairs from a phantom; a natural next step is to evaluate fine-tuned models on real patient bronchoscopy videos with depth ground truth obtained from CT registration, which would directly test clinical transfer.
  • The large performance gap between generic depth models and fine-tuned models on ROOM data suggests that the simulator's noise and lighting modeling, not just airway geometry, carries most of the transfer signal — an assumption that could be tested by ablating the noise-modeling stage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents ROOM, a pipeline for generating photorealistic multi-modal bronchoscopy datasets from patient CT scans. The pipeline combines airway segmentation with a modified 3D U-Net, medial-axis trajectory extraction, PyBullet-based continuum robot simulation, and Blender path tracing with Principled BSDF materials and a data-driven sensor noise model. The authors evaluate the generated data on multi-view pose estimation (Table I), monocular depth estimation (Table II), and fine-tuning depth models on an external phantom dataset (Table III), and they include a qualitative navigation demonstration. The central quantitative claim is that fine-tuning BREA-Depth on ROOM data raises delta-1 accuracy from 65.39% to 67.70% on external phantom data, suggesting that synthetic ROOM data can help bridge the sim-to-real gap for bronchoscopy.

Significance. If validated, ROOM would be a useful open-source contribution to bronchoscopy robotics, addressing data scarcity by automating CT-to-dataset generation and releasing synchronized RGB, depth, normal, flow, point-cloud, and pose data. The paper's strengths include the open-source release, the multi-modal output format, the physics-based continuum robot model, and the decision to evaluate fine-tuning on an external phantom dataset rather than only on ROOM's own renderings. However, the external transfer evidence is currently thin, the sensor-noise model is fitted from unspecified data, and there are multiple quantitative inconsistencies that need correction before the results can be fully assessed.

major comments (4)
  1. [Section IV-C, Table III] The central claim that fine-tuning on ROOM data transfers to real settings is supported only by ten 'selected representative image-depth pairs' from an external phantom dataset. No selection protocol, per-pair or per-sequence breakdown, confidence intervals, or statistical test is given. With n=10, the 2.31-point delta-1 gain (65.39 to 67.70) is within plausible sampling variability, and 'selected' invites selection bias. Moreover, the only real bronchoscopy evidence (Fig. 7) is qualitative and has no ground truth. Please report results on all frames or on a random/defined subset, include per-pair results with paired bootstrap confidence intervals or a permutation test, specify the fine-tuning data size and hyperparameters, and clearly separate quantitative from qualitative claims.
  2. [Section III-C, Eq. (6)-(7)] The sensor noise model is fitted from unspecified 'real endoscopic data'; no source, sample size, capture setup, or validation of the fitted PSD is provided. Since the photorealistic appearance claim and the depth-estimation results in Section IV depend on matching clinical noise statistics, this is a load-bearing unverified input. Please identify the data used, describe the fitting procedure and the normalization of P(omega), and add a quantitative validation or ablation, for example comparing noise statistics on held-out real frames or evaluating depth models with and without the shaped noise.
  3. [Section IV-A / IV-B vs Tables I and II] Several reported numbers do not match the tables. Section IV-A gives DUSt3R RTA as 0.37% and VGGT as 0.5%, but Table I lists 0.21 and 0.25. Section IV-B states UniDepth achieves the best L1 error (0.0103 m) and RMSE (0.0160 m), but Table II shows UniDepth at 0.106 and 0.166, with BREA-Depth lowest at 0.091 and 0.141. The caption in Table II also claims absolute relative errors of 0.44-0.49 and delta-1 of 26-28%, while the table contains 0.421-0.486 and 27.1-30.8. These inconsistencies must be reconciled before the quantitative evaluation can be assessed.
  4. [Section III-C and Section I] The paper's claim of a 'fully automated pipeline' that preserves 'geometric constraints' is not directly validated. The accuracy of the CT segmentation, the robustness of the medial-axis extraction, and the geometric fidelity of the reconstructed airway tree are not quantified, for example with Dice overlap against reference segmentations or centerline error. Because the downstream pose and depth tasks inherit any reconstruction errors, a quantitative check of the anatomical reconstruction stage would substantially strengthen the geometric-validity claim.
minor comments (6)
  1. [Section III-B, Eq. (1)] The notation vec(R) and the hat operator on u are not defined. Please clarify the derivative convention and the relation between u and its skew-symmetric matrix.
  2. [Section III-C, Medial Axis Extraction] The condition d/dt [grad phi(x(t))] dot n = 0 uses n ambiguously; earlier n(t) is the inward normal. Please specify the sign convention and define the gradient-transition criterion more precisely.
  3. [Table II and Fig. 6 captions] Table II includes seven methods, and Fig. 6's caption says 'Five state-of-the-art models' while enumerating six or seven. Correct the counts.
  4. [Table I] RTA values appear in the table as 0.07, 0.17, 0.25 without an explicit percent sign, while the text writes 0.07%. Please standardize the units and notation.
  5. [Section IV-C and Fig. 7] The statement that fine-tuned models improve 'even when tested on real bronchoscopy images' is only qualitative because the real image in Fig. 7 has no depth ground truth. Please adjust the wording to avoid implying a quantitative result.
  6. [References [26]-[28]] Several references use 'et al.' in the author lists rather than complete author names. Please use a consistent citation format.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: ROOM's fine-tuning transfer is evaluated on an external phantom dataset, and the fitted noise PSD is an input to rendering, not a derived prediction.

full rationale

The paper's load-bearing quantitative claim is that fine-tuning depth models on ROOM-generated synthetic data improves performance on an external phantom-based bronchoscope dataset with ground truth (Table III), explicitly avoiding evaluation on in-distribution ROOM renderings. The fine-tuned models are compared against the same models before fine-tuning, and the external dataset [36] is not produced by ROOM and is not used to fit any ROOM parameter. The sensor noise PSD fitted in Eqs. (6)-(7) is a rendering input derived from unspecified real endoscopic data; it is not fitted to the downstream depth-estimation outcomes, so the downstream evaluation does not reduce to this fit. The self-citations (BREA-Depth [35], [14], [20]) provide baselines or standard modeling assumptions but do not carry the circularity burden: the fine-tuning improvement is measured rather than imported from those papers. The main substantive weaknesses, such as the small number of selected external test pairs, lack of confidence intervals, and qualitative-only real-image evaluation, are evidence-quality limitations rather than circular derivation steps. No equation in the paper reduces to a target result by construction, and no central claim is justified solely by a self-citation chain. Therefore the paper is self-contained against its central transfer claim, and no circular step is present.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The pipeline rests on standard geometric, rendering, and robotics tools, but the key realism and transfer claims rely on unvalidated domain assumptions about tissue appearance, lighting, and sensor noise. No new physical entities are introduced.

free parameters (5)
  • static and dynamic friction coefficients mu_s, mu_d = 0.3, 0.25
    Chosen to reproduce stick-slip at airway bifurcations in Section III.B; no measurement source or uncertainty is provided.
  • actuator noise scale and delay = 0.05; U(0,0.1)s
    Parameters in Eq. (4)-(5) are hand-set to mimic tendon stretching and backlash; no clinical calibration data is given.
  • sensor noise amplitude beta = not stated
    Controls noise strength in I_synth = I_rendered + beta * n_synth; no fitting procedure is specified, and it directly affects photorealism.
  • BSDF material parameters = not stated
    Base color, metallic, and roughness maps are procedural choices for tissue appearance; no calibration against measured tissue BRDF is reported.
  • robot geometric constants l and gamma = 50 mm; 1.75e-3 m
    Used in Eq. (3) for the continuum robot model; these appear to be robot design inputs, but the paper gives no citation or uncertainty.
assumptions (6)
  • domain assumption Constant curvature / Cosserat rod model describes the bronchoscope's shape and tendon transmission
    Eq. (1)-(3); standard for continuum robots but simplified relative to the compliance and hysteresis of real flexible bronchoscopes.
  • domain assumption Coulomb friction with calibrated coefficients models bronchoscope-tissue contact
    Section III.B; contact and deformable dynamics are simplified, a limitation the authors acknowledge in the Discussion.
  • domain assumption SDF grassfire / medial axis extraction gives valid collision-free airway centerlines
    Section III.C; assumes the segmented airway geometry is accurate enough for clinically representative navigation trajectories.
  • ad hoc to paper Principled BSDF materials and point light with exponential falloff approximate wet mucosal tissue and endoscopic lighting
    Section III.C and Fig. 3; this visual realism assumption underpins sim-to-real transfer but is not quantitatively validated.
  • domain assumption Bilateral filtering separates real endoscopic signal from noise, and PSD shaping reproduces sensor noise
    Eq. (6)-(7); assumes noise is additive, stationary, and transferable from unspecified real endoscopic images to all bronchoscopy sensors.
  • domain assumption Modified 3D U-Net segmentation produces anatomically accurate airway lumen masks
    Section III.C; no dataset name, training details, or segmentation accuracy numbers are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation." pith.science (2026). https://pith.science/paper/TX7DQXKH

@misc{pith2026250913177,
  author       = {Pith},
  title        = {Pith review of: ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TX7DQXKH}},
  note         = {Machine review of arXiv:2509.13177}
}
read the original abstract

Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their development is limited by the lack of realistic training and test environments: Real data is difficult to collect due to ethical constraints and patient safety concerns, and developing autonomy algorithms requires realistic imaging and physical feedback. We present ROOM (Realistic Optical Observation in Medicine), a comprehensive simulation framework designed for generating photorealistic bronchoscopy training data. By leveraging patient CT scans, our pipeline renders multi-modal sensor data including RGB images with realistic noise and light specularities, metric depth maps, surface normals, optical flow and point clouds at medically relevant scales. We validate the data generated by ROOM in two canonical tasks for medical robotics: multi-view pose estimation and monocular depth estimation, demonstrating diverse challenges that state-of-the-art methods must overcome to transfer to these medical settings. Furthermore, we show that the data produced by ROOM can be used to fine-tune existing depth estimation models to overcome these challenges, also enabling other downstream applications such as navigation. We expect that ROOM will enable large-scale data generation across diverse patient anatomies and procedural scenarios that are challenging to capture in clinical settings. Code and data: https://github.com/iamsalvatore/room.

Figures

Figures reproduced from arXiv: 2509.13177 by the authors.

Figure 1
Figure 1. ROOM framework overview. Given patient CT scans (left), our pipeline reconstructs accurate 3D lung models and extracts medial axis trajectories, enabling physics-based continuum robot simulation to generate photorealistic multi-modal sensor data (right). This includes RGB images with realistic noise and lighting, metric depth maps, surface normals, optical flow, point clouds, and ground-truth poses, for different me… view at source ↗
Figure 2
Figure 2. ROOM data generation pipeline. The system consists of four main stages: (1) Medial Axis Extraction from segmented CT lung models, (2) Automated Sampling along skeletal branches with higher density at bifurcations and high-curvature regions, (3) Data Synthesis generating synchronized multi-modal sensor streams from t0 to tn timesteps, and (4) Sensor Noise Modeling applying realistic noise characteristics matching rea… view at source ↗
Figure 4
Figure 4. Continuum robot model used in ROOM simulation. The bronchoscope is modelled as a flexible, cable-driven continuum robot with constant curvature bending and three degrees of freedom: tendon actuation for bending curvature (q1), axial rotation for bending plane (q2), and linear insertion depth (q3). The physics-based simulation incorporates realistic friction models, actuator noise, and collision dynamics calibrated t… view at source ↗
Figures from the paper (4 more)
Figure 5
Figure 5. Figure 5: ROOM pipeline output folder structure. The framework generates synchronized multi-modal sensor data organized by patient anatomy and sequence. Each sequence contains RGB images (600×600), metric depth maps, surface normals, optical flow fields, point clouds, ground-tru…
Figure 6
Figure 6. Figure 6: Comparative monocular depth estimation results on ROOM synthetic bronchoscopy sequences. Top rows show L1 error maps between predicted depth estimation and ground truth depth, where warmer colours indicate higher absolute errors, while bottom rows display corresponding…
Figure 7
Figure 7. Figure 7: Monocular depth estimation examples of pre-trained models and fine-tuned on ROOM. We show examples on a phantom-based dataset with ground truth [36] as well as real images. Please note that the real image does not have depth ground truth available. monocular depth esti…
Figure 8
Figure 8. Figure 8: Vision-based navigation examples. We demonstrate qualitative results of the relative monocular depth predictions (scaled with ground￾truth scale), as input for a sampling-based local planner. Left: projection of the collision-free path. Right: 3D visualisation of the p…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

37 extracted references · 4 linked inside Pith

  1. [1]

    Continuum Robots for Medical Interventions,

    P. E. Dupont, N. Simaan, H. Choset, and C. D. Rucker, “Continuum Robots for Medical Interventions,”Proceedings of the IEEE, 2022

  2. [2]

    Kubric: A Scalable Dataset Generator,

    K. Greff, F. Belletti, L. Beyer, C. Doersch, Y . Du, D. Duckworth, D. J. Fleet, D. Gnanapragasam, F. Golemo, C. Herrmann, T. Kipf, A. Kundu, D. Lagun, I. Laradji, H.-T. D. Liu, H. Meyer, Y . Miao, D. Nowrouzezahrai, C. Oztireli, E. Pot, N. Radwan, D. Rebain, S. Sabour, M. S. M. Sajjadi, M. Sela, V . Sitzmann, A. Stone, D. Sun, S. V ora, Z. Wang, T. Wu, K....

  3. [3]

    TartanAir: A Dataset to Push the Limits of Visual SLAM,

    W. Wang, D. Zhu, X. Wang, Y . Hu, Y . Qiu, C. Wang, Y . Hu, A. Kapoor, and S. Scherer, “TartanAir: A Dataset to Push the Limits of Visual SLAM,” inIEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2020. Fig. 8:Vision-based navigation examples.We demonstrate qualitative results of the relative monocular depth predictions (scaled with ground- t...

  4. [4]

    SimCol3D — 3D reconstruction during colonoscopy challenge,

    A. Rau, S. Bano, Y . Jin, P. Azagra, J. Morlana, R. Kader, E. Sander- son, B. J. Matuszewski, J. Y . Lee, D.-J. Lee, E. Posner, N. Frank, V . Elangovan, S. Raviteja, Z. Li, J. Liu, S. Lalithkumar, M. Islam, H. Ren, L. B. Lovat, J. M. Montiel, and D. Stoyanov, “SimCol3D — 3D reconstruction during colonoscopy challenge,”Medical Image Analysis

  5. [5]

    ORBIT-Surgical: An Open-Simulation Framework for Learning Sur- gical Augmented Dexterity,

    Q. Yu, M. Moghani, K. Dharmarajan, V . Schorp, W. C.-H. Panitch, J. Liu, K. Hari, H. Huang, M. Mittal, K. Goldberg, and A. Garg, “ORBIT-Surgical: An Open-Simulation Framework for Learning Sur- gical Augmented Dexterity,” inIEEE Intl. Conf. on Robotics and Automation (ICRA), 2024

  6. [6]

    Software toolkit for modeling, simulation and control of soft robots,

    E. Coevoet, T. Morales-Bieze, F. Largilli `ere, Z. Zhang, M. Thieffry, M. Sanz-Lopez, B. Carrez, D. Marchal, O. Goury, J. Dequidt, and C. Duriez, “Software toolkit for modeling, simulation and control of soft robots,”Advanced Robotics, 2017

  7. [7]

    TMTDyn: A Matlab package for modeling and control of hybrid rigid–continuum robots based on discretized lumped systems and reduced-order models,

    S. M. H. Sadati, S. E. Naghibi, A. Shiva, B. Michael, L. Renson, M. Howard, C. D. Rucker, K. Althoefer, T. Nanayakkara, S. Zschaler, C. Bergeles, H. Hauser, and I. D. Walker, “TMTDyn: A Matlab package for modeling and control of hybrid rigid–continuum robots based on discretized lumped systems and reduced-order models,”Intl. J. of Robot. Res., 2021

  8. [8]

    Surgical Simulation: A Systematic Review,

    L. M. Sutherland, P. W. Middleton, A. Russell, M. Wijenayake, N. Maddern, and G. J. Maddern, “Surgical Simulation: A Systematic Review,”Annals of Surgery, 2006

Show all 37 references
  1. [9]

    Simulation in Endoscopy: Practical Educational Strategies to Improve Learning,

    C. H. Park, M. J. Ryou, and C. C. Thompson, “Simulation in Endoscopy: Practical Educational Strategies to Improve Learning,” World Journal of Gastroenterology, 2019

  2. [10]

    EndoGaussian: Real-time Gaussian Splatting for Dynamic Endoscopic Scene Reconstruction,

    Y . Liu, C. Li, C. Yang, and Y . Yuan, “EndoGaussian: Real-time Gaussian Splatting for Dynamic Endoscopic Scene Reconstruction,” arXiv preprint arXiv:2401.12561, 2024

  3. [11]

    Endora: Video Generation Models as Endoscopy Sim- ulators,

    C. Li, H. Liu, Y . Liu, B. Y . Feng, W. Li, X. Liu, Z. Chen, J. Shao, and Y . Yuan, “Endora: Video Generation Models as Endoscopy Sim- ulators,” inMed. Image Comput. Comput. Assist. Interv. (MICCAI), 2024

  4. [12]

    Design optimization of a contact-aided continuum robot for endobronchial interventions based on anatomical constraints,

    L. Ros-Freixedes, A. Gao, N. Liu, M. Shen, and G.-Z. Yang, “Design optimization of a contact-aided continuum robot for endobronchial interventions based on anatomical constraints,”International Journal of Computer Assisted Radiology and Surgery, 2019

  5. [13]

    PANS: Probabilistic Airway Navigation System for Real-time Robust Bronchoscope Localization,

    Q. Tian, Z. Chen, H. Liao, X. Huang, B. Yang, L. Li, and H. Liu, “PANS: Probabilistic Airway Navigation System for Real-time Robust Bronchoscope Localization,” inMed. Image Comput. Comput. Assist. Interv. (MICCAI), 2024

  6. [14]

    Feature-based Visual Odometry for Bronchoscopy: A Dataset and Benchmark,

    J. Deng, P. Li, K. Dhaliwal, C. X. Lu, and M. Khadem, “Feature-based Visual Odometry for Bronchoscopy: A Dataset and Benchmark,” in IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2023

  7. [15]

    BronchoPose: an analysis of data and model configuration for vision-based bronchoscopy pose estimation,

    J. Borrego-Carazo, C. S ´anchez, D. Castells-Rufas, J. Carrabina, and D. Gil, “BronchoPose: an analysis of data and model configuration for vision-based bronchoscopy pose estimation,”Computer Methods and Programs in Biomedicine, 2023

  8. [16]

    BM-BronchoLC: A rich bronchoscopy dataset for anatomical landmarks and lung cancer lesion recognition,

    V . e. a. Vu, “BM-BronchoLC: A rich bronchoscopy dataset for anatomical landmarks and lung cancer lesion recognition,”Scientific Data, 2024

  9. [17]

    UAAL Dataset: Upper Airway Anatomical Landmark Dataset for Automated Bronchoscopy and Intubation,

    R. e. a. Hao, “UAAL Dataset: Upper Airway Anatomical Landmark Dataset for Automated Bronchoscopy and Intubation,”Figshare, 2024

  10. [18]

    AI Co-Pilot Bronchoscope Robot,

    J. Zhang, L. Liu, P. Xiang, Q. Fang, X. Nie, H. Ma, J. Hu, R. Xiong, Y . Wang, and H. Lu, “AI Co-Pilot Bronchoscope Robot,”Nature Communications, 2024

  11. [19]

    Bron- choCopilot: Towards Autonomous Robotic Bronchoscopy via Mul- timodal Reinforcement Learning,

    J. Zhao, H. Chen, Q. Tian, J. Chen, B. Yang, and H. Liu, “Bron- choCopilot: Towards Autonomous Robotic Bronchoscopy via Mul- timodal Reinforcement Learning,”arXiv preprint arXiv:2403.01483, 2024

  12. [20]

    On the Benefits of Hysteresis in Tendon Driven Continuum Robots,

    D. Hanley, F. Alambeigi, and M. Khadem, “On the Benefits of Hysteresis in Tendon Driven Continuum Robots,” inIEEE Intl. Conf. on Robotics and Automation (ICRA), 2025

  13. [21]

    3D skeletons: A state-of-the-art report,

    A. Tagliasacchi, T. Delame, M. Spagnuolo, N. Amenta, and A. Telea, “3D skeletons: A state-of-the-art report,”Computer Graphics Forum, 2016

  14. [22]

    Structure-from-Motion Revis- ited,

    J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-Motion Revis- ited,” inIEEE Int. Conf. Computer Vision and Pattern Recognition, 2016, pp. 4104–4113

  15. [23]

    ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual- Inertial and Multi-Map SLAM

    “ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual- Inertial and Multi-Map SLAM.”

  16. [24]

    DUSt3R: Geometric 3D Vision Made Easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “DUSt3R: Geometric 3D Vision Made Easy,” inIEEE Int. Conf. Computer Vision and Pattern Recognition, 2024

  17. [25]

    VGGT: Visual Geometry Grounded Transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “VGGT: Visual Geometry Grounded Transformer,” in IEEE Int. Conf. Computer Vision and Pattern Recognition, 2025

  18. [26]

    Endomapper dataset of complete calibrated en- doscopy procedures,

    P. Azagraet al., “Endomapper dataset of complete calibrated en- doscopy procedures,”Scientific Data, 2023

  19. [27]

    CudaSIFT-SLAM: multiple-map visual SLAM for full procedure mapping in real human endoscopy,

    R. Elvira, J. D. Tard ´os, and J. M. Montiel, “CudaSIFT-SLAM: multiple-map visual SLAM for full procedure mapping in real human endoscopy,”arXiv preprint arXiv:2405.16932, 2024

  20. [28]

    Pose estimation via structure-depth information from monocular endoscopy images sequence,

    Z. Liet al., “Pose estimation via structure-depth information from monocular endoscopy images sequence,”Optica Publishing Group, 2024

  21. [29]

    Bronchoscopy,

    J. Klapper, S. Raja, N. Ninan, and S. Shofer, “Bronchoscopy,” TSRA Primer in Cardiothoracic Surgery, The American Association for Thoracic Surgery, 2024

  22. [30]

    Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation,

    M. Hu, W. Yin, C. Zhang, Z. Cai, X. Long, H. Chen, K. Wang, G. Yu, C. Shen, and S. Shen, “Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation,”IEEE Trans. Pattern Anal. Mach. Intell., 2024

  23. [31]

    Depth Anything V2,

    L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth Anything V2,” inAdvances in Neural Information Processing Systems, 2024

  24. [32]

    UniDepth: Universal Monocular Metric Depth Estimation,

    L. Piccinelli, Y .-H. Yang, C. Sakaridis, M. Segu, S. Li, L. Van Gool, and F. Yu, “UniDepth: Universal Monocular Metric Depth Estimation,” inIEEE Int. Conf. Computer Vision and Pattern Recognition, 2024

  25. [33]

    Endodac: Efficient adapting foundation model for self-supervised depth estimation from any endoscopic camera,

    B. Cui, M. Islam, L. Bai, A. Wang, and H. Ren, “Endodac: Efficient adapting foundation model for self-supervised depth estimation from any endoscopic camera,” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2024, pp. 208–218

  26. [34]

    EndoOmni: Zero-shot cross-dataset depth estimation in endoscopy by robust self-learning from noisy labels,

    Q. Tian, Z. Chen, H. Liao, X. Huang, L. Li, S. Ourselin, and H. Liu, “EndoOmni: Zero-shot cross-dataset depth estimation in endoscopy by robust self-learning from noisy labels,”arXiv preprint arXiv:2409.05442, 2024

  27. [35]

    BREA-Depth: Bronchoscopy Realistic Airway- geometric Depth Estimation,

    F. X. Zhang, E. Mackute, M. Kasaei, K. Dhaliwal, R. Thomson, and M. Khadem, “BREA-Depth: Bronchoscopy Realistic Airway- geometric Depth Estimation,” inMedical Image Computing and Computer-Assisted Intervention – MICCAI 2025, 2025

  28. [36]

    Deep monocular 3D reconstruction for assisted navigation in bronchoscopy,

    M. Visentini-Scarzanella, T. Sugiura, T. Kaneko, and S. Koto, “Deep monocular 3D reconstruction for assisted navigation in bronchoscopy,” International journal of computer assisted radiology and surgery, vol. 12, pp. 1089–1099, 2017

  29. [37]

    VP-STO: Via-point-based Stochastic Trajectory Optimization for Reactive Robot Behavior,

    J. Jankowski, L. Bruderm ¨uller, N. Hawes, and S. Calinon, “VP-STO: Via-point-based Stochastic Trajectory Optimization for Reactive Robot Behavior,” inIEEE Intl. Conf. on Robotics and Automation (ICRA), 2023, pp. 10 125–10 131

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.