Pith. sign in

REVIEW 2 cited by

BodySLAM: A Generalized Monocular Visual SLAM Framework for Surgical Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.03078 v2 pith:WFXMDEBK submitted 2024-08-06 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords monocularendoscopicestimationbodyslamchallengesdepthmvslamapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Endoscopic surgery relies on two-dimensional views, posing challenges for surgeons in depth perception and instrument manipulation. While Monocular Visual Simultaneous Localization and Mapping (MVSLAM) has emerged as a promising solution, its implementation in endoscopic procedures faces significant challenges due to hardware limitations, such as the use of a monocular camera and the absence of odometry sensors. This study presents BodySLAM, a robust deep learning-based MVSLAM approach that addresses these challenges through three key components: CycleVO, a novel unsupervised monocular pose estimation module; the integration of the state-of-the-art Zoe architecture for monocular depth estimation; and a 3D reconstruction module creating a coherent surgical map. The approach is rigorously evaluated using three publicly available datasets (Hamlyn, EndoSLAM, and SCARED) spanning laparoscopy, gastroscopy, and colonoscopy scenarios, and benchmarked against four state-of-the-art methods. Results demonstrate that CycleVO exhibited competitive performance with the lowest inference time among pose estimation methods, while maintaining robust generalization capabilities, whereas Zoe significantly outperformed existing algorithms for depth estimation in endoscopy. BodySLAM's strong performance across diverse endoscopic scenarios demonstrates its potential as a viable MVSLAM solution for endoscopic applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new 7.1-million-frame RGB-D dataset of real cataract surgery with auto-generated 3D hand meshes and instrument poses, plus two baseline models that set benchmarks.

  2. Unifying Scale-Aware Depth Prediction and Perceptual Priors for Monocular Endoscope Pose Estimation and Tissue Reconstruction

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A monocular endoscopy framework fuses Depth Pro and Depth Anything depth with RAFT-LPIPS temporal refinement and dog-leg pose optimization to reconstruct tissue surfaces and camera trajectories.

Pith tools