Pith. sign in

REVIEW 2 cited by

CudaSIFT-SLAM: multiple-map visual SLAM for full procedure mapping in real human endoscopy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16932 v1 pith:NYUEUCZN submitted 2024-05-27 cs.RO cs.CV

classification cs.ROcs.CV
keywords datasetfullhumanmappingorb-slam3realreal-timesub-maps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Monocular visual simultaneous localization and mapping (V-SLAM) is nowadays an irreplaceable tool in mobile robotics and augmented reality, where it performs robustly. However, human colonoscopies pose formidable challenges like occlusions, blur, light changes, lack of texture, deformation, water jets or tool interaction, which result in very frequent tracking losses. ORB-SLAM3, the top performing multiple-map V-SLAM, is unable to recover from them by merging sub-maps or relocalizing the camera, due to the poor performance of its place recognition algorithm based on ORB features and DBoW2 bag-of-words. We present CudaSIFT-SLAM, the first V-SLAM system able to process complete human colonoscopies in real-time. To overcome the limitations of ORB-SLAM3, we use SIFT instead of ORB features and replace the DBoW2 direct index with the more computationally demanding brute-force matching, being able to successfully match images separated in time for relocation and map merging. Real-time performance is achieved thanks to CudaSIFT, a GPU implementation for SIFT extraction and brute-force matching. We benchmark our system in the C3VD phantom colon dataset, and in a full real colonoscopy from the Endomapper dataset, demonstrating the capabilities to merge sub-maps and relocate in them, obtaining significantly longer sub-maps. Our system successfully maps in real-time 88 % of the frames in the C3VD dataset. In a real screening colonoscopy, despite the much higher prevalence of occluded and blurred frames, the mapping coverage is 53 % in carefully explored areas and 38 % in the full sequence, a 70 % improvement over ORB-SLAM3.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. C3VDv2 -- Colonoscopy 3D video dataset with enhanced realism

    eess.IV 2025-06 conditional novelty 6.0 of 10

    C3VDv2 releases 169 registered colonoscopy videos with depth, normals, optical flow, occlusion, pose, and 3D model ground truth, plus eight full-colon screening videos and fifteen deformation videos with enhanced realism.

  2. ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    ROOM is an open simulation pipeline that generates photorealistic, multimodal synthetic bronchoscopy data from CT scans, and fine-tuning depth models on this data improves their performance on an external phantom-base...

Pith tools