Pith. sign in

REVIEW 4 cited by

EndoOmni: Zero-Shot Cross-Dataset Depth Estimation in Endoscopy by Robust Self-Learning from Noisy Labels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05442 v4 pith:LJVAV2XK submitted 2024-09-09 cs.CV

classification cs.CV
keywords depthestimationmodeltrainingendoomnidataendoscopylabels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Single-image depth estimation is essential for endoscopy tasks such as localization, reconstruction, and augmented reality. Most existing methods in surgical scenes focus on in-domain depth estimation, limiting their real-world applicability. This constraint stems from the scarcity and inferior labeling quality of medical data for training. In this work, we present EndoOmni, the first foundation model for zero-shot cross-domain depth estimation for endoscopy. To harness the potential of diverse training data, we refine the advanced self-learning paradigm that employs a teacher model to generate pseudo-labels, guiding a student model trained on large-scale labeled and unlabeled data. To address training disturbance caused by inherent noise in depth labels, we propose a robust training framework that leverages both depth labels and estimated confidence from the teacher model to jointly guide the student model training. Moreover, we propose a weighted scale-and-shift invariant loss to adaptively adjust learning weights based on label confidence, thus imposing learning bias towards cleaner label pixels while reducing the influence of highly noisy pixels. Experiments on zero-shot relative depth estimation show that our EndoOmni improves state-of-the-art methods in medical imaging for 33\% and existing foundation models for 34\% in terms of absolute relative error on specific datasets. Furthermore, our model provides strong initialization for fine-tuning metric depth estimation, maintaining superior performance in both in-domain and out-of-domain scenarios. The source code is publicly available at https://github.com/TianCuteQY/EndoOmni.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new 7.1-million-frame RGB-D dataset of real cataract surgery with auto-generated 3D hand meshes and instrument poses, plus two baseline models that set benchmarks.

  2. C3VDv2 -- Colonoscopy 3D video dataset with enhanced realism

    eess.IV 2025-06 conditional novelty 6.0 of 10

    C3VDv2 releases 169 registered colonoscopy videos with depth, normals, optical flow, occlusion, pose, and 3D model ground truth, plus eight full-colon screening videos and fifteen deformation videos with enhanced realism.

  3. ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    ROOM is an open simulation pipeline that generates photorealistic, multimodal synthetic bronchoscopy data from CT scans, and fine-tuning depth models on this data improves their performance on an external phantom-base...

  4. Harnessing Foundation Models for Robust and Generalizable 6-DOF Bronchoscopy Localization

    cs.CV 2025-05 conditional novelty 4.0 of 10

    PANSv2, a combination of foundation-model depth and landmark cues with centerline constraints and failure re-initialization, reports SR-5 of 46.5% on filtered frames, 18.1 points above the previous best on a private dataset.

Pith tools