SurgAtlas is a new dataset of 15,291 surgical videos totaling 2,391 hours with multi-level annotations that supports finetuning models to competitive performance on surgical benchmarks.
Cholecseg8k: a semantic segmen- tation dataset for laparoscopic cholecystectomy based on cholec80
13 Pith papers cite this work, alongside 43 external citations. Polarity classification is still indexing.
abstract
Computer-assisted surgery has been developed to enhance surgery correctness and safety. However, researchers and engineers suffer from limited annotated data to develop and train better algorithms. Consequently, the development of fundamental algorithms such as Simultaneous Localization and Mapping (SLAM) is limited. This article elaborates on the efforts of preparing the dataset for semantic segmentation, which is the foundation of many computer-assisted surgery mechanisms. Based on the Cholec80 dataset [3], we extracted 8,080 laparoscopic cholecystectomy image frames from 17 video clips in Cholec80 and annotated the images. The dataset is named CholecSeg8K and its total size is 3GB. Each of these images is annotated at pixel-level for thirteen classes, which are commonly founded in laparoscopic cholecystectomy surgery. CholecSeg8k is released under the license CC BY- NC-SA 4.0.
citation-role summary
citation-polarity summary
representative citing papers
A hierarchical prompt tree with self-reflection graph propagation enables positive forward and backward knowledge transfer in incremental surgical instrument segmentation, improving over baselines by more than 5% and 11% on two benchmarks.
SAM 3 outperforms SAM 2 under click prompting for zero-shot 3D medical segmentation across 16 datasets and 54 structures, with fewer failure modes in prompt-frame over-segmentation and prediction retention.
A consensus-based catalog of 18 validation pitfalls, with evidence that common practices understate uncertainty, hide failures, and flip algorithm rankings in surgical video AI.
Introduces the triplet segmentation task, CholecTriplet-Seg dataset with over 30,000 frames, and TargetFusionNet architecture extending Mask2Former for instance-level grounding of surgical <instrument, verb, target> triplets.
ReferEndoscopy plus attribute-retrieval and frequency-aware fusion yields open-vocabulary compositional referring segmentation that outperforms natural-image RIS baselines on endoscopic data and generalizes to an unseen robotic prostatectomy set.
Introduces OR-Action benchmark for multi-role fine-grained actions in OR videos and a vision-only temporal model with multi-to-single view alignment that outperforms graph-based approaches.
EndoGSim integrates MLLM-guided material initialization with 4D Gaussian Splatting and differentiable Material Point Method to achieve physics-aware 4D reconstruction and simulation of endoscopic scenes.
A multi-frame network with SAM 3-derived mask priors achieves 72.4% F1 tip and 58.0% F1 anchor localization in surgical videos without manual mask annotations for training.
Presents ATLAS-120k dataset and ATLAS model for context-aware surgical anatomy segmentation using foundation representations and temporal cues.
DEX is a modular network using dynamically activated experts and a group-EMA director to learn emergent modular representations for multi-modality medical vision foundation models, evaluated on a new 4M-image benchmark across 10 modalities and 26 downstream tasks.
DenseTRF adapts texture-aware representations via slot attention for unsupervised improvement of cross-domain generalization in surgical dense prediction tasks.
The paper summarizes results from the SurgToolLoc and SurgVU challenges held at MICCAI conferences from 2022 to 2025.
citing papers explorer
-
SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery
SurgAtlas is a new dataset of 15,291 surgical videos totaling 2,391 hours with multi-level annotations that supports finetuning models to competitive performance on surgical benchmarks.
-
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
A hierarchical prompt tree with self-reflection graph propagation enables positive forward and backward knowledge transfer in incremental surgical instrument segmentation, improving over baselines by more than 5% and 11% on two benchmarks.
-
Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data
SAM 3 outperforms SAM 2 under click prompting for zero-shot 3D medical segmentation across 16 datasets and 54 structures, with fewer failure modes in prompt-frame over-segmentation and prediction retention.
-
Current validation practice undermines surgical AI development
A consensus-based catalog of 18 validation pitfalls, with evidence that common practices understate uncertainty, hide failures, and flip algorithm rankings in surgical video AI.
-
Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
Introduces the triplet segmentation task, CholecTriplet-Seg dataset with over 30,000 frames, and TargetFusionNet architecture extending Mask2Former for instance-level grounding of surgical <instrument, verb, target> triplets.
-
Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation
ReferEndoscopy plus attribute-retrieval and frequency-aware fusion yields open-vocabulary compositional referring segmentation that outperforms natural-image RIS baselines on endoscopic data and generalizes to an unseen robotic prostatectomy set.
-
OR-Action: Multi-Role Video Understanding with Fine-Grained Actions
Introduces OR-Action benchmark for multi-role fine-grained actions in OR videos and a vision-only temporal model with multi-to-single view alignment that outperforms graph-based approaches.
-
EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting
EndoGSim integrates MLLM-guided material initialization with 4D Gaussian Splatting and differentiable Material Point Method to achieve physics-aware 4D reconstruction and simulation of endoscopic scenes.
-
Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos
A multi-frame network with SAM 3-derived mask priors achieves 72.4% F1 tip and 58.0% F1 anchor localization in surgical videos without manual mask annotations for training.
-
Surgical Anatomy Recognition with Context Learning using Foundation Representations
Presents ATLAS-120k dataset and ATLAS model for context-aware surgical anatomy segmentation using foundation representations and temporal cues.
-
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
DEX is a modular network using dynamically activated experts and a group-EMA director to learn emergent modular representations for multi-modality medical vision foundation models, evaluated on a new 4M-image benchmark across 10 modalities and 26 downstream tasks.
-
DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
DenseTRF adapts texture-aware representations via slot attention for unsupervised improvement of cross-domain generalization in surgical dense prediction tasks.
-
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
The paper summarizes results from the SurgToolLoc and SurgVU challenges held at MICCAI conferences from 2022 to 2025.