Pith. sign in

REVIEW 16 cited by

2018 Robotic Scene Segmentation Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.11190 v3 pith:DOWWJZ5F submitted 2020-01-30 cs.CV cs.RO

2018 Robotic Scene Segmentation Challenge

classification cs.CV cs.RO
keywords challengesegmentationtissueinstrumentbackgrounddatasetmotionporcine
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In 2015 we began a sub-challenge at the EndoVis workshop at MICCAI in Munich using endoscope images of ex-vivo tissue with automatically generated annotations from robot forward kinematics and instrument CAD models. However, the limited background variation and simple motion rendered the dataset uninformative in learning about which techniques would be suitable for segmentation in real surgery. In 2017, at the same workshop in Quebec we introduced the robotic instrument segmentation dataset with 10 teams participating in the challenge to perform binary, articulating parts and type segmentation of da Vinci instruments. This challenge included realistic instrument motion and more complex porcine tissue as background and was widely addressed with modifications on U-Nets and other popular CNN architectures. In 2018 we added to the complexity by introducing a set of anatomical objects and medical devices to the segmented classes. To avoid over-complicating the challenge, we continued with porcine data which is dramatically simpler than human tissue due to the lack of fatty tissue occluding many organs.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

    cs.CV 2026-06 unverdicted novelty 7.0

    SurgAtlas is a new dataset of 15,291 surgical videos totaling 2,391 hours with multi-level annotations that supports finetuning models to competitive performance on surgical benchmarks.

  2. Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework

    cs.CV 2026-04 conditional novelty 7.0

    A hierarchical prompt tree with self-reflection graph propagation enables positive forward and backward knowledge transfer in incremental surgical instrument segmentation, improving over baselines by more than 5% and ...

  3. SurgiSR4K: A High-Resolution Endoscopic Video Dataset for Robotic-Assisted Minimally Invasive Procedures

    eess.IV 2025-06 unverdicted novelty 7.0

    Introduces the first publicly accessible native 4K resolution endoscopic video dataset for robotic-assisted minimally invasive procedures.

  4. Stitch-Inferencer: Enhance Endoscopic Video Segmentation and Tracking via Panoramic Reconstruction

    cs.CV 2026-07 conditional novelty 6.0

    An inference-time framework builds an explicit panoramic memory from endoscopic video and uses it to improve off-the-shelf segmentation and tracking models without retraining.

  5. SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy

    cs.RO 2026-07 conditional novelty 6.0

    SurgAM fuses DINOv2 semantic features with Stable Diffusion spatial features plus hierarchical prompts to predict surgical affordance maps that enable autonomous phantom tasks.

  6. Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

    cs.CV 2026-06 unverdicted novelty 6.0

    FAROS uses flow-guided propagation from zero-shot masks and optical flow to create dense temporally consistent labels from sparse keyframes, improving joint multi-task learning across temporal and spatial surgical tas...

  7. RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

    cs.CV 2026-05 unverdicted novelty 6.0

    RoboSurg-VQA is a new segmentation-aware VQA benchmark created by repurposing public surgical datasets with fixed clinically motivated questions and closed answer sets.

  8. On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

    cs.CV 2026-01 conditional novelty 6.0

    RGB-D pre-training with explicit cross-modal objectives (MultiMAE) improves surgical detection, segmentation, pose, and depth estimation over RGB-only pre-training, with gains persisting when fine-tuned on 25% of labe...

  9. SAM 2: Segment Anything in Images and Videos

    cs.CV 2024-08 conditional novelty 6.0

    SAM 2 delivers more accurate video segmentation with 3x fewer user interactions and 6x faster image segmentation than the original SAM by training a streaming-memory transformer on the largest video segmentation datas...

  10. DeGenseGS: Geometrically and Semantically Decoupled Surgical Scene Understanding in 4D Gaussian Splatting

    cs.CV 2026-07 conditional novelty 5.5

    Decoupling geometry and semantics in 4DGS via HexPlane kinematic latents and rasterization-native extraction raises surgical semantic mIoU from 53.46% to 68.20% on CholecSeg8k.

  11. Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

    cs.CV 2026-06 unverdicted novelty 5.0

    A multi-frame network with SAM 3-derived mask priors achieves 72.4% F1 tip and 58.0% F1 anchor localization in surgical videos without manual mask annotations for training.

  12. SurfSurg6D: Geometry Consistent Dense Correspondence for Textureless Surgical Instrument Pose Estimation

    cs.CV 2026-05 unverdicted novelty 5.0

    A new synthetic dataset and geometry-consistent dense correspondence framework improve RGB-only pose estimation accuracy for surgical instruments on three evaluation datasets.

  13. SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking

    cs.CV 2025-11 conditional novelty 5.0

    SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.

  14. Surgical Visual Understanding (SurgVU) Dataset

    cs.CV 2025-01 unverdicted novelty 5.0

    Releases the SurgVU dataset of surgical videos and labels to enable machine learning research in surgical data science.

  15. SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge

    cs.CV 2024-07 accept novelty 5.0

    SegSTRONG-C provides a new benchmark where top models reach 0.9394 DSC and 0.9301 NSD on corrupted surgical tool segmentation tests, showing conventional techniques help but calling for more innovative robustness methods.

  16. Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

    cs.CV 2026-07 conditional novelty 4.0

    LoRA on SAM3’s prompt encoder, detector, and tracker (0.98% of parameters) raises surgical concept-segmentation mIoU over zero-shot SAM3 and Medical SAM3 while fitting in ~9 GB GPU memory.