Pith. sign in

REVIEW 2 cited by

Pixel-Wise Recognition for Holistic Surgical Scene Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11174 v3 pith:EICQF46G submitted 2024-01-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords surgicalscenetasksunderstandingbenchmarkholisticinstrumentsegmentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents the Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies (GraSP) dataset, a curated benchmark that models surgical scene understanding as a hierarchy of complementary tasks with varying levels of granularity. Our approach encompasses long-term tasks, such as surgical phase and step recognition, and short-term tasks, including surgical instrument segmentation and atomic visual actions detection. To exploit our proposed benchmark, we introduce the Transformers for Actions, Phases, Steps, and Instrument Segmentation (TAPIS) model, a general architecture that combines a global video feature extractor with localized region proposals from an instrument segmentation model to tackle the multi-granularity of our benchmark. Through extensive experimentation in ours and alternative benchmarks, we demonstrate TAPIS's versatility and state-of-the-art performance across different tasks. This work represents a foundational step forward in Endoscopic Vision, offering a novel framework for future research towards holistic surgical scene understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TEMSET-24K: Densely Annotated Dataset for Indexing Multipart Endoscopic Videos using Surgical Timeline Segmentation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A new densely labeled dataset of 24,000+ TEMS surgical video clips with phase, task, and action labels, plus a benchmark model for automatic timeline indexing.

  2. EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A training-free Lagrangian motion magnification framework with periodic reference resetting and tissue-aware dual-mask control improves vascular pulsation visibility in endoscopic surgery videos.

Pith tools