Pith. sign in

REVIEW 1 cited by

HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10075 v2 pith:7Y2BQOLV submitted 2024-05-16 cs.CV cs.AI

classification cs.CVcs.AI
keywords surgicalhierarchicalmodelhecvlphaserecognitiontextsacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Natural language could play an important role in developing generalist surgical models by providing a broad source of supervision from raw texts. This flexible form of supervision can enable the model's transferability across datasets and tasks as natural language can be used to reference learned visual concepts or describe new ones. In this work, we present HecVL, a novel hierarchical video-language pretraining approach for building a generalist surgical model. Specifically, we construct a hierarchical video-text paired dataset by pairing the surgical lecture video with three hierarchical levels of texts: at clip-level, atomic actions using transcribed audio texts; at phase-level, conceptual text summaries; and at video-level, overall abstract text of the surgical procedure. Then, we propose a novel fine-to-coarse contrastive learning framework that learns separate embedding spaces for the three video-text hierarchies using a single model. By disentangling embedding spaces of different hierarchical levels, the learned multi-modal representations encode short-term and long-term surgical concepts in the same model. Thanks to the injected textual semantics, we demonstrate that the HecVL approach can enable zero-shot surgical phase recognition without any human annotation. Furthermore, we show that the same HecVL model for surgical phase recognition can be transferred across different surgical procedures and medical centers. The code is available at https://github.com/CAMMA-public/SurgVLP

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

    cs.CV 2025-01 conditional novelty 6.0 of 10

    RASO recognizes surgical instruments and anatomy in images and video using a weakly supervised training pipeline built from automatically generated tag-image-text pairs from surgical lecture videos.

Pith tools