Pith. sign in

REVIEW 3 cited by

Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.15220 v5 pith:4FF4RCNW submitted 2023-07-27 cs.CV cs.AI

classification cs.CVcs.AI
keywords surgicallearningmulti-modalsurgvlpvideodownstreamlanguagelectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The [training code](https://github.com/CAMMA-public/PeskaVLP) and [weights](https://github.com/CAMMA-public/SurgVLP) are public.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment

    cs.CR 2025-02 conditional novelty 7.0 of 10

    An adversarial domain alignment method steals a medical multimodal LLM's radiology report generation using natural images and an oracle LLM, without medical data.

  2. SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SurgVLM, a family of surgical vision-language models trained on 1.81M frames and 7.79M conversations, outperforms 14 commercial VLMs on a six-dataset surgical benchmark.

  3. SurgX: Neuron-Concept Association for Explainable Surgical Phase Recognition

    cs.CV 2025-07 conditional novelty 5.0 of 10

    SurgX associates neurons in surgical phase recognition models with surgical concepts and uses the concepts of high-contribution neurons to explain predictions on Cholec80.

Pith tools