Pith. sign in

REVIEW 3 cited by

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.15421 v2 pith:TSSVTOWK submitted 2024-11-23 cs.CV

classification cs.CV
keywords surgicalophclippretraininghierarchicalretrieval-augmentedvideosadvancedframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Surgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A student CLIP model distilled from nine medical CLIP teachers outperforms its teachers across most of 58 biomedical benchmarks.

  2. Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SPA uses few-shot spatial adaptation, a diffusion model trained on a hand-provided procedure graph, and test-time mutual agreement to achieve strong surgical phase recognition with minimal labeled data.

  3. MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Applying GRPO reinforcement learning with spatial-semantic rewards and a Chain-of-Box prompt achieves state-of-the-art medical image grounding without chain-of-thought annotations.

Pith tools