A single multi-task CLIP model trained with one positive label per image matches task-specific surgical benchmarks on phase, CVS, and triplet recognition.
Medical I mage Analysis 78, 102433 (2022) 10 S
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
A single multi-task CLIP model trained with one positive label per image matches task-specific surgical benchmarks on phase, CVS, and triplet recognition.