A single multi-task CLIP model trained with one positive label per image matches task-specific surgical benchmarks on phase, CVS, and triplet recognition.
In: Proceedings of th e IEEE/CVF Conference on Computer Vision and Pattern Recognition
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
A single multi-task CLIP model trained with one positive label per image matches task-specific surgical benchmarks on phase, CVS, and triplet recognition.