Fine-tuning a CLIP image encoder with learned prompts improves surgical phase recognition by 1 to 4 accuracy points on three datasets, though the comparison baseline does not control for CLIP pretraining.
In: 2009 IEEE conferenc e on computer vision and pattern recognition
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model
Fine-tuning a CLIP image encoder with learned prompts improves surgical phase recognition by 1 to 4 accuracy points on three datasets, though the comparison baseline does not control for CLIP pretraining.