An integrated structural prompt learning method for CLIP reports a state-of-the-art average harmonic mean of 80.70 on 11 base-to-new few-shot classification benchmarks using self- and cross-modal prompt-token interactions and difficulty-based loss weighting.
In: 2009 IEEE Conference on Computer Vision and Pattern Recognition
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Integrated Structural Prompt Learning for Vision-Language Models
An integrated structural prompt learning method for CLIP reports a state-of-the-art average harmonic mean of 80.70 on 11 base-to-new few-shot classification benchmarks using self- and cross-modal prompt-token interactions and difficulty-based loss weighting.