A visual prompt tuning method that injects fixed color, texture, and shape features plus cascaded self-attention maps into a frozen ViT improves accuracy across 34 datasets with 0.74% trainable parameters.
Color is one of the most expressive and invariant visual cues, robust to rotation and scaling (Swain & Ballard, 1991)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adapting Vision Foundation Models with Cascaded Semantics
A visual prompt tuning method that injects fixed color, texture, and shape features plus cascaded self-attention maps into a frozen ViT improves accuracy across 34 datasets with 0.74% trainable parameters.