Vision-language models show stronger shape-over-color in-context inductive bias from images than from text, and are biased toward the first-mentioned adjective in text.
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The in-context inductive biases of vision-language models differ across modalities
Vision-language models show stronger shape-over-color in-context inductive bias from images than from text, and are biased toward the first-mentioned adjective in text.