A study of 61 prompt variants across 10 vision-language models and 3 benchmarks finds accuracy swings of up to 15 points, with proprietary models more sensitive than open-source ones.
**Question:** What design element best describes the visuals? **Options:** A
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
A study of 61 prompt variants across 10 vision-language models and 3 benchmarks finds accuracy swings of up to 15 points, with proprietary models more sensitive than open-source ones.