In a controlled image-text task, Claude 3.5 integrated visual and question cues to make pragmatic speaker-ignorance inferences, while GPT-4o and Gemini 1.5 Pro relied more on literal meanings.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
In a controlled image-text task, Claude 3.5 integrated visual and question cues to make pragmatic speaker-ignorance inferences, while GPT-4o and Gemini 1.5 Pro relied more on literal meanings.