On a new Polish medical VQA benchmark from board certification exams, vision-language models perform better from question text alone than from images alone and score above chance from answer choices alone, indicating weak visual grounding.
Towards Developing a Multilingual and Code-Mixed Visual Question Answering System by Knowledge Distillation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
On a new Polish medical VQA benchmark from board certification exams, vision-language models perform better from question text alone than from images alone and score above chance from answer choices alone, indicating weak visual grounding.