A new 65M-image, 200M-QA dataset for museum exhibits lets fine-tuned vision-language models beat general-purpose VLMs on museum attribute questions, especially on questions requiring background knowledge.
Language bias in Visual Question Answering: A Survey and Taxonomy
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Visual question answering (VQA) is a challenging task, which has attracted more and more attention in the field of computer vision and natural language processing. However, the current visual question answering has the problem of language bias, which reduces the robustness of the model and has an adverse impact on the practical application of visual question answering. In this paper, we conduct a comprehensive review and analysis of this field for the first time, and classify the existing methods according to three categories, including enhancing visual information, weakening language priors, data enhancement and training strategies. At the same time, the relevant representative methods are introduced, summarized and analyzed in turn. The causes of language bias are revealed and classified. Secondly, this paper introduces the datasets mainly used for testing, and reports the experimental results of various existing methods. Finally, we discuss the possible future research directions in this field.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Understanding Museum Exhibits using Vision-Language Reasoning
A new 65M-image, 200M-QA dataset for museum exhibits lets fine-tuned vision-language models beat general-purpose VLMs on museum attribute questions, especially on questions requiring background knowledge.