BLaVe-CoT predicts whether candidate answers to a visual question refer to the same or different image regions, improving F1 on the VQA-AnswerTherapy benchmark.
Vqask: a multimodal android gpt- based application to help blind users visualize pictures
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
BLaVe-CoT predicts whether candidate answers to a visual question refer to the same or different image regions, improving F1 on the VQA-AnswerTherapy benchmark.