Training a visual question answering model on LLM-generated subtask rationales with bounding boxes improves GQA accuracy from 64.0 to 65.1 percent and adds grounded explanations.
Cric: A vqa dataset for compositional reasoning on vision and commonsense
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Visually Interpretable Subtask Reasoning for Visual Question Answering
Training a visual question answering model on LLM-generated subtask rationales with bounding boxes improves GQA accuracy from 64.0 to 65.1 percent and adds grounded explanations.