REVIEW 1 cited by
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a question. Thus, we introduce the VQA-HAT (Human ATtention) dataset. We evaluate attention maps generated by state-of-the-art VQA models against human attention both qualitatively (via visualizations) and quantitatively (via rank-order correlation). Overall, our experiments show that current attention models in VQA do not seem to be looking at the same regions as humans.
Forward citations
Cited by 1 Pith paper
-
Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams
A 250M-parameter vision-language model fine-tuned on Skew-T diagrams achieves CSI comparable to IFS-HRES for 3-hour precipitation probability in South Korean summer.
Discussion (0). Sign in to comment.