Pith. sign in

REVIEW 1 cited by

Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1606.03556 v2 pith:IIGRTC22 submitted 2016-06-11 cs.CV cs.CL

classification cs.CVcs.CL
keywords attentionhumanhumansquestionregionsansweransweringlook
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a question. Thus, we introduce the VQA-HAT (Human ATtention) dataset. We evaluate attention maps generated by state-of-the-art VQA models against human attention both qualitatively (via visualizations) and quantitatively (via rank-order correlation). Overall, our experiments show that current attention models in VQA do not seem to be looking at the same regions as humans.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams

    physics.ao-ph 2025-08 conditional novelty 6.0 of 10

    A 250M-parameter vision-language model fine-tuned on Skew-T diagrams achieves CSI comparable to IFS-HRES for 3-hour precipitation probability in South Korean summer.

Pith tools