Pith. sign in

REVIEW 3 cited by

Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1704.03162 v2 pith:7ZTYASTF submitted 2017-04-11 cs.CV

classification cs.CV
keywords modelquestionansweringresultsvisualbaselineimagereported
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents a new baseline for visual question answering task. Given an image and a question in natural language, our model produces accurate answers according to the content of the image. Our model, while being architecturally simple and relatively small in terms of trainable parameters, sets a new state of the art on both unbalanced and balanced VQA benchmark. On VQA 1.0 open ended challenge, our model achieves 64.6% accuracy on the test-standard set without using additional data, an improvement of 0.4% over state of the art, and on newly released VQA 2.0, our model scores 59.7% on validation set outperforming best previously reported results by 0.5%. The results presented in this paper are especially interesting because very similar models have been tried before but significantly lower performance were reported. In light of the new results we hope to see more meaningful research on visual question answering in the future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AdaCM2 adaptively evicts low-attention video tokens using cross-modality attention scores, keeping memory bounded while improving accuracy on long-video QA, captioning, and classification.

  2. Answering Questions about Data Visualizations using Efficient Bimodal Fusion

    cs.CV 2019-08 accept novelty 6.0 of 10

    PReFIL combines LSTM question embeddings with two levels of convolutional features via 1x1 convolutions and recurrent spatial aggregation, setting new state-of-the-art accuracy on FigureQA and DVQA.

  3. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.

Pith tools