Pith. sign in

REVIEW 1 cited by

Graph-Structured Representations for Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1609.05600 v2 pith:5QPIUCT7 submitted 2016-09-19 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords questionrepresentationsscenestructurevisualaccuracyansweringapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes to improve visual question answering (VQA) with structured representations of both scene contents and questions. A key challenge in VQA is to require joint reasoning over the visual and text domains. The predominant CNN/LSTM-based approach to VQA is limited by monolithic vector representations that largely ignore structure in the scene and in the form of the question. CNN feature vectors cannot effectively capture situations as simple as multiple object instances, and LSTMs process questions as series of words, which does not reflect the true complexity of language structure. We instead propose to build graphs over the scene objects and over the question words, and we describe a deep neural network that exploits the structure in these representations. This shows significant benefit over the sequential processing of LSTMs. The overall efficacy of our approach is demonstrated by significant improvements over the state-of-the-art, from 71.2% to 74.4% in accuracy on the "abstract scenes" multiple-choice benchmark, and from 34.7% to 39.1% in accuracy over pairs of "balanced" scenes, i.e. images with fine-grained differences and opposite yes/no answers to a same question.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cognitively-Inspired Emergent Communication via Knowledge Graphs for Assisting the Visually Impaired

    cs.AI 2025-05 reject novelty 5.0 of 10

    VAG-EC, a graph-based emergent communication method, reports higher TopSim and Context Independence scores than a baseline EC model on synthetic dining scenes, though the evaluation is incomplete.

Pith tools