Pith. sign in

REVIEW 6 cited by

A Focused Dynamic Attention Model for Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1604.01485 v1 pith:QXUIY7F4 submitted 2016-04-06 cs.CV cs.CLcs.NE

classification cs.CVcs.CLcs.NE
keywords questionfeaturesimagevisualregionsansweringanswersattention
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Visual Question and Answering (VQA) problems are attracting increasing interest from multiple research disciplines. Solving VQA problems requires techniques from both computer vision for understanding the visual contents of a presented image or video, as well as the ones from natural language processing for understanding semantics of the question and generating the answers. Regarding visual content modeling, most of existing VQA methods adopt the strategy of extracting global features from the image or video, which inevitably fails in capturing fine-grained information such as spatial configuration of multiple objects. Extracting features from auto-generated regions -- as some region-based image recognition methods do -- cannot essentially address this problem and may introduce some overwhelming irrelevant features with the question. In this work, we propose a novel Focused Dynamic Attention (FDA) model to provide better aligned image content representation with proposed questions. Being aware of the key words in the question, FDA employs off-the-shelf object detector to identify important regions and fuse the information from the regions and global features via an LSTM unit. Such question-driven representations are then combined with question representation and fed into a reasoning unit for generating the answers. Extensive evaluation on a large-scale benchmark dataset, VQA, clearly demonstrate the superior performance of FDA over well-established baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal Unified Attention Networks for Vision-and-Language Interactions

    cs.CV 2019-08 conditional novelty 5.0 of 10

    MUAN applies a stacked gated self-attention block to concatenated visual and textual tokens, jointly modeling intra-modal and inter-modal attention, and achieves top results on VQA and visual grounding benchmarks.

  2. Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    Introduces GRIT, LTMI, and a hierarchical attention framework claiming performance gains on image captioning, visual dialog, and ALFRED instruction following.

  3. A Better Way to Attend: Attention with Trees for Video Question Answering

    cs.CV 2019-09 conditional novelty 4.0 of 10

    A tree-structured memory network that uses parse trees and distinguishes visual from verbal words improves video question answering accuracy over flat sequence attention baselines.

  4. Visual question answering: from early developments to recent advances -- a survey

    cs.CV 2025-01 conditional novelty 2.0 of 10

    A survey that classifies VQA architectures by encoder, fusion, and decoder, reviews datasets and metrics, and discusses applications and future directions.

  5. Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A survey that organizes VQA methods from feature extraction through MLLM reasoning, datasets, and metrics, without introducing new experimental results.

  6. A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

    cs.CV 2024-11 unverdicted

    A broad but dated survey of VQA datasets and algorithms that organizes the pre-2021 literature into four dataset categories and six model paradigms.

Pith tools