Pith. sign in

REVIEW 3 cited by

MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.18491 v3 pith:N2EPT4AH submitted 2025-03-24 cs.CL

classification cs.CL
keywords knowledgecommonsensemagic-vqareasoninginferencelvlmsvisualanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual Question Answering (VQA) requires reasoning across visual and textual modalities, yet Large Vision-Language Models (LVLMs) often lack integrated commonsense knowledge, limiting their robustness in real-world scenarios. To address this, we introduce MAGIC-VQA, a novel framework that enhances VQA by systematically integrating commonsense knowledge with LVLMs. MAGIC-VQA employs a three-stage process: (1) Explicit Knowledge Integration from external sources, (2) By-Type Post-Processing for contextual refinement, and (3) Implicit Knowledge Augmentation using a Graph Neural Network (GNN) for structured reasoning. While GNNs bring greater depth to structured inference, they enable superior relational inference beyond LVLMs. MAGIC-VQA bridges a key gap by unifying commonsensse knowledge with LVLM-driven reasoning, eliminating the need for extensive pre-training or complex prompt tuning. Our framework achieves state-of-the-art performance on benchmark datasets, significantly improving commonsense reasoning in VQA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A self-evolving multimodal model using continuous self-consistency rewards improves math reasoning by about 2–3% using only raw images, without labels or external reward models.

  2. Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Aspect ratio of weight matrices biases heavy-tail spectral metrics; the new FARMS subsampling method removes this bias and improves downstream layer-wise tuning.

  3. VisionTrap: Unanswerable Questions On Visual Data

    cs.CV 2025-07 conditional novelty 5.0 of 10

    VisionTrap shows that GPT-4o, GPT-4.1, Gemini Flash 2.5, and LLaVA tend to answer unanswerable visual questions rather than abstain, especially when given multiple-choice options.

Pith tools