Pith. sign in

REVIEW 1 cited by

A Dataset and Baselines for Visual Question Answering on Art

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.12520 v1 pith:7VOL3K4K submitted 2020-08-28 cs.CV cs.CL

classification cs.CVcs.CL
keywords answeringquestionvisualdatasetknowledgequestionsbaselinecorrectness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Answering questions related to art pieces (paintings) is a difficult task, as it implies the understanding of not only the visual information that is shown in the picture, but also the contextual knowledge that is acquired through the study of the history of art. In this work, we introduce our first attempt towards building a new dataset, coined AQUA (Art QUestion Answering). The question-answer (QA) pairs are automatically generated using state-of-the-art question generation methods based on paintings and comments provided in an existing art understanding dataset. The QA pairs are cleansed by crowdsourcing workers with respect to their grammatical correctness, answerability, and answers' correctness. Our dataset inherently consists of visual (painting-based) and knowledge (comment-based) questions. We also present a two-branch model as baseline, where the visual and knowledge questions are handled independently. We extensively compare our baseline model against the state-of-the-art models for question answering, and we provide a comprehensive study about the challenges and potential future directions for visual question answering on art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MMArt offers 74,234 paintings with four complementary perspectives and a unified caption, and shows the perspectives are task-asymmetric.

Pith tools