Pith. sign in

REVIEW 2 cited by

Scene Text Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.13648 v2 pith:OLFDN42D submitted 2019-05-31 cs.CV

classification cs.CV
keywords textdatasetinformationscenevisualansweringfurtherquestion
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process. We use this dataset to define a series of tasks of increasing difficulty for which reading the scene text in the context provided by the visual information is necessary to reason and generate an appropriate answer. We propose a new evaluation metric for these tasks to account both for reasoning errors as well as shortcomings of the text recognition module. In addition we put forward a series of baseline methods, which provide further insight to the newly released dataset, and set the scene for further research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    DashboardQA is a new benchmark of 405 question-answer pairs over 112 interactive dashboards; the strongest tested GUI agent reaches only 38.69% accuracy.

  2. A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

    cs.CV 2024-11 unverdicted

    A broad but dated survey of VQA datasets and algorithms that organizes the pre-2021 literature into four dataset categories and six model paradigms.

Pith tools