Pith. sign in

REVIEW 15 cited by

DocVQA: A Dataset for VQA on Document Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.00398 v3 pith:UDBRHHUC submitted 2020-07-01 cs.CV cs.IR

classification cs.CVcs.IR
keywords datasetdocumentdocvqaimagesmodelsquestionscomprehensionexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at docvqa.org

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RADIO1D: Elastic Representations for Condensed Vision Modeling

    cs.CV 2026-07 accept novelty 7.0 of 10

    RADIO1D produces elastic hierarchical 1D visual tokens via multi-teacher distillation that match or beat fixed 2D encoders in VLMs at lower token counts.

  2. Representation Forcing for Bottleneck-Free Unified Multimodal Models

    cs.CV 2026-05 unverdicted novelty 6.5 of 10

    Representation Forcing lets a UMM decoder autoregressively predict its own understanding representations as in-context tokens that guide pixel-space diffusion, matching VAE-based generation without an external latent space.

  3. Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Persistent agent skill evolution is sparse, validation-filtered search whose gains depend strongly on model, benchmark, and which feedback (failures versus successes) is shown.

  4. VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality

    cs.CV 2025-09 conditional novelty 6.0 of 10

    VLM-in-the-Wild provides an enterprise-focused benchmark and the BlockWeaver OCR matching algorithm, reporting that a small fine-tuned model can rival a 32B model on some tasks.

  5. VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VDInstruct achieves strong zero-shot key-information extraction by combining a region detector with content-aware vision tokenization, using about 500 image tokens per page.

  6. EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EffiVLM-Bench is a benchmark study showing token compression is task- and model-dependent, KV cache methods are more loyal, and parameter compression preserves accuracy better at typical ratios.

  7. Spoken question answering for visual queries

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A LLaVA-style model with an added Whisper speech encoder answers spoken questions about images, trained on TTS-synthesized speech and reaching near the text-input baseline.

  8. Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Griffon-R generates its own grounding hints and rationale before answering, achieving state-of-the-art visual reasoning on VSR and CLEVR while improving MMBench, ScienceQA, and TextVQA.

  9. CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Multimodal context-learning benchmark CLBench-V separates grounding, information application, and knowledge acquisition; the best evaluated model scores 0.2847.

  10. Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

    cs.IR 2026-03 conditional novelty 5.0 of 10

    CoCoA forces an MLLM to reconstruct masked text through a single EOS token, improving multimodal embedding quality on MMEB-V1 and matching MoCa at 3B with far less pretraining data.

  11. On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

    cs.IR 2025-06 conditional novelty 4.0 of 10

    Preprocessing financial PDFs into text, tables, and chart data with existing tools improves LLM question-answering accuracy over direct GPT-4o image input in a small private evaluation.

  12. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

  13. DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation

    cs.IR 2026-05 conditional novelty 3.0 of 10

    DocAnnot combines an LVLM, OCR, and a spatial matching heuristic to auto-annotate KIE documents at F1 0.68–0.85, and models trained on that data reach roughly 0.68 F1 on CORD.

  14. Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning

    cs.CV 2025-08 conditional novelty 3.0 of 10

    A prompt-engineered Claude 3.7, guided by GPT-4o-generated prompts and few-shot examples, reaches near-ceiling accuracy on most of the 18 MIRAGE multi-image reasoning tasks.

  15. Differential Multimodal Transformers

    cs.AI 2025-07 reject novelty 2.0 of 10

    A proposed multimodal extension of differential attention turns out to be equivalent to a learned scalar multiplier on standard attention, so the reported retrieval gains do not support the claimed mechanism.

Pith tools