Pith. sign in

REVIEW 2 cited by

Document Understanding Dataset and Evaluation (DUDE)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.08455 v3 pith:7P576ZQI submitted 2023-05-15 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords documentdatasetdudeevaluationunderstandingcommunitycreatingcurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We call on the Document AI (DocAI) community to reevaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks. Document Understanding Dataset and Evaluation (DUDE) seeks to remediate the halted research progress in understanding visually-rich documents (VRDs). We present a new dataset with novelties related to types of questions, answers, and document layouts based on multi-industry, multi-domain, and multi-page VRDs of various origins, and dates. Moreover, we are pushing the boundaries of current methods by creating multi-task and multi-domain evaluation setups that more accurately simulate real-world situations where powerful generalization and adaptation under low-resource settings are desired. DUDE aims to set a new standard as a more practical, long-standing benchmark for the community, and we hope that it will lead to future extensions and contributions that address real-world challenges. Finally, our work illustrates the importance of finding more efficient ways to model language, images, and layout in DocAI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Internalized Reasoning for Long-Context Visual Document Understanding

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    Synthetic page-ranked reasoning traces plus low-strength model merging give a 32B VLM 58.3 on MMLongBenchDoc, beating a 235B teacher while cutting output tokens ~12× versus explicit reasoning.

  2. Hierarchical Evidence-Driven Reasoning for Long Document Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical multimodal RAG pipeline with GRPO-trained multi-page evidence verification and memory-guided iteration improves long-document QA accuracy by ~8% over prior open-source baselines.

Pith tools