PureDocBench shows document parsing is far from solved, with top models at ~74/100, small specialists competing with large VLMs, and ranking reversals under real degradation.
arXiv preprint arXiv:2111.08609 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
Document AI, or Document Intelligence, is a relatively new research topic that refers to the techniques for automatically reading, understanding, and analyzing business documents. It is an important research direction for natural language processing and computer vision. In recent years, the popularity of deep learning technology has greatly advanced the development of Document AI, such as document layout analysis, visual information extraction, document visual question answering, document image classification, etc. This paper briefly reviews some of the representative models, tasks, and benchmark datasets. Furthermore, we also introduce early-stage heuristic rule-based document analysis, statistical machine learning algorithms, and deep learning approaches especially pre-training methods. Finally, we look into future directions for Document AI research.
citation-role summary
citation-polarity summary
years
2026 3roles
background 2polarities
background 2representative citing papers
The paper systematises India's EV compliance-document lifecycle into a two-layer evidence model, a six-stage lifecycle with four failure loci, an exergy-destruction analytic lens, and a six-problem research agenda.
CC-OCR V2 reveals that state-of-the-art large multimodal models substantially underperform on challenging real-world document processing tasks.
citing papers explorer
-
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings
PureDocBench shows document parsing is far from solved, with top models at ~74/100, small specialists competing with large VLMs, and ranking reversals under real degradation.
-
The Documentation and Traceability Burden of the Indian EV Transition
The paper systematises India's EV compliance-document lifecycle into a two-layer evidence model, a six-stage lifecycle with four failure loci, an exergy-destruction analytic lens, and a six-problem research agenda.
-
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
CC-OCR V2 reveals that state-of-the-art large multimodal models substantially underperform on challenging real-world document processing tasks.