Pith. sign in

REVIEW 3 major objections 4 minor 40 references

A new 100-item benchmark of professional PDF tasks reports that no frontier multimodal model answers even a third of the items correctly, with the best passing 30.7%.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-02 06:56 UTC pith:EJ567ELZ

load-bearing objection A genuinely useful benchmark for professional PDF reasoning, with a robust headline finding even if the judge validation and the exact 30.7% threshold need closer scrutiny. the 3 major comments →

arxiv 2607.11192 v3 pith:EJ567ELZ submitted 2026-07-13 cs.CV

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

classification cs.CV
keywords benchmarkPDF documentsmultimodal reasoningdocument understandinggroundingprofessional workflowsevaluationvisual question answering
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that standard document-AI benchmarks measure reading skills in isolation and therefore overstate how well models handle real professional documents. It introduces GDP.pdf, a benchmark of 100 question–document pairs authored by working professionals in ten fields, where every item was kept only after at least two frontier models failed it in a substantive way. Using atomic rubrics graded by an LLM judge, the benchmark reports that the best of seventeen frontier models passes only 30.7% of items and the worst 2%. The paper's central claim is that current models are not reliable for grounded professional document work such as reading exclusions, matching floor-plan symbols, or detecting superseded amendments.

Core claim

On the paper's own terms, the central discovery is that frontier multimodal models, despite strong scores on standard visual QA suites, fail most expert-authored professional document tasks when success requires grounding the answer in the right evidence from the original PDF. The failures are systematic rather than scattered: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text. The benchmark's strict pass rate, which credits an item only when every atomic rubric criterion is satisfied, produces a leaderboard range of 2% to 30.7% across seventeen models. The paper also establishes a reusabl

What carries the argument

The central object is the GDP.pdf benchmark itself, whose construction makes the difficulty load-bearing. Each item must satisfy an adversarial screening rule: a candidate task is admitted only if at least two frontier models commit a major failure on it (a wrong answer, dropped decisive evidence, or a fabricated claim). Each item carries an expert-written rubric of atomic yes/no criteria; grading is done by an LLM judge that sees only the response, not the PDF or the gold answer, and the benchmark reports both a graded rubric score and a strict pass rate. This machinery is what lets the paper attribute low scores to grounded, evidence-level failure rather than to superficial stylistic misma

Load-bearing premise

The results depend on the LLM judge reliably grading each atomic rubric criterion without seeing the source PDF or the gold answer; if the judge is biased toward certain response styles or against valid paraphrases, the reported pass rates could be wrong.

What would settle it

Re-grade a random sample of model responses from the released benchmark with independent expert human raters, using the same rubrics, and compare the human strict pass rates to the LLM-judge pass rates; if agreement is low (for example, more than a few percentage points of difference), the reported leaderboard numbers would not survive.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the benchmark is representative, leaderboard rankings on broad multimodal suites substantially overstate readiness for professional document workflows.
  • A model that passes a GDP.pdf item has satisfied every requirement in the rubric, so strict pass rate is a floor-style reliability measure suitable for deployment decisions.
  • Because the decisive evidence is often a footnote, legend, or superseding amendment, improving performance likely requires treating fine print as content, not noise, rather than merely extending context windows.
  • The benchmark provides a reusable harness: new models can be scored against the released items via the public rubric, with the strict pass rate as the head-to-head metric.
  • Error patterns concentrated in tables, charts, spatial reasoning, and abstention define concrete targets for model development over the next generation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because items were selected for being hard for frontier models, the 30.7% figure likely underestimates performance on typical, unselected professional PDF tasks; it is a measurement of the hard tail, not of average work.
  • The strict pass rate may penalize partially correct answers that a human would find useful, so the benchmark's headline number is not the same as the share of tasks a model can assist with.
  • One testable extension is to run the same items through tool-using or retrieval-augmented pipelines to see whether grounding failures can be repaired by giving models access to structured page extracts.
  • The benchmark's abstention items suggest a concrete path: models with calibrated uncertainty could earn credit by declining to answer unsupported queries, which many current models fail.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. GDP.pdf is a 100-item benchmark for grounded multimodal reasoning over professional PDFs. Items are expert-authored questions from ten professional domains, each paired with an original PDF, an atomic rubric, and capability tags; a candidate item is admitted only if at least two frontier models commit a major failure on it. The authors evaluate seventeen frontier models, reporting a graded rubric score and a strict task-level pass rate. Their headline result is that no model passes more than 30.7% of items, with the worst models at 2%, and the failure analysis attributes most errors to a small set of recurring patterns: table misalignment, chart misreads, dropped footnotes/exclusions, spatial miscounts, scan noise, and superseding amendments. The benchmark, rubrics, evaluation harness, and leaderboard are publicly released.

Significance. If the benchmark and evaluation protocol are sound, GDP.pdf makes a useful, distinctive contribution: unlike single-capability document-AI suites, it measures whether a model can complete realistic professional tasks on the original PDFs, with the decisive evidence often in a footnote, legend, exclusion, or amendment. The released dataset, atomic rubric design, capability taxonomy, and publicly available evaluation harness are concrete assets, and the qualitative failure analysis is valuable for model development. The paper is also appropriately transparent about the main caveat of its construction: because every item was adversarially screened, absolute pass rates describe performance on a deliberately hard subset, not on representative professional document work at large.

major comments (3)
  1. [Section 4, 'Judging and manual verification'] The load-bearing assumption is that the LLM judge (Gemini 3.5 Flash) grades atomic rubric criteria accurately, but the paper reports no agreement statistic, no validation sample size, and no confusion matrix. The text only says the judge was 'calibrated to ensure high agreement with expert human raters.' This is not a reporting nicety: the central claim 'no model passes a third of the items' rests on a margin of about 2.6 items for GPT-5.6 Sol (30.7% vs. 33.3%), and strict pass requires every criterion to pass, so per-criterion judge errors compound. The manual verification in Section 5.3 covers only the failures discussed qualitatively, not a random sample of passes, so a judge that incorrectly marks a wrong answer as passing would go undetected. Because the judge's own model family is also evaluated (Gemini 3.5 Flash appears in Table 6), a family-specific grading bias is not merely hyp
  2. [Section 5.1, 'Evaluated Models'] Pass rates are reported as point estimates averaged over five runs per model, but no error bars or confidence intervals are given. With 100 items and a strict binary pass indicator, the difference between the top two models (30.7% vs. 29.8%) is well within sampling noise, and even the gap between 30.7% and the 33.3% headline threshold is not statistically meaningful without variance estimates. The authors should report per-model variance or binomial confidence intervals, and ideally per-item stability across the five runs, before drawing conclusions about model rankings or the 'no model passes a third' claim.
  3. [Section 6, 'Discussion'] The paper correctly states that 'because every item defeated at least two frontier models at collection time, absolute pass rates measure performance on adversarially selected professional tasks, not on professional document work at large.' However, this caveat appears only in the Discussion, while the Abstract, Introduction, and Conclusion present the 30.7% figure without this qualification, and the closing claim that this 'does not support unsupervised use' is a generalization beyond the benchmark. The authors should either move this caveat into the Abstract and Conclusions, or add a non-adversarially sampled reference subset to ground the broader practical claim.
minor comments (4)
  1. [Abstract] Typographical spacing issue: 'GDP .pdf' should be 'GDP.pdf.'
  2. [Section 5.1, Table 6] Model configuration labels such as 'Adaptive Max' and 'xHigh reasoning' are not defined; a brief footnote describing what these configuration choices mean would improve comparability across providers.
  3. [Section 4] It is stated that the judge is given 'neither the source PDF nor the gold answer.' This is a sensible anti-anchoring choice, but it means the judge cannot verify factual grounding independently; it can only check surface consistency with rubric criteria. The authors should clarify how rubrics encode document-dependent facts so that a response repeating a false but plausible claim cannot pass.
  4. [Section 5.2] The sentence 'The hardest task slices were the spatial ones...' would be more informative if accompanied by quantitative slice scores in Table 6 or a separate table, rather than only qualitative summary.

Circularity Check

0 steps flagged

No significant circularity: the low pass-rate result is a measured evaluation on an explicitly caveated adversarial benchmark, not a prediction derived from its own construction.

full rationale

The paper's central claim is an empirical measurement, not a derivation. Items are screened so that at least two frontier models fail them (§3.2, §6), which makes absolute pass rates low partly by construction; however, the paper states this caveat explicitly in Section 6: "because every item defeated at least two frontier models at collection time, absolute pass rates measure performance on adversarially selected professional tasks, not on professional document work at large." The reported 30.7% and 2% figures are then obtained by running seventeen models on the released benchmark with an external evaluation harness, not by plugging the screening failures into an equation. The rubric grading uses an LLM judge from a family that is also evaluated, but the judge sees neither the source PDF nor the gold answer, and criterion-level gradeability is a test-protocol concern, not a circularity of the claimed result. No self-citations are load-bearing; all references are prior external benchmarks. The main limitations (unquantified judge agreement, adversarial selection) are acknowledged or disclosed in the paper, and the benchmark is publicly released for independent verification. Under the rule that only explicit equation-level or citation-level reductions count, no circular step is exhibited.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The benchmark's numeric outputs depend on design decisions (screening threshold, rubric strictness, judge model) rather than physical axioms. No free parameters are fitted to data in the theory sense; the listed items are evaluation-design controls that directly influence the reported numbers.

free parameters (2)
  • screening threshold = at least 2 frontier model failures
    Item admission depends on this threshold; changing it would directly change the difficulty and pass rates. It is a design choice rather than a fitted constant, but it controls the headline 30.7% figure.
  • strict pass threshold = all atomic rubric criteria satisfied
    The strict pass rate requires every rubric element; a more lenient threshold would produce higher pass rates, so this design decision shapes the reported rankings.
axioms (4)
  • domain assumption Domain experts' questions and reference answers are accurate and representative of real professional workflows.
    Section 3.5, 'Collection and Curation Workflow': tasks are submitted by contributors from their own work; if their judgment is unrepresentative, the benchmark does not measure professional work.
  • domain assumption Atomic rubric criteria fully capture task correctness, including equivalent formulations and prohibited claims.
    Section 4, 'Atomic rubric design': the paper assumes rubrics are complete and unambiguous; erroneous rubrics would make pass rates unreliable.
  • domain assumption The LLM judge (Gemini 3.5 Flash) reliably evaluates rubric criteria without seeing the source PDF or gold answer.
    Section 4, 'Judging and manual verification': the paper states the judge is calibrated for high agreement with human raters but does not provide inter-rater statistics.
  • domain assumption An item's failure by at least two frontier models identifies a meaningful, non-superficial difficulty.
    Section 3.2, 'Adversarial construction': this screening rule assumes model failures correspond to substantive errors, not stylistic differences.

reviewed 2026-08-02 · how reviews work

0 comments
Cite this review

Pith. "Pith review of GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents." pith.science (2026). https://pith.science/paper/EJ567ELZ

@misc{pith2026260711192,
  author       = {Pith},
  title        = {Pith review of: GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EJ567ELZ}},
  note         = {Machine review of arXiv:2607.11192}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any one of them does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF. GDP_pdf is a benchmark built to measure this directly. It consists of question-document pairs authored by working professionals in ten fields, and a candidate question was kept only when at least two frontier multimodal models failed it in a way that mattered: a wrong answer, missed decisive evidence, or a fabricated claim, rather than a superficial difference such as style. Each item comes with a rubric of atomic criteria, so we can report a graded rubric score as well as a strict task-level pass rate, and each item is tagged against a taxonomy of eleven capabilities in three tiers, spanning text extraction and grounding, table and chart comprehension, cross-referencing, spatial reasoning, and abstention on unsupported queries. We report results for seventeen frontier models on the 100-item benchmark: the best model passes only 30.7% of the items and the worst passes 2%. Most errors trace back to a small set of recurring loss patterns: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 4 linked inside Pith

  1. [1]

    Ali Furkan Biten, Rub`en Tito, Andres Mafla, Lluis Gomez, Marc ¸al Rusi˜nol, Ernest Valveny, C. V . Jawahar, and Dimos- thenis Karatzas. Scene text visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019. 5

  2. [2]

    MEGA-Bench: Scaling multimodal evaluation to over 500 real-world tasks

    Jiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang, Yubo Wang, Yuansheng Ni, Wang Zhu, Ziyan Jiang, Bohan Lyu, Dongfu Jiang, Xuan He, Yuan Liu, Hexiang Hu, Xiang Yue, and Wenhu Chen. MEGA-Bench: Scaling multimodal evaluation to over 500 real-world tasks. InInternational Conference on Learning Representations,

  3. [3]

    HybridQA: A dataset of multi-hop question answering over tabular and textual data

    Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. HybridQA: A dataset of multi-hop question answering over tabular and textual data. InFindings of the Association for Computational Linguistics: EMNLP 2020, 2020. 2

  4. [4]

    RoDLA: Benchmarking the robustness of document layout analysis models

    Yufan Chen, Jiaming Zhang, Kunyu Peng, Junwei Zheng, Ruiping Liu, Philip Torr, and Rainer Stiefelhagen. RoDLA: Benchmarking the robustness of document layout analysis models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 2

  5. [5]

    FinQA: A dataset of numerical reasoning over financial data

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. FinQA: A dataset of numerical reasoning over financial data. InProceedings of the Conference on Empirical Methods in Natural Language Processing, 2021. 2

  6. [6]

    M-LongDoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework

    Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, and Lidong Bing. M-LongDoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework. InProceedings of the Conference on Empirical Methods in Natural Language Processing, 2025. 2

  7. [7]

    LongDocURL: a comprehensive multimodal long document benchmark integrating understanding, reason- ing, and locating

    Chao Deng, Jiale Yuan, Pi Bu, Peijie Wang, Zhong-Zhi Li, Jian Xu, Xiao-Hui Li, Yuan Gao, Jun Song, Bo Zheng, and Cheng-Lin Liu. LongDocURL: a comprehensive multimodal long document benchmark integrating understanding, reason- ing, and locating. InProceedings of the Annual Meeting of the Association for Computational Linguistics, 2025. 2

  8. [8]

    Benchmarking retrieval- augmented multimodal generation for document question answering

    Kuicai Dong, Yujing Chang, Shijie Huang, Yasheng Wang, Ruiming Tang, and Yong Liu. Benchmarking retrieval- augmented multimodal generation for document question answering. InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track, 2025. 2

  9. [9]

    OCRBench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reason- ing

    Ling Fu, Zhebin Kuang, Jiajun Song, Mingxin Huang, Biao Yang, Yuzhe Li, Linghao Zhu, Qidi Luo, Xinyu Wang, Hao Lu, Zhang Li, Guozhi Tang, Bin Shan, Chunhui Lin, Qi Liu, Binghong Wu, Hao Feng, Hao Liu, Can Huang, Jingqun Tang, Wei Chen, Lianwen Jin, Yuliang Liu, and Xiang Bai. OCRBench v2: An improved benchmark for evaluating large multimodal models on vis...

  10. [10]

    Ho, Christopher R ´e, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N

    Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher R ´e, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zam- brano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H....

  11. [11]

    FinanceBench: A new benchmark for financial question answering.arXiv preprint arXiv:2311.11944, 2023

    Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen. FinanceBench: A new benchmark for financial question answering.arXiv preprint arXiv:2311.11944, 2023. 2

  12. [12]

    FUNSD: A dataset for form understanding in noisy scanned documents

    Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. FUNSD: A dataset for form understanding in noisy scanned documents. InICDAR Workshop on Open Services and Tools for Document Analysis, 2019. 2

  13. [13]

    Fenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng, Haoming Jiang, Yifan Gao, Priyanka Nigam, Sreyashi Nag, Bing Yin, Yin- ing Hua, Xuan Zhou, Omid Rohanian, Anshul Thakur, Lei Clifton, and David A. Clifton. Large language models are poor clinical decision-makers: A comprehensive benchmark. InProceedings of...

  14. [14]

    OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai. OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 67(12):220102, 2024. 2

  15. [15]

    MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts. In International Conference on Learning Representations, 2024. 2

  16. [16]

    MMLongBench-Doc: Bench- marking long-context document understanding with visualiza- tions

    Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, and Aixin Sun. MMLongBench-Doc: Bench- marking long-context document understanding with visualiza- tions. InAdvances in Neural Information Processing Systems Datasets and Benchmark...

  17. [17]

    OK-VQA: A visual question answering benchmark requiring external knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. 2

  18. [18]

    ChartQA: A benchmark for question answering about charts with visual and logical reasoning

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, 2022. 2

  19. [19]

    ChartQAPro: A more diverse and challenging benchmark for chart question answering

    Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj, Firoz Kabir, Aaryaman Kartha, Md Tah- mid Rahman Laskar, Mizanur Rahman, Shadikur Rah- man, Mehrad Shahmohammadi, Megh Thakkar, Md Rizwan Parvez, Enamul Hoque, and Shafiq Joty. ChartQAPro: A more diverse and challenging benchmark for chart question answering. InFindings of the Association for ...

  20. [20]

    Minesh Mathew, Dimosthenis Karatzas, and C. V . Jawahar. DocVQA: A dataset for VQA on document images. InPro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021. 2, 5

  21. [21]

    Minesh Mathew, Viraj Bagal, Rub `en Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V . Jawahar. Infograph- icVQA. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022. 2

  22. [22]

    Khapra, and Pratyush Kumar

    Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. PlotQA: Reasoning over scientific plots. InProceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision, 2020. 2

  23. [23]

    OmniDocBench: Benchmarking di- verse PDF document parsing with comprehensive annotations

    Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, and Conghui He. OmniDocBench: Benchmarking di- verse PDF document parsing with comprehensive annotations. InProceedings of the IEEE/CVF C...

  24. [24]

    CORD: A consolidated receipt dataset for post-OCR parsing

    Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jae- heung Surh, Minjoon Seo, and Hwalsuk Lee. CORD: A consolidated receipt dataset for post-OCR parsing. InWork- shop on Document Intelligence at NeurIPS 2019, 2019. 2

  25. [25]

    ANLS* – a universal document processing metric for generative large language models.arXiv preprint arXiv:2402.03848, 2024

    David Peer, Philemon Sch¨opf, V olckmar Nebendahl, Alexan- der Rietzler, and Sebastian Stabinger. ANLS* – a universal document processing metric for generative large language models.arXiv preprint arXiv:2402.03848, 2024. 5

  26. [26]

    Nassar, and Peter W

    Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S. Nassar, and Peter W. J. Staar. DocLayNet: A large human- annotated dataset for document-layout analysis. InProceed- ings of the ACM SIGKDD Conference on Knowledge Discov- ery and Data Mining, 2022. 2

  27. [27]

    Rossi, and Franck Dernoncourt

    Jon Saad-Falcon, Joe Barrow, Alexa Siu, Ani Nenkova, Se- unghyun Yoon, Ryan A. Rossi, and Franck Dernoncourt. PDF- Triage: Question answering over long, structured documents. InProceedings of the Conference on Empirical Methods in Natural Language Processing: Industry Track, 2024. 2

  28. [28]

    A-OKVQA: A benchmark for visual question answering using world knowl- edge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowl- edge. InEuropean Conference on Computer Vision, 2022. 2

  29. [29]

    DocILE benchmark for document information localization and extraction

    ˇStˇep´an ˇSimsa, Milan ˇSulc, Michal Uˇriˇc´aˇr, Yash Patel, Ahmed Hamdi, Mat ˇej Koci ´an, Maty ´aˇs Skalick ´y, Ji ˇr´ı Matas, An- toine Doucet, Micka¨el Coustaty, and Dimosthenis Karatzas. DocILE benchmark for document information localization and extraction. InInternational Conference on Document Analysis and Recognition, 2023. 2

  30. [30]

    Hi- erarchical multimodal transformers for Multipage DocVQA

    Rub`en Tito, Dimosthenis Karatzas, and Ernest Valveny. Hi- erarchical multimodal transformers for Multipage DocVQA. Pattern Recognition, 144:109834, 2023. 2

  31. [31]

    Document understanding dataset and evaluation (DUDE)

    Jordy Van Landeghem, Rub `en Tito, Łukasz Borchmann, Michał Pietruszka, Paweł J ´oziak, Rafał Powalski, Dawid Jurkiewicz, Micka ¨el Coustaty, Bertrand Anckaert, Ernest Valveny, Matthew Blaschko, Sien Moens, and Tomasz Sta- nisławek. Document understanding dataset and evaluation (DUDE). InProceedings of the IEEE/CVF International Conference on Computer Vis...

  32. [32]

    CharXiv: Charting gaps in realistic chart understand- ing in multimodal LLMs

    Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sad- hika Malladi, Alexis Chevalier, Sanjeev Arora, and Danqi Chen. CharXiv: Charting gaps in realistic chart understand- ing in multimodal LLMs. InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 2

  33. [33]

    MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for exp...

  34. [34]

    MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark

    Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, and Graham Neu- big. MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark. InProceedings of the Annual Meeting of the Association for Computational Linguistics,

  35. [35]

    Pub- LayNet: Largest dataset ever for document layout analysis

    Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. Pub- LayNet: Largest dataset ever for document layout analysis. In International Conference on Document Analysis and Recog- nition, 2019. 2

  36. [36]

    Real5-OmniDocBench: A full-scale physical reconstruction benchmark for robust docu- ment parsing in the wild.arXiv preprint arXiv:2603.04205,

    Changda Zhou, Ziyue Gao, Xueqing Wang, Tingquan Gao, Cheng Cui, Jing Tang, and Yi Liu. Real5-OmniDocBench: A full-scale physical reconstruction benchmark for robust docu- ment parsing in the wild.arXiv preprint arXiv:2603.04205,

  37. [37]

    TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance

    Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. InProceedings of the Annual Meeting of the Association for Computational Linguistics, 2021. 2

  38. [38]

    Towards complex document understanding by discrete reasoning

    Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. Towards complex document understanding by discrete reasoning. InProceedings of the ACM International Conference on Multimedia, 2022. 2

  39. [39]

    MMDocBench: Benchmarking large vision-language models for fine-grained visual document understanding and grounding

    Fengbin Zhu, Ziyang Liu, Xiang Yao Ng, Haohui Wu, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, and Tat-Seng Chua. MMDocBench: Benchmarking large vision-language models for fine-grained visual document understanding and grounding. InProceedings of the International Conference on Multimedia Modeling, 2026. 2

  40. [40]

    DOCBENCH: A benchmark for evaluating LLM-based doc- ument reading systems.arXiv preprint arXiv:2407.10701,

    Anni Zou, Wenhao Yu, Hongming Zhang, Kaixin Ma, Deng Cai, Zhuosheng Zhang, Hai Zhao, and Dong Yu. DOCBENCH: A benchmark for evaluating LLM-based doc- ument reading systems.arXiv preprint arXiv:2407.10701,

This paper was first reviewed by deepseek-v4-flash on August 2, 2026.