REVIEW 3 major objections 4 minor 40 references
A new 100-item benchmark of professional PDF tasks reports that no frontier multimodal model answers even a third of the items correctly, with the best passing 30.7%.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-08-02 06:56 UTC pith:EJ567ELZ
load-bearing objection A genuinely useful benchmark for professional PDF reasoning, with a robust headline finding even if the judge validation and the exact 30.7% threshold need closer scrutiny. the 3 major comments →
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is that frontier multimodal models, despite strong scores on standard visual QA suites, fail most expert-authored professional document tasks when success requires grounding the answer in the right evidence from the original PDF. The failures are systematic rather than scattered: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text. The benchmark's strict pass rate, which credits an item only when every atomic rubric criterion is satisfied, produces a leaderboard range of 2% to 30.7% across seventeen models. The paper also establishes a reusabl
What carries the argument
The central object is the GDP.pdf benchmark itself, whose construction makes the difficulty load-bearing. Each item must satisfy an adversarial screening rule: a candidate task is admitted only if at least two frontier models commit a major failure on it (a wrong answer, dropped decisive evidence, or a fabricated claim). Each item carries an expert-written rubric of atomic yes/no criteria; grading is done by an LLM judge that sees only the response, not the PDF or the gold answer, and the benchmark reports both a graded rubric score and a strict pass rate. This machinery is what lets the paper attribute low scores to grounded, evidence-level failure rather than to superficial stylistic misma
Load-bearing premise
The results depend on the LLM judge reliably grading each atomic rubric criterion without seeing the source PDF or the gold answer; if the judge is biased toward certain response styles or against valid paraphrases, the reported pass rates could be wrong.
What would settle it
Re-grade a random sample of model responses from the released benchmark with independent expert human raters, using the same rubrics, and compare the human strict pass rates to the LLM-judge pass rates; if agreement is low (for example, more than a few percentage points of difference), the reported leaderboard numbers would not survive.
If this is right
- If the benchmark is representative, leaderboard rankings on broad multimodal suites substantially overstate readiness for professional document workflows.
- A model that passes a GDP.pdf item has satisfied every requirement in the rubric, so strict pass rate is a floor-style reliability measure suitable for deployment decisions.
- Because the decisive evidence is often a footnote, legend, or superseding amendment, improving performance likely requires treating fine print as content, not noise, rather than merely extending context windows.
- The benchmark provides a reusable harness: new models can be scored against the released items via the public rubric, with the strict pass rate as the head-to-head metric.
- Error patterns concentrated in tables, charts, spatial reasoning, and abstention define concrete targets for model development over the next generation.
Where Pith is reading between the lines
- Because items were selected for being hard for frontier models, the 30.7% figure likely underestimates performance on typical, unselected professional PDF tasks; it is a measurement of the hard tail, not of average work.
- The strict pass rate may penalize partially correct answers that a human would find useful, so the benchmark's headline number is not the same as the share of tasks a model can assist with.
- One testable extension is to run the same items through tool-using or retrieval-augmented pipelines to see whether grounding failures can be repaired by giving models access to structured page extracts.
- The benchmark's abstention items suggest a concrete path: models with calibrated uncertainty could earn credit by declining to answer unsupported queries, which many current models fail.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GDP.pdf is a 100-item benchmark for grounded multimodal reasoning over professional PDFs. Items are expert-authored questions from ten professional domains, each paired with an original PDF, an atomic rubric, and capability tags; a candidate item is admitted only if at least two frontier models commit a major failure on it. The authors evaluate seventeen frontier models, reporting a graded rubric score and a strict task-level pass rate. Their headline result is that no model passes more than 30.7% of items, with the worst models at 2%, and the failure analysis attributes most errors to a small set of recurring patterns: table misalignment, chart misreads, dropped footnotes/exclusions, spatial miscounts, scan noise, and superseding amendments. The benchmark, rubrics, evaluation harness, and leaderboard are publicly released.
Significance. If the benchmark and evaluation protocol are sound, GDP.pdf makes a useful, distinctive contribution: unlike single-capability document-AI suites, it measures whether a model can complete realistic professional tasks on the original PDFs, with the decisive evidence often in a footnote, legend, exclusion, or amendment. The released dataset, atomic rubric design, capability taxonomy, and publicly available evaluation harness are concrete assets, and the qualitative failure analysis is valuable for model development. The paper is also appropriately transparent about the main caveat of its construction: because every item was adversarially screened, absolute pass rates describe performance on a deliberately hard subset, not on representative professional document work at large.
major comments (3)
- [Section 4, 'Judging and manual verification'] The load-bearing assumption is that the LLM judge (Gemini 3.5 Flash) grades atomic rubric criteria accurately, but the paper reports no agreement statistic, no validation sample size, and no confusion matrix. The text only says the judge was 'calibrated to ensure high agreement with expert human raters.' This is not a reporting nicety: the central claim 'no model passes a third of the items' rests on a margin of about 2.6 items for GPT-5.6 Sol (30.7% vs. 33.3%), and strict pass requires every criterion to pass, so per-criterion judge errors compound. The manual verification in Section 5.3 covers only the failures discussed qualitatively, not a random sample of passes, so a judge that incorrectly marks a wrong answer as passing would go undetected. Because the judge's own model family is also evaluated (Gemini 3.5 Flash appears in Table 6), a family-specific grading bias is not merely hyp
- [Section 5.1, 'Evaluated Models'] Pass rates are reported as point estimates averaged over five runs per model, but no error bars or confidence intervals are given. With 100 items and a strict binary pass indicator, the difference between the top two models (30.7% vs. 29.8%) is well within sampling noise, and even the gap between 30.7% and the 33.3% headline threshold is not statistically meaningful without variance estimates. The authors should report per-model variance or binomial confidence intervals, and ideally per-item stability across the five runs, before drawing conclusions about model rankings or the 'no model passes a third' claim.
- [Section 6, 'Discussion'] The paper correctly states that 'because every item defeated at least two frontier models at collection time, absolute pass rates measure performance on adversarially selected professional tasks, not on professional document work at large.' However, this caveat appears only in the Discussion, while the Abstract, Introduction, and Conclusion present the 30.7% figure without this qualification, and the closing claim that this 'does not support unsupervised use' is a generalization beyond the benchmark. The authors should either move this caveat into the Abstract and Conclusions, or add a non-adversarially sampled reference subset to ground the broader practical claim.
minor comments (4)
- [Abstract] Typographical spacing issue: 'GDP .pdf' should be 'GDP.pdf.'
- [Section 5.1, Table 6] Model configuration labels such as 'Adaptive Max' and 'xHigh reasoning' are not defined; a brief footnote describing what these configuration choices mean would improve comparability across providers.
- [Section 4] It is stated that the judge is given 'neither the source PDF nor the gold answer.' This is a sensible anti-anchoring choice, but it means the judge cannot verify factual grounding independently; it can only check surface consistency with rubric criteria. The authors should clarify how rubrics encode document-dependent facts so that a response repeating a false but plausible claim cannot pass.
- [Section 5.2] The sentence 'The hardest task slices were the spatial ones...' would be more informative if accompanied by quantitative slice scores in Table 6 or a separate table, rather than only qualitative summary.
Circularity Check
No significant circularity: the low pass-rate result is a measured evaluation on an explicitly caveated adversarial benchmark, not a prediction derived from its own construction.
full rationale
The paper's central claim is an empirical measurement, not a derivation. Items are screened so that at least two frontier models fail them (§3.2, §6), which makes absolute pass rates low partly by construction; however, the paper states this caveat explicitly in Section 6: "because every item defeated at least two frontier models at collection time, absolute pass rates measure performance on adversarially selected professional tasks, not on professional document work at large." The reported 30.7% and 2% figures are then obtained by running seventeen models on the released benchmark with an external evaluation harness, not by plugging the screening failures into an equation. The rubric grading uses an LLM judge from a family that is also evaluated, but the judge sees neither the source PDF nor the gold answer, and criterion-level gradeability is a test-protocol concern, not a circularity of the claimed result. No self-citations are load-bearing; all references are prior external benchmarks. The main limitations (unquantified judge agreement, adversarial selection) are acknowledged or disclosed in the paper, and the benchmark is publicly released for independent verification. Under the rule that only explicit equation-level or citation-level reductions count, no circular step is exhibited.
Axiom & Free-Parameter Ledger
free parameters (2)
- screening threshold =
at least 2 frontier model failures
- strict pass threshold =
all atomic rubric criteria satisfied
axioms (4)
- domain assumption Domain experts' questions and reference answers are accurate and representative of real professional workflows.
- domain assumption Atomic rubric criteria fully capture task correctness, including equivalent formulations and prohibited claims.
- domain assumption The LLM judge (Gemini 3.5 Flash) reliably evaluates rubric criteria without seeing the source PDF or gold answer.
- domain assumption An item's failure by at least two frontier models identifies a meaningful, non-superficial difficulty.
Cite this review
Pith. "Pith review of GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents." pith.science (2026). https://pith.science/paper/EJ567ELZ
@misc{pith2026260711192,
author = {Pith},
title = {Pith review of: GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents},
year = {2026},
howpublished = {\url{https://pith.science/paper/EJ567ELZ}},
note = {Machine review of arXiv:2607.11192}
}
read the original abstract
A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any one of them does not necessarily reveal whether a model can answer a realistic question that someone in the field would actually ask about a specific PDF. GDP_pdf is a benchmark built to measure this directly. It consists of question-document pairs authored by working professionals in ten fields, and a candidate question was kept only when at least two frontier multimodal models failed it in a way that mattered: a wrong answer, missed decisive evidence, or a fabricated claim, rather than a superficial difference such as style. Each item comes with a rubric of atomic criteria, so we can report a graded rubric score as well as a strict task-level pass rate, and each item is tagged against a taxonomy of eleven capabilities in three tiers, spanning text extraction and grounding, table and chart comprehension, cross-referencing, spatial reasoning, and abstention on unsupported queries. We report results for seventeen frontier models on the 100-item benchmark: the best model passes only 30.7% of the items and the worst passes 2%. Most errors trace back to a small set of recurring loss patterns: misaligned tables, misread charts, skipped footnotes and exclusions, miscounted floor-plan symbols, scan noise, and amendments that supersede earlier text.
Reference graph
Works this paper leans on
-
[1]
Ali Furkan Biten, Rub`en Tito, Andres Mafla, Lluis Gomez, Marc ¸al Rusi˜nol, Ernest Valveny, C. V . Jawahar, and Dimos- thenis Karatzas. Scene text visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019. 5
2019
-
[2]
MEGA-Bench: Scaling multimodal evaluation to over 500 real-world tasks
Jiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang, Yubo Wang, Yuansheng Ni, Wang Zhu, Ziyan Jiang, Bohan Lyu, Dongfu Jiang, Xuan He, Yuan Liu, Hexiang Hu, Xiang Yue, and Wenhu Chen. MEGA-Bench: Scaling multimodal evaluation to over 500 real-world tasks. InInternational Conference on Learning Representations,
-
[3]
HybridQA: A dataset of multi-hop question answering over tabular and textual data
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. HybridQA: A dataset of multi-hop question answering over tabular and textual data. InFindings of the Association for Computational Linguistics: EMNLP 2020, 2020. 2
2020
-
[4]
RoDLA: Benchmarking the robustness of document layout analysis models
Yufan Chen, Jiaming Zhang, Kunyu Peng, Junwei Zheng, Ruiping Liu, Philip Torr, and Rainer Stiefelhagen. RoDLA: Benchmarking the robustness of document layout analysis models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 2
2024
-
[5]
FinQA: A dataset of numerical reasoning over financial data
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. FinQA: A dataset of numerical reasoning over financial data. InProceedings of the Conference on Empirical Methods in Natural Language Processing, 2021. 2
2021
-
[6]
M-LongDoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, and Lidong Bing. M-LongDoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework. InProceedings of the Conference on Empirical Methods in Natural Language Processing, 2025. 2
2025
-
[7]
LongDocURL: a comprehensive multimodal long document benchmark integrating understanding, reason- ing, and locating
Chao Deng, Jiale Yuan, Pi Bu, Peijie Wang, Zhong-Zhi Li, Jian Xu, Xiao-Hui Li, Yuan Gao, Jun Song, Bo Zheng, and Cheng-Lin Liu. LongDocURL: a comprehensive multimodal long document benchmark integrating understanding, reason- ing, and locating. InProceedings of the Annual Meeting of the Association for Computational Linguistics, 2025. 2
2025
-
[8]
Benchmarking retrieval- augmented multimodal generation for document question answering
Kuicai Dong, Yujing Chang, Shijie Huang, Yasheng Wang, Ruiming Tang, and Yong Liu. Benchmarking retrieval- augmented multimodal generation for document question answering. InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track, 2025. 2
2025
-
[9]
OCRBench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reason- ing
Ling Fu, Zhebin Kuang, Jiajun Song, Mingxin Huang, Biao Yang, Yuzhe Li, Linghao Zhu, Qidi Luo, Xinyu Wang, Hao Lu, Zhang Li, Guozhi Tang, Bin Shan, Chunhui Lin, Qi Liu, Binghong Wu, Hao Feng, Hao Liu, Can Huang, Jingqun Tang, Wei Chen, Lianwen Jin, Yuliang Liu, and Xiang Bai. OCRBench v2: An improved benchmark for evaluating large multimodal models on vis...
2025
-
[10]
Ho, Christopher R ´e, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher R ´e, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zam- brano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H....
2023
-
[11]
FinanceBench: A new benchmark for financial question answering.arXiv preprint arXiv:2311.11944, 2023
Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen. FinanceBench: A new benchmark for financial question answering.arXiv preprint arXiv:2311.11944, 2023. 2
Pith/arXiv arXiv 2023
-
[12]
FUNSD: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. FUNSD: A dataset for form understanding in noisy scanned documents. InICDAR Workshop on Open Services and Tools for Document Analysis, 2019. 2
2019
-
[13]
Fenglin Liu, Zheng Li, Hongjian Zhou, Qingyu Yin, Jingfeng Yang, Xianfeng Tang, Chen Luo, Ming Zeng, Haoming Jiang, Yifan Gao, Priyanka Nigam, Sreyashi Nag, Bing Yin, Yin- ing Hua, Xuan Zhou, Omid Rohanian, Anshul Thakur, Lei Clifton, and David A. Clifton. Large language models are poor clinical decision-makers: A comprehensive benchmark. InProceedings of...
2024
-
[14]
OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 67(12):220102, 2024
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai. OCRBench: On the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 67(12):220102, 2024. 2
2024
-
[15]
MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts. In International Conference on Learning Representations, 2024. 2
2024
-
[16]
MMLongBench-Doc: Bench- marking long-context document understanding with visualiza- tions
Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, and Aixin Sun. MMLongBench-Doc: Bench- marking long-context document understanding with visualiza- tions. InAdvances in Neural Information Processing Systems Datasets and Benchmark...
2024
-
[17]
OK-VQA: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. 2
2019
-
[18]
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, 2022. 2
2022
-
[19]
ChartQAPro: A more diverse and challenging benchmark for chart question answering
Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj, Firoz Kabir, Aaryaman Kartha, Md Tah- mid Rahman Laskar, Mizanur Rahman, Shadikur Rah- man, Mehrad Shahmohammadi, Megh Thakkar, Md Rizwan Parvez, Enamul Hoque, and Shafiq Joty. ChartQAPro: A more diverse and challenging benchmark for chart question answering. InFindings of the Association for ...
2025
-
[20]
Minesh Mathew, Dimosthenis Karatzas, and C. V . Jawahar. DocVQA: A dataset for VQA on document images. InPro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021. 2, 5
2021
-
[21]
Minesh Mathew, Viraj Bagal, Rub `en Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V . Jawahar. Infograph- icVQA. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022. 2
2022
-
[22]
Khapra, and Pratyush Kumar
Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. PlotQA: Reasoning over scientific plots. InProceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision, 2020. 2
2020
-
[23]
OmniDocBench: Benchmarking di- verse PDF document parsing with comprehensive annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, and Conghui He. OmniDocBench: Benchmarking di- verse PDF document parsing with comprehensive annotations. InProceedings of the IEEE/CVF C...
2025
-
[24]
CORD: A consolidated receipt dataset for post-OCR parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jae- heung Surh, Minjoon Seo, and Hwalsuk Lee. CORD: A consolidated receipt dataset for post-OCR parsing. InWork- shop on Document Intelligence at NeurIPS 2019, 2019. 2
2019
-
[25]
David Peer, Philemon Sch¨opf, V olckmar Nebendahl, Alexan- der Rietzler, and Sebastian Stabinger. ANLS* – a universal document processing metric for generative large language models.arXiv preprint arXiv:2402.03848, 2024. 5
Pith/arXiv arXiv 2024
-
[26]
Nassar, and Peter W
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S. Nassar, and Peter W. J. Staar. DocLayNet: A large human- annotated dataset for document-layout analysis. InProceed- ings of the ACM SIGKDD Conference on Knowledge Discov- ery and Data Mining, 2022. 2
2022
-
[27]
Rossi, and Franck Dernoncourt
Jon Saad-Falcon, Joe Barrow, Alexa Siu, Ani Nenkova, Se- unghyun Yoon, Ryan A. Rossi, and Franck Dernoncourt. PDF- Triage: Question answering over long, structured documents. InProceedings of the Conference on Empirical Methods in Natural Language Processing: Industry Track, 2024. 2
2024
-
[28]
A-OKVQA: A benchmark for visual question answering using world knowl- edge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowl- edge. InEuropean Conference on Computer Vision, 2022. 2
2022
-
[29]
DocILE benchmark for document information localization and extraction
ˇStˇep´an ˇSimsa, Milan ˇSulc, Michal Uˇriˇc´aˇr, Yash Patel, Ahmed Hamdi, Mat ˇej Koci ´an, Maty ´aˇs Skalick ´y, Ji ˇr´ı Matas, An- toine Doucet, Micka¨el Coustaty, and Dimosthenis Karatzas. DocILE benchmark for document information localization and extraction. InInternational Conference on Document Analysis and Recognition, 2023. 2
2023
-
[30]
Hi- erarchical multimodal transformers for Multipage DocVQA
Rub`en Tito, Dimosthenis Karatzas, and Ernest Valveny. Hi- erarchical multimodal transformers for Multipage DocVQA. Pattern Recognition, 144:109834, 2023. 2
2023
-
[31]
Document understanding dataset and evaluation (DUDE)
Jordy Van Landeghem, Rub `en Tito, Łukasz Borchmann, Michał Pietruszka, Paweł J ´oziak, Rafał Powalski, Dawid Jurkiewicz, Micka ¨el Coustaty, Bertrand Anckaert, Ernest Valveny, Matthew Blaschko, Sien Moens, and Tomasz Sta- nisławek. Document understanding dataset and evaluation (DUDE). InProceedings of the IEEE/CVF International Conference on Computer Vis...
2023
-
[32]
CharXiv: Charting gaps in realistic chart understand- ing in multimodal LLMs
Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sad- hika Malladi, Alexis Chevalier, Sanjeev Arora, and Danqi Chen. CharXiv: Charting gaps in realistic chart understand- ing in multimodal LLMs. InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 2
2024
-
[33]
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for exp...
2024
-
[34]
MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, and Graham Neu- big. MMMU-Pro: A more robust multi-discipline multimodal understanding benchmark. InProceedings of the Annual Meeting of the Association for Computational Linguistics,
-
[35]
Pub- LayNet: Largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. Pub- LayNet: Largest dataset ever for document layout analysis. In International Conference on Document Analysis and Recog- nition, 2019. 2
2019
-
[36]
Changda Zhou, Ziyue Gao, Xueqing Wang, Tingquan Gao, Cheng Cui, Jing Tang, and Yi Liu. Real5-OmniDocBench: A full-scale physical reconstruction benchmark for robust docu- ment parsing in the wild.arXiv preprint arXiv:2603.04205,
-
[37]
TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance
Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. InProceedings of the Annual Meeting of the Association for Computational Linguistics, 2021. 2
2021
-
[38]
Towards complex document understanding by discrete reasoning
Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. Towards complex document understanding by discrete reasoning. InProceedings of the ACM International Conference on Multimedia, 2022. 2
2022
-
[39]
MMDocBench: Benchmarking large vision-language models for fine-grained visual document understanding and grounding
Fengbin Zhu, Ziyang Liu, Xiang Yao Ng, Haohui Wu, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, and Tat-Seng Chua. MMDocBench: Benchmarking large vision-language models for fine-grained visual document understanding and grounding. InProceedings of the International Conference on Multimedia Modeling, 2026. 2
2026
-
[40]
Anni Zou, Wenhao Yu, Hongming Zhang, Kaixin Ma, Deng Cai, Zhuosheng Zhang, Hai Zhao, and Dong Yu. DOCBENCH: A benchmark for evaluating LLM-based doc- ument reading systems.arXiv preprint arXiv:2407.10701,
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.