Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey of visually rich document question answering concludes that explicit 2D layout encoding, not image resolution, separates the top-scoring models.

desk verdict A genuinely useful survey map of VRD question answering, but the Section 6 ranking claim contradicts its own Table 3 and should be softened. read the letter →

arxiv 2501.02235 v2 pith:BNUODZGZ submitted 2025-01-04 cs.CL

classification cs.CL
keywords visuallyrichdocumentunderstandingvisualquestionansweringlayout-awarelanguagemodelslargevision-languagemulti-pagepositionalencodingANLSevaluationretrieval-augmentedgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey maps the design space of question answering over visually rich documents—scanned or born-digital pages whose meaning lives in text, tables, layout, and figures—and draws a comparative conclusion from the benchmark numbers it collects. The authors argue that models that explicitly encode 2D layout, such as ERNIE-Layout and Arctic-TILT, achieve the strongest results, and that this shows text and layout information are essential for document question answering, even when the question concerns a chart or figure. They also contend that vision-only large vision-language models, which treat the page as a single image, are ill-suited to multi-page documents unless paired with a retriever or heavy page compression, and they recommend multimodal fusion guided by cross-attention as the way forward. The survey matters because it gives practitioners a structured comparison of encoding strategies, fusion mechanisms, and multi-page techniques, while explicitly cautioning that its cross-paper score comparisons are not controlled experiments.

What carries the argument

The object that carries the argument is the document representation itself, decomposed into three modalities: text tokens, bounding boxes (layout), and the page image. The survey's comparison is organised around how models fuse these modalities: absolute 2D positional embeddings, relative 2D attention biases, disentangled attention that separates a token's semantic meaning from its horizontal and vertical distance to other tokens, cross-attention between visual and textual tokens, and page-level compression tokens. The evaluation machinery is ANLS (Average Normalized Levenshtein Similarity), the metric used in Table 3 to rank single- and multi-page VQA systems. The taxonomy's pivot is the contrast between structured encoders that consume layout explicitly and vision-only LVLMs that see the whole page as an image; the authors argue that this distinction, not resolution, separates the top performers.

What would settle it

Run a matched experiment: take one backbone, train a vision-only variant that sees page images at high resolution and a layout-aware variant that consumes text tokens with 2D positional biases, on the same VQA datasets, and compare ANLS on DocVQA, DUDE, and MMLongBench-Doc; if the vision-only variant matches or beats the layout-aware one on multi-page questions, the survey's central claim that text and layout are essential would be falsified.

Watch

Extended reading notes

Core claim

The central claim is that, in current document VQA, how a model represents the spatial layout of a page matters more than how many pixels it can see. The authors read Table 3 as showing that models making extensive use of positional features—ERNIE-Layout's disentangled attention over sequential, horizontal, and vertical relative distances, and Arctic-TILT's blockwise attention with a role bias for text tokens—have the best results, and they infer from this that text and layout information are essential for answering questions, even in complex charts and figures. They argue that structured multimodal approaches combining text, layout, and vision are more efficient for multi-page understanding than vision-only LVLMs, which either compress each page so heavily that performance degrades or must depend on a retriever to select relevant pages. They therefore recommend that the community prioritise layout handling and explicit 2D position encoding, and that visual features be injected through cross-attention with text tokens as queries rather than through self-attention over concatenated visual and textual tokens.

Load-bearing premise

The load-bearing premise is that the ANLS scores in Table 3 can be compared across models even though each model was trained and evaluated by a different team under different protocols; the authors themselves flag this in Section 7, noting that it is challenging to draw definitive conclusions.

Editorial extensions

If this is right

  • Architecture choices should favour explicit 2D position encoding, whether absolute embeddings, relative biases, or 2D rotary positions, over treating document pages as generic images.
  • For multi-page QA, sparse-attention designs such as global-local or blockwise attention are the paper's recommended direction, ahead of page-by-page compression and retrieval-dependent pipelines.
  • Pretraining on document parsing tasks that turn page screenshots into structured text (HTML, Markdown, CSV/JSON) should be a standard step, since it aligns text, layout, and vision and makes visual features partially redundant.
  • Cross-attention visual injection with text tokens as queries is preferred over self-attention over concatenated token lists, because it keeps visual features separate while letting the LLM interrogate them.
  • Because the paper's score table is not a controlled comparison, the rankings should be re-validated with uniform training protocols before architectural conclusions are treated as settled.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the correlation between positional-feature use and top scores is causal, a matched ablation should show a larger layout advantage on table-heavy and multi-page benchmarks than on plain-text documents; the survey's cross-paper table cannot demonstrate this directly.
  • The paper's evidence implies that for text-dense pages the visual modality is largely redundant; a concrete extension would route text-heavy pages through layout-aware text encoders and reserve vision encoders for figures, cutting compute with little accuracy loss.
  • The recommended text-guided cross-attention fusion could be made adaptive by switching the query modality (text vs. vision) based on detected document type, generalising the paper's binary suggestion into a testable mechanism.
  • Because the survey restricts itself to transformers, its conclusion that layout encoding is essential is scoped; graph-based layout models would be the natural comparison to see whether the finding survives outside attention architectures.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript is a survey of question answering over visually rich documents. It organizes recent work into three encoding families: structured multi-modal encoders that combine text, layout bounding boxes, and visual features (Section 2); vision-only large vision-language models that treat pages as images (Section 3); and multi-page strategies based on retrieval, per-page compression/query tokens, or sparse/recurrent attention (Section 4). Section 5 compares self-attention and cross-attention mechanisms for injecting visual features into an LLM decoder. The survey includes comparative tables of models and datasets (Tables 1-4) and concludes in Section 6 that models making extensive use of positional features, such as ERNIE-Layout and Arctic-TILT, achieve the best results and that text and layout are essential, while Section 7 acknowledges that cross-paper evaluation conditions differ.

Significance. The paper is a useful structured map of a fast-moving area, with broad coverage of encoder designs, multi-page strategies, and dataset characteristics. Its principal value is taxonomic: it collects a large set of recent models into a clear two-step pipeline and provides an appendix of VQA datasets. To the authors' credit, the limitations of the cross-paper comparison are explicitly acknowledged in Section 7. However, the central comparative conclusion in Section 6 is not established by the evidence in Table 3, and as written the conclusion is internally inconsistent with the caveat in Section 7. The significance of the paper as a guide for architecture choice therefore depends on a revision that either presents the conclusion as a hypothesis or supplies a controlled comparison.

major comments (3)
  1. [Section 6, Table 3] Section 6 states that 'models that make extensive use of positional features—such as ERNIE-Layout and Arctic-TILT—have the best results,' but the DocVQA column of Table 3 shows GPT-4o at 92.8 and InternLMXComposer2-4KHD at 90.0, both above ERNIE-Layout's 88.4 and with InternLMXComposer2-4KHD essentially tied with Arctic-TILT's 90.2. GPT-4o is a vision-only commercial LVLM in the survey's own taxonomy, and InternLMXComposer2-4KHD is also a vision-only model; the claimed ranking is therefore contradicted by the table's own numbers.
  2. [Section 6 vs Section 7] The causal conclusion in Section 6 ('This indicates that text and layout information are essential') directly conflicts with the limitation stated in Section 7: because methods are 'evaluated in their original experimental setups, which differ in terms of model architecture, training protocols, and datasets,' the authors themselves say it is 'challenging to draw definitive conclusions.' Observational cross-paper ANLS values cannot establish that layout information is essential; that would require a controlled ablation in which the same base model, training data, and protocol are evaluated with and without layout features.
  3. [Table 3] Table 3 mixes incomparable conditions: rows marked * use retrievers (e.g., InternLMXComposer2-4KHD with PDF-Wukong, Pix2Struct with Naidu et al., QwenVL and Idefics2 with M3DocRAG), rows marked ² concatenate page representations rather than performing true multi-page reasoning, and several cells are empty. Rankings based on such heterogeneous scores are not robust. The top-3 bold marking should be disclosed per column and restricted to comparable settings, or the table should be relabeled as a compilation of reported scores without ranking claims.
minor comments (5)
  1. [References] References Huang et al. 2024a and 2024b are identical ('From detection to application...'), and Xu et al. 2024a and 2024b are identical (LLaVA-UHD); please merge or disambiguate them.
  2. [Table 3] Table 3 model names are inconsistent: 'mPLUGDoc', 'mPLUGDoc1.5', and 'ILMXC24KHD' should be spelled as in the main text (mPLUG-DocOwl, mPLUG-DocOwl1.5, InternLMXComposer2-4KHD).
  3. [Table 4] Table 4 header says '#Pages' per document but BoundingDocs reports 237k, which appears to be a total rather than a per-document average; please clarify the units in that column.
  4. [Section 2.1] The notation \hat{V} = V ∪ [BBOX] should use a set of special tokens, e.g., V ∪ {[BBOX]}, and clarify how the marker is tokenized and inserted into the sequence.
  5. [Table 3] The bold top-3 markers are not visible in the text and the basis for choosing top-3 across a heterogeneous score matrix should be stated explicitly in the caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's claims are interpretive summaries of external cited results, not derivations from its own scaffolding.

full rationale

This is a survey paper, so its claims are literature-synthesis statements rather than results derived from equations or fitted parameters. The main load-bearing claim in Section 6, that layout-aware models such as ERNIE-Layout and Arctic-TILT perform best, is presented as a reading of Table 3, which reports ANLS scores taken from external papers. The survey does not fit any parameter, define a quantity in terms of its own conclusion, or invoke a uniqueness theorem; the Table 3 numbers are independent evidence, albeit gathered under heterogeneous setups. The authors' own Section 7 explicitly disclaims the comparability of these numbers, which weakens the strength of the Section 6 inference but does not make it circular. The claim is a contestable interpretation of external observations, not a reduction of the survey's conclusion to its own inputs. Self-citation is not used as load-bearing support, and no cited prior work by the authors is invoked to forbid alternatives or supply an unverified premise. The correct verdict is therefore no significant circularity, with a score of 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey rests on the assumption that the external papers it cites are accurate and that its qualitative selection of works is representative. No free parameters or invented entities are introduced. The comparability assumption is the most fragile and is explicitly acknowledged by the authors.

assumptions (2)
  • domain assumption The reported benchmark scores in the cited papers are accurate and directly comparable across models.
    The survey relies on numbers from the original papers without re-running experiments. If these numbers are wrong or measured under different conditions, the comparisons in Table 3 and the conclusions in Section 6 are invalid.
  • domain assumption The selection of papers is representative of the field.
    The survey selects papers based on the authors' judgment, not a systematic search, and the scope is limited to transformer-based VQA approaches, as admitted in Section 7.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends." pith.science (2026). https://pith.science/paper/BNUODZGZ

@misc{pith2026250102235,
  author       = {Pith},
  title        = {Pith review of: Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BNUODZGZ}},
  note         = {Machine review of arXiv:2501.02235}
}
read the original abstract

The field of visually-rich document understanding, which involves interacting with visually-rich documents (whether scanned or born-digital), is rapidly evolving and still lacks consensus on several key aspects of the processing pipeline. In this work, we provide a comprehensive overview of state-of-the-art approaches, emphasizing their strengths and limitations, pointing out the main challenges in the field, and proposing promising research directions.

Figures

Figures reproduced from arXiv: 2501.02235 by the authors.

Figure 1
Figure 1. Illustration of the datasets listed in this survey [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Scene-aware multi-agent document synthesis plus error-driven hard-example expansion improves compact Qwen3-VL models on constrained and open-category KIE, topping reported on-device baselines.

  2. DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

    cs.LG 2026-05 conditional novelty 5.0 of 10

    OCR tools can be ranked without ground-truth labels by measuring how much a multimodal LLM must correct each tool's output.

Reference graph

Works this paper leans on

110 extracted references · 15 canonical work pages · cited by 2 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Joshua Ainslie, Santiago Onta \ n \'o n, Chris Alberti, Philip Pham, Anirudh Ravula, Sumit Sanghai, Zhuyun Meng, and Lana Hou. 2020. https://arxiv.org/abs/2004.08483 Etc: Encoding long and structured data in transformers . arXiv preprint arXiv:2004.08483

  4. [4]

    Mirna Al-Shetairy, Hanan Hindy, Dina Khattab, and Mostafa M. Aref. 2024. https://arxiv.org/abs/2410.13883 Transformers utilization in chart understanding: A review of recent advances and future trends . Preprint, arXiv:2410.13883

  5. [5]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...

  6. [6]

    Manmatha

    Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R. Manmatha. 2021. https://arxiv.org/abs/2106.11539 Docformer: End-to-end transformer for document understanding . Preprint, arXiv:2106.11539

  7. [7]

    Manmatha

    Srikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran, Yichu Zhou, and R. Manmatha. 2023. https://arxiv.org/abs/2306.01733 Docformerv2: Local features for document understanding . Preprint, arXiv:2306.01733

  8. [8]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. https://arxiv.org/abs/2308.12966 Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond . Preprint, arXiv:2308.12966

Show all 110 references
  1. [9]

    Tsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts, Roy Ganz, Elad Ben Avraham, Aviad Aberdam, Shahar Tsiper, and Ron Litman. 2024. https://arxiv.org/abs/2401.03411 Gram: Global reasoning for multi-page vqa . Preprint, arXiv:2401.03411

  2. [10]

    Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. 2023. https://arxiv.org/abs/2308.13418 Nougat: Neural optical understanding for academic documents . Preprint, arXiv:2308.13418

  3. [11]

    Borchmann, Michal Pietruszka, Wojciech Ja'skowski, Dawid Jurkiewicz, Piotr Halama, Pawel J'oziak, Lukasz Garncarek, Pawel Liskowski, Karolina Szyndler, Andrzej Gretkowski, Julita Oltusek, Gabriela Nowakowska, Artur Zawlocki, Lukasz Duhr, Pawel Dyda, and Michal Turski. 2024. ht...

  4. [12]

    Haoyu Cao, Changcun Bao, Chaohu Liu, Huang Chen, Kun Yin, Hao Liu, Yinsong Liu, Deqiang Jiang, and Xing Sun. 2023. https://arxiv.org/abs/2309.01131 Attention where it matters: Rethinking visual document understanding with selective region concentration . Preprint, arXiv:2309.01131

  5. [13]

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. https://arxiv.org/abs/2005.12872 End-to-end object detection with transformers . Preprint, arXiv:2005.12872

  6. [14]

    Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh. 2024. https://arxiv.org/abs/2312.06742 Honeybee: Locality-enhanced projector for multimodal llm . Preprint, arXiv:2312.06742

  7. [15]

    Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. 2024. https://arxiv.org/abs/2403.06764 An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models . Preprint, arXiv:2403.06764

  8. [16]

    Pu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung, Yin-Wen Chang, and Chun-Sung Ferng. 2021. https://arxiv.org/abs/2104.08698 A simple and effective positional encoding for transformers . Preprint, arXiv:2104.08698

  9. [17]

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. https://arxiv.org/abs/1909.11740 Uniter: Universal image-text representation learning . Preprint, arXiv:1909.11740

  10. [18]

    Yew Ken Chia, Liying Cheng, Hou Pong Chan, Chaoqun Liu, Maojia Song, Sharifah Mahani Aljunied, Soujanya Poria, and Lidong Bing. 2024. https://arxiv.org/abs/2411.06176 M-longdoc: A benchmark for multimodal super-long document understanding and a retrieval-aware tuning framework...

  11. [19]

    Jaemin Cho, Debanjan Mahata, Ozan Irsoy, Yujie He, and Mohit Bansal. 2024. https://arxiv.org/abs/2411.04952 M3docrag: Multi-modal retrieval is what you need for multi-page multi-document understanding . Preprint, arXiv:2411.04952

  12. [20]

    He-Sen Dai, Xiao-Hui Li, Fei Yin, Xudong Yan, Shuqi Mei, and Cheng-Lin Liu. 2024. https://api.semanticscholar.org/CorpusID:272694741 Graphmllm: A graph-based multi-level layout language-independent model for document understanding . In IEEE International Conference on Document...

  13. [21]

    Brian Davis, Bryan Morse, Bryan Price, Chris Tensmeyer, Curtis Wigington, and Vlad Morariu. 2022. https://arxiv.org/abs/2203.16618 End-to-end document recognition and understanding with dessurt . Preprint, arXiv:2203.16618

  14. [22]

    Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020. https://arxiv.org/abs/2006.14806 Turl: Table understanding through representation learning . Preprint, arXiv:2006.14806

  15. [23]

    Mohamed Dhouib, Ghassen Bettaieb, and Aymen Shabou. 2023. https://arxiv.org/abs/2304.12484 Docparser: End-to-end ocr-free information extraction from visually rich documents . Preprint, arXiv:2304.12484

  16. [24]

    Yihao Ding, Jean Lee, and Soyeon Caren Han. 2024. https://arxiv.org/abs/2408.01287 Deep learning based visually rich document content understanding: A survey . Preprint, arXiv:2408.01287

  17. [25]

    Qi Dong, Lei Kang, and Dimosthenis Karatzas. 2024 a . https://api.semanticscholar.org/CorpusID:272701570 Multi-page document vqa with recurrent memory transformer . In International Workshop on Document Analysis Systems

  18. [26]

    Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, a...

  19. [27]

    Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo. 2024. https://arxiv.org/abs/2407.01449 Colpali: Efficient document retrieval with vision language models . Preprint, arXiv:2407.01449

  20. [28]

    Hao Feng, Qi Liu, Hao Liu, Jingqun Tang, Wengang Zhou, Houqiang Li, and Can Huang. 2024. https://arxiv.org/abs/2311.11810 Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding . Preprint, arXiv:2311.11810

  21. [29]

    Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang. 2023. https://arxiv.org/abs/2308.11592 Unidoc: A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding . Preprint, arXiv:2308.11592

  22. [30]

    Masato Fujitake. 2024. https://arxiv.org/abs/2403.14252 Layoutllm: Large language model instruction tuning for visually rich document understanding . Preprint, arXiv:2403.14252

  23. [31]

    Lukasz Garncarek, Rafal Powalski, Tomasz Stanislawek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Grali'nski. 2020. https://api.semanticscholar.org/CorpusID:235262539 Lambert: Layout-aware language modeling for information extraction . In IEEE International Confer...

  24. [32]

    Simone Giovannini, Fabio Coppini, Andrea Gemelli, and Simone Marinai. 2025. https://arxiv.org/abs/2501.03403 Boundingdocs: a unified dataset for document question answering with spatial annotations . Preprint, arXiv:2501.03403

  25. [33]

    Prashant Gupta, Daniel Borchmann, Alvaro Dossantos, and Umapada Pal. 2022. https://arxiv.org/abs/2207.06881 Recurrent memory transformer . arXiv preprint arXiv:2207.06881

  26. [34]

    Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://arxiv.org/abs/2006.03654 Deberta: Decoding-enhanced bert with disentangled attention . Preprint, arXiv:2006.03654

  27. [35]

    Teakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2022. https://arxiv.org/abs/2108.04539 Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents . Preprint, arXiv:2108.04539

  28. [36]

    Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxuan Zhang, Juanzi Li, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang. 2024. https://arxiv.org/abs/2312.08914 Cogagent: A visual language model for gui agents . Preprint, arXiv:2312.08914

  29. [37]

    Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024 a . https://arxiv.org/abs/2403.12895 mplug-docowl 1.5: Unified structure learning for ocr-free document understanding . Preprint, arXiv:2403.12895

  30. [38]

    Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024 b . https://arxiv.org/abs/2409.03420 mplug-docowl2: High-resolution compressing for ocr-free multi-page document understanding . Preprint, arXiv:2409.03420

  31. [39]

    Pengfei Hu, Zhenrong Zhang, Jiefeng Ma, Shuhang Liu, Jun Du, and Jianshu Zhang. 2025. https://arxiv.org/abs/2409.11887 Docmamba: Efficient document pre-training with state space model . Preprint, arXiv:2409.11887

  32. [41]

    Jiani Huang, Haihua Chen, Fengchang Yu, and Wei Lu. 2024 b . https://doi.org/10.1145/3657285 From detection to application: Recent advances in understanding scientific tables and figures . ACM Comput. Surv., 56(10)

  33. [42]

    Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022. https://arxiv.org/abs/2204.08387 Layoutlmv3: Pre-training for document ai with unified text and image masking . Preprint, arXiv:2204.08387

  34. [43]

    Geewook Kim, Teakgyu Hong, Moonbin Yim, Jeongyeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. 2022. https://arxiv.org/abs/2111.15664 Ocr-free document understanding transformer . Preprint, arXiv:2111.15664

  35. [44]

    Jordy Van Landeghem, Rubén Tito, Łukasz Borchmann, Michał Pietruszka, Paweł Józiak, Rafał Powalski, Dawid Jurkiewicz, Mickaël Coustaty, Bertrand Ackaert, Ernest Valveny, Matthew Blaschko, Sien Moens, and Tomasz Stanisławek. 2023. https://arxiv.org/abs/2305.08455 Document under...

  36. [45]

    Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. 2024. https://arxiv.org/abs/2405.02246 What matters when building vision-language models? Preprint, arXiv:2405.02246

  37. [46]

    Junlong Lee, Yiheng Xu, Yang Xiao, Huan Wang, Jinlong Zhao, Pengchuan Xie, Miao Xu, Baolin Shi, and Lei Xu. 2022. https://arxiv.org/abs/2203.08411 Formnet: Structural encoding beyond sequential modeling in form document information extraction . arXiv preprint arXiv:2203.08411

  38. [47]

    Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. https://arxiv.org/abs/2210.03347 Pix2struct: Screenshot parsing as pretraining for visual language understanding . Pr...

  39. [48]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. https://arxiv.org/abs/2005.11401 Retrieval-augmented generation for knowledge-int...

  40. [49]

    Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2021 a . https://arxiv.org/abs/2105.11210 Structurallm: Structural pre-training for form understanding . Preprint, arXiv:2105.11210

  41. [50]

    Jia-Nan Li, Jian Guan, Wei Wu, Zhengtao Yu, and Rui Yan. 2024 a . https://arxiv.org/abs/2409.19700 2d-tpe: Two-dimensional positional encoding enhances table understanding for large language models . Preprint, arXiv:2409.19700

  42. [51]

    Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. 2022. https://arxiv.org/abs/2203.02378 Dit: Self-supervised pre-training for document image transformer . Preprint, arXiv:2203.02378

  43. [52]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 a . https://arxiv.org/abs/2301.12597 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models . Preprint, arXiv:2301.12597

  44. [53]

    Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu

    Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. 2021 b . https://arxiv.org/abs/2106.03331 Selfdoc: Self-supervised document representation learning . Preprint, arXiv:2106.03331

  45. [54]

    Peng Li, Xiaotang Zhao, Wei Fang, et al. 2024 b . https://arxiv.org/pdf/2407.02392 Tokenpacker: Efficient visual projector for multimodal llm . arXiv preprint arXiv:2407.02392

  46. [55]

    Li, Xiantao Cai, Bo Du, and Hai Zhao

    Qiwei Li, Z. Li, Xiantao Cai, Bo Du, and Hai Zhao. 2023 b . https://api.semanticscholar.org/CorpusID:260899841 Enhancing visually-rich document understanding via layout structure modeling . Proceedings of the 31st ACM International Conference on Multimedia

  47. [56]

    Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. 2024 c . https://arxiv.org/abs/2403.18814 Mini-gemini: Mining the potential of multi-modality vision language models . Preprint, arXiv:2403.18814

  48. [57]

    Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2024 d . https://arxiv.org/abs/2311.06607 Monkey: Image resolution and text label are important things for large multi-modal models . Preprint, arXiv:2311.06607

  49. [58]

    Manmatha, and Vijay Mahadevan

    Haofu Liao, Aruni RoyChowdhury, Weijian Li, Ankan Bansal, Yuting Zhang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, and Vijay Mahadevan. 2023. https://arxiv.org/abs/2307.07929 Doctr: Document transformer for structured information extraction in documents . Preprint, arXiv:2307.07929

  50. [59]

    Wenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang, Jun Huang, and Lianwen Jin. 2024. https://arxiv.org/abs/2408.15045 Doclayllm: An efficient and effective multi-modal extension of large language models for text-rich document understanding . Preprint, arXiv:2408.15045

  51. [60]

    Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang, Yichi Zhang, Xuming He, Hongsheng Li, and Yu Qiao. 2023. https://arxiv.org/abs/2311.07575 Sphinx: The joint mixing of weights, tasks, and ...

  52. [61]

    Chaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun, and Linli Xu. 2024 a . https://arxiv.org/abs/2404.06918 Hrvda: High-resolution visual document assistant . Preprint, arXiv:2404.06918

  53. [62]

    Fangyu Liu, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, and Yasemin Altun. 2023. https://arxiv.org/abs/2212.10505 Deplot: One-shot visual language reasoning by plot-to-table translation . Pre...

  54. [63]

    Hao Liu, Xinghua Jiang, Xin Li, Antai Guo, Deqiang Jiang, and Bo Ren. 2022 a . https://arxiv.org/abs/2204.08227 The devil is in the frequency: Geminated gestalt autoencoder for self-supervised visual pre-training . Preprint, arXiv:2204.08227

  55. [64]

    Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024 b . https://arxiv.org/abs/2403.04473 Textmonkey: An ocr-free large multimodal model for understanding document . Preprint, arXiv:2403.04473

  56. [65]

    Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. 2022 b . https://arxiv.org/abs/2111.09883 Swin transformer v2: Scaling up capacity and resolution . Preprint, arXiv:2111.09883

  57. [66]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. https://arxiv.org/abs/2103.14030 Swin transformer: Hierarchical vision transformer using shifted windows . Preprint, arXiv:2103.14030

  58. [67]

    Junyu Lu, Dixiang Zhang, Songxin Zhang, Zejian Xie, Zhuoyang Song, Cong Lin, Jiaxing Zhang, Bingyi Jing, and Pingjian Zhang. 2024. https://arxiv.org/abs/2312.05278 Lyrics: Boosting fine-grained language-vision alignment and comprehension via semantic-aware visual objects . Pre...

  59. [68]

    Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, and Rongrong Ji. 2024. https://arxiv.org/abs/2403.03003 Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models . Preprint, arXiv:2403.03003

  60. [69]

    Tengchao Lv, Yupan Huang, Jingye Chen, Yuzhong Zhao, Yilin Jia, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, Shaoxiang Wu, Guoxin Wang, Cha Zhang, and Furu Wei. 2024. https://arxiv.org/abs/2309.11419 Kosmos-2.5: A multimodal literate mode...

  61. [70]

    Feipeng Ma, Yizhou Zhou, Hebei Li, Zilong He, Siying Wu, Fengyun Rao, Yueyi Zhang, and Xiaoyan Sun. 2024 a . https://ar5iv.labs.arxiv.org/html/2408.11795v1 Ee-mllm: A data-efficient and compute-efficient multimodal large language model . arXiv preprint arXiv:2408.11795

  62. [71]

    Xueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen, and Jimmy Lin. 2024 b . https://arxiv.org/abs/2406.11251 Unifying multimodal retrieval via document screenshot embedding . Preprint, arXiv:2406.11251

  63. [72]

    Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, and Aixin Sun. 2024 c . https://arxiv.org/abs/2407.01523 MMLongBench-Doc: Benchmarking Long-context ...

  64. [73]

    Zhiming Mao, Haoli Bai, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024. https://arxiv.org/abs/2403.16516 Visually guided generative text-layout pre-training for document intelligence . Preprint, arXiv:2403.16516

  65. [74]

    V Jawahar

    Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V Jawahar. 2021 a . https://arxiv.org/abs/2104.12756 Infographicvqa . Preprint, arXiv:2104.12756

  66. [75]

    Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021 b . https://arxiv.org/abs/2007.00398 Docvqa: A dataset for vqa on document images . Preprint, arXiv:2007.00398

  67. [76]

    Chaitanya Naidu, Mohammad Khan, and C. V. Jawahar. 2024. https://arxiv.org/abs/2404.19024 Multi-page document visual question answering using self-attention scoring mechanism . arXiv preprint arXiv:2404.19024

  68. [77]

    Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2022. https://arxiv.org/abs/2210.06155 Ernie-layout: Layout knowledge enhanced pre-training for...

  69. [78]

    Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka. 2021. https://arxiv.org/abs/2102.09550 Going full-tilt boogie on document understanding with text-image-layout transformer . Preprint, arXiv:2102.09550

  70. [79]

    Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel. 2022. https://arxiv.org/abs/2009.14457 Towards a multi-modal, multi-task learning based pre-training framework for document representation learning . Preprint, arXiv:2009.14457

  71. [80]

    Smith, and Mike Lewis

    Ofir Press, Noah A. Smith, and Mike Lewis. 2022. https://arxiv.org/abs/2108.12409 Train short, test long: Attention with linear biases enables input length extrapolation . Preprint, arXiv:2108.12409

  72. [81]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. https://arxiv.org/abs/1910.10683 Exploring the limits of transfer learning with a unified text-to-text transformer . Preprint, arXiv:1910.10683

  73. [82]

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. https://arxiv.org/abs/1506.01497 Faster r-cnn: Towards real-time object detection with region proposal networks . Preprint, arXiv:1506.01497

  74. [83]

    Imanol Schlag, Paul Smolensky, Roland Fernandez, Nebojsa Jojic, Jürgen Schmidhuber, and Jianfeng Gao. 2020. https://arxiv.org/abs/1910.06611 Enhancing the transformer with explicit relational encoding for math problem solving . Preprint, arXiv:1910.06611

  75. [84]

    Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. 2024. https://arxiv.org/abs/2403.15388 Llava-prumerge: Adaptive token reduction for efficient large multimodal models . Preprint, arXiv:2403.15388

  76. [85]

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. https://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding . Preprint, arXiv:2104.09864

  77. [86]

    Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, and Jun Suzuki. 2024. https://arxiv.org/abs/2401.13313 Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions . Preprint, arXiv:2401.13313

  78. [87]

    Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, and Kuniko Saito. 2023. https://arxiv.org/abs/2301.04883 Slidevqa: A dataset for document visual question answering on multiple images . Preprint, arXiv:2301.04883

  79. [88]

    Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. 2021. https://arxiv.org/abs/2101.11272 Visualmrc: Machine reading comprehension on document images . Preprint, arXiv:2101.11272

  80. [89]

    Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, and Can Huang. 2024. https://arxiv.org/abs/2404.12803 Textsquare: Scaling up text-centric visual instruction tuning ....

  81. [90]

    Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal. 2023. https://arxiv.org/abs/2212.02623 Unifying vision, text, and layout for universal document processing . Preprint, arXiv:2212.02623

  82. [91]

    Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny. 2023. https://arxiv.org/abs/2212.05935 Hierarchical multimodal transformers for multi-page docvqa . Preprint, arXiv:2212.05935

  83. [92]

    Zilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun, Jingbo Shang, and Vlad I. Morariu. 2022. https://arxiv.org/abs/2211.14958 Mgdoc: Pre-training with multi-granular hierarchy for document image understanding . Preprint, arXiv:2211.14958

  84. [93]

    Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2023. https://arxiv.org/abs/2312.06109 Vary: Scaling up the vision vocabulary for large vision-language models . Preprint, arXiv:2312.06109

  85. [94]

    Xudong Xie, Liang Yin, Hao Yan, Yang Liu, Jing Ding, Minghui Liao, Yuliang Liu, Wei Chen, and Xiang Bai. 2024. https://doi.org/10.48550/arXiv.2410.05970 Pdf-wukong: A large multimodal model for efficient long pdf reading with end-to-end sparse sampling . arXiv preprint arXiv:2...

  86. [96]

    Ruyi Xu, Yuan Yao, Zonghao Guo, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, Maosong Sun, and Gao Huang. 2024 b . https://arxiv.org/abs/2403.11703 Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images . Preprint, arXiv:2403.11703

  87. [97]

    Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2022. https://arxiv.org/abs/2012.14740 Layoutlmv2: Multi-modal pre-training for visually-rich document understanding . Preprint, ar...

  88. [98]

    Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2019. https://api.semanticscholar.org/CorpusID:209515395 Layoutlm: Pre-training of text and layout for document image understanding . Proceedings of the 26th ACM SIGKDD International Conference on Knowledg...

  89. [99]

    Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. 2021. https://arxiv.org/abs/2104.08836 Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding . Preprint, arXiv:2104.08836

  90. [100]

    Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023 a . https://arxiv.org/abs/2307.02499 mplug-docowl: Modularized multimodal large language model for document understandin...

  91. [101]

    Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, Qin Jin, Liang He, Xin Alex Lin, and Fei Huang. 2023 b . https://arxiv.org/abs/2310.05126 Ureader: Universal ocr-free visually-situated language understanding with m...

  92. [102]

    Pengcheng Yin, Graham Neubig, Wen tau Yih, and Sebastian Riedel. 2020. https://arxiv.org/abs/2005.08314 Tabert: Pretraining for joint understanding of textual and tabular data . Preprint, arXiv:2005.08314

  93. [103]

    Ya-Qi Yu, Minghui Liao, Jihao Wu, Yongxin Liao, Xiaoyu Zheng, and Wei Zeng. 2024. https://arxiv.org/abs/2404.09204 Texthawk: Exploring efficient fine-grained perception of multimodal large language models . Preprint, arXiv:2404.09204

  94. [104]

    Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2025. https://openreview.net/forum?id=DgaY5mDdmT MLLM s know where to look: Training-free perception of small visual details with multimodal LLM s . In The Thirteenth International Conference on Learning R...

  95. [105]

    Jiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie, and Lianwen Jin. 2024 a . https://arxiv.org/abs/2406.19101 Dockylin: A large multimodal model for visual document understanding with efficient visual slimming . arXiv preprint arXiv:2406.19101

  96. [106]

    Liang Zhang, Anwen Hu, Haiyang Xu, Ming Yan, Yichen Xu, Qin Jin, Ji Zhang, and Fei Huang. 2024 b . https://arxiv.org/abs/2404.16635 Tinychart: Efficient chart understanding with visual token merging and program-of-thoughts learning . Preprint, arXiv:2404.16635

  97. [107]

    Qintong Zhang, Victor Shea-Jay Huang, Bin Wang, Junyuan Zhang, Zhengren Wang, Hao Liang, Shawn Wang, Matthieu Lin, Conghui He, and Wentao Zhang. 2024 c . https://arxiv.org/abs/2410.21169 Document parsing unveiled: Techniques, challenges, and prospects for structured informatio...

  98. [108]

    Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun. 2024 d . https://arxiv.org/abs/2306.17107 Llavar: Enhanced visual instruction tuning for text-rich image understanding . Preprint, arXiv:2306.17107

  99. [109]

    Zhenrong Zhang, Jiefeng Ma, Jun Du, Licheng Wang, and Jianshu Zhang. 2022. https://api.semanticscholar.org/CorpusID:247748605 Multimodal pre-training based on graph attention network for document understanding . IEEE Transactions on Multimedia, 25:6743--6755

  100. [110]

    Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. 2022. https://doi.org/10.1145/3503161.3548422 Towards complex document understanding by discrete reasoning . In Proceedings of the 30th ACM International Conference on Multimedia, page 4857–4866. ACM

  101. [111]

    Fengbin Zhu, Ziyang Liu, Xiang Yao Ng, Haohui Wu, Wenjie Wang, Fuli Feng, Chao Wang, Huanbo Luan, and Tat Seng Chua. 2024. https://arxiv.org/abs/2410.21311 Mmdocbench: Benchmarking large vision-language models for fine-grained visual document understanding . Preprint, arXiv:2410.21311

  102. [112]

    Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2021. https://arxiv.org/abs/2010.04159 Deformable detr: Deformable transformers for end-to-end object detection . Preprint, arXiv:2010.04159

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.