Pith. sign in

REVIEW 3 major objections 6 minor 137 references

A document-native vision encoder, trained with text generation plus pixel reconstruction, transfers across OCR, parsing, and understanding better than natural-image foundations.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 04:44 UTC pith:NDLC6ZRV

load-bearing objection Solid document-native encoder with clean multi-task transfer; the MDPBench SOTA headline is real but not cleanly encoder-attributable. the 3 major comments →

arxiv 2607.11562 v1 pith:NDLC6ZRV submitted 2026-07-13 cs.CV

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

classification cs.CV
keywords document AIvisual pretrainingOCRdocument parsingdocument understandingimage reconstructionmultilingual documentsvision foundation models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Mainstream vision encoders learn object and scene semantics from natural photos, so they miss the stroke-level detail that document images depend on. This paper argues that a dedicated visual foundation for documents is possible: pretrain an encoder on a large multilingual document corpus so that it both reads text and reconstructs pixels. The generation objective ties features to words; the reconstruction objective keeps character strokes, glyphs, and layout that pure language supervision would throw away. Swapping that encoder into existing systems improves text recognition, formula recognition, detection, tampering localization, and overlapping-text segmentation. Kept frozen and paired with a small language model, it also drives competitive multilingual document parsing and stronger document understanding than CLIP-, DINO-, or SAM-style backbones under matched training. The practical claim is that document intelligence can rest on its own visual pretraining rather than borrowed natural-image features.

Core claim

Document-oriented pretraining that jointly optimizes image-to-text generation and pixel-level reconstruction produces transferable character-level visual representations: as a backbone swap it improves five document analysis tasks, and as a frozen encoder with a lightweight language model it yields a 0.7B parser that reaches open-source state of the art on multilingual MDPBench while also outperforming CLIP, DINO, and SAM counterparts on eight document-understanding benchmarks under identical training.

What carries the argument

Dual-objective document pretraining on MonkeyDoc v2 (113M images, 17 languages): image-to-text generation aligns visual tokens with textual content, while pixel-level reconstruction (MSE, optionally edge- and distance-aware) forces the encoder to retain strokes, glyphs, and layout that text supervision alone can discard.

Load-bearing premise

The large training labels from multi-expert agreement and automatic layout filters are clean enough that measured gains reflect better visual features rather than shared errors between those labels and the evaluation stack.

What would settle it

Under fully matched data, optimization, and decoding, freeze MonkeyOCRv2 and a strong natural-image encoder of similar size, train the same lightweight document parser, and check whether the document-native encoder still wins on photographed non-Latin pages of MDPBench and on scrambled-text recognition at low resolution; a clear loss would falsify the claim that the dual objective, not scale or task setup, is doing the work.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Document systems can replace ImageNet, CLIP, DINO, or SAM backbones with a compact document encoder and expect gains without rewriting the rest of the pipeline.
  • A frozen ~0.1B document vision encoder plus a ~0.6B language model is enough for competitive multilingual parsing, so large general VLMs are not required for that task class.
  • Pixel reconstruction should reduce reliance on language priors when text is scrambled, low-resolution, or deliberately conflicted with linguistic expectations.
  • Future document foundation work can treat text strokes and layout as first-class visual targets rather than side effects of semantic alignment.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same dual recipe may help other dense-symbol domains (sheet music, circuit diagrams, engineering drawings) where global semantic encoders discard local marks.
  • If reconstruction is what narrows the semantic–scrambled gap, progressive post-training of the language head alone may not close remaining gaps on saturated parsing benches without stronger visual evidence.
  • Balancing MonkeyDoc-style supervision toward low-resource and historical scripts would be a direct test of whether the method generalizes beyond high-resource languages that dominate the current mix.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. MonkeyOCRv2 proposes a document-native visual encoder pretrained on MonkeyDoc v2 (113M images, 17 languages) with a joint objective of image-to-text generation and pixel-level reconstruction (Eqs. 1–11). The encoder is evaluated as a backbone substitution on five document analysis tasks (text recognition, formula recognition, text detection, tampering detection, overlapping text segmentation) and, frozen, as the vision tower of lightweight VLMs for document parsing and understanding. The paper reports consistent gains from encoder replacement (Tabs. 2–5, Fig. 4), open-source SOTA on MDPBench for a 0.7B frozen-encoder parser (+2.8 over 3B dots.mocr with a much smaller ViT; Tab. 6), and superior document-understanding scores versus CLIP/DINO/SAM/OpenVision under matched LLM, data, and training (Tab. 8). Reconstruction is further supported by scrambled-text and CHAOS-Bench analyses (Sec. 5.1–5.2).

Significance. If the transfer results hold under the stated controls, the work is a substantial contribution to document AI: it argues, with multi-task evidence, that character-level document pretraining can serve as a foundation rather than a domain adaptation of natural-image encoders. Strengths include a large multilingual corpus, a dual-objective recipe with explicit reconstruction ablations (MSE and structure-aware variants in Tab. 8; scrambled-text gap narrowing in Fig. 5), controlled frozen-encoder VLM comparisons on eight benchmarks (Tab. 8), and honest system-level caveats for OmniDocBench (Sec. 4.6, Tab. 7). The five-task backbone-swap protocol is particularly useful for the community. Code and data release is promised, which would further raise impact.

major comments (3)
  1. Sec. 4.6 and Tab. 6 present MonkeyOCRv2-Parsing as open-source SOTA on MDPBench (+2.8 over dots.mocr, ~11× smaller vision encoder). The same section correctly notes that OmniDocBench (Tab. 7) is system-level and that encoder attribution should use Tab. 8. That caveat is not applied to MDPBench: the parser couples a frozen encoder to autoregressive layout prediction, per-element re-crop recognition, and assembly, with no matched experiment that freezes a strong alternative encoder (e.g., OpenVision-B, RADIOv2.5-B, SigLIP 2) inside the identical parsing pipeline, data, and training recipe. Without that control, the headline +2.8 and size framing can credit pipeline design and training as much as document-native pretraining. Please either add the matched encoder swap in the parsing stack or reframe Tab. 6 as a system result and rest the encoder claim primarily on Tabs. 2–5 and Tab. 8.
  2. Sec. 3.1 (Expert Model Labeling; Data Filtering) relies on multi-expert OCR agreement and LLM layout/reading-order filters for large-scale supervision. The residual error structure of those automatic labels is not quantified against human gold or against the evaluation stacks (including MDPBench, which shares research-lineage with prior MonkeyOCR work). If label failure modes correlate with the pretraining objective or with same-lineage benchmarks, measured gains partly reflect label–model correlation. A short audit—e.g., human agreement rates on a stratified sample, or performance when pretraining only on fully public human-annotated subsets—would strengthen the claim that gains are representation quality rather than supervision artifacts.
  3. Tab. 8 is the cleanest encoder-level comparison, but input configurations differ substantially (App. D: CLIP 196 tokens vs SAM 4096 vs MonkeyOCRv2 ~1082). The paper states each encoder uses its native setting, which is reasonable, yet token budget and resolution are known confounders for document VQA. A sensitivity check that equalizes approximate visual-token count (or reports a fixed-token budget ablation for the top baselines) would make the 13.2-point gap over OpenVision-B more attributable to pretraining rather than resolution policy.
minor comments (6)
  1. Abstract and Fig. 2(a) lead with the MDPBench SOTA and 11× smaller encoder; after addressing the major comment on attribution, align abstract wording with the revised claim so abstract and Sec. 4.6 do not over-promise encoder-only causality.
  2. Eq. (11) sets λ=1.0 with α, β, T, τ fixed without tuning (Sec. 3.2). A brief sensitivity note (even one-dimensional in λ) would help readers assess robustness of the dual-objective balance.
  3. Table 1 lists MonkeyDoc v2 as 113M multi-type documents / 17 languages; App. A shows strong English/Chinese skew. Mentioning this imbalance earlier (not only in Limitations) would set expectations for low-resource scripts.
  4. Fig. 4 uses bar charts without numeric tables in the main text; adding exact F-measures in a small table or appendix would aid citation and reproducibility.
  5. Typo/consistency: abstract says “previous best 3B dots.mocr” while related work and Tab. 6 use “dots.mocr”; unify naming. Also “UniMERNet-T to outperform the 325M UniMERNet-B” is clear in Tab. 3 but the abstract’s “enabling the 110M” phrasing could state ExpRate/CDM explicitly.
  6. Sec. 5.1 correctly treats the semantic–scrambled gap as an operational proxy; consider moving one sentence of that caveat into the figure caption of Fig. 5 so casual readers do not over-interpret the gap as a pure hallucination metric.

Circularity Check

1 steps flagged

No derivation-by-construction circularity; mild same-lineage risk on MDPBench SOTA framing, while backbone swaps and matched VLM controls remain independent.

specific steps
  1. self citation load bearing [Sec. 4.6 Document Parsing; Tab. 6 MDPBench; abstract SOTA claim]
    "Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench... surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11× smaller. ... While MDPBench originates from the same research line as our prior benchmarks, the encoder, the LLM, and the data pipeline evaluated here are independent of its construction"

    The headline open-source SOTA and 11×-smaller-encoder framing rest on MDPBench, which the paper itself notes comes from the same research line. The paper does not freeze a strong alternative encoder inside the exact autoregressive-layout + re-crop parsing pipeline, so the load-bearing +2.8 claim partly leans on same-lineage benchmark construction rather than a fully independent external test. This is mild self-citation risk, not a definitional reduction of the pretraining objective.

full rationale

MonkeyOCRv2 is an empirical CV foundation-model paper, not a first-principles derivation. The dual objective L_pretrain = L_text + λ L_rec is a training recipe, not a claim that one quantity is mathematically forced by another. Gains are measured by encoder substitution into CRNN/PARSeq, UniMERNet-T, DBNet/PSENet/DPText-DETR, FFDN, Mask2Former/MOTS (Tabs. 2–5, Fig. 4) and by frozen-encoder VLMs under fixed LLM/data/optimization (Tab. 8). Reconstruction ablations (Fig. 5, Tab. 8 baseline vs MSE vs structure-aware, Tab. 9 CHAOS-Bench) compare trained variants rather than renaming a fit as a prediction. MDPBench shares authorship lineage with prior MonkeyOCR work, and Sec. 4.6/Tab. 7 explicitly warn that OmniDocBench is system-level, so the +2.8 open-source SOTA headline is not cleanly encoder-attributed without a matched alternative-encoder ablation inside the same parsing pipeline—this is attribution softness, not circular reduction of a claimed derivation. No self-definitional equations, fitted-input-as-prediction, uniqueness theorems, or ansatz-via-citation chains appear. Score 2 for mild self-citation load on the parsing headline only.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

This is an empirical systems paper. Load-bearing premises are standard deep-learning transfer assumptions plus specific data-labeling and loss-weight choices, not new physical entities. Free parameters are the dual-loss weights and training knobs fixed without extensive search. Axioms cover the claim that reconstruction preserves character-level evidence and that automatic multi-expert labels are adequate at 113M scale.

free parameters (4)
  • λ (reconstruction weight in L_pretrain)
    Balances text generation vs reconstruction; fixed at 1.0 without reported sweep beyond the presence/absence ablation.
  • α, β (structure-aware reconstruction weights)
    Fixed at 0.5 and 0.25 for edge/distance terms used in document-understanding variants; not systematically tuned.
  • T, τ (distance-to-edge iterations and edge temperature)
    Fixed at 16 and 0.08 for structure-aware loss; ad hoc numerical choices.
  • peak learning rate and batch size for pretraining
    1e-3 and global batch 256 on 64 A800s; standard but claim-dependent training hyperparameters.
axioms (4)
  • domain assumption Pixel reconstruction (MSE and optional edge/distance matching) forces the encoder to retain character strokes and layout that pure text supervision discards.
    Core justification for the dual objective (Sec. 3.2 and Sec. 5.1); supported by ablations but not proven as the unique mechanism.
  • domain assumption Multi-expert agreement among OCR systems plus LLM layout/order filters yields sufficiently accurate labels for large-scale pretraining.
    Data Engine (Sec. 3.1) relies on this for real-document supervision quality.
  • domain assumption Frozen-encoder transfer under matched LLM, data, and optimization isolates visual representation quality.
    Used for document understanding (Tab. 8) and parsing claims; residual differences in native resolution/token counts remain (App. D).
  • standard math Standard transformer/ViT training dynamics and cross-entropy + MSE optimization are valid for learning transferable document features.
    Background ML practice assumed throughout pretraining and fine-tuning.
invented entities (2)
  • MonkeyDoc v2 corpus no independent evidence
    purpose: Provide 113M multilingual document images with dense text supervision for document-native pretraining.
    Constructed resource, not a physical postulate; independent value depends on public release and external reuse.
  • MonkeyOCRv2 dual-objective encoder family (S/B/AS) no independent evidence
    purpose: Serve as a reusable document vision backbone for recognition, detection, parsing, and understanding.
    New model family defined by architecture + pretraining recipe; evidence is internal benchmarks pending external reimplementation.

pith-pipeline@v1.1.0-grok45 · 44414 in / 3332 out tokens · 38181 ms · 2026-07-14T04:44:35.070340+00:00 · methodology

0 comments
read the original abstract

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.

Figures

Figures reproduced from arXiv: 2607.11562 by Dongliang Luo, Handong Zheng, Jiajun Song, Jiarui Zhang, Qiang Liu, Shuo Zhang, Xiang Bai, Xinhan Wang, Yang Liu, Yuliang Liu, Zhang Li, Zhiyin Ma, Zidun Guo, Ziyang Zhang.

Figure 1
Figure 1. Figure 1: Overview of MonkeyOCRv2. Existing vision foundation models are primarily designed for natural images and emphasize object semantics, global alignment, semantic features, or region boundaries. MonkeyOCRv2 addresses the resulting representation mismatch by jointly learning text generation and pixel-level reconstruction, producing document-native visual representations that transfer across diverse document AI… view at source ↗
Figure 2
Figure 2. Figure 2: Performance overview of MonkeyOCRv2. (a) Performance versus vision-encoder size on MDPBench [56], a challenging multilingual document parsing benchmark. MonkeyOCRv2 achieves 83.3%, outperforming the previous best open-source model, dots.mocr, with a vision encoder roughly 11× smaller. Bubble area indicates the total number of model parameters. (b) Absolute performance improvements across seven document ana… view at source ↗
Figure 3
Figure 3. Figure 3: MonkeyOCRv2 is pretrained on a large-scale corpus of multilingual, multi-type document [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: For text detection, MonkeyOCRv2 consistently delivers robust performance gains on [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Left: scrambled text recognition accuracy. Right: the accuracy gap between semantically [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Visualization comparisons with leading document parsing models on an Arabic document [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Visualization comparisons with leading document parsing models on a photographed [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Visualization comparisons with popular vision foundation models on document understand [PITH_FULL_IMAGE:figures/full_fig_p018_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Detailed data distribution of MonkeyDoc v2. [PITH_FULL_IMAGE:figures/full_fig_p031_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Visualization of Arabic, German, English, Spanish, French, and Hindi images in Monkey [PITH_FULL_IMAGE:figures/full_fig_p033_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Visualization of Indonesian, Italian, Japanese, Korean, Dutch, and Portuguese images in [PITH_FULL_IMAGE:figures/full_fig_p034_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Visualization of Russian, Thai, Vietnamese, Simplified Chinese, and Traditional Chinese [PITH_FULL_IMAGE:figures/full_fig_p035_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

137 extracted references · 20 linked inside Pith

  1. [1]

    Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...

  2. [2]

    Qwen2.5-vl technical report.arXiv preprint arXiv:2502.13923, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, 19 Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report...

  3. [3]

    BEiT: BERT pre-training of image transformers

    Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. BEiT: BERT pre-training of image transformers. InInternational Conference on Learning Representations, 2022

  4. [4]

    Scene text recognition with permuted autoregressive sequence models

    Darwin Bautista and Rowel Atienza. Scene text recognition with permuted autoregressive sequence models. InProceedings of the European Conference on Computer Vision, pages 178–196, 2022

  5. [5]

    Nougat: Neu- ral optical understanding for academic documents

    Lukas Blecher, Guillem Cucurull Preixens, Thomas Scialom, and Robert Stojnic. Nougat: Neu- ral optical understanding for academic documents. InInternational Conference on Learning Representations, 2024

  6. [6]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9650–9660, 2021

  7. [7]

    Encoder-decoder with atrous separable convolution for semantic image segmentation

    Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision, pages 801–818, 2018

  8. [8]

    Enhancing tampered text detection through frequency feature fusion and decomposition

    Zhongxi Chen, Shen Chen, Taiping Yao, Ke Sun, Shouhong Ding, Xianming Lin, Liujuan Cao, and Rongrong Ji. Enhancing tampered text detection through frequency feature fusion and decomposition. InProceedings of the European Conference on Computer Vision, pages 200–217, 2024

  9. [9]

    Masked-attention mask transformer for universal image segmentation

    Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022

  10. [10]

    Per-pixel classification is not all you need for semantic segmentation.Advances in Neural Information Processing Systems, 34:17864–17875, 2021

    Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation.Advances in Neural Information Processing Systems, 34:17864–17875, 2021

  11. [11]

    M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis

    Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang, Qiyuan Zhu, Zecheng Xie, Jing Li, Kai Ding, and Lianwen Jin. M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 151...

  12. [12]

    Reproducible scaling laws for contrastive language-image learning

    Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2818–2829, 2023

  13. [13]

    Total-text: A comprehensive dataset for scene text detection and recognition

    Chee Kheng Ch’ng and Chee Seng Chan. Total-text: A comprehensive dataset for scene text detection and recognition. InProceedings of the International Conference on Document Analysis and Recognition, pages 935–942, 2017

  14. [14]

    Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art

    Chee Kheng Chng, Yuliang Liu, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, and Lianwen Jin. Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art. InProceedings of the International Conference on Document Analysis and Recognition, pag...

  15. [15]

    Paddleocr-vl-1.5: Towards a multi-task 0.9b vlm for robust in-the-wild document parsing.arXiv preprint arXiv:2601.21957, 2026

    Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Yi Liu, Dianhai Yu, and Yanjun Ma. Paddleocr-vl-1.5: Towards a multi-task 0.9b vlm for robust in-the-wild document parsing.arXiv preprint arXiv:2601.21957, 2026

  16. [16]

    Boosting document parsing efficiency and performance with coarse-to-fine visual processing

    Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Jing Zhang, Jun Zhang, Xing Wei, Yi Liu, Dianhai Yu, and Yanjun Ma. Boosting document parsing efficiency and performance with coarse-to-fine visual processing. InProceedings of the IEEE/CVF Conference on Comp...

  17. [17]

    Vision grid transformer for document layout analysis

    Cheng Da, Chuwei Luo, Qi Zheng, and Cong Yao. Vision grid transformer for document layout analysis. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 19462–19472, 2023

  18. [18]

    Imagenet: A large- scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009

  19. [19]

    Decaf: A deep convolutional activation feature for generic visual recognition

    Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. InProceedings of the International Conference on Machine Learning, pages 647–655, 2014

  20. [20]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InInternational Conference on Learning Representations, 2021

  21. [21]

    Out of length text recognition with sub-string matching

    Yongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao, and Yu-Gang Jiang. Out of length text recognition with sub-string matching. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 2798–2806, 2025

  22. [22]

    Context perception parallel decoder for scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

    Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Chenxia Li, Yuning Du, and Yu-Gang Jiang. Context perception parallel decoder for scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  23. [23]

    Instruction-guided scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4):2723–2738, 2025

    Yongkun Du, Zhineng Chen, Yuchen Su, Caiyan Jia, and Yu-Gang Jiang. Instruction-guided scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4):2723–2738, 2025

  24. [24]

    Svtrv2: Ctc beats encoder-decoder models in scene text recognition

    Yongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia, and Yu-Gang Jiang. Svtrv2: Ctc beats encoder-decoder models in scene text recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 20147–20156, 2025

  25. [25]

    Unirec-0.1 b: Unified text and formula recognition with 0.1 b parameters.arXiv preprint arXiv:2512.21095, 2025

    Yongkun Du, Zhineng Chen, Yazhen Xie, Weikang Bai, Hao Feng, Wei Shi, Yuchen Su, Can Huang, and Yu-Gang Jiang. Unirec-0.1 b: Unified text and formula recognition with 0.1 b parameters.arXiv preprint arXiv:2512.21095, 2025

  26. [26]

    Glm-ocr technical report.arXiv preprint arXiv:2603.10910, 2026

    Shuaiqi Duan, Yadong Xue, Weihan Wang, Zhe Su, Huan Liu, Sheng Yang, Guobing Gan, Guo Wang, Zihan Wang, Shengdong Yan, Dexin Jin, Yuxuan Zhang, Guohong Wen, Yanfeng Wang, Yutao Zhang, Xiaohan Zhang, Wenyi Hong, Yukuo Cen, Da Yin, Bin Chen, Wenmeng Yu, Xiaotao Gu, and Jie Tang. Glm-ocr technical report.arXiv preprint arXiv:2603.10910, 2026

  27. [27]

    Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition

    Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang. Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7098–7107, 2021

  28. [28]

    Mathwriting: A dataset for handwrit- ten mathematical expression recognition

    Philippe Gervais, Anastasiia Fadeeva, and Andrii Maksai. Mathwriting: A dataset for handwrit- ten mathematical expression recognition. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V . 2, pages 5459–5469, 2025

  29. [29]

    White, Silvia C

    Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J Szostkiewicz, Dmytro Shved, Gavin J Gyimesi, Jon M Laurent, Samantha M Wright, Muhammed T Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. A multi-agent system for automating scientific discovery.Nature, 655(8122):1–3, 2026

  30. [30]

    Rich feature hierarchies for accurate object detection and semantic segmentation

    Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 580–587, 2014

  31. [31]

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomašev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Riman...

  32. [32]

    Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

    Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. InProceedings of the 23rd international conference on Machine learning, pages 369–376, 2006

  33. [33]

    Speech recognition with deep recurrent neural networks

    Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In2013 IEEE international conference on acoustics, speech and signal processing, pages 6645–6649. Ieee, 2013

  34. [34]

    Unimernet: A universal network for real-world mathematical expression recognition

    Zhuangcheng Gu, Guang Liang, Bin Wang, Zhiyuan Zhao, Qintong Zhang, Weijia Li, Chao Xu, Bo Zhang, Botian Shi, Jiang Wu, Wentao Zhang, and Conghui He. Unimernet: A universal network for real-world mathematical expression recognition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 34106–34115, 2026

  35. [35]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  36. [36]

    Icpr2018 contest on robust reading for multi-type web images

    Mengchao He, Yuliang Liu, Zhibo Yang, Sheng Zhang, Canjie Luo, Feiyu Gao, Qi Zheng, Yongpan Wang, Xin Zhang, and Lianwen Jin. Icpr2018 contest on robust reading for multi-type web images. InProceedings of the International Conference on Pattern Recognition, pages 7–12, 2018

  37. [37]

    Radiov2.5: Improved baselines for agglomerative vision founda- tion models

    Greg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catan- zaro, and Pavlo Molchanov. Radiov2.5: Improved baselines for agglomerative vision founda- tion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22487–22497, 2025

  38. [38]

    mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

    Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. mplug-docowl 1.5: Unified structure learning for ocr-free document understanding. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 3096–3120, 2024

  39. [39]

    Layoutlmv3: Pre-training for document ai with unified text and image masking

    Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. Layoutlmv3: Pre-training for document ai with unified text and image masking. InProceedings of the 30th ACM International Conference on Multimedia, pages 4083–4091, 2022

  40. [40]

    Revisiting scene text recognition: A data perspective

    Qing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu, and Lianwen Jin. Revisiting scene text recognition: A data perspective. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 20543–20554, 2023

  41. [41]

    Icdar 2015 competition on robust reading

    Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny. Icdar 2015 competition on robust reading. InProceedings of the International Conference on Document Analysis and Recognition...

  42. [42]

    Icdar 2013 robust reading competition

    Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. Icdar 2013 robust reading competition. InProceedings of the International Conference on Document Analysis and Recognition, pages 1484–1493, 2013

  43. [43]

    Ocr-free document understanding transformer

    Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. InProceedings of the European Conference on Computer Vision, pages 498–517, 2022

  44. [44]

    Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023

  45. [45]

    Open images v5 text annotation and yet another mask text spotter

    Ilya Krylov, Sergei Nosov, and Vladislav Sovrasov. Open images v5 text annotation and yet another mask text spotter. InAsian Conference on Machine Learning, pages 379–389, 2021

  46. [46]

    Cat-net: Compression artifact tracing network for detection and localization of image splicing

    Myung-Joon Kwon, In-Jae Yu, Seung-Hun Nam, and Heung-Kyu Lee. Cat-net: Compression artifact tracing network for detection and localization of image splicing. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 375–384, 2021

  47. [47]

    Towards better structured and less noisy web data: Oscar with register annotations

    Veronika Laippala, Anna Salmela, Samuel Rönnqvist, Alham Fikri Aji, Li-Hsin Chang, Asma Dhifallah, Larissa Goulart, Henna Kortelainen, Marc Pàmies, Deise Prina Dutra, Valtteri Skantsi, Lintang Sutawika, and Sampo Pyysalo. Towards better structured and less noisy web data: Oscar with register annotations. InProceedings of the Eighth Workshop on Noisy User-...

  48. [48]

    Pix2struct: Screenshot parsing as pretraining for visual language understanding

    Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisen- schlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. Pix2struct: Screenshot parsing as pretraining for visual language understanding. InProceedings of the International Conference on Machine Learning, pages 18893–18912, 2023

  49. [49]

    Building a test collection for complex document information processing

    David Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David Grossman, and Jefferson Heard. Building a test collection for complex document information processing. InProceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval, pages 665–666, 2006

  50. [50]

    Hunyuanocr-1.5: Making lightweight ocr vlms faster and better.arXiv preprint arXiv:2607.04884, 2026

    Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, and Yu Zhou. Hunyuanocr-1.5: Making lightweight ocr vlms faster and better.arXiv prepri...

  51. [51]

    Dit: Self-supervised pre-training for document image transformer

    Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. Dit: Self-supervised pre-training for document image transformer. InProceedings of the 30th ACM International Conference on Multimedia, pages 3530–3539, 2022

  52. [52]

    Trocr: Transformer-based optical character recognition with pre- trained models

    Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei. Trocr: Transformer-based optical character recognition with pre- trained models. InProceedings of the AAAI conference on artificial intelligence, volume 37, pages 13094–13102, 2023

  53. [53]

    Openvision: A fully-open, cost- effective family of advanced vision encoders for multimodal learning

    Xianhang Li, Yanqing Liu, Haoqin Tu, and Cihang Xie. Openvision: A fully-open, cost- effective family of advanced vision encoders for multimodal learning. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 3977–3987, 2025

  54. [54]

    Exploring plain vision transformer backbones for object detection

    Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. InProceedings of the European Conference on Computer Vision, pages 280–296, 2022

  55. [55]

    dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025

    Yumeng Li, Guang Yang, Hao Liu, Bowen Wang, and Colin Zhang. dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025

  56. [56]

    Mdpbench: A benchmark for multilingual document parsing in real-world scenarios.arXiv preprint arXiv:2603.28130, 2026

    Zhang Li, Zhibo Lin, Qiang Liu, Ziyang Zhang, Shuo Zhang, Zidun Guo, Jiajun Song, Jiarui Zhang, Xiang Bai, and Yuliang Liu. Mdpbench: A benchmark for multilingual document parsing in real-world scenarios.arXiv preprint arXiv:2603.28130, 2026

  57. [57]

    Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm.arXiv preprint arXiv:2506.05218, 2025

    Zhang Li, Yuliang Liu, Qiang Liu, Zhiyin Ma, Ziyang Zhang, Shuo Zhang, Biao Yang, Zidun Guo, Jiarui Zhang, Xinyu Wang, and Xiang Bai. Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm.arXiv preprint arXiv:2506.05218, 2025

  58. [58]

    Monkey: Image resolution and text label are important things for large multi-modal models

    Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Monkey: Image resolution and text label are important things for large multi-modal models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26763–26773, 2024

  59. [59]

    Real-time scene text detection with differentiable binarization

    Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen, and Xiang Bai. Real-time scene text detection with differentiable binarization. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11474–11481, 2020

  60. [60]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024

  61. [61]

    Xiaohong Liu, Yaojie Liu, Jun Chen, and Xiaoming Liu. Pscc-net: Progressive spatio-channel correlation network for image manipulation detection and localization.IEEE Transactions on Circuits and Systems for Video Technology, 32(11):7505–7517, 2022

  62. [62]

    Multi-scenario overlapping text segmen- tation with depth awareness

    Yang Liu, Xudong Xie, Yuliang Liu, and Xiang Bai. Multi-scenario overlapping text segmen- tation with depth awareness. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 17454–17463, 2025

  63. [63]

    Openvision 2: A family of generative pretrained visual encoders for multimodal learning

    Yanqing Liu, Xianhang Li, Letian Zhang, Zirui Wang, Zeyu Zheng, Yuyin Zhou, and Cihang Xie. Openvision 2: A family of generative pretrained visual encoders for multimodal learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 39164–39174, 2026

  64. [64]

    Multilingual denoising pre-training for neural machine translation.Transactions of the Association for Computational Linguistics, 8:726–742, 2020

    Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. Multilingual denoising pre-training for neural machine translation.Transactions of the Association for Computational Linguistics, 8:726–742, 2020

  65. [65]

    Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019

    Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang. Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019

  66. [66]

    Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai. Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

  67. [67]

    Textmonkey: An ocr-free large multimodal model for understanding document.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 48(5):6008–6019, 2026

    Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. Textmonkey: An ocr-free large multimodal model for understanding document.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 48(5):6008–6019, 2026

  68. [68]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021

  69. [69]

    A convnet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022

  70. [70]

    Towards end-to-end unified scene text detection and layout analysis

    Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. Towards end-to-end unified scene text detection and layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022

  71. [71]

    Toward real text manipulation detection: New dataset and new solution.Pattern Recognition, 157:110828, 2025

    Dongliang Luo, Yuliang Liu, Rui Yang, Xianjin Liu, Jishen Zeng, Yu Zhou, and Xiang Bai. Toward real text manipulation detection: New dataset and new solution.Pattern Recognition, 157:110828, 2025

  72. [72]

    Chartqa: A benchmark for question answering about charts with visual and logical reasoning

    Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguistics: ACL 2022, pages 2263–2279, 2022

  73. [73]

    Infographicvqa

    Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawa- har. Infographicvqa. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022

  74. [74]

    Docvqa: A dataset for vqa on document images

    Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2200–2209, 2021

  75. [75]

    Scene text recognition using higher order language priors

    Anand Mishra, Karteek Alahari, and CV Jawahar. Scene text recognition using higher order language priors. InBMVC-British Machine Vision Conference, 2012

  76. [76]

    Mineru2.5: A decoupled vision-language model for efficient high-resolution document parsing

    Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang, Linke Ouyang, Zhiyuan Zhao, Tao Chu, Tianyao He, Fan Wu, Qintong Zhang, Zhenjiang Jin, Guang Liang, Rui Zhang, Wen- zheng Zhang, Yuan Qu, Zhifei Ren, Yuefeng Sun, Zirui Tang, Boyu Niu, Yuanhong Zheng, Dongsheng Ma, Ziyang Miao, Hejun Dong, Siyi Qian, Junyuan Zhang, Fangdong Wang, Jingzhou Chen, Xiaomeng Zhao...

  77. [77]

    Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jégou, Julien Mairal, Patrick La...

  78. [78]

    Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations

    Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, and Conghui He. Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations. InProceedings of the IEEE/CVF Con...

  79. [79]

    Compositional semantic parsing on semi-structured tables

    Panupong Pasupat and Percy Liang. Compositional semantic parsing on semi-structured tables. InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1470–1480, 2015

  80. [80]

    Doclaynet: A large human-annotated dataset for document-layout segmentation

    Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S Nassar, and Peter Staar. Doclaynet: A large human-annotated dataset for document-layout segmentation. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3743–3751, 2022

Showing first 80 references.