REVIEW 3 major objections 6 minor 137 references
A document-native vision encoder, trained with text generation plus pixel reconstruction, transfers across OCR, parsing, and understanding better than natural-image foundations.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 04:44 UTC pith:NDLC6ZRV
load-bearing objection Solid document-native encoder with clean multi-task transfer; the MDPBench SOTA headline is real but not cleanly encoder-attributable. the 3 major comments →
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Document-oriented pretraining that jointly optimizes image-to-text generation and pixel-level reconstruction produces transferable character-level visual representations: as a backbone swap it improves five document analysis tasks, and as a frozen encoder with a lightweight language model it yields a 0.7B parser that reaches open-source state of the art on multilingual MDPBench while also outperforming CLIP, DINO, and SAM counterparts on eight document-understanding benchmarks under identical training.
What carries the argument
Dual-objective document pretraining on MonkeyDoc v2 (113M images, 17 languages): image-to-text generation aligns visual tokens with textual content, while pixel-level reconstruction (MSE, optionally edge- and distance-aware) forces the encoder to retain strokes, glyphs, and layout that text supervision alone can discard.
Load-bearing premise
The large training labels from multi-expert agreement and automatic layout filters are clean enough that measured gains reflect better visual features rather than shared errors between those labels and the evaluation stack.
What would settle it
Under fully matched data, optimization, and decoding, freeze MonkeyOCRv2 and a strong natural-image encoder of similar size, train the same lightweight document parser, and check whether the document-native encoder still wins on photographed non-Latin pages of MDPBench and on scrambled-text recognition at low resolution; a clear loss would falsify the claim that the dual objective, not scale or task setup, is doing the work.
If this is right
- Document systems can replace ImageNet, CLIP, DINO, or SAM backbones with a compact document encoder and expect gains without rewriting the rest of the pipeline.
- A frozen ~0.1B document vision encoder plus a ~0.6B language model is enough for competitive multilingual parsing, so large general VLMs are not required for that task class.
- Pixel reconstruction should reduce reliance on language priors when text is scrambled, low-resolution, or deliberately conflicted with linguistic expectations.
- Future document foundation work can treat text strokes and layout as first-class visual targets rather than side effects of semantic alignment.
Where Pith is reading between the lines
- The same dual recipe may help other dense-symbol domains (sheet music, circuit diagrams, engineering drawings) where global semantic encoders discard local marks.
- If reconstruction is what narrows the semantic–scrambled gap, progressive post-training of the language head alone may not close remaining gaps on saturated parsing benches without stronger visual evidence.
- Balancing MonkeyDoc-style supervision toward low-resource and historical scripts would be a direct test of whether the method generalizes beyond high-resource languages that dominate the current mix.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MonkeyOCRv2 proposes a document-native visual encoder pretrained on MonkeyDoc v2 (113M images, 17 languages) with a joint objective of image-to-text generation and pixel-level reconstruction (Eqs. 1–11). The encoder is evaluated as a backbone substitution on five document analysis tasks (text recognition, formula recognition, text detection, tampering detection, overlapping text segmentation) and, frozen, as the vision tower of lightweight VLMs for document parsing and understanding. The paper reports consistent gains from encoder replacement (Tabs. 2–5, Fig. 4), open-source SOTA on MDPBench for a 0.7B frozen-encoder parser (+2.8 over 3B dots.mocr with a much smaller ViT; Tab. 6), and superior document-understanding scores versus CLIP/DINO/SAM/OpenVision under matched LLM, data, and training (Tab. 8). Reconstruction is further supported by scrambled-text and CHAOS-Bench analyses (Sec. 5.1–5.2).
Significance. If the transfer results hold under the stated controls, the work is a substantial contribution to document AI: it argues, with multi-task evidence, that character-level document pretraining can serve as a foundation rather than a domain adaptation of natural-image encoders. Strengths include a large multilingual corpus, a dual-objective recipe with explicit reconstruction ablations (MSE and structure-aware variants in Tab. 8; scrambled-text gap narrowing in Fig. 5), controlled frozen-encoder VLM comparisons on eight benchmarks (Tab. 8), and honest system-level caveats for OmniDocBench (Sec. 4.6, Tab. 7). The five-task backbone-swap protocol is particularly useful for the community. Code and data release is promised, which would further raise impact.
major comments (3)
- Sec. 4.6 and Tab. 6 present MonkeyOCRv2-Parsing as open-source SOTA on MDPBench (+2.8 over dots.mocr, ~11× smaller vision encoder). The same section correctly notes that OmniDocBench (Tab. 7) is system-level and that encoder attribution should use Tab. 8. That caveat is not applied to MDPBench: the parser couples a frozen encoder to autoregressive layout prediction, per-element re-crop recognition, and assembly, with no matched experiment that freezes a strong alternative encoder (e.g., OpenVision-B, RADIOv2.5-B, SigLIP 2) inside the identical parsing pipeline, data, and training recipe. Without that control, the headline +2.8 and size framing can credit pipeline design and training as much as document-native pretraining. Please either add the matched encoder swap in the parsing stack or reframe Tab. 6 as a system result and rest the encoder claim primarily on Tabs. 2–5 and Tab. 8.
- Sec. 3.1 (Expert Model Labeling; Data Filtering) relies on multi-expert OCR agreement and LLM layout/reading-order filters for large-scale supervision. The residual error structure of those automatic labels is not quantified against human gold or against the evaluation stacks (including MDPBench, which shares research-lineage with prior MonkeyOCR work). If label failure modes correlate with the pretraining objective or with same-lineage benchmarks, measured gains partly reflect label–model correlation. A short audit—e.g., human agreement rates on a stratified sample, or performance when pretraining only on fully public human-annotated subsets—would strengthen the claim that gains are representation quality rather than supervision artifacts.
- Tab. 8 is the cleanest encoder-level comparison, but input configurations differ substantially (App. D: CLIP 196 tokens vs SAM 4096 vs MonkeyOCRv2 ~1082). The paper states each encoder uses its native setting, which is reasonable, yet token budget and resolution are known confounders for document VQA. A sensitivity check that equalizes approximate visual-token count (or reports a fixed-token budget ablation for the top baselines) would make the 13.2-point gap over OpenVision-B more attributable to pretraining rather than resolution policy.
minor comments (6)
- Abstract and Fig. 2(a) lead with the MDPBench SOTA and 11× smaller encoder; after addressing the major comment on attribution, align abstract wording with the revised claim so abstract and Sec. 4.6 do not over-promise encoder-only causality.
- Eq. (11) sets λ=1.0 with α, β, T, τ fixed without tuning (Sec. 3.2). A brief sensitivity note (even one-dimensional in λ) would help readers assess robustness of the dual-objective balance.
- Table 1 lists MonkeyDoc v2 as 113M multi-type documents / 17 languages; App. A shows strong English/Chinese skew. Mentioning this imbalance earlier (not only in Limitations) would set expectations for low-resource scripts.
- Fig. 4 uses bar charts without numeric tables in the main text; adding exact F-measures in a small table or appendix would aid citation and reproducibility.
- Typo/consistency: abstract says “previous best 3B dots.mocr” while related work and Tab. 6 use “dots.mocr”; unify naming. Also “UniMERNet-T to outperform the 325M UniMERNet-B” is clear in Tab. 3 but the abstract’s “enabling the 110M” phrasing could state ExpRate/CDM explicitly.
- Sec. 5.1 correctly treats the semantic–scrambled gap as an operational proxy; consider moving one sentence of that caveat into the figure caption of Fig. 5 so casual readers do not over-interpret the gap as a pure hallucination metric.
Circularity Check
No derivation-by-construction circularity; mild same-lineage risk on MDPBench SOTA framing, while backbone swaps and matched VLM controls remain independent.
specific steps
-
self citation load bearing
[Sec. 4.6 Document Parsing; Tab. 6 MDPBench; abstract SOTA claim]
"Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench... surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11× smaller. ... While MDPBench originates from the same research line as our prior benchmarks, the encoder, the LLM, and the data pipeline evaluated here are independent of its construction"
The headline open-source SOTA and 11×-smaller-encoder framing rest on MDPBench, which the paper itself notes comes from the same research line. The paper does not freeze a strong alternative encoder inside the exact autoregressive-layout + re-crop parsing pipeline, so the load-bearing +2.8 claim partly leans on same-lineage benchmark construction rather than a fully independent external test. This is mild self-citation risk, not a definitional reduction of the pretraining objective.
full rationale
MonkeyOCRv2 is an empirical CV foundation-model paper, not a first-principles derivation. The dual objective L_pretrain = L_text + λ L_rec is a training recipe, not a claim that one quantity is mathematically forced by another. Gains are measured by encoder substitution into CRNN/PARSeq, UniMERNet-T, DBNet/PSENet/DPText-DETR, FFDN, Mask2Former/MOTS (Tabs. 2–5, Fig. 4) and by frozen-encoder VLMs under fixed LLM/data/optimization (Tab. 8). Reconstruction ablations (Fig. 5, Tab. 8 baseline vs MSE vs structure-aware, Tab. 9 CHAOS-Bench) compare trained variants rather than renaming a fit as a prediction. MDPBench shares authorship lineage with prior MonkeyOCR work, and Sec. 4.6/Tab. 7 explicitly warn that OmniDocBench is system-level, so the +2.8 open-source SOTA headline is not cleanly encoder-attributed without a matched alternative-encoder ablation inside the same parsing pipeline—this is attribution softness, not circular reduction of a claimed derivation. No self-definitional equations, fitted-input-as-prediction, uniqueness theorems, or ansatz-via-citation chains appear. Score 2 for mild self-citation load on the parsing headline only.
Axiom & Free-Parameter Ledger
free parameters (4)
- λ (reconstruction weight in L_pretrain)
- α, β (structure-aware reconstruction weights)
- T, τ (distance-to-edge iterations and edge temperature)
- peak learning rate and batch size for pretraining
axioms (4)
- domain assumption Pixel reconstruction (MSE and optional edge/distance matching) forces the encoder to retain character strokes and layout that pure text supervision discards.
- domain assumption Multi-expert agreement among OCR systems plus LLM layout/order filters yields sufficiently accurate labels for large-scale pretraining.
- domain assumption Frozen-encoder transfer under matched LLM, data, and optimization isolates visual representation quality.
- standard math Standard transformer/ViT training dynamics and cross-entropy + MSE optimization are valid for learning transferable document features.
invented entities (2)
-
MonkeyDoc v2 corpus
no independent evidence
-
MonkeyOCRv2 dual-objective encoder family (S/B/AS)
no independent evidence
read the original abstract
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.
Figures
Reference graph
Works this paper leans on
-
[1]
Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan ...
Pith/arXiv arXiv 2025
-
[2]
Qwen2.5-vl technical report.arXiv preprint arXiv:2502.13923, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, 19 Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report...
Pith/arXiv arXiv 2025
-
[3]
BEiT: BERT pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. BEiT: BERT pre-training of image transformers. InInternational Conference on Learning Representations, 2022
2022
-
[4]
Scene text recognition with permuted autoregressive sequence models
Darwin Bautista and Rowel Atienza. Scene text recognition with permuted autoregressive sequence models. InProceedings of the European Conference on Computer Vision, pages 178–196, 2022
2022
-
[5]
Nougat: Neu- ral optical understanding for academic documents
Lukas Blecher, Guillem Cucurull Preixens, Thomas Scialom, and Robert Stojnic. Nougat: Neu- ral optical understanding for academic documents. InInternational Conference on Learning Representations, 2024
2024
-
[6]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9650–9660, 2021
2021
-
[7]
Encoder-decoder with atrous separable convolution for semantic image segmentation
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision, pages 801–818, 2018
2018
-
[8]
Enhancing tampered text detection through frequency feature fusion and decomposition
Zhongxi Chen, Shen Chen, Taiping Yao, Ke Sun, Shouhong Ding, Xianming Lin, Liujuan Cao, and Rongrong Ji. Enhancing tampered text detection through frequency feature fusion and decomposition. InProceedings of the European Conference on Computer Vision, pages 200–217, 2024
2024
-
[9]
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1290–1299, 2022
2022
-
[10]
Per-pixel classification is not all you need for semantic segmentation.Advances in Neural Information Processing Systems, 34:17864–17875, 2021
Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation.Advances in Neural Information Processing Systems, 34:17864–17875, 2021
2021
-
[11]
M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis
Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang, Qiyuan Zhu, Zecheng Xie, Jing Li, Kai Ding, and Lianwen Jin. M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 151...
2023
-
[12]
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2818–2829, 2023
2023
-
[13]
Total-text: A comprehensive dataset for scene text detection and recognition
Chee Kheng Ch’ng and Chee Seng Chan. Total-text: A comprehensive dataset for scene text detection and recognition. InProceedings of the International Conference on Document Analysis and Recognition, pages 935–942, 2017
2017
-
[14]
Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art
Chee Kheng Chng, Yuliang Liu, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, and Lianwen Jin. Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art. InProceedings of the International Conference on Document Analysis and Recognition, pag...
2019
-
[15]
Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Yi Liu, Dianhai Yu, and Yanjun Ma. Paddleocr-vl-1.5: Towards a multi-task 0.9b vlm for robust in-the-wild document parsing.arXiv preprint arXiv:2601.21957, 2026
Pith/arXiv arXiv 2026
-
[16]
Boosting document parsing efficiency and performance with coarse-to-fine visual processing
Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, Yue Zhang, Yubo Zhang, Jing Zhang, Jun Zhang, Xing Wei, Yi Liu, Dianhai Yu, and Yanjun Ma. Boosting document parsing efficiency and performance with coarse-to-fine visual processing. InProceedings of the IEEE/CVF Conference on Comp...
2026
-
[17]
Vision grid transformer for document layout analysis
Cheng Da, Chuwei Luo, Qi Zheng, and Cong Yao. Vision grid transformer for document layout analysis. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 19462–19472, 2023
2023
-
[18]
Imagenet: A large- scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009
2009
-
[19]
Decaf: A deep convolutional activation feature for generic visual recognition
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. InProceedings of the International Conference on Machine Learning, pages 647–655, 2014
2014
-
[20]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. InInternational Conference on Learning Representations, 2021
2021
-
[21]
Out of length text recognition with sub-string matching
Yongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao, and Yu-Gang Jiang. Out of length text recognition with sub-string matching. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 2798–2806, 2025
2025
-
[22]
Context perception parallel decoder for scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Chenxia Li, Yuning Du, and Yu-Gang Jiang. Context perception parallel decoder for scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
2025
-
[23]
Instruction-guided scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4):2723–2738, 2025
Yongkun Du, Zhineng Chen, Yuchen Su, Caiyan Jia, and Yu-Gang Jiang. Instruction-guided scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4):2723–2738, 2025
2025
-
[24]
Svtrv2: Ctc beats encoder-decoder models in scene text recognition
Yongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia, and Yu-Gang Jiang. Svtrv2: Ctc beats encoder-decoder models in scene text recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 20147–20156, 2025
2025
-
[25]
Yongkun Du, Zhineng Chen, Yazhen Xie, Weikang Bai, Hao Feng, Wei Shi, Yuchen Su, Can Huang, and Yu-Gang Jiang. Unirec-0.1 b: Unified text and formula recognition with 0.1 b parameters.arXiv preprint arXiv:2512.21095, 2025
Pith/arXiv arXiv 2025
-
[26]
Glm-ocr technical report.arXiv preprint arXiv:2603.10910, 2026
Shuaiqi Duan, Yadong Xue, Weihan Wang, Zhe Su, Huan Liu, Sheng Yang, Guobing Gan, Guo Wang, Zihan Wang, Shengdong Yan, Dexin Jin, Yuxuan Zhang, Guohong Wen, Yanfeng Wang, Yutao Zhang, Xiaohan Zhang, Wenyi Hong, Yukuo Cen, Da Yin, Bin Chen, Wenmeng Yu, Xiaotao Gu, and Jie Tang. Glm-ocr technical report.arXiv preprint arXiv:2603.10910, 2026
arXiv 2026
-
[27]
Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition
Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang. Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7098–7107, 2021
2021
-
[28]
Mathwriting: A dataset for handwrit- ten mathematical expression recognition
Philippe Gervais, Anastasiia Fadeeva, and Andrii Maksai. Mathwriting: A dataset for handwrit- ten mathematical expression recognition. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V . 2, pages 5459–5469, 2025
2025
-
[29]
White, Silvia C
Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J Szostkiewicz, Dmytro Shved, Gavin J Gyimesi, Jon M Laurent, Samantha M Wright, Muhammed T Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. A multi-agent system for automating scientific discovery.Nature, 655(8122):1–3, 2026
2026
-
[30]
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 580–587, 2014
2014
-
[31]
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomašev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Riman...
2026
-
[32]
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. InProceedings of the 23rd international conference on Machine learning, pages 369–376, 2006
2006
-
[33]
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In2013 IEEE international conference on acoustics, speech and signal processing, pages 6645–6649. Ieee, 2013
2013
-
[34]
Unimernet: A universal network for real-world mathematical expression recognition
Zhuangcheng Gu, Guang Liang, Bin Wang, Zhiyuan Zhao, Qintong Zhang, Weijia Li, Chao Xu, Bo Zhang, Botian Shi, Jiang Wu, Wentao Zhang, and Conghui He. Unimernet: A universal network for real-world mathematical expression recognition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 34106–34115, 2026
2026
-
[35]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016
2016
-
[36]
Icpr2018 contest on robust reading for multi-type web images
Mengchao He, Yuliang Liu, Zhibo Yang, Sheng Zhang, Canjie Luo, Feiyu Gao, Qi Zheng, Yongpan Wang, Xin Zhang, and Lianwen Jin. Icpr2018 contest on robust reading for multi-type web images. InProceedings of the International Conference on Pattern Recognition, pages 7–12, 2018
2018
-
[37]
Radiov2.5: Improved baselines for agglomerative vision founda- tion models
Greg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catan- zaro, and Pavlo Molchanov. Radiov2.5: Improved baselines for agglomerative vision founda- tion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22487–22497, 2025
2025
-
[38]
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. mplug-docowl 1.5: Unified structure learning for ocr-free document understanding. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 3096–3120, 2024
2024
-
[39]
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. Layoutlmv3: Pre-training for document ai with unified text and image masking. InProceedings of the 30th ACM International Conference on Multimedia, pages 4083–4091, 2022
2022
-
[40]
Revisiting scene text recognition: A data perspective
Qing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu, and Lianwen Jin. Revisiting scene text recognition: A data perspective. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 20543–20554, 2023
2023
-
[41]
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny. Icdar 2015 competition on robust reading. InProceedings of the International Conference on Document Analysis and Recognition...
2015
-
[42]
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. Icdar 2013 robust reading competition. InProceedings of the International Conference on Document Analysis and Recognition, pages 1484–1493, 2013
2013
-
[43]
Ocr-free document understanding transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. InProceedings of the European Conference on Computer Vision, pages 498–517, 2022
2022
-
[44]
Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023
2023
-
[45]
Open images v5 text annotation and yet another mask text spotter
Ilya Krylov, Sergei Nosov, and Vladislav Sovrasov. Open images v5 text annotation and yet another mask text spotter. InAsian Conference on Machine Learning, pages 379–389, 2021
2021
-
[46]
Cat-net: Compression artifact tracing network for detection and localization of image splicing
Myung-Joon Kwon, In-Jae Yu, Seung-Hun Nam, and Heung-Kyu Lee. Cat-net: Compression artifact tracing network for detection and localization of image splicing. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 375–384, 2021
2021
-
[47]
Towards better structured and less noisy web data: Oscar with register annotations
Veronika Laippala, Anna Salmela, Samuel Rönnqvist, Alham Fikri Aji, Li-Hsin Chang, Asma Dhifallah, Larissa Goulart, Henna Kortelainen, Marc Pàmies, Deise Prina Dutra, Valtteri Skantsi, Lintang Sutawika, and Sampo Pyysalo. Towards better structured and less noisy web data: Oscar with register annotations. InProceedings of the Eighth Workshop on Noisy User-...
2022
-
[48]
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisen- schlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. Pix2struct: Screenshot parsing as pretraining for visual language understanding. InProceedings of the International Conference on Machine Learning, pages 18893–18912, 2023
2023
-
[49]
Building a test collection for complex document information processing
David Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David Grossman, and Jefferson Heard. Building a test collection for complex document information processing. InProceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval, pages 665–666, 2006
2006
-
[50]
Hunyuanocr-1.5: Making lightweight ocr vlms faster and better.arXiv preprint arXiv:2607.04884, 2026
Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang, Hao Feng, Yongkun Du, Binghong Wu, Zheng Ruan, Zhiqiong Lu, Liang Wu, Pengyuan Lyu, Huawen Shen, Zibin Lin, Shijing Hu, Jieneng Yang, Hongbing Wen, Guanghua Yu, Hong Liu, Bochao Wang, Can Ma, Han Hu, Chengquan Zhang, and Yu Zhou. Hunyuanocr-1.5: Making lightweight ocr vlms faster and better.arXiv prepri...
Pith/arXiv arXiv 2026
-
[51]
Dit: Self-supervised pre-training for document image transformer
Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. Dit: Self-supervised pre-training for document image transformer. InProceedings of the 30th ACM International Conference on Multimedia, pages 3530–3539, 2022
2022
-
[52]
Trocr: Transformer-based optical character recognition with pre- trained models
Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei. Trocr: Transformer-based optical character recognition with pre- trained models. InProceedings of the AAAI conference on artificial intelligence, volume 37, pages 13094–13102, 2023
2023
-
[53]
Openvision: A fully-open, cost- effective family of advanced vision encoders for multimodal learning
Xianhang Li, Yanqing Liu, Haoqin Tu, and Cihang Xie. Openvision: A fully-open, cost- effective family of advanced vision encoders for multimodal learning. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 3977–3987, 2025
2025
-
[54]
Exploring plain vision transformer backbones for object detection
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. InProceedings of the European Conference on Computer Vision, pages 280–296, 2022
2022
-
[55]
Yumeng Li, Guang Yang, Hao Liu, Bowen Wang, and Colin Zhang. dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025
arXiv 2025
-
[56]
Zhang Li, Zhibo Lin, Qiang Liu, Ziyang Zhang, Shuo Zhang, Zidun Guo, Jiajun Song, Jiarui Zhang, Xiang Bai, and Yuliang Liu. Mdpbench: A benchmark for multilingual document parsing in real-world scenarios.arXiv preprint arXiv:2603.28130, 2026
arXiv 2026
-
[57]
Zhang Li, Yuliang Liu, Qiang Liu, Zhiyin Ma, Ziyang Zhang, Shuo Zhang, Biao Yang, Zidun Guo, Jiarui Zhang, Xinyu Wang, and Xiang Bai. Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm.arXiv preprint arXiv:2506.05218, 2025
arXiv 2025
-
[58]
Monkey: Image resolution and text label are important things for large multi-modal models
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. Monkey: Image resolution and text label are important things for large multi-modal models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26763–26773, 2024
2024
-
[59]
Real-time scene text detection with differentiable binarization
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen, and Xiang Bai. Real-time scene text detection with differentiable binarization. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11474–11481, 2020
2020
-
[60]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024
2024
-
[61]
Xiaohong Liu, Yaojie Liu, Jun Chen, and Xiaoming Liu. Pscc-net: Progressive spatio-channel correlation network for image manipulation detection and localization.IEEE Transactions on Circuits and Systems for Video Technology, 32(11):7505–7517, 2022
2022
-
[62]
Multi-scenario overlapping text segmen- tation with depth awareness
Yang Liu, Xudong Xie, Yuliang Liu, and Xiang Bai. Multi-scenario overlapping text segmen- tation with depth awareness. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 17454–17463, 2025
2025
-
[63]
Openvision 2: A family of generative pretrained visual encoders for multimodal learning
Yanqing Liu, Xianhang Li, Letian Zhang, Zirui Wang, Zeyu Zheng, Yuyin Zhou, and Cihang Xie. Openvision 2: A family of generative pretrained visual encoders for multimodal learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 39164–39174, 2026
2026
-
[64]
Multilingual denoising pre-training for neural machine translation.Transactions of the Association for Computational Linguistics, 8:726–742, 2020
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. Multilingual denoising pre-training for neural machine translation.Transactions of the Association for Computational Linguistics, 8:726–742, 2020
2020
-
[65]
Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang. Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019
2019
-
[66]
Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai. Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024
2024
-
[67]
Textmonkey: An ocr-free large multimodal model for understanding document.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 48(5):6008–6019, 2026
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. Textmonkey: An ocr-free large multimodal model for understanding document.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 48(5):6008–6019, 2026
2026
-
[68]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021
2021
-
[69]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11976–11986, 2022
2022
-
[70]
Towards end-to-end unified scene text detection and layout analysis
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. Towards end-to-end unified scene text detection and layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022
2022
-
[71]
Toward real text manipulation detection: New dataset and new solution.Pattern Recognition, 157:110828, 2025
Dongliang Luo, Yuliang Liu, Rui Yang, Xianjin Liu, Jishen Zeng, Yu Zhou, and Xiang Bai. Toward real text manipulation detection: New dataset and new solution.Pattern Recognition, 157:110828, 2025
2025
-
[72]
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguistics: ACL 2022, pages 2263–2279, 2022
2022
-
[73]
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawa- har. Infographicvqa. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022
2022
-
[74]
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2200–2209, 2021
2021
-
[75]
Scene text recognition using higher order language priors
Anand Mishra, Karteek Alahari, and CV Jawahar. Scene text recognition using higher order language priors. InBMVC-British Machine Vision Conference, 2012
2012
-
[76]
Mineru2.5: A decoupled vision-language model for efficient high-resolution document parsing
Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang, Linke Ouyang, Zhiyuan Zhao, Tao Chu, Tianyao He, Fan Wu, Qintong Zhang, Zhenjiang Jin, Guang Liang, Rui Zhang, Wen- zheng Zhang, Yuan Qu, Zhifei Ren, Yuefeng Sun, Zirui Tang, Boyu Niu, Yuanhong Zheng, Dongsheng Ma, Ziyang Miao, Hejun Dong, Siyi Qian, Junyuan Zhang, Fangdong Wang, Jingzhou Chen, Xiaomeng Zhao...
2026
-
[77]
Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jégou, Julien Mairal, Patrick La...
Pith/arXiv arXiv 2023
-
[78]
Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, and Conghui He. Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations. InProceedings of the IEEE/CVF Con...
2025
-
[79]
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang. Compositional semantic parsing on semi-structured tables. InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1470–1480, 2015
2015
-
[80]
Doclaynet: A large human-annotated dataset for document-layout segmentation
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S Nassar, and Peter Staar. Doclaynet: A large human-annotated dataset for document-layout segmentation. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 3743–3751, 2022
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.