Pith. sign in

REVIEW 2 major objections 6 minor 1 cited by

A lightweight OCR vision-language model can be made much faster on long structured outputs and stronger on long-tail tasks without redesigning its backbone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 11:57 UTC pith:MPNCYOT6

load-bearing objection Solid systems upgrade: DFlash speedups plus agentic long-tail data on a frozen 1B backbone, with one real gap on AR-vs-DFlash accuracy parity. the 2 major comments →

arxiv 2607.04884 v1 pith:MPNCYOT6 submitted 2026-07-06 cs.CV

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

classification cs.CV
keywords OCRvision-language modelsspeculative decodingdocument parsingmultilingual OCRlightweight VLMagentic data constructiontext spotting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that end-to-end OCR does not need a new architecture to become practical at scale. Starting from an already compact OCR-specialized vision-language model, it targets two bottlenecks: slow decoding of long structured outputs such as dense pages, tables, and formulas, and weak coverage of long-tail skills like ancient scripts, low-resource languages, multi-image document questions, and faithful seen-text parsing. Speed comes from adapting DFlash speculative decoding so a small draft model proposes token blocks that the frozen target model verifies in one pass, cutting latency while aiming to keep the same answers. Capability growth comes from Agentic Data Flow, which turns measured model failures into executable data pipelines for search, cleaning, synthesis, and iteration, plus upgraded pretraining and post-training that raise resolution and context length. The reported result is a single lightweight model that is both the fastest in its class on the authors’ latency tests and competitive or leading across broad OCR benchmarks, with weights and training code planned for release.

Core claim

HunyuanOCR-1.5 keeps the validated lightweight end-to-end backbone and still becomes the fastest lightweight OCR VLM in the authors’ comparisons—6.37× under Transformers and 2.14× under vLLM via DFlash—while ranking among top-tier end-to-end systems on OmniDocBench v1.6 and setting new performance milestones on ancient-script OCR, fine-grained chart and table parsing, multi-image QA, low-resource multilingual parsing, and document hallucination evaluation.

What carries the argument

DFlash: a block-diffusion draft model that proposes multiple candidate tokens in one parallel step, then has the frozen target model verify them while preserving the target output distribution, which is especially useful for regular structured OCR text. Agentic Data Flow: an agent loop that converts model weaknesses into data requirements and automates material collection, quality checks, and pipeline writing for long-tail training data.

Load-bearing premise

That DFlash’s draft-and-verify process really leaves the model’s output distribution unchanged on the OCR tasks where accuracy is reported, so the speed numbers and the quality numbers can be claimed together.

What would settle it

Rerun OmniDocBench and the long-tail suites with DFlash on and off under identical hardware, prompts, and decoding settings; check whether page scores and accepted token sequences stay effectively identical and whether the published latency and throughput ratios still appear.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Long structured OCR generation—tables, formulas, dense Markdown—can be decoded several times faster without switching to a multi-stage layout pipeline.
  • A ~1B-class end-to-end OCR model can expand into ancient scripts, low-resource languages, and multi-page document QA without a full backbone redesign.
  • Agent-driven weakness-to-data loops can systematically supply the training material that closes long-tail OCR gaps.
  • Both high-throughput server serving and local PC-side OCR deployment become more realistic for production use of a single unified model.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Drafting methods that exploit local regularity in HTML, LaTeX, and Markdown may transfer to other long structured generation tasks outside OCR.
  • If agentic data construction stays cheaper than architecture churn, specialized VLMs may improve mainly through closed-loop data rather than new backbones.
  • The still-low absolute recall on seen-but-meaningless word preservation suggests language priors remain a hard failure mode even when overall parsing scores look strong.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. HunyuanOCR-1.5 upgrades the lightweight end-to-end OCR VLM HunyuanOCR-1.0 without redesigning the backbone. Efficiency is improved by adapting DFlash block-diffusion speculative decoding (B=16 draft model, ~90.7M params, initialized from the last five target decoder layers; Eqs. 1–2) to long structured OCR outputs, reporting 6.37× Transformers and 2.14× vLLM speedups (Tables 2–6) while claiming distribution preservation. Capability is extended via Agentic Data Flow (agent-driven material search, cleaning, and pipeline iteration for low-resource, ancient-script, and multi-image data), Stage3 pretraining upgrades (4K resolution, 128K context), refined SFT, and capability-routed RL (IcePop/GRPO-style with parse, judge, and degeneration rewards). Evaluation uses a capability tree spanning OmniDocBench v1.6, Chronicles-OCR, ChartArena, TableVerse-5K, DUDE, MORE, CHAOS-Bench, and several inherited/in-house tasks, with planned weight and code release.

Significance. If the speed and accuracy claims hold under the stated deployment settings, the paper is a strong systems contribution: a compact (~1B) end-to-end OCR VLM that is simultaneously among the fastest reported lightweight OCR VLMs and competitive or leading on open document-parsing and long-tail OCR benchmarks. The DFlash adaptation is well-motivated for HTML/LaTeX-style OCR outputs; the multi-axis speed analysis (length, content type, concurrency) is unusually thorough for this area. Agentic Data Flow and CHAOS-Bench are useful methodological additions for long-tail data construction and seen-text faithfulness. Explicit plans to release weights and training code further increase community value. The work is incremental relative to HunyuanOCR-1.0 but practically significant for deployable OCR VLMs.

major comments (2)
  1. [Sec. 3.2; Tables 2–6] Sec. 3.2 and Tables 2–6: The central “faster and better” claim treats DFlash as distribution-preserving by construction of draft-then-verify, yet all accuracy tables (OmniDocBench, Chronicles-OCR, ChartArena, MORE, CHAOS-Bench, etc.) appear to report a single decoding mode without an AR vs DFlash side-by-side on the same 930-sample OmniDocBench split (or on long-tail sets). Speculative decoding preserves the target distribution only if verification is implemented correctly and acceptance does not interact with decoding hyperparameters in practice. Please add a compact accuracy parity table (AR vs DFlash) on OmniDocBench overall/per-dimension scores and at least one long structured subset (tables/formulas), plus any decoding settings that differ.
  2. [Secs. 4–5; Sec. 7.2] Secs. 4–5 and 7.2–7.3: Capability gains (ancient scripts, MORE, ChartArena, TableVerse-5K, DUDE, OmniDocBench) are attributed jointly to Agentic Data Flow, Stage3 re-planning, SFT cleanup, and RL, but the manuscript lacks controlled ablations that isolate these factors (e.g., Stage3 without new agentic data; SFT-only vs SFT+RL on the same held-out hard set). Without this, the load-bearing claim that Agentic Data Flow “substantially improves” long-tail capabilities remains only partially supported. A minimal ablation suite on 2–3 representative benchmarks would make the causal story defensible.
minor comments (6)
  1. [Table 12; Sec. 7.3] Table 12 / OmniDocBench formula discussion: The authors note that multi-line formula GT matching may underestimate end-to-end models. This caveat should be flagged in the table caption or main text near the formula CDM number so readers do not over-interpret the formula gap vs. some competitors.
  2. [Sec. 6.2; Table 11] CHAOS-Bench (Sec. 6.2, Table 11): Absolute page-average recall remains low (14.15). The construction protocol is clear, but please state dataset size N, perturbation rate, and whether the benchmark will be released with the code so the “milestone” claim is reproducible.
  3. [Tables 13, 15] In-house Spotting / IE / Video Subtitle benchmarks (Tables 13, 15) are useful for application claims but opaque. A short appendix with sample counts, domain mix, and metric definitions would improve external interpretability.
  4. [Sec. 3.2; Eqs. (1)–(2)] Eqs. (1)–(2) and Fig. 2: Define γ’s units/role more explicitly in the main text (decay over draft positions) and state whether the draft model is used only at inference or also affects any reported accuracy runs.
  5. [Abstract; Introduction] Abstract/intro wording “the top-tier” vs. “among the top-tier” is slightly inconsistent with Table 12’s careful end-to-end expert ranking; align the claim language with the table scope (end-to-end expert OCR models).
  6. [Sec. 2; Tables 7–12] Related Work and comparisons mix open and proprietary systems of very different sizes; when claiming SOTA among lightweight models, a size-filtered subtable would reduce ambiguity.

Circularity Check

0 steps flagged

No significant circularity; empirical claims rest on external open benchmarks and standard draft-verify guarantees, not self-definitional reductions.

full rationale

HunyuanOCR-1.5 is an empirical systems paper whose load-bearing claims (6.37×/2.14× speedups via DFlash, top-tier OmniDocBench v1.6 among end-to-end models, and gains on Chronicles-OCR/ChartArena/TableVerse-5K/MORE/DUDE/CHAOS-Bench) are measured against open external benchmarks and third-party systems, not derived from fitted parameters or self-referential equations. DFlash (Sec. 3.2, Eqs. 1–2, B=16 draft from last 5 frozen target layers) inherits the standard speculative-decoding acceptance rule that preserves the target output distribution by construction of parallel verification; the paper reports latency/TPS only and does not rename a fit as an accuracy prediction. Agentic Data Flow (Sec. 4) seeds synthetic data from HunyuanOCR-1.0 failure modes and tool-assisted filtering, but the resulting model is scored on held-out external sets, so the evaluation is not tautological with the seeds. Self-comparisons to HunyuanOCR-1.0 [28] are transparent versioning of the same architecture, not uniqueness theorems or load-bearing self-citations that force the result. No self-definitional loops, no fitted-input-called-prediction, no ansatz smuggled via overlapping-author uniqueness claims, and no renaming of known empirical patterns appear in the derivation chain. The paper is therefore self-contained against external evidence.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 2 invented entities

Empirical systems paper. Load-bearing free parameters are the DFlash and RL hyperparameters chosen by the authors; domain assumptions are standard speculative-decoding correctness and the validity of the chosen open + in-house benchmarks. Invented entities are the named systems and the new benchmark introduced in this work.

free parameters (4)
  • DFlash block size B = 16
    Set to 16; controls draft length and acceptance statistics; chosen by authors, not derived.
  • DFlash anchors n and decay γ = n=16, γ=7.0
    n=16 anchors per sequence, γ=7.0 exponential position weight in the draft loss (Eqs. 1–2); hand-chosen training knobs.
  • RL group size G, IcePop ratio bounds, PPO clip ε, KL weight γ
    Policy-optimization hyperparameters that shape which tokens receive gradient; not fixed by theory.
  • Parse reward weights λ1, λ2 and table content/structure mix 0.5/0.5 = 0.5 / 0.5 for table content/structure
    Balance plain-text edit distance vs. element-specific scores; chosen for training stability.
axioms (3)
  • domain assumption Speculative decoding with a draft model verified by the target model preserves the exact output distribution of the target model.
    Standard claim of the draft-then-verify paradigm (Sec. 3.2); required for accuracy numbers to transfer from AR to DFlash.
  • domain assumption Open-source OCR benchmarks (OmniDocBench, Chronicles-OCR, ChartArena, MORE, DUDE, etc.) plus the authors’ in-house Spotting/IE/Video sets are sufficiently representative of real-world OCR workloads.
    Underpins all SOTA and long-tail claims; evaluation tree in Sec. 6.
  • domain assumption Agent-generated synthetic data (fonts, backgrounds, multi-page QA) after tool-based cleaning transfers to real ancient-script, low-resource, and multi-image documents.
    Core premise of Agentic Data Flow (Sec. 4); if false, long-tail gains would not generalize.
invented entities (2)
  • Agentic Data Flow no independent evidence
    purpose: Closed-loop agent system that turns model weaknesses into material search, cleaning, pipeline code, and training data.
    Named contribution of the paper; no independent prior existence outside this work.
  • CHAOS-Bench no independent evidence
    purpose: Controlled WYSIWYG hallucination benchmark that inserts meaningless perturbed words and measures page-average recall of those words.
    Introduced in this paper (Sec. 6.2) to quantify seen-text faithfulness.

pith-pipeline@v1.1.0-grok45 · 36109 in / 3151 out tokens · 29368 ms · 2026-07-11T11:57:15.465868+00:00 · methodology

0 comments
read the original abstract

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image document understanding within a single end-to-end VLM. Building upon the lightweight architecture of HunyuanOCR-1.0, HunyuanOCR-1.5 does not redesign the backbone, but systematically improves both efficiency and capability. For efficiency, we adapt DFlash to OCR decoding, significantly reducing the latency of long structured outputs such as dense documents, tables, and formulas while preserving output distribution. Powered by DFlash, HunyuanOCR-1.5 achieves a 6.37x Transformer inference speedup and a 2.14x speedup under vLLM, delivering the fastest inference among lightweight OCR VLMs. For capability, we propose Agentic Data Flow, an agent-driven data construction system that transforms model weaknesses into executable data requirements and autonomously performs material search, quality verification, and pipeline development. It substantially improves long-tail capabilities in ancient-script OCR, fine-grained chart and table parsing, multi-image text-centric QA, low-resource multilingual parsing, and document hallucination evaluation. HunyuanOCR-1.5 ranks among the top-tier end-to-end OCR solutions on OmniDocBench v1.6 while achieving new performance milestones across these long-tail tasks. Combined with an upgraded pretraining and post-training recipe, HunyuanOCR-1.5 further extends its capability in high-resolution, long-context, and multi-task scenarios. Experiments demonstrate faster inference, broader OCR capability coverage, and the deployment advantages of a lightweight end-to-end model. We will release the model weights and training code to support future research and real-world OCR applications.

Figures

Figures reproduced from arXiv: 2607.04884 by Binghong Wu, Bochao Wang, Can Ma, Chengquan Zhang, Gengluo Li, Guanghua Yu, Han Hu, Hao Feng, Hongbing Wen, Hong Liu, Huawen Shen, Jieneng Yang, Liang Wu, Pengyuan Lyu, Shangpin Peng, Shijing Hu, Weinong Wang, Xingyu Wan, Yongkun Du, Yu Zhou, Zheng Ruan, Zhiqiong Lu, Zibin Lin.

Figure 1
Figure 1. Figure 1: Overview of the HunyuanOCR-1.5 architecture. A compact end-to-end model that unifies diverse OCR-centric capabilities, including document parsing, text spotting, information extraction, OCR-aware QA, ancient-script recognition, chart parsing, image translation, and video subtitle extraction. 3 Model Design 3.1 Model Architecture HunyuanOCR-1.5 follows the compact, fully end-to-end architecture of HunyuanOC… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of DFlash training with a joint FlexAttention mask. One target forward is performed, K anchors are sampled at random positions, and all K blocks attend in a single pass. Rows denote Query tokens, and columns denote Key/Value tokens. block-diagonal mask, where each block can attend to the target hidden states before its anchor and to the mask tokens within the same block, while different blocks rem… view at source ↗
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Overview of the RL framework. The RL framework that optimizes the general OCR model toward more faithful, stronger, and more comprehensive behavior through three complementary reward components. or human preferences [66, 67]. HunyuanOCR [28] has already validated this potential in the OCR domain: through high-quality RL data and an ability-adaptive reward design, it achieves stable and effective training, … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

    cs.CV 2026-07 conditional novelty 6.5

    A document-oriented ViT family pretrained with text generation plus pixel reconstruction on 113M images transfers across recognition, detection, parsing, and understanding, setting open-source SOTA on MDPBench with a ...

Reference graph

Works this paper leans on

114 extracted references · 42 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Scene text detection and recognition: The deep learning era.International Journal of Computer Vision, 129(1):161–184, 2021

    Shangbang Long, Xin He, and Cong Yao. Scene text detection and recognition: The deep learning era.International Journal of Computer Vision, 129(1):161–184, 2021

  2. [2]

    Document parsing unveiled: Techniques, challenges, and prospects for structured information extraction.arXiv preprint arXiv:2410.21169, 2024

    Qintong Zhang, Bin Wang, Victor Shea-Jay Huang, Junyuan Zhang, Zhengren Wang, Hao Liang, Conghui He, and Wentao Zhang. Document parsing unveiled: Techniques, challenges, and prospects for structured information extraction.arXiv preprint arXiv:2410.21169, 2024

  3. [3]

    OmniDocBench: Benchmarking diverse PDF document parsing with comprehensive annotations

    Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, et al. OmniDocBench: Benchmarking diverse PDF document parsing with comprehensive annotations. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2025

  4. [4]

    Chronicles-OCR: A cross-temporal perception benchmark for the evolutionary trajectory of chinese characters.arXiv preprint arXiv:2605.11960, 2026

    Gengluo Li, Shangpin Peng, Xingyu Wan, Chengquan Zhang, Hao Feng, Xin Xu, Pian Wu, Bang Li, Zengmao Ding, Yongge Liu, et al. Chronicles-OCR: A cross-temporal perception benchmark for the evolutionary trajectory of chinese characters.arXiv preprint arXiv:2605.11960, 2026

  5. [5]

    ChartArena: Benchmarking chart parsing across languages, scenarios, and formats.arXiv preprint arXiv:2606.01348, 2026

    Shangpin Peng, Gengluo Li, Xingyu Wan, Chengquan Zhang, Hao Feng, Binghong Wu, Huawen Shen, Weinong Wang, Ziyi Cai, Zhuotao Tian, Han Hu, Can Ma, and Yu Zhou. ChartArena: Benchmarking chart parsing across languages, scenarios, and formats.arXiv preprint arXiv:2606.01348, 2026

  6. [6]

    OCRBench: on the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 67(12), 2024

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xiang Bai. OCRBench: on the hidden mystery of OCR in large multimodal models.Science China Information Sciences, 67(12), 2024

  7. [7]

    Document image machine translation with dynamic multi-pre-trained models assembling

    Yupu Liang, Yaping Zhang, Cong Ma, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, and Yu Zhou. Document image machine translation with dynamic multi-pre-trained models assembling. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pa...

  8. [8]

    MMTIT-Bench: A multilingual and multi-scenario benchmark with cognition- perception-reasoning guided text-image machine translation

    Gengluo Li, Chengquan Zhang, Yupu Liang, Huawen Shen, Yaping Zhang, Pengyuan Lyu, Weinong Wang, Xingyu Wan, Gangyan Zeng, Han Hu, et al. MMTIT-Bench: A multilingual and multi-scenario benchmark with cognition- perception-reasoning guided text-image machine translation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages ...

  9. [9]

    Document understanding dataset and evaluation (DUDE)

    Jordy Van Landeghem, Rubèn Tito, Łukasz Borchmann, Michał Pietruszka, Pawel Joziak, Rafal Powalski, Dawid Jurkiewicz, Mickaël Coustaty, Bertrand Anckaert, Ernest Valveny, et al. Document understanding dataset and evaluation (DUDE). InProceedings of the IEEE International Conference on Computer Vision, pages 19528–19540, 2023

  10. [10]

    MMLongBench-Doc: Benchmarking long-context document understanding with visualizations

    Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, et al. MMLongBench-Doc: Benchmarking long-context document understanding with visualizations. In Proceedings of Advances in Neural Information Processing Systems, volume 37, pages 95963–96010, 2024

  11. [11]

    GPT-4 Technical Report.arXiv preprint arXiv:2303.08774, 2023

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 Technical Report.arXiv preprint arXiv:2303.08774, 2023

  12. [12]

    Gemini: A family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: A family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023

  13. [13]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.arXiv preprint arXiv:2403.05530, 2024

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.arXiv preprint arXiv:2403.05530, 2024

  14. [14]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025

  15. [15]

    Gemini 3 Pro Model Card

    Google. Gemini 3 Pro Model Card. https://storage.googleapis.com/deepmind-media/Model-Cards/ Gemini-3-Pro-Model-Card.pdf, 2026

  16. [16]

    Gemini 3.1 Pro: A smarter model for your most complex tasks

    Google. Gemini 3.1 Pro: A smarter model for your most complex tasks. https://blog.google/ innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/, 2026

  17. [17]

    Qwen-VL: A versatile vision-language model for understanding, localization.Text Reading, and Beyond, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for understanding, localization.Text Reading, and Beyond, 2023

  18. [18]

    Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution.arXiv 21 preprint arXiv:2409.12191, 2024

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution.arXiv 21 preprint arXiv:2409.12191, 2024

  19. [19]

    Qwen2.5-VL Technical Report.arXiv preprint arXiv:2502.13923, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, et al. Qwen2.5-VL Technical Report.arXiv preprint arXiv:2502.13923, 2025

  20. [20]

    Qwen3-VL Technical Report.arXiv preprint arXiv:2511.21631, 2025

    Shuai Bai, Yuxuan Cai, et al. Qwen3-VL Technical Report.arXiv preprint arXiv:2511.21631, 2025

  21. [21]

    InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 24185–24198, 2024

  22. [22]

    How far are we to GPT-4V? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024

    Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to GPT-4V? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024

  23. [23]

    Mini-InternVL: A flexible-transfer pocket multimodal model with 5% parameters and 90% performance.Visual Intelligence, 2(1):1–17, 2024

    Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren, Weiyun Wang, Jinguo Zhu, Hao Tian, Shenglong Ye, Junjun He, Xizhou Zhu, et al. Mini-InternVL: A flexible-transfer pocket multimodal model with 5% parameters and 90% performance.Visual Intelligence, 2(1):1–17, 2024

  24. [24]

    Enhancing the reasoning ability of multimodal large language models via mixed preference optimization.arXiv preprint arXiv:2411.10442, 2024

    Weiyun Wang, Zhe Chen, Wenhai Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Jinguo Zhu, Xizhou Zhu, Lewei Lu, Yu Qiao, and Jifeng Dai. Enhancing the reasoning ability of multimodal large language models via mixed preference optimization.arXiv preprint arXiv:2411.10442, 2024

  25. [25]

    Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271, 2024

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271, 2024

  26. [26]

    InternVL3: Exploring advanced training and test-time recipes for open-source multimodal models

    Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. InternVL3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479, 2025

  27. [27]

    InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency.arXiv preprint arXiv:2508.18265, 2025

    Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency.arXiv preprint arXiv:2508.18265, 2025

  28. [28]

    HunyuanOCR Technical Report.arXiv preprint arXiv:2511.19575, 2025

    Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan, Gengluo Li, Shangpin Peng, Weinong Wang, Liang Wu, Huawen Shen, Yu Zhou, Canhui Tang, et al. HunyuanOCR Technical Report.arXiv preprint arXiv:2511.19575, 2025

  29. [29]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InProceedings of Advances in Neural Information Processing Systems, volume 30, 2017

  30. [30]

    MinerU-Diffusion: Rethinking document OCR as inverse rendering via diffusion decoding.arXiv preprint arXiv:2603.22458, 2026

    Hejun Dong, Junbo Niu, Bin Wang, Weijun Zeng, Wentao Zhang, and Conghui He. MinerU-Diffusion: Rethinking document OCR as inverse rendering via diffusion decoding.arXiv preprint arXiv:2603.22458, 2026

  31. [31]

    Fast inference from transformers via speculative decoding

    Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In Proceedings of the International Conference on Machine Learning, pages 19274–19286, 2023

  32. [32]

    EAGLE: Speculative sampling requires rethinking feature uncertainty

    Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE: Speculative sampling requires rethinking feature uncertainty. InProceedings of the International Conference on Machine Learning, pages 28935–28948, 2024

  33. [33]

    EAGLE-2: Faster inference of language models with dynamic draft trees

    Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE-2: Faster inference of language models with dynamic draft trees. InProceedings of the Conference on Empirical Methods in Natural Language Processing, pages 7421–7432, 2024

  34. [34]

    EAGLE-3: Scaling up inference acceleration of large language models via training-time test

    Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE-3: Scaling up inference acceleration of large language models via training-time test. InProceedings of Advances in Neural Information Processing Systems, volume 38, pages 136737–136756, 2026

  35. [35]

    Medusa: Simple LLM inference acceleration framework with multiple decoding heads

    Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. InProceedings of the 41st International Conference on Machine Learning, pages 5209–5235, 2024

  36. [36]

    DFlash: Block diffusion for flash speculative decoding.arXiv preprint arXiv:2602.06036, 2026

    Jian Chen, Yesheng Liang, and Zhijian Liu. DFlash: Block diffusion for flash speculative decoding.arXiv preprint arXiv:2602.06036, 2026

  37. [37]

    Block Diffusion: Interpolating between autoregressive and diffusion language models

    Marianne Arriola, Aaron Gokaslan, Justin Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sahoo, and V olodymyr Kuleshov. Block Diffusion: Interpolating between autoregressive and diffusion language models. In Proceedings of the International Conference on Learning Representations, pages 50726–50753, 2025

  38. [38]

    SDAR: A synergistic diffusion-autoregression paradigm for scalable sequence generation

    Shuang Cheng, Yihan Bian, Dawei Liu, Linfeng Zhang, Qian Yao, Zhongbo Tian, Wenhai Wang, Qipeng Guo, Kai Chen, Biqing Qi, et al. SDAR: A synergistic diffusion-autoregression paradigm for scalable sequence generation. 22 arXiv preprint arXiv:2510.06303, 2025

  39. [39]

    Fast-dLLM v2: Efficient block-diffusion LLM.arXiv preprint arXiv:2509.26328, 2025

    Chengyue Wu, Hao Zhang, Shuchen Xue, Shizhe Diao, Yonggan Fu, Zhijian Liu, Pavlo Molchanov, Ping Luo, Song Han, and Enze Xie. Fast-dLLM v2: Efficient block-diffusion LLM.arXiv preprint arXiv:2509.26328, 2025

  40. [40]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023

  41. [41]

    Llama.cpp – run LLM inference in C/C++

    Georgi Gerganov and llama.cpp contributors. Llama.cpp – run LLM inference in C/C++. https://github.com/ ggml-org/llama.cpp, 2023. GitHub repository

  42. [42]

    AgentInstruct: Toward generative teaching with agentic flows

    Arindam Mitra, Luciano Del Corro, Guoqing Zheng, Shweti Mahajan, Dany Rouhana, Andres Codas, Yadong Lu, Wei-ge Chen, Olga Vrousgos, Corby Rosset, et al. AgentInstruct: Toward generative teaching with agentic flows. arXiv preprint arXiv:2407.03502, 2024

  43. [43]

    TaskCraft: Automated generation of agentic tasks.arXiv preprint arXiv:2506.10055, 2025

    Dingfeng Shi, Jingyi Cao, Qianben Chen, Weichen Sun, Weizhen Li, Hongxuan Lu, Fangchen Dong, Tianrui Qin, King Zhu, Minghao Liu, et al. TaskCraft: Automated generation of agentic tasks.arXiv preprint arXiv:2506.10055, 2025

  44. [44]

    MetaSynth: Meta-prompting-driven agentic scaffolds for diverse synthetic data generation

    Haris Riaz, Sourav Sanjukta Bhabesh, Vinayak Arannil, Miguel Ballesteros, and Graham Horwood. MetaSynth: Meta-prompting-driven agentic scaffolds for diverse synthetic data generation. InFindings of the Association for Computational Linguistics: ACL, pages 18770–18803, 2025

  45. [45]

    PaddleOCR-VL: Boosting multilingual document parsing via a 0.9B ultra-compact vision-language model.arXiv preprint arXiv:2510.14528, 2025

    Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, et al. PaddleOCR-VL: Boosting multilingual document parsing via a 0.9B ultra-compact vision-language model.arXiv preprint arXiv:2510.14528, 2025

  46. [46]

    PaddleOCR-VL-1.5: Towards a multi-task 0.9B VLM for robust in-the-wild document parsing.arXiv preprint arXiv:2601.21957, 2026

    Cheng Cui, Ting Sun, Suyin Liang, Tingquan Gao, Zelun Zhang, Jiaxuan Liu, Xueqing Wang, Changda Zhou, Hongen Liu, Manhui Lin, et al. PaddleOCR-VL-1.5: Towards a multi-task 0.9B VLM for robust in-the-wild document parsing.arXiv preprint arXiv:2601.21957, 2026

  47. [47]

    Infinity Parser: Layout aware reinforcement learning for scanned document parsing.arXiv preprint arXiv:2506.03197, 2025

    Baode Wang, Biao Wu, Weizhen Li, Meng Fang, Zuming Huang, Jun Huang, Haozhe Wang, Yanjie Liang, Ling Chen, Wei Chu, et al. Infinity Parser: Layout aware reinforcement learning for scanned document parsing.arXiv preprint arXiv:2506.03197, 2025

  48. [48]

    HSD: Training-free acceleration for document parsing vision-language model with hierarchical speculative decoding.arXiv preprint arXiv:2602.12957, 2026

    Wenhui Liao, Hongliang Li, Pengyu Xie, Xinyu Cai, Yufan Shen, Yi Xin, Qi Qin, Shenglong Ye, Tianbin Li, Ming Hu, et al. HSD: Training-free acceleration for document parsing vision-language model with hierarchical speculative decoding.arXiv preprint arXiv:2602.12957, 2026

  49. [49]

    Logics-Parsing Technical Report.arXiv preprint arXiv:2509.19760, 2025

    Xiangyang Chen, Shuzhao Li, Xiuwen Zhu, Yongfan Chen, Fan Yang, Cheng Fang, Lin Qu, Xiaoxiao Xu, Hu Wei, and Minggang Wu. Logics-Parsing Technical Report.arXiv preprint arXiv:2509.19760, 2025

  50. [50]

    Logics-Parsing-Omni Technical Report.arXiv preprint arXiv:2603.09677, 2026

    Xin An, Jingyi Cai, Xiangyang Chen, Huayao Liu, Peiting Liu, Peng Wang, Bei Yang, Xiuwen Zhu, Yongfan Chen, Yan Gao, et al. Logics-Parsing-Omni Technical Report.arXiv preprint arXiv:2603.09677, 2026

  51. [51]

    Unirec-0.1b: Unified text and formula recognition with 0.1b parameters.arXiv preprint arXiv:2512.21095, 2025

    Yongkun Du, Zhineng Chen, Yazhen Xie, Weikang Bai, Hao Feng, Wei Shi, Yuchen Su, Can Huang, and Yu-Gang Jiang. Unirec-0.1b: Unified text and formula recognition with 0.1b parameters.arXiv preprint arXiv:2512.21095, 2025

  52. [52]

    dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025

    Yumeng Li, Guang Yang, Hao Liu, Bowen Wang, and Colin Zhang. dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025

  53. [53]

    DeepSeek-OCR: Contexts optical compression.arXiv preprint arXiv:2510.18234, 2025

    Haoran Wei, Yaofeng Sun, and Yukun Li. DeepSeek-OCR: Contexts optical compression.arXiv preprint arXiv:2510.18234, 2025

  54. [54]

    Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding

    Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. In Findings of the Association for Computational Linguistics: ACL, pages 7655–7671, 2024

  55. [55]

    Better & faster large language models via multi-token prediction.arXiv preprint arXiv:2404.19737, 2024

    Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve. Better & faster large language models via multi-token prediction.arXiv preprint arXiv:2404.19737, 2024

  56. [56]

    An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020

  57. [57]

    Patch n’ Pack: NaViT, a Vision Transformer for any aspect ratio and resolution

    Mostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek, Matthias Minderer, Mathilde Caron, Andreas Steiner, Joan Puigcerver, Robert Geirhos, Ibrahim M Alabdulmohsin, et al. Patch n’ Pack: NaViT, a Vision Transformer for any aspect ratio and resolution. InProceedings of Advances in Neural Information Processing Systems, volume 36, pages 2252–2274, 2023

  58. [58]

    Hunyuan-0.5B.https://github.com/Tencent-Hunyuan/Hunyuan-0.5B, 2025

    Tencent Hunyuan. Hunyuan-0.5B.https://github.com/Tencent-Hunyuan/Hunyuan-0.5B, 2025. 23

  59. [59]

    Flex Attention: A programming model for generating optimized attention kernels.arXiv preprint arXiv:2412.05496, 2024

    Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang, and Horace He. Flex Attention: A programming model for generating optimized attention kernels.arXiv preprint arXiv:2412.05496, 2024

  60. [60]

    Qwen3.5: Towards native multimodal agents, February 2026

    Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URL https://qwen.ai/blog?id= qwen3.5

  61. [61]

    Synthetic data for text localisation in natural images

    Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. Synthetic data for text localisation in natural images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2315–2324, 2016

  62. [62]

    OCR-Free document understanding transformer

    Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. OCR-Free document understanding transformer. InProceedings of the European Conference on Computer Vision, 2022

  63. [63]

    DeepSeekMath: Pushing the limits of mathematical reasoning in open language models

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024

  64. [64]

    Seg-Zero: Reasoning-chain guided segmentation via cognitive reinforcement.arXiv preprint arXiv:2503.06520, 2025

    Yuqi Liu, Bohao Peng, Zhisheng Zhong, Zihao Yue, Fanbin Lu, Bei Yu, and Jiaya Jia. Seg-Zero: Reasoning-chain guided segmentation via cognitive reinforcement.arXiv preprint arXiv:2503.06520, 2025

  65. [65]

    Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs.arXiv preprint arXiv:2506.14245, 2025

    Xumeng Wen, Zihan Liu, Shun Zheng, Zhijian Xu, Shengyu Ye, Zhirong Wu, Xiao Liang, Yang Wang, Junjie Li, Ziming Miao, et al. Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs.arXiv preprint arXiv:2506.14245, 2025

  66. [66]

    Uni-DPO: A unified paradigm for dynamic preference optimization of LLMs

    Shangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang, Xing Wu, Haotian Xu, Chengquan Zhang, Takashi Isobe, Baotian Hu, and Min Zhang. Uni-DPO: A unified paradigm for dynamic preference optimization of LLMs. arXiv preprint arXiv:2506.10054, 2025

  67. [67]

    Mitigating object hallucinations via sentence-level early intervention

    Shangpin Peng, Senqiao Yang, Li Jiang, and Zhuotao Tian. Mitigating object hallucinations via sentence-level early intervention. InProceedings of the IEEE International Conference on Computer Vision, 2025

  68. [68]

    Every Step Evolves: Scaling reinforcement learning for trillion-scale thinking model.arXiv preprint arXiv:2510.18855, 2025

    Ling Team, Anqi Shen, Baihui Li, Bin Hu, Bin Jing, Cai Chen, Chao Huang, Chao Zhang, Chaokun Yang, Cheng Lin, et al. Every Step Evolves: Scaling reinforcement learning for trillion-scale thinking model.arXiv preprint arXiv:2510.18855, 2025

  69. [69]

    DAPO: An open-source LLM reinforcement learning system at scale

    Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. DAPO: An open-source LLM reinforcement learning system at scale. InProceedings of Advances in Neural Information Processing Systems, volume 38, pages 113222–113244, 2026

  70. [70]

    Image-Based Table Recognition: Data, model, and evaluation

    Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes. Image-Based Table Recognition: Data, model, and evaluation. InProceedings of the European Conference on Computer Vision, 2020

  71. [71]

    StrucTab: A structured optimization framework for table parsing.arXiv preprint arXiv:2606.29905, 2026

    Gengluo Li, Shangpin Peng, Chengquan Zhang, Binghong Wu, Hao Feng, Weinong Wang, Pengyuan Lyu, Huawen Shen, Xingyu Wan, Zhuotao Tian, Han Hu, Can Ma, and Yu Zhou. StrucTab: A structured optimization framework for table parsing.arXiv preprint arXiv:2606.29905, 2026

  72. [72]

    StructChart: On the schema, metric, and augmentation for visual chart understanding.arXiv preprint arXiv:2309.11268, 2023

    Renqiu Xia, Bo Zhang, Haoyang Peng, Hancheng Ye, Xiangchao Yan, Peng Ye, Botian Shi, Yu Qiao, and Junchi Yan. StructChart: On the schema, metric, and augmentation for visual chart understanding.arXiv preprint arXiv:2309.11268, 2023

  73. [73]

    MORE: A multilingual document parsing benchmark and evaluation

    Long Xu, Binghong Wu, Tinghao Yu, Hao Feng, Zhenyu Huang, Haoqing Jiang, Yunhao Wang, Shuo Huang, and Feng Zhang. MORE: A multilingual document parsing benchmark and evaluation. InProceedings of the International Conference on Machine Learning, 2026

  74. [74]

    Transformers: State-of-the-art natural language processing

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. Transformers: State-of-the-art natural language processing. InPro- ceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, 2020

  75. [75]

    GLM-OCR Technical Report.arXiv preprint arXiv:2603.10910, 2026

    Shuaiqi Duan, Yadong Xue, Weihan Wang, Zhe Su, Huan Liu, Sheng Yang, Guobing Gan, Guo Wang, Zihan Wang, Shengdong Yan, et al. GLM-OCR Technical Report.arXiv preprint arXiv:2603.10910, 2026

  76. [76]

    PaddleOCR-VL-1.6: Expanding the frontier of document parsing with under-optimized region refinement and progressive post-training.arXiv preprint arXiv:2606.03264, 2026

    Zelun Zhang, Hongen Liu, Suyin Liang, Yubo Zhang, Yiqing Xiang, Jiaxuan Liu, Ting Sun, Manhui Lin, Yue Zhang, Changda Zhou, et al. PaddleOCR-VL-1.6: Expanding the frontier of document parsing with under-optimized region refinement and progressive post-training.arXiv preprint arXiv:2606.03264, 2026

  77. [77]

    Unlimited OCR works.arXiv preprint arXiv:2606.23050, 2026

    Youyang Yin, Huanhuan Liu, Qunyi Xie, Chaorun Liu, Shiqi Yang, Shaohua Wang, Zhanlong Liu, Hao Zou, Jinyue Chen, Shu Wei, et al. Unlimited OCR works.arXiv preprint arXiv:2606.23050, 2026

  78. [78]

    DeepSeek-OCR 2: Visual causal flow.arXiv preprint arXiv:2601.20552, 2026

    Haoran Wei, Yaofeng Sun, and Yukun Li. DeepSeek-OCR 2: Visual causal flow.arXiv preprint arXiv:2601.20552, 2026. 24

  79. [79]

    Gemma: Open models based on gemini research and technology.arXiv preprint arXiv:2403.08295, 2024

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology.arXiv preprint arXiv:2403.08295, 2024

  80. [80]

    MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe

    Tianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang, Wenshuo Ma, Zhihui He, Tianchi Cai, Weize Chen, Yuxiang Huang, Ranchi Zhao, et al. MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 11704–11715, 2026

Showing first 80 references.