Pith. sign in

REVIEW 12 cited by

DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.11810 v4 pith:PEULKRWT submitted 2023-11-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords docpediamodeldocumentlargevisualcomprehensiondomainfrequency
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

This work presents DocPedia, a novel large multimodal model (LMM) for versatile OCR-free document understanding, capable of parsing images up to 2,560$\times$2,560 resolution. Unlike existing work either struggle with high-resolution documents or give up the large language model thus vision or language ability constrained, our DocPedia directly processes visual input in the frequency domain rather than the pixel space. The unique characteristic enables DocPedia to capture a greater amount of visual and textual information using a limited number of visual tokens. To consistently enhance both perception and comprehension abilities of our model, we develop a dual-stage training strategy and enrich instructions/annotations of all training tasks covering multiple document types. Extensive quantitative and qualitative experiments conducted on various publicly available benchmarks confirm the mutual benefits of jointly learning perception and comprehension tasks. The results provide further evidence of the effectiveness and superior performance of our DocPedia over other methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Constraining visual token budgets during SFT and RL forces VLMs to learn functional active perception, yielding ~5% relative gains and strong transfer to unconstrained evaluation.

  2. MFH: Marrying Frequency Domain with Handwritten Mathematical Expression Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MFH fuses high-frequency DCT features with spatial features from standard HMER encoders, improving recognition accuracy by about 1 to 2 points on CROHME 2014/2016/2019.

  3. Document Image Rectification Bases on Self-Adaptive Multitask Fusion

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A multitask fusion network with inter-task feature aggregation and gating reports state-of-the-art document dewarping results on DIR300, DocUNet, and DocReal, subject to comparison caveats.

  4. Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Granite Vision is a ~3B parameter open-weights vision-language model that reaches state-of-the-art scores on document understanding benchmarks despite its small size.

  5. EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition

    cs.CV 2025-02 reject novelty 6.0 of 10

    The paper presents the first event-camera dataset for scene text recognition and an LLM-based recognizer, but test-set tuning and contradictory data filtering weaken the evaluation.

  6. Docopilot: Improving Multimodal Models for Document-Level Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A new academic-paper dataset and a retrieval-free fine-tuned InternVL2 model improve multi-page document QA accuracy and latency on several benchmarks.

  7. ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An event-stream scene text recognizer trained with LLM-generated chain-of-thought rationales improves BLEU-1 on EventSTR from 0.638 to 0.648 and accuracy on WordArt* and IC15* by about half a point.

  8. Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    CAR routes each query to either a short answer or full reasoning based on the perplexity of the model's draft answer, improving accuracy and cutting token use on VQA, KIE, and math/common sense benchmarks.

  9. Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Converting document images into markup-language representations before answering questions improves visual document understanding, and the released DocMark datasets enable an adaptive markup generation pipeline.

  10. Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs

    cs.CV 2024-12 conditional novelty 5.0 of 10

    SAAO improves transferability of 3D point cloud adversarial attacks by performing Admix-style mixing in the graph Fourier domain with learnable weights and gradient-based path selection, yielding higher transfer attac...

  11. Visual Large Language Models for Generalized and Specialized Applications

    cs.CV 2025-01 conditional novelty 3.0 of 10

    This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.

  12. Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

    cs.CL 2025-01 conditional novelty 3.0 of 10

    A structured overview of question answering over visually rich documents, comparing encoding, vision-only, and multi-page methods, and highlighting comparability issues in existing benchmarks.

Pith tools