Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2307.02499.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:23.176102Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
17
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 552b7a20-9103-4138-af8e-9e7ac74207fe · inbound
A Survey on Multimodal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d4b8cd34-6934-476d-8aab-f6f555e1387c · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ebf46cb7-67d1-423c-a210-057833a6f6f0 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1c2852a6-e767-40dd-90b9-a368384dc822 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ca4339bc-d3f9-4dc1-81b3-4fc30fb7f0ce · inbound
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cd5b616a-9db4-4433-8e40-f861a0281528 · inbound
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f82e1aa3-ea2a-405e-8e62-5e676a63ff1d · inbound
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3725caea-2055-4b6d-afd2-24a2cf74229d · inbound
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 135
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea1d49c-3ff7-4efb-81f6-a1562d4e110c · inbound
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275ad23c-29ba-4d86-ab9c-d17dcc0f2072 · inbound
Medical Multimodal Foundation Models in Clinical Diagnosis and Treatment: Applications, Challenges, and Future Directions mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e16f75-503e-48f9-861a-2a58642a65e1 · inbound
Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 294
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e974b614-b4b3-4f7f-bdfa-88217caad9ab · inbound
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da39107-c5a1-4406-a17c-8cc8c739aa9d · inbound
From Elements to Design: A Layered Approach for Automatic Graphic Design Composition mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb693c1-9c97-4b31-b968-addcf92fd827 · inbound
Slow Perception: Let's Perceive Geometric Figures Step-by-step mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b457a58-b9ae-41d7-981f-999f974956ae · inbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd1ab135-9013-424b-9169-dbc158c19fcc · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation defa629e-8143-497f-8995-6a7095da0f6b · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40ba76a-3d62-49c7-8a50-52de43bfd51e · inbound
Visual Large Language Models for Generalized and Specialized Applications mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 290
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c431f3-29a7-40a3-8006-5dbebc659c34 · inbound
Ocean-OCR: Towards General OCR Application via a Vision-Language Model mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e866717-1d12-4ccf-bf1b-c775db334c8d · inbound
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793e0fc3-89bc-46e8-bca9-fba460646b40 · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4b35a914-4d46-41a5-8e7f-974d3018cb85 · inbound
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f7d9831-2036-445d-acc1-bd80ebf19868 · inbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922212e8-41a2-47e2-8247-2eb451015deb · inbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f853f63-5d03-4a70-b316-80e57e437fed · inbound
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1152ff-f775-4bbb-8ec6-57a697ca8c0f · inbound
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6123c3b3-0e54-484e-b7dd-3352b9d48e95 · inbound
Clapper: Compact Learning and Video Representation in VLMs mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f2b3ba-ba36-44b4-85d1-72b640d28a8d · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffcb8151-50dd-4afe-9c37-6c0203142738 · inbound
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924717a1-24cd-45db-99c1-8602fcd9fe8b · inbound
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d2e2b8-8e95-45c1-8ff7-cd7cace1d5df · inbound
Structured Attention Matters to Multimodal LLMs in Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b668490-1497-4a9e-ba9e-36a94dd2cd0e · inbound
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2454a794-357c-430c-9d0d-367ccdaf14d4 · inbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a032bf00-e823-4686-9959-b392ac51eed5 · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72b07ee6-91d8-4aa3-8d59-51d67e29215a · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0deda2f6-0968-456e-a19b-f59e3ea0d9f2 · inbound
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d97af87a-dda4-48a4-ab70-dfe2e3584110 · inbound
InstructTable: Improving Table Structure Recognition Through Instructions mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 55a49441-a1a8-4ad4-9196-273bfdafe020 · inbound
CPT: Controllable and Editable Design Variations with Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5b04b076-e704-4649-8935-650c895e0e26 · inbound
Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 66a90a39-25bf-43da-836c-38ff35a6b6e3 · inbound
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eeaf4af3-5576-49c6-9670-71318ade87d9 · inbound
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e9aa1bb3-8649-49e7-a2b6-482da0b6da85 · inbound
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 260e67b2-aaaa-497b-8ef6-f37e8d9c6d05 · inbound
LLM Agents Can See Code Repositories mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 65af3fe7-43ec-47db-8b30-253a1489eb4d · inbound
LLM Agents Can See Code Repositories mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bbb07e4-ee82-4607-9762-d4ee6dd4254a · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c091df3d-a442-4ea3-82ba-d232b2938d98 · inbound
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 66a0e3c9-5836-4a18-908b-d3b9bdbd9293 · inbound
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 28144c9c-5287-4997-b2ff-b722feec3549 · inbound
Qwen-Audio-VAE Technical Report mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd04d064-6e10-4f74-8517-0dc5168f0875 · inbound
DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f55687b-0776-4368-9f17-49dfe5daaebf · inbound
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · inbound
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f114d9c3-8704-4518-bafa-06da7d6b8d6d · inbound
InSight-doc: Agentic Visual Perception for Long-Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.