Pith. sign in

Paper Citation Record · LEDGER

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2307.02499.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.02499 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:23.176102Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

17
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 552b7a20-9103-4138-af8e-9e7ac74207fe · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.407056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:9246ac67b3f4fd01a7a5a7d11fb4c65c59b5b8ad377123523a5005faedd4f41d

Observation d4b8cd34-6934-476d-8aab-f6f555e1387c · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.774332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:c3278b9054e1e53092f2a3801014529757e2b44aa0794d9ef1b24d3edb6e510f

Observation ebf46cb7-67d1-423c-a210-057833a6f6f0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 123

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.354084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:309dffc60d3cd694f7f5bab38d7fd4b424c2a2a951576f08948961d8d115f52f

Observation 1c2852a6-e767-40dd-90b9-a368384dc822 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.308851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:c38432a147eff964fd18c2df774f955dae9819503772a1653fa74b84dc2492aa

Observation ca4339bc-d3f9-4dc1-81b3-4fc30fb7f0ce · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.899919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d368a819d03eb89c0ea37fa16d57a7f49b051e5a034bb8b0141246533a2f3501

Observation cd5b616a-9db4-4433-8e40-f861a0281528 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.728217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:b96408e356f95a196d98d2d6479dba29f304a9491b7704688ee5162227da2f80

Observation f82e1aa3-ea2a-405e-8e62-5e676a63ff1d · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:37:25.863817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:05f80ab59e8131fe85ddafcd71294a50204f20f5f2d33cf7a6517f94981aa907

Observation 3725caea-2055-4b6d-afd2-24a2cf74229d · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:01.046270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:01.046270Z digest=sha256:b82c52d8fbb5e258a2beac7950ede3c3b37d582444d6ba1531f4ce1f58d9112a

Observation 3ea1d49c-3ff7-4efb-81f6-a1562d4e110c · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.314292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.314292Z digest=sha256:aa239459dd948c4f5661bdb3e95814114cd39f62396a39be3cc728823d7fe57b

Observation 275ad23c-29ba-4d86-ab9c-d17dcc0f2072 · inbound

Medical Multimodal Foundation Models in Clinical Diagnosis and Treatment: Applications, Challenges, and Future Directions cites this paper.

Medical Multimodal Foundation Models in Clinical Diagnosis and Treatment: Applications, Challenges, and Future Directions mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:17:44.741480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:17:44.741480Z digest=sha256:458ff464377ad22a5d6caac7c84ba715fced764f378216e855033f6bf1d11504

Observation 05e16f75-503e-48f9-861a-2a58642a65e1 · inbound

Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges cites this paper.

Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 294

Resolution
unresolved
no resolver link, observed 2026-08-11T22:41:22.638715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:41:22.638715Z digest=sha256:ebbb819df49cc9f4e251bb887a148bdc2ae5269547ed3b8425a90df131934f45

Observation e974b614-b4b3-4f7f-bdfa-88217caad9ab · inbound

LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining cites this paper.

LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:09:03.961608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:09:03.961608Z digest=sha256:66c5203e09946bc6ed54611802969e633c2e4f18a543346234611ef5db0475af

Observation 8da39107-c5a1-4406-a17c-8cc8c739aa9d · inbound

From Elements to Design: A Layered Approach for Automatic Graphic Design Composition cites this paper.

From Elements to Design: A Layered Approach for Automatic Graphic Design Composition mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:06:12.390006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:06:12.390006Z digest=sha256:313c67cc337e7712112b1ec09f86a03dba04185233b5f136a1a92bbe0205362d

Observation cdb693c1-9c97-4b31-b968-addcf92fd827 · inbound

Slow Perception: Let's Perceive Geometric Figures Step-by-step cites this paper.

Slow Perception: Let's Perceive Geometric Figures Step-by-step mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:42.690920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:21:42.690920Z digest=sha256:707f3ceedb89c5ebc89983757683667142891009e2f150e8d77769d6974da697

Observation 1b457a58-b9ae-41d7-981f-999f974956ae · inbound

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking cites this paper.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.889462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.889462Z digest=sha256:6582893a256ea42b794a092b38af8c57ac1b99ba79eae1fe26cecedca3240535

Observation fd1ab135-9013-424b-9169-dbc158c19fcc · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.778534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:19db7fcba580c1ae916281031417d83dac389a210d0fa611302af4c43493af59

Observation defa629e-8143-497f-8995-6a7095da0f6b · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.084106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.084106Z digest=sha256:ca004126165e6803138b3127209761e4af0e736236d01b54c385c5f5b4a541ce

Observation b40ba76a-3d62-49c7-8a50-52de43bfd51e · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 290

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:10.024946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:10.024946Z digest=sha256:a0f8f5af4ebc1547ada4b8ffe168de9e01472c616b031d342c6a2c8e6ecdd424

Observation b9c431f3-29a7-40a3-8006-5dbebc659c34 · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.796461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.796461Z digest=sha256:58a146fb8943c574ac590dde31131c3805d5a9bc73333ae4dc399ea184c6b57c

Observation 2e866717-1d12-4ccf-bf1b-c775db334c8d · inbound

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers cites this paper.

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:18.879391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:18.879391Z digest=sha256:55810fb2b0f282a5170887323240254125a1dffc2e0207a2781e89949cb30cb6

Observation 793e0fc3-89bc-46e8-bca9-fba460646b40 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.784774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:ec3601389207da7a342cf5405fb40bdc2cbad4ffdd97418221328d47d71e1671

Observation 4b35a914-4d46-41a5-8e7f-974d3018cb85 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.769640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:ea15d43fa2200869720050fe9208a1c3bc0a4be6abf6b7a14598c0164b473b99

Observation 4f7d9831-2036-445d-acc1-bd80ebf19868 · inbound

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning cites this paper.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.176102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.176102Z digest=sha256:8350182cb0ac7d93e33650a4ae5cd78eb7a74602d11f0df28c4a5b3deb021e65

Observation 922212e8-41a2-47e2-8247-2eb451015deb · inbound

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling cites this paper.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.578722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.578722Z digest=sha256:28e28c729bb1e932484a2ea5b029d759138f2599a1bcc26889f1ba4abadcf426

Observation 8f853f63-5d03-4a70-b316-80e57e437fed · inbound

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? cites this paper.

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:42.005297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:42.005297Z digest=sha256:d9f26e30c51f60154327b21ff443227e7ce9c309ec86182a2209aa1367c751a5

Observation 2d1152ff-f775-4bbb-8ec6-57a697ca8c0f · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.576365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.576365Z digest=sha256:cb60424ba3eb40b383ca7bcd99383a0b787eddd240dfadc914871f2c0ae48d4d

Observation 6123c3b3-0e54-484e-b7dd-3352b9d48e95 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.848486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.848486Z digest=sha256:ce3ad213b261fb72f535c9812ad5dd5c50a5c64b163cdc0cb3ae2f6ddc43c9d2

Observation a3f2b3ba-ba36-44b4-85d1-72b640d28a8d · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.457998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.457998Z digest=sha256:6c5e2e4ef684f736e250eb01dac802e111f99fcd4a25c39ec4ab1bfce995b5b8

Observation ffcb8151-50dd-4afe-9c37-6c0203142738 · inbound

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning cites this paper.

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:40.379294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:40.379294Z digest=sha256:1d65e726be7ad0dbac39689a00b84790a6e0a6bbca424a9d19ae6a6d470d464a

Observation 924717a1-24cd-45db-99c1-8602fcd9fe8b · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:50.252890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:50.252890Z digest=sha256:ceb6b664221293e340fff3f7235b1a1537d36c6711c3158160a1364fb69aaff4

Observation f6d2e2b8-8e95-45c1-8ff7-cd7cace1d5df · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.656467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.656467Z digest=sha256:79ecc16b70c2d83d858f32232039111ce5d3cf49585b795617571cfb79ea1ffb

Observation 4b668490-1497-4a9e-ba9e-36a94dd2cd0e · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:21.088382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:21.088382Z digest=sha256:9df72ba62ea4e8c1ec55cf5a73f5f023404ce27011eeb363015fb9664181c0f4

Observation 2454a794-357c-430c-9d0d-367ccdaf14d4 · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.342834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.342834Z digest=sha256:a25a0df4c4d1df60061cdd8a51f1feac881d97944a8d5e1996bf8b581349f0aa

Observation a032bf00-e823-4686-9959-b392ac51eed5 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.350623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:7a3feac2a72066152a041342913986dbc41f7a0a4f9bcf43b8b7a7c5922e7282

Observation 72b07ee6-91d8-4aa3-8d59-51d67e29215a · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.975519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.975519Z digest=sha256:b951a7697a5a6ba59bdcd58fa61db198d8ba9f27184132581d41ef82bbff3829

Observation 0deda2f6-0968-456e-a19b-f59e3ea0d9f2 · inbound

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding cites this paper.

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.635607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T06:54:03.390655Z digest=sha256:7ba4db5f9d5c8b7d9d854e5099a681739b7c4ff3dc8710e4c14c1b6d90c79be1

Observation d97af87a-dda4-48a4-ab70-dfe2e3584110 · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:16.038985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:a3d5efd04a290a42017b0332c4cf60f59a96bd954015fc5d353eb0fcef7ad0a5

Observation 55a49441-a1a8-4ad4-9196-273bfdafe020 · inbound

CPT: Controllable and Editable Design Variations with Language Models cites this paper.

CPT: Controllable and Editable Design Variations with Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:51.622018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T20:01:39.648794Z digest=sha256:5107aabb8bd87a11857bb59d3df650e8008da45e85f54e51cd90f1c3d263b1f8

Observation 5b04b076-e704-4649-8935-650c895e0e26 · inbound

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization cites this paper.

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.300577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:28:16.315767Z digest=sha256:e1a032f9a4429a6f5dd799fed44354aba5365075cbc32c02c1f814895a7160a9

Observation 66a90a39-25bf-43da-836c-38ff35a6b6e3 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.486108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:9ebb41bf833d31063fd64e6a4b61f8a11e9b7fd5ab6d68fa9b472b49d1d5db77

Observation eeaf4af3-5576-49c6-9670-71318ade87d9 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.507739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:807b2ac6bb014350cea5df1762a6f945b3e118b141cab9d56b7097dc1ff23775

Observation e9aa1bb3-8649-49e7-a2b6-482da0b6da85 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.416164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:961240f96873dff460f91cfa9b0be88f722acb6aae4bfad0eb7dc372d986b963

Observation 260e67b2-aaaa-497b-8ef6-f37e8d9c6d05 · inbound

LLM Agents Can See Code Repositories cites this paper.

LLM Agents Can See Code Repositories mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-27T05:30:36.093677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T05:22:45.210932Z digest=sha256:9821e25ac2cee0fee9fcf0eda8c75715f479f33d35f2ea1bb158329ae15c3ad0

Observation 65af3fe7-43ec-47db-8b30-253a1489eb4d · inbound

LLM Agents Can See Code Repositories cites this paper.

LLM Agents Can See Code Repositories mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:46:26.190461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:46:26.190461Z digest=sha256:1e80374493ab258b4a1a0b5931eecd4865e8963a84b001067c547077e6107330

Observation 3bbb07e4-ee82-4607-9762-d4ee6dd4254a · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.032383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:82ac4850f3291c2c62f0f046464df04e53b4dc7d5906fcd88c459b5232f0c2e9

Observation c091df3d-a442-4ea3-82ba-d232b2938d98 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.783753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:de7f39539d98c43295cbb51400152d9a00037c2d1db81ef34fabc86b15cd1005

Observation 66a0e3c9-5836-4a18-908b-d3b9bdbd9293 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:51.235675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:987ca154f0c8ba548085bf12f73a06b0021e95dd7cbc963a82f5d3c1a886b47c

Observation 28144c9c-5287-4997-b2ff-b722feec3549 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 114

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:744413a4f672ef57c41d365cc56db2cfb6af852400fbeb30887bf472dbb236d9

Observation dd04d064-6e10-4f74-8517-0dc5168f0875 · inbound

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding cites this paper.

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T11:50:16.975786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:50:16.975786Z digest=sha256:3de7394755a03195f7493831b1abd3cc820b690dc1171cc68c6033b8c4f1a496

Observation 5f55687b-0776-4368-9f17-49dfe5daaebf · inbound

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding cites this paper.

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:07.816209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:07.816209Z digest=sha256:bf95845318d8752e56950a8a8f3367068f355c0420419ecb1687e798d69041de

Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.347023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.347023Z digest=sha256:75e497cd5ea9c3e9daf7a2a85290244c42b994cff30e8c4262e86188bd419fe4

Observation f114d9c3-8704-4518-bafa-06da7d6b8d6d · inbound

InSight-doc: Agentic Visual Perception for Long-Document Understanding cites this paper.

InSight-doc: Agentic Visual Perception for Long-Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T20:43:51.739388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:43:51.739388Z digest=sha256:02453465a9d61ae8d804cc9dc57edee531424b5e0c6fe4dc552c7e134e515597