Pith. sign in

Paper Citation Record · LEDGER

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2307.02499.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.02499 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:46:26.190461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 552b7a20-9103-4138-af8e-9e7ac74207fe · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.407056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:f0dd8968b424e898d58e820a90cd31be33495be7f60e8ba8e6f100607a63973a

Observation d4b8cd34-6934-476d-8aab-f6f555e1387c · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.774332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:1d9b711a19d6f164f586de5aa5229c5abb139a89ba5d446a995c14e096fd6921

Observation ebf46cb7-67d1-423c-a210-057833a6f6f0 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 123

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.354084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:0981e3e8d5fceba27b2d435b1813dce49efe4c8a5fdf065245d9819c38fff98e

Observation 1c2852a6-e767-40dd-90b9-a368384dc822 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.308851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:912eec3f660a45f08c0bf90a092b1d504ef96e674a9b8a169b77079df7321338

Observation ca4339bc-d3f9-4dc1-81b3-4fc30fb7f0ce · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.899919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:cd8aeb47c0357d941965d837d629d4420c514f868e5d4eca8ad7fc5d993977c6

Observation cd5b616a-9db4-4433-8e40-f861a0281528 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.728217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:38b76e437ff6942287e2f3077bea43bd14378fac070a69e4a285a87743047048

Observation f82e1aa3-ea2a-405e-8e62-5e676a63ff1d · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:37:25.863817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:59f9b28c3e01c0230bcf199b10483a67984fa93beb475c06afd56fad04013d20

Observation fd1ab135-9013-424b-9169-dbc158c19fcc · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.778534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:43df4dcdd21b02a5e2632d2c40312be65726e2f16b188da88dfb52a302d755a0

Observation 793e0fc3-89bc-46e8-bca9-fba460646b40 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.784774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:d840b564bdcdf2e744a5498c8989a5feaba36ff1b0d091c9621a7e6c5d74a4ff

Observation 4b35a914-4d46-41a5-8e7f-974d3018cb85 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.769640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:0b82bceb7d6b99b0090d0241e7488cb63d08bde63ee1ae53529ac206dc3e8af9

Observation a032bf00-e823-4686-9959-b392ac51eed5 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.350623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:b17a13c02bbb8f4ca039443eefd127f7ac0a633c9193130ed97e22534391fb60

Observation 0deda2f6-0968-456e-a19b-f59e3ea0d9f2 · inbound

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding cites this paper.

Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.635607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T06:54:03.390655Z digest=sha256:329eadaf579cecff92517d7b53402cfd18d69b071ff8ccb912a74d7190b6a99f

Observation d97af87a-dda4-48a4-ab70-dfe2e3584110 · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:16.038985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:37f5ef98a5155911a976ed9cb5d74274e1f0664beb2e95cd4fe4419e62950a1e

Observation 55a49441-a1a8-4ad4-9196-273bfdafe020 · inbound

CPT: Controllable and Editable Design Variations with Language Models cites this paper.

CPT: Controllable and Editable Design Variations with Language Models mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:51.622018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:01:39.648794Z digest=sha256:04b704044f69a635cf2c9a531910915e7b0d52244949fa1641586c9ca4b5a0a7

Observation 5b04b076-e704-4649-8935-650c895e0e26 · inbound

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization cites this paper.

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.300577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:28:16.315767Z digest=sha256:4444e170f237450c732efb7bfc3bc0561b5fca04330e3011fc25138419ef0a78

Observation 66a90a39-25bf-43da-836c-38ff35a6b6e3 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.486108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:85899c281a50c2e8945171e4713f5bc447978479486f3151cfd1ace977eb4c8b

Observation eeaf4af3-5576-49c6-9670-71318ade87d9 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.507739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:7221f1e42c687b9147e4130a472271f67d3c6e99704951cf6a973397d78deb1a

Observation e9aa1bb3-8649-49e7-a2b6-482da0b6da85 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.416164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:5733dfd0f73b65824b526c1253070789e7baa3c891bb10fdbbf84098d2f88e22

Observation 260e67b2-aaaa-497b-8ef6-f37e8d9c6d05 · inbound

LLM Agents Can See Code Repositories cites this paper.

LLM Agents Can See Code Repositories mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-27T05:30:36.093677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T05:22:45.210932Z digest=sha256:599c3549d8b5a33e9b41ab51786b85ebfffe2e215520fa1561f5c337af2e0fc4

Observation 65af3fe7-43ec-47db-8b30-253a1489eb4d · inbound

LLM Agents Can See Code Repositories cites this paper.

LLM Agents Can See Code Repositories mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:46:26.190461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:46:26.190461Z digest=sha256:fd150a82f9f130cb40e63be6bbe1b7349a86a03db6e8d1bd1ca138bd8f9a42df

Observation 3bbb07e4-ee82-4607-9762-d4ee6dd4254a · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.032383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:b61d2a11da5a6deabd5dfd2dc74a5cad13d75b784ceb198f0caf9597f53ac49b

Observation c091df3d-a442-4ea3-82ba-d232b2938d98 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.783753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:f5356197317492fd85842f8fc9f3bc36ae8df029b80abdaf675d826541f88a06

Observation 66a0e3c9-5836-4a18-908b-d3b9bdbd9293 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:51.235675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:9e474362b75b2a81fea97eff975d458374b86a45704ddd4c2c15673e5d626fd9

Observation 28144c9c-5287-4997-b2ff-b722feec3549 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 114

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:fee59ff79fc8ba69fa95998750211e4f1cfdebc30dba28edd57933e27505805e

Observation dd04d064-6e10-4f74-8517-0dc5168f0875 · inbound

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding cites this paper.

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T11:50:16.975786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:50:16.975786Z digest=sha256:cc4e6dd3cca6f3351d82c6a7690e754e7c585c85531a94d8140fd78d8164461d

Observation 5f55687b-0776-4368-9f17-49dfe5daaebf · inbound

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding cites this paper.

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:07.816209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:07.816209Z digest=sha256:6082dcce7422e404cdb562f4c520626ba2880ede804d0b8c3a02383d22eaaea7