Pith. sign in

Paper Citation Record · LEDGER

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2403.12895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12895 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:30.234729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T17:07:25.743722Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9636f990-2707-4fc7-ac5b-fcd8b73293c8 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:56:41.782309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:abd1e0f6f85295e56f9ca1ac4af61afadedd80897f98fd11354a7d6175addc18

Observation 87856638-81ac-452d-8d4d-5731764e1050 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T20:58:58.997584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:fe1c6b610deb9dffa44ef62c5db2ab01ee4e4f1980202f0a0e5de1bc9e9f2ee2

Observation 24372ced-89e2-4eb8-9ceb-ec8101d89e10 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T10:46:28.806610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:c7114d1de2edfe5519933e45d05cddf5c1d72f87b15508f499829f43c7359903

Observation 20d6da08-e96d-4de7-90c8-56bcb0da7189 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:59:32.748963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:699f79bda2a0be1d658819e8873532eb30fc08b64ac7be85b2ccf6a8a4923a2a

Observation 29680334-4140-4af1-ae06-56f51675dc11 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.864055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:8b98eaab83aef4ff070d70bab1fb2ed55f9c1e5703ec9cd2608a3b948eb18ade

Observation c2efe267-4983-441c-a96f-21bb893de652 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:00:25.807671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:682dbb00b0b53a5f88821772663e8540cb81221abd0c38c49fd6f3a72a797562

Observation 20b7ed55-3765-49f7-9512-ba4426200797 · inbound

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction cites this paper.

PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:12:14.657006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:12:14.613620Z digest=sha256:4bdeffebdd592e90329c9112bb060bf303e51f08289a5f11bc9d580bc5f3f249

Observation 0bb3dfa2-7fa4-495d-bb7d-4f3df067811d · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.181515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:3e4f4a8e207a72eb159b62a469f41116f51090f2601dd36b1d6871205be4d834

Observation 36eb7878-7b97-440b-86f7-8d44812fc120 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.855416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:8c549ee083a7a78c0ed66c2f388de7292d03219e38e37679b1903dfd773d1df8

Observation a94e4643-1970-4d16-8da9-39b3ed21684b · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.789867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:338892d9027aaa8f34abd049e0757a81d3e0804c62ab1defa1013e0433dfca11

Observation 13ce348c-6f4e-4ff0-a75c-36eed2ec4bee · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.234729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.234729Z digest=sha256:393ec831980c8d035150ab66a35a1d317ee0f0ef3dbf2c18c0e14735e5151a2f

Observation aa7ec550-3dc2-4674-a241-f86be58475fe · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:23.097635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:23.097635Z digest=sha256:f6bb2a6ebf1a38388048b890cadb63ce4db0b3e96191d83ec9692b59da33a1e4

Observation 57b3f667-02d7-41d7-844d-fc169ac4be3f · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.065585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.065585Z digest=sha256:0ed6f10ab2caeb691788eef11316ef1bdb85dfe448a6db75f12bb3019a003c07

Observation 29ef0fcd-2bc9-4672-916c-740cfec8368e · inbound

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models cites this paper.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:08.059391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:08.059391Z digest=sha256:384429cdda734ccc44bcdf0482a2a8ee6a525c04cf4effe4041912754a6b5831

Observation 28ac8605-6fd0-46f2-bb60-72bcd742971a · inbound

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning cites this paper.

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:37.113140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:37.113140Z digest=sha256:6a09e4f0e43df61b5624783239b729defbdea5894a12bb0dc75053957d60473d

Observation 310cdb34-bc96-4d48-ab61-5808c1d8a2b8 · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:17.602697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:17.602697Z digest=sha256:3d3a00116c36e9547b92956939e84ed3b2365c2b48533a600bd6c7b9406a5aa8

Observation 3cb95d74-63eb-4d1c-bd59-1439d6c85896 · inbound

A document is worth a structured record: Principled inductive bias design for document recognition cites this paper.

A document is worth a structured record: Principled inductive bias design for document recognition mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 39

Resolution
malformed identifier
arxiv_id, observed 2026-05-19T05:02:05.215068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T04:57:35.441758Z digest=sha256:cffba5881c84ede4422d4de8d58b262a297edc77ac54d37df7c095beb85e4646

Observation 2921b240-0064-4586-9758-4014f25367d4 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:49.641010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:49.641010Z digest=sha256:5cd62a186d22a18e18d4bbad80ab08227666ba398b29469474e1024c92a485ad

Observation 7dd154c6-d455-4066-8564-f221c43666c4 · inbound

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering cites this paper.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.796214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.796214Z digest=sha256:1b82384b2ec366bd2e09cfcd7373d80acbbd7f30675dc5cdd5efa10a383b78e2

Observation acea7df4-fe20-4d75-9810-9ed67ee89d3c · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:32:36.576223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:85f5d78b85183973ea955d8edaf1d6480385c5c18404658f76f6cb2e7679452b

Observation dfca068a-ce1c-4348-87f7-f1a55734cadc · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:14.038805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:14.038805Z digest=sha256:369162e36ec3390f02212bba17b9eca809a4b46281b9d3add740d89d5f031c9a

Observation 2ba7f8f2-2ea6-4e31-99fd-aa52ca51d8e0 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.692458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:36076f55e1c250552af57187ec04eaec322c39ffd40f99999600d899859ff19c

Observation 8ed5f5c4-6e72-483c-9ba1-909a80bff665 · inbound

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models cites this paper.

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.438914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:56:05.583390Z digest=sha256:45eeba2ef99905a202a0820ca43a880f92b224b04273d1b858955b6ecd0cb37d

Observation 9c6b683e-1544-444d-9dd6-40278a803d47 · inbound

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence cites this paper.

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:37:58.163050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:37:36.144960Z digest=sha256:ad3281c65df0b53b5edf4473d58608edc0b65402905dccc8694b22455b2aa2bc

Observation 6492d638-4007-49ba-a24b-3196fe499c0d · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-10T17:07:25.745651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T17:02:28.089092Z digest=sha256:1c4c8efeea31c9e1a48e92e5d6aba947b3d8d11065cb8e3a58584b97a0cbcb7a

Observation ccfee072-f5e6-40a4-be92-bdac0ad074c8 · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T08:03:51.775355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:03:51.775355Z digest=sha256:05ed70f649b50403c4b6e51c7681518e10d6d2208588e0d8347c89b84a8d3495