Pith. sign in

Paper Citation Record · LEDGER

TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.04473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.04473 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:36.782313Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T19:43:23.773160Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 541007d1-b508-4c6d-8895-89630c02dfa3 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.414279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:2b0756ca9cf99f6f0142b7a47494267041f8526bd046ac7531939c00b2267952

Observation dce1f58e-15c5-4d58-9261-e90500800966 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.111646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:685fabdcb421818dbe14f38898568b3516b0dd59b7fa01cd719c120988702772

Observation 77c8df6e-1d1c-461e-89b0-5599283f556b · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.774762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:9010a7aee84c0ad9ccff750aa2af1843bbb408075ea759f6634038fcbaf0c51d

Observation d0d22a99-2dca-4228-b9e2-14d19b081665 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.096671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:eec580dfce9a90ba63ce823355ec8ab45fa175168279d3bd3794d5d760ea5397

Observation 9f91f9c5-3858-4714-b421-252add1f24f6 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.827628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:3f8e517c1effd13bc929d545fe7b38029e505054c1130cfd49fcef7b733640fc

Observation 3032a2af-cde3-4694-8376-af8e8cfababc · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.878737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d95f244f69d0aa63094d223cbf45b5ad29fa4cbf461d032bc645ad22a23097a0

Observation 8a11461b-8a19-4f64-b893-9cb44861c518 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.776809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:eeea753060fdba80e5c63639c269c75d3f6e67dbfd5e5d4cbd949f2b86d99260

Observation fe01c260-d247-4af5-b785-ac057b13bc24 · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:37:25.911108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:74c8da7e1536d2880658395378800ef11054eaa8477dccadc3ffa09aa03cee5b

Observation ca16e56a-e147-4132-b09d-d7c651eb04ce · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.281080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:3718ae02402b2ef9018a5c4f3d7269c7cce7a62886621dfb2c1b433b27ddeae9

Observation 7731a63b-db2d-49a7-ae06-be6343d86ea3 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.724787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:969988898eea506989c75a391dafdaa32f5bbdc69e010c87280a8b8483c995a9

Observation 331caa8d-4bf1-46d3-9584-531062fe2fce · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.851273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:a206b3a83465377ec6b16764881f55d0bdbb85e5619202053940ed7491027c9e

Observation cdb269cb-b2e4-4c84-8496-dac5bfbdaa65 · inbound

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning cites this paper.

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:36.782313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:36.782313Z digest=sha256:eb2a19e4f385aad566d7504f04b35cdace10445ed59dbb0ceb099684f06ded95

Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.762351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.762351Z digest=sha256:385b03219b31d043253dc6ef1c290517b0ff578abd588c80dacf870eb9a97ea3

Observation 836f1d85-7cb8-4389-ad5a-7a2d6669fde8 · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:18.671551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:18.671551Z digest=sha256:38e66401452e1e47050ee4fb8b693dd8379eb83f6c4480c8c1275cae880c999b

Observation 2195bfaf-1295-49b4-9434-3315e423845d · inbound

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency cites this paper.

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:11.824348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:11.824348Z digest=sha256:9ab48a0e0fe10ebfb72cc90a4f68848ec502164d7c7fa25a3ca94fc7de245908

Observation cc6f6b2f-255c-40f5-b1df-25125e92583b · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.251861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.251861Z digest=sha256:e3c6ede40bcf98893aa2d27709eafe44da39645223165c5ec5c31ace643935cc

Observation 146b0fc2-3f5b-499c-9f4e-1565187e2bd5 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.414900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:2aa6864ccd60d14e174b2b268996bb210f8ad8ed5d5cc364217392f4d31bbeef

Observation b2244a65-2661-4e5a-9c3e-68e7cf52ac17 · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.051188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.051188Z digest=sha256:30bdd9a31f9dd4925eab4ae2bc7dc81898a040e62329e73cc843224f04346eb8

Observation ece6fc9a-60b6-4784-89f6-ca02fa63b724 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.725020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.725020Z digest=sha256:1fc37e5b6764bbeda9f3d72ae52e977a6f9c1a82328109786e44a32fc95f1687

Observation 1a118deb-9a5b-4c1f-b5ab-fbba38a6e8da · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.710762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.710762Z digest=sha256:d087ad851e6c0ef84750a4fa47d906e0528a0f4efb6e6cbc850b9640244674b2

Observation b1e09e38-907f-449a-b072-d1c13ad8f18f · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.017809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:2542afd18bc0e100b90a78fc6c3a8718a12ffe94f14cc060ae22966d57a5730f

Observation 97769deb-fd86-43e5-8f63-1cee51cff3d4 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.788856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.788856Z digest=sha256:aad12cd4411326eda414321b63251b756e8e90c5b1694a415a58ece403a711a2

Observation a2564aa5-037b-4aff-a5fc-20746eea6fa1 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:14.895125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:14.895125Z digest=sha256:28c83a342544267c0f21846a7c6f9edf77bcd9ee8ac24e52b261a5bdcd538b4a

Observation 1e380ce8-469e-41cf-ad6b-d6c707f8fde4 · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:21.113889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:b5966acade68a1ae0100a2a11db012ea402db24868952356fd3b2231c916661d

Observation 44327a09-f177-4137-8c73-f66c3ad7f413 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.702294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:8d252d8045d97a26da8803e56975263ed237d796a11f4a844525cc1e39b4ea47

Observation e2959da3-271d-403f-a559-ddea62488bf4 · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.250374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:b8c8ec92f57659f29971f65a8e52cfcbf46da0cb7fcd30f7f463bbbff781782e

Observation ddd1ef0b-72b4-4ab3-999f-4f7910dcaa41 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.472624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:5359b409ae7d99453e00ba752ebe2211087ce037fdf34335da63dd83154fa002

Observation 16a8f0b1-41c3-4e79-8933-f7731feeafea · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.600822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:a9bea269850aeb387f018ce27fba64e920eef7aa237cc1dff00aa00da8aa0be0