Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.04473.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:36.782313Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T19:43:23.773160Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 541007d1-b508-4c6d-8895-89630c02dfa3 · inbound
A Survey on Multimodal Large Language Models TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dce1f58e-15c5-4d58-9261-e90500800966 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 77c8df6e-1d1c-461e-89b0-5599283f556b · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0d22a99-2dca-4228-b9e2-14d19b081665 · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f91f9c5-3858-4714-b421-252add1f24f6 · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3032a2af-cde3-4694-8376-af8e8cfababc · inbound
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a11461b-8a19-4f64-b893-9cb44861c518 · inbound
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe01c260-d247-4af5-b785-ac057b13bc24 · inbound
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ca16e56a-e147-4132-b09d-d7c651eb04ce · inbound
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7731a63b-db2d-49a7-ae06-be6343d86ea3 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 331caa8d-4bf1-46d3-9584-531062fe2fce · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cdb269cb-b2e4-4c84-8496-dac5bfbdaa65 · inbound
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · inbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836f1d85-7cb8-4389-ad5a-7a2d6669fde8 · inbound
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2195bfaf-1295-49b4-9434-3315e423845d · inbound
Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6f6b2f-255c-40f5-b1df-25125e92583b · inbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146b0fc2-3f5b-499c-9f4e-1565187e2bd5 · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2244a65-2661-4e5a-9c3e-68e7cf52ac17 · inbound
Describe Anything Model for Visual Question Answering on Text-rich Images TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece6fc9a-60b6-4784-89f6-ca02fa63b724 · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a118deb-9a5b-4c1f-b5ab-fbba38a6e8da · inbound
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e09e38-907f-449a-b072-d1c13ad8f18f · inbound
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97769deb-fd86-43e5-8f63-1cee51cff3d4 · inbound
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2564aa5-037b-4aff-a5fc-20746eea6fa1 · inbound
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e380ce8-469e-41cf-ad6b-d6c707f8fde4 · inbound
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44327a09-f177-4137-8c73-f66c3ad7f413 · inbound
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2959da3-271d-403f-a559-ddea62488bf4 · inbound
ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ddd1ef0b-72b4-4ab3-999f-4f7910dcaa41 · inbound
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16a8f0b1-41c3-4e79-8933-f7731feeafea · inbound
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.