Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2310.05126.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.600852Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T16:18:37.543028Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7359fe39-2afa-448a-b8bd-ed7f0739d030 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e34c82b-4a20-42d7-bde2-f41ceb3c4f19 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 959f748c-3963-4c39-b326-c35b441f9251 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 159
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bf3a85f8-994d-4172-92fc-4ad13c8c5fd8 · inbound
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dc054965-971e-4bc6-872d-ed373c5e9791 · inbound
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3c622375-6f48-4b08-86a0-7efd163a6d25 · inbound
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3187935f-0a0e-43a9-be6f-1c4c72f363f7 · inbound
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66bab495-d4ab-4ccd-b68f-a59783f6139f · inbound
Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c95798-9e2b-41f5-b04c-686960fd8bd6 · inbound
FILA: Fine-Grained Vision Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f687888-e6cb-4ea7-8640-113138fd9b41 · inbound
DocVLM: Make Your VLM an Efficient Reader UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea32079-4f42-49f9-8e56-720633a0d8b3 · inbound
DocFusion: A Unified Framework for Document Parsing Tasks UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00f6c14-9e11-4af5-a241-2cfa952461de · inbound
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 111
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d0c6cb-54eb-499f-adce-c11763938a7a · inbound
InstructOCR: Instruction Boosting Scene Text Spotting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf827e25-0aed-4782-8627-8dfd9377e165 · inbound
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b97aaed-ddd9-4dcb-b263-fdce21bd4bb3 · inbound
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bbae7ba5-2786-4931-9df5-154052e9c283 · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 375c2eee-b73a-4cbc-8f92-e376e4caa9b2 · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6ce3d4-4ac3-4fef-8fcc-5f7d52ab49e8 · inbound
Visual Large Language Models for Generalized and Specialized Applications UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e807b4-543c-4b27-8c86-056d1dee2615 · inbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f27e3e-d5aa-4e3c-98fc-813fd950510a · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be00772f-ab01-4e96-9645-29dd506a6267 · inbound
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e12b97-07d3-408d-9d50-b077d70af03e · inbound
OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2354dc82-88e1-4cb8-ac30-81994c8141c4 · inbound
Return of the Encoder: Maximizing Parameter Efficiency for SLMs UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a489b425-3008-48d3-9604-c616a7e11f3f · inbound
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797929cd-e1e4-421a-ba37-31d011647556 · inbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a8e5eb-970f-4e92-ad66-8d2e24d4f79a · inbound
Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14b27b0-8113-4d7f-8a65-b09e36415b5a · inbound
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0190d4b3-deae-47a2-960d-9ef4e1052993 · inbound
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853d8cd9-fce2-459e-a3ef-e6e69e8209b9 · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9c12fa-4573-4023-a931-b729e6d14542 · inbound
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d962ee53-3e50-4bc5-864a-9d9354d1a2b4 · inbound
Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 074c23a4-80a0-40e5-8eb7-63c922968231 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.