Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2404.09204.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.501180Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:13:15.344020Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 03c3bda6-d15d-4433-b521-9eefbf563a73 · inbound
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4ce4de7-7a45-411c-a2a9-754906917c48 · inbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c9bc405-f105-4bd6-abc5-9ae0a315a36c · inbound
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45f6205c-b1ea-48cc-969b-71dca94e7c7a · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f857a2fb-c246-49ff-8285-92827b57a347 · inbound
Visual Large Language Models for Generalized and Specialized Applications TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4d8a90-9c17-4568-987c-dd50b5a74b1f · inbound
RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab3b555-d38c-47f6-afe0-f98bd8f69ca1 · inbound
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd675666-0726-40a7-ada7-dcd14b2e0ccb · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da945551-70eb-4509-a555-6964945401cb · inbound
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf2c4e41-add8-435f-831c-2cee03d3ce58 · inbound
Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ded1b3-bba6-442c-9ca4-9370fca55ed5 · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4dd574f-1f6e-4058-a219-a9e9662bea66 · inbound
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b5d9ef0d-62c2-4bed-9758-5cc8ab1a50a9 · inbound
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f2017440-6d04-4e6e-98ab-582e3ae656eb · inbound
ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c1e3a305-7ffb-4e9a-9c0a-8ffd881c53d2 · inbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 619ea8c1-1590-46c9-a539-750845217293 · inbound
Grounded 3D-Aware Spatial Vision-Language Modeling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.