Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:51:41.504443Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2412.18327.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:51:41.504443Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b13f1780-b852-4c66-81f7-29c62b1203d3 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Textcaps: a dataset for image captioning with reading compre- hension,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4d15c7f0-4683-4d32-a25d-048bd19f9614 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Towards vqa models that can read,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fce03ac-160f-474e-be42-57b4f41b65a5 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Scene text visual question answering,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 947b8f29-7617-4143-81be-1827ba2f137e · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6845cd48-f542-4935-a45d-6b161b81db68 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Visualmrc: Machine reading comprehension on document images,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b9f4e125-59c9-4e21-a69d-a6509f3abd84 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Textbook question answering under instructor guidance with memory networks,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9641ece4-cc8e-448f-8941-46004f050281 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad4ae96-5f80-4f18-82d8-fd7c1b2137ee · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Docvqa: A dataset for vqa on document images,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4685e783-8a7c-43ea-9679-20d2ad89f37c · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Infographicvqa,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cdc5accd-ca60-4c84-92e9-bf5a8ae19529 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Iterative answer prediction with pointer-augmented multimodal trans- formers for textvqa,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 90cfb49b-6902-4e11-9324-50a9f6011988 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Tap: Text- aware pre-training for text-vqa and text-caption,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 764ffbff-1bd1-4b1c-a658-d51cfe931b0b · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Vilt: Vision-and-language transformer without convolution or region supervision,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c417c1f9-1530-4def-a164-59ff8f68653c · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41fc879e-a193-4ee4-a796-5aae5b8012bf · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee8ae27-66d3-44e4-af24-911f9d44c533 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b162b25a-c8ac-4f19-b3c6-a53157779eba · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Pix2struct: Screenshot parsing as pretraining for visual language understanding,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e2eea61b-064d-4b38-ba10-864913a2286c · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce0abf7-27f0-4f9c-bbbc-b94bc4fb3574 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Improved baselines with visual instruction tuning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf4b0ec-05e8-4315-96cb-dd3096ae5d26 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f5a895-1ada-4aa3-95e1-5dca3a1ef7b4 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b65599-22bc-4802-b0fd-8dc5b7e743b1 · outbound
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Unresolved cited work
Reference 500
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.