Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:17:28.096510Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 4 inbound Pith citation observations for arXiv:2501.02235.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:17:28.096510Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:55:42.901109Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T04:42:04.421908Z
100 of 110 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a87a521c-e834-4c86-8d73-1b7fc2f1b674 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97ef716-8afa-4527-ae6c-5a1e1a521ce9 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f028033-ebc7-4af0-9d3b-00e8cafeae53 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends ETC: Encoding Long and Structured Inputs in Transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4fe70d7-730d-44ec-a9d7-0bbe43d0ba8c · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fdf7032-bfc8-4664-b3fa-fdda4914506e · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Flamingo: a Visual Language Model for Few-Shot Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31782bb-a13f-4b1f-b46a-5df1dbbe8e02 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocFormer: End-to-End Transformer for Document Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc4dcb05-74b7-407e-ab3b-6b50f37df76f · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocFormerv2: Local Features for Document Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5522fe15-c9fd-4eb3-839e-a640778e2d68 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce5017d-ae49-4033-b08b-d5fe46fbd72d · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends GRAM: Global Reasoning for Multi-Page VQA
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ceb1194-763e-4333-a493-58cae499d38e · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Nougat: Neural Optical Understanding for Academic Documents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7365d8-7ece-469f-9786-363106f73bed · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Arctic-TILT. Business Document Understanding at Sub-Billion Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d80531d-da38-4266-b14c-f2c504a58cc2 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ebb039-5d0e-46ed-b1b0-eaee596f8ea9 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends End-to-End Object Detection with Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce00d56-c995-4d7b-b8cc-998ca2f0ef98 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Honeybee: Locality-enhanced Projector for Multimodal LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d19ddcd-8134-4d6a-bcfb-06d07fa22729 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f844ec4a-2413-4aff-b8b1-5606877cd5ca · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends A Simple and Effective Positional Encoding for Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671a0c85-5eee-4c68-bcdb-5b3249cb232b · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UNITER: UNiversal Image-TExt Representation Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc1a2e8-291e-4aa2-ab54-d3a8adb086c5 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d9c5f0f-aba5-42e8-8937-f1436a47afb6 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbbd4caf-45b3-4356-a7cc-f6619b9dd0ed · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f5e7bc3-c235-4d08-a821-657d948ca559 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends End-to-end Document Recognition and Understanding with Dessurt
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c1ef9c-8ae0-4295-b6c2-ab2e0237dd5a · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TURL: Table Understanding through Representation Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb73613-b01f-495c-bbd8-aa0f4b58c866 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocParser: End-to-end OCR-free Information Extraction from Visually Rich Documents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aedf4986-3e8a-4395-982f-8252b9bd0f4d · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Deep Learning based Visually Rich Document Content Understanding: A Survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f07980a-0123-4c16-817d-08848457040d · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08e4b647-bd1e-4562-a1c9-ced97488f4b5 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 212c6106-97e8-4c8f-a7b7-e2aa6176e97a · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends ColPali: Efficient Document Retrieval with Vision Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfec1f6-86e8-4190-8d26-71614137f408 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933f771a-b857-4313-bd4a-d0c0c2c23601 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb67e5b8-9347-4fe9-a764-99d2ead83b82 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0674f496-1aea-4dd0-874e-205e988f1887 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52252601-a075-4735-ab8b-826497012417 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c25711d-786e-48b7-8621-cfa1da1a74b2 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Recurrent Memory Transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d31d31-fe16-4d7f-9b15-afaae4507a37 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4841cf7d-fd52-4d7c-9dd1-52dbefd08ce2 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217d370a-9e1e-4a49-8939-8182104f7c3c · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends CogAgent: A Visual Language Model for GUI Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293a213d-5bfb-4193-a461-b7e04c19e7f7 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2130ea05-3bd0-4f24-b5f2-f834062031b3 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473e2b7f-a936-4fdf-a9d9-8f2cc0945bfd · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocMamba: Efficient Document Pre-training with State Space Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bb04131f-41b7-41b4-b2cb-fe9ee3b324ac · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 921d78d4-9d5c-4ed5-af6c-1612d4efa661 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be746ae-776c-4683-94bf-deb8101e035e · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends OCR-free Document Understanding Transformer
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d07c202-f27d-46ba-9984-8437a160391e · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Document Understanding Dataset and Evaluation (DUDE)
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa021069-4a9b-48bb-9cf6-9aebacad6cc4 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends What matters when building vision-language models?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff5fb33-c7ed-4ee1-a598-68b3bff6e210 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d343a2fb-2145-4c82-ab37-cec424bc1cec · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7c3aee-0a59-4cdb-b524-4588676e8914 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a276cd64-bbde-48a1-914f-20e41eea28b8 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends StructuralLM: Structural Pre-training for Form Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8090ed47-0e84-4d48-bbbd-307b8dd3ceb1 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends 2D-TPE: Two-Dimensional Positional Encoding Enhances Table Understanding for Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 940e10ee-762a-4aff-93d6-79c8f2f618d9 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DiT: Self-supervised Pre-training for Document Image Transformer
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78e5d51d-411e-48f0-8235-6f0ef20d85f5 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca0d8ba-6634-4cc8-b16f-9c3b3959f829 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends SelfDoc: Self-Supervised Document Representation Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71909ff6-8976-435a-8fde-f3e3fed2aac0 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c89e9c-8933-49ab-be6e-bb6d86d622ff · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Li, Xiantao Cai, Bo Du, and Hai Zhao
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7933008f-4111-4d9d-b21e-4b86bde448c7 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c0552b-b785-414c-8216-a17b76ff1262 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017efb57-c07a-4094-9927-3e3876af7e55 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocTr: Document Transformer for Structured Information Extraction in Documents
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 195cf54e-b498-4da7-9066-a0ab28b3dd3f · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678ba77d-103b-4b67-9a82-425fee3c4328 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8278004f-44fd-44d6-b2b9-f1d83de8f5ca · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends HRVDA: High-Resolution Visual Document Assistant
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a8ca76-d402-4570-861e-7ce36df88bbb · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DePlot: One-shot visual language reasoning by plot-to-table translation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df51213-601d-4b76-9bcd-678678c3e131 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends The Devil is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b40beacd-a01f-4471-96be-11bfc57346a6 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1a6122-37c8-4037-813a-31867920994a · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Swin Transformer V2: Scaling Up Capacity and Resolution
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d20a12-8dbd-4391-a271-9cee8c121a02 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbfde75-2dbb-48b7-94af-bed52354636e · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e5db2af-59d7-494f-b743-b43cde59b289 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f1f4dd-0b04-4f43-867f-71dad31aed6a · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends KOSMOS-2.5: A Multimodal Literate Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 059369a6-f95f-460d-a178-cdab4bdbbb61 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ff63e4-f94b-4245-9fff-2ff617cce2dc · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unifying Multimodal Retrieval via Document Screenshot Embedding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efecb521-6231-4109-a8cc-985719cd36f6 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14cde736-0904-48f8-873d-b081ece20785 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Visually Guided Generative Text-Layout Pre-training for Document Intelligence
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbbfbcab-553b-455a-a67e-1aded7501099 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends InfographicVQA
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a4441b-ee2c-4614-a220-41721203a39b · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocVQA: A Dataset for VQA on Document Images
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f4addf-03e1-4410-a769-b221d0fe2b0b · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9951ccd0-0ec7-44b1-912b-eb258b88f30f · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e3ecd4bf-a7a8-41e4-8219-0997b3b4dea3 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7febd7-dad6-4552-99d3-6312fd4e8aba · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38f74c5-ecdf-4564-8cea-abeb30d608d6 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6f2ecf-7b9b-4cc9-9406-567752ae3855 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b20da880-08c3-4c1b-a77d-e83546f01594 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7d0f5b-b5ea-4b20-9ba1-b6849bfa5cff · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a7c3ed-5353-4cde-af76-5b8a6d5426db · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14d7e133-ac42-42e1-a486-6c9174a5d576 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9677bfae-716b-49be-a318-1b2a5fb6300c · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cdb31871-084f-40e8-b7a7-650e2162565d · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 445f9482-1cdc-4302-9486-6527b0ef5b59 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends VisualMRC: Machine Reading Comprehension on Document Images
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 57dc6faa-2bb4-4db8-87f5-71ffba39ffa9 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d8a585-a1e6-483d-b932-cf459dcb0afb · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unifying Vision, Text, and Layout for Universal Document Processing
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a378fd44-d891-4df3-b5a1-949fee1ebd35 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Hierarchical multimodal transformers for Multi-Page DocVQA
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53799283-9fef-4ff8-93b2-826388d0ad77 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e74542a-568d-4e1c-8f20-96874b93fa0a · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31be4889-c881-4221-b3df-9f70fc11967f · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfcd838-d87c-49f2-bbd7-cf1f88539ebb · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0501f79-02b3-4fef-b600-b3d363282b86 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc0dc13-f1e2-433b-a0d7-11645661b412 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beead3ab-b136-44b2-b759-f9101ed01b89 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defa629e-8143-497f-8995-6a7095da0f6b · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375c2eee-b73a-4cbc-8f92-e376e4caa9b2 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609e6b1e-9562-442f-b398-b04fa4296387 · outbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b6da64-0b84-489c-b57c-7971bed9bdb7 · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a011b765-b42d-4abb-b59e-16d2341360cb · inbound
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 18963410-65f3-4049-80f0-aa3e7a05fe89 · inbound
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4f9d30-3599-415f-be88-369f081809d4 · inbound
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends
Reference 221
Source-reported events for the cited work
Unavailable: canonical work link unavailable.