Pith. sign in

Paper Citation Record · LEDGER

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

As of 11 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 4 inbound Pith citation observations for arXiv:2501.02235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02235 v2

Coverage vector

measured 100 of 110 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:17:28.096510Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:55:42.901109Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:42:04.421908Z

Reference resolution

100 of 110 outbound references displayed

  • verified exact13
  • verified fuzzy0
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a87a521c-e834-4c86-8d73-1b7fc2f1b674 · outbound

This paper cites online" 'onlinestring :=.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.366319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.366319Z digest=sha256:7a4d96a9ecb3beabf129877b852b11943e44b6963f270a9c26737968cfcdcccf

Observation f97ef716-8afa-4527-ae6c-5a1e1a521ce9 · outbound

This paper cites write newline.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.374570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.374570Z digest=sha256:2f4c13b5965183ea361b83921b2a5a7eb482956769bc725207dec3429b052e63

Observation 2f028033-ebc7-4af0-9d3b-00e8cafeae53 · outbound

This paper cites ETC: Encoding Long and Structured Inputs in Transformers.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends ETC: Encoding Long and Structured Inputs in Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.383485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.383485Z digest=sha256:7c200a291da86f753adf052dd57ade67aa31098b442f546a80b91ecd2aee5cb5

Observation b4fe70d7-730d-44ec-a9d7-0bbe43d0ba8c · outbound

This paper cites Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.390834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.390834Z digest=sha256:070353ac15edfc2237a4059545da65d8c167f7fb9fd72f503b5127c818adf932

Observation 8fdf7032-bfc8-4664-b3fa-fdda4914506e · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Flamingo: a Visual Language Model for Few-Shot Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.398595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.398595Z digest=sha256:7e51f9694490885ea7aadb337f9ba773e5726066f76c88f14e3b5c25b0d083b8

Observation b31782bb-a13f-4b1f-b46a-5df1dbbe8e02 · outbound

This paper cites DocFormer: End-to-End Transformer for Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocFormer: End-to-End Transformer for Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.407100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.407100Z digest=sha256:9f9fbb6d0ca7013672e12eab8561efeb6c4df0e7788ccea4cb9cfa04054c2410

Observation cc4dcb05-74b7-407e-ab3b-6b50f37df76f · outbound

This paper cites DocFormerv2: Local Features for Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocFormerv2: Local Features for Document Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.414864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.414864Z digest=sha256:27bc5589b6634d2f76b742f637cda0be5e744ffc1bf077d4926339c2d20e6796

Observation 5522fe15-c9fd-4eb3-839e-a640778e2d68 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.422891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.422891Z digest=sha256:46380b3c76fa3088e48392bc7ed2b3467c205bdca0111bec0dd4b0cfd96a6176

Observation 9ce5017d-ae49-4033-b08b-d5fe46fbd72d · outbound

This paper cites GRAM: Global Reasoning for Multi-Page VQA.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends GRAM: Global Reasoning for Multi-Page VQA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.431622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.431622Z digest=sha256:458230bd3eadc435994134ce5170da5fce443e6ea432dd0d932b5acdf3922cd5

Observation 3ceb1194-763e-4333-a493-58cae499d38e · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Nougat: Neural Optical Understanding for Academic Documents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.438492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.438492Z digest=sha256:0bb145387d440c34ce6ef1561f5a5ed1e386c7522d52d8932e26a37e21fcb837

Observation 3e7365d8-7ece-469f-9786-363106f73bed · outbound

This paper cites Arctic-TILT. Business Document Understanding at Sub-Billion Scale.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Arctic-TILT. Business Document Understanding at Sub-Billion Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.445253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.445253Z digest=sha256:a1442f19bbc7e8dde6654b3a585ae3f617833bd0d1a5bb086ee7ff6210157e11

Observation 1d80531d-da38-4266-b14c-f2c504a58cc2 · outbound

This paper cites Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region Concentration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.452347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.452347Z digest=sha256:17ab9be903ae5b7afb460065c0a6073cacf4994eacdcc7d59484d6c9955cf78c

Observation b9ebb039-5d0e-46ed-b1b0-eaee596f8ea9 · outbound

This paper cites End-to-End Object Detection with Transformers.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends End-to-End Object Detection with Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.459874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.459874Z digest=sha256:ae88dacbff6e659655eb37dec57aca5589937fa3a96f6b7e9530576fa64cf6bd

Observation 1ce00d56-c995-4d7b-b8cc-998ca2f0ef98 · outbound

This paper cites Honeybee: Locality-enhanced Projector for Multimodal LLM.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Honeybee: Locality-enhanced Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.466019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.466019Z digest=sha256:c6ee562d96bcfa5ab465ea3f94c171962cd30e0d9c9f6d61a257369fe9bc8c65

Observation 5d19ddcd-8134-4d6a-bcfb-06d07fa22729 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.472844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.472844Z digest=sha256:04073b3bff7afcdb73728650658fe42a66bc039f585364818aed168ebbe14804

Observation f844ec4a-2413-4aff-b8b1-5606877cd5ca · outbound

This paper cites A Simple and Effective Positional Encoding for Transformers.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends A Simple and Effective Positional Encoding for Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.479336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.479336Z digest=sha256:a5fff155d4abb7a5c9fc10d38478e5b65c346b9eca5898048248e57672ba3a5c

Observation 671a0c85-5eee-4c68-bcdb-5b3249cb232b · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UNITER: UNiversal Image-TExt Representation Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.486497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.486497Z digest=sha256:a5384614b6d0883fdcc0b2af33426d7460dce1e4bab22fe3c72524b6c18dd2fd

Observation 0cc1a2e8-291e-4aa2-ab54-d3a8adb086c5 · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.493239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.493239Z digest=sha256:e4b9c1811bee92059887e8acef830f34efa1bd7d2eac5d4cc1c5b2a46d67a8a5

Observation 3d9c5f0f-aba5-42e8-8937-f1436a47afb6 · outbound

This paper cites M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.499998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.499998Z digest=sha256:ab328ec35d923b270938d06bde97da066dc64376a2440230c9bdfb36009ee2c1

Observation dbbd4caf-45b3-4356-a7cc-f6619b9dd0ed · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.510035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.510035Z digest=sha256:2d41c9e1bb41e64a4d95f534d51bb2f6cea0c66c16e014369aa0efdb269e0a8a

Observation 1f5e7bc3-c235-4d08-a821-657d948ca559 · outbound

This paper cites End-to-end Document Recognition and Understanding with Dessurt.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends End-to-end Document Recognition and Understanding with Dessurt

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.519732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.519732Z digest=sha256:68e8462768dee3f0519ae2b3fd3725e3b1b5182bab745f9a4d1141fe3f63a58b

Observation 45c1ef9c-8ae0-4295-b6c2-ab2e0237dd5a · outbound

This paper cites TURL: Table Understanding through Representation Learning.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TURL: Table Understanding through Representation Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.527458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.527458Z digest=sha256:73fecaf73fa96282edba8ac0af400c15e8bab77b9382e1ec109c246478e4b991

Observation 7cb73613-b01f-495c-bbd8-aa0f4b58c866 · outbound

This paper cites DocParser: End-to-end OCR-free Information Extraction from Visually Rich Documents.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocParser: End-to-end OCR-free Information Extraction from Visually Rich Documents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.536554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.536554Z digest=sha256:d41a98c4f56e1b863632c3f04afbcdff428a5a723d843ff698b94db131db7466

Observation aedf4986-3e8a-4395-982f-8252b9bd0f4d · outbound

This paper cites Deep Learning based Visually Rich Document Content Understanding: A Survey.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Deep Learning based Visually Rich Document Content Understanding: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.543243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.543243Z digest=sha256:4ec1cc24b8725a6804f3d75b8b070724c279bd6eda201adbd179be3ad96a3e84

Observation 2f07980a-0123-4c16-817d-08848457040d · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.551174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.551174Z digest=sha256:44ba7dd35cb540bc6677e75997f7ddcc55dd19cf77e9b78e7c0f4ae270174e96

Observation 08e4b647-bd1e-4562-a1c9-ced97488f4b5 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.557702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.557702Z digest=sha256:8e06bd5a4656e4805ac0f226721ee035e4572932595e24718ba2d60dc2335ce3

Observation 212c6106-97e8-4c8f-a7b7-e2aa6176e97a · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends ColPali: Efficient Document Retrieval with Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.564498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.564498Z digest=sha256:a09cce688a2e6e9f654f945a0642cd48bae97e3a7dfed7cbef26fe5c26715f44

Observation 6cfec1f6-86e8-4190-8d26-71614137f408 · outbound

This paper cites DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.570746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.570746Z digest=sha256:4548d475567ccacb8c74119347c43117d9aaf6d7c8abb571d487007c738bd0fa

Observation 933f771a-b857-4313-bd4a-d0c0c2c23601 · outbound

This paper cites UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.577212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.577212Z digest=sha256:3a63dfef5cd9a1018f6ba0c175f34023fbb6a3c4188cc5cd3097533e0d4c772a

Observation cb67e5b8-9347-4fe9-a764-99d2ead83b82 · outbound

This paper cites LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:30.786704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.584075Z digest=sha256:bb5e3f4b1ec88a53a064c13a06c601ad8f1bbd9fd7d5bf9bfc6cefde04388449

Observation 0674f496-1aea-4dd0-874e-205e988f1887 · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.591368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.591368Z digest=sha256:16bb3f3a1b3faac1201616784510706eb46f6b9ab74dbf521a571a00d8d291e7

Observation 52252601-a075-4735-ab8b-826497012417 · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-10T22:17:30.750564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.599532Z digest=sha256:ee6f9a110896b6a18fdd633d91ce435664836940527d5bc982e12d0f80f4de74

Observation 6c25711d-786e-48b7-8621-cfa1da1a74b2 · outbound

This paper cites Recurrent Memory Transformer.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Recurrent Memory Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.607101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.607101Z digest=sha256:4ff0092fc71242646d551883e6e644c3f84dc1348ca20ee7ce2a0542fc08f249

Observation 67d31d31-fe16-4d7f-9b15-afaae4507a37 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.615379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.615379Z digest=sha256:c635129989ec971c12ef49d2a52c2c694044bc38a45af1fb7bbd171e3207f276

Observation 4841cf7d-fd52-4d7c-9dd1-52dbefd08ce2 · outbound

This paper cites BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.622203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.622203Z digest=sha256:f44c7687af8906d56188fc26c4803c6c03d31480505f57e507f5f22fd2def597

Observation 217d370a-9e1e-4a49-8939-8182104f7c3c · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends CogAgent: A Visual Language Model for GUI Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.630693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.630693Z digest=sha256:844b8e5ea5de56093c6d75633fd5eb491ac3af9110a0ce94f0cc16651a8bf241

Observation 293a213d-5bfb-4193-a461-b7e04c19e7f7 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.639579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.639579Z digest=sha256:6471889f7594eb2c71990ee3df19c99cf68ab788b8aaf4db60f12140e108ebe3

Observation 2130ea05-3bd0-4f24-b5f2-f834062031b3 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.645608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.645608Z digest=sha256:3ac98c0fbd9b86554901a0447f01f83a2d12c063b5db6d87d36e55986291a9ea

Observation 473e2b7f-a936-4fdf-a9d9-8f2cc0945bfd · outbound

This paper cites DocMamba: Efficient Document Pre-training with State Space Model.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocMamba: Efficient Document Pre-training with State Space Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:30.424784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.653498Z digest=sha256:504ed6cc883f6bababd2b0fcee07b7e548f8cfcfe442006451bb26d5418349ab

Observation bb04131f-41b7-41b4-b2cb-fe9ee3b324ac · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.668883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.668883Z digest=sha256:993cea2e5e0f23a9351af2cff813558303d0e85fa0cf153878aefa160c711d74

Observation 921d78d4-9d5c-4ed5-af6c-1612d4efa661 · outbound

This paper cites LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.678099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.678099Z digest=sha256:0dd4315a57e5b23b3b107e82e9b522ac859c4047dd2d87079a5d324d05686ad3

Observation 9be746ae-776c-4683-94bf-deb8101e035e · outbound

This paper cites OCR-free Document Understanding Transformer.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends OCR-free Document Understanding Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.686894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.686894Z digest=sha256:3bafa2c708230c55fc8d37694213b6318a3f0b6dbb73bb88b6c2201db4d98147

Observation 3d07c202-f27d-46ba-9984-8437a160391e · outbound

This paper cites Document Understanding Dataset and Evaluation (DUDE).

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Document Understanding Dataset and Evaluation (DUDE)

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.698732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.698732Z digest=sha256:4001cb54f1f043d226054d080105386042658114f87864e7811fce43e6f85b9d

Observation fa021069-4a9b-48bb-9cf6-9aebacad6cc4 · outbound

This paper cites What matters when building vision-language models?.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends What matters when building vision-language models?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.705612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.705612Z digest=sha256:e58e47fd361b222f8c40a81b07af8bdc8ecf5659ab3cd72f6f6d3c9fec3cc13b

Observation aff5fb33-c7ed-4ee1-a598-68b3bff6e210 · outbound

This paper cites FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:30.247387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.714235Z digest=sha256:53712a7c32b357adfb119587c124cc2d60f4ee5231f898761a888cff9264a0ae

Observation d343a2fb-2145-4c82-ab37-cec424bc1cec · outbound

This paper cites Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.721321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.721321Z digest=sha256:bc13b6e0b78a837f42498eebb950809b33186d790943230ff5eeee2b3774739d

Observation 2a7c3aee-0a59-4cdb-b524-4588676e8914 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.727769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.727769Z digest=sha256:00cc971886c7c8fda92cc2bb9f995af590b4a5c0963208b60a627f67399334ed

Observation a276cd64-bbde-48a1-914f-20e41eea28b8 · outbound

This paper cites StructuralLM: Structural Pre-training for Form Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends StructuralLM: Structural Pre-training for Form Understanding

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:30.150455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.734958Z digest=sha256:c104c8c89e33bc1016f4a2677b5991afeb35c3be7e09a97bb57dc2d8db80859a

Observation 8090ed47-0e84-4d48-bbbd-307b8dd3ceb1 · outbound

This paper cites 2D-TPE: Two-Dimensional Positional Encoding Enhances Table Understanding for Large Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends 2D-TPE: Two-Dimensional Positional Encoding Enhances Table Understanding for Large Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:30.102896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.742135Z digest=sha256:45d1ab87f33e0fc20cef9ee51b83b6169877b1a8bc71760a1a094d13a7f42f78

Observation 940e10ee-762a-4aff-93d6-79c8f2f618d9 · outbound

This paper cites DiT: Self-supervised Pre-training for Document Image Transformer.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DiT: Self-supervised Pre-training for Document Image Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.749896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.749896Z digest=sha256:f08bb74945ac726accbb95db5aa259da7798a84ed1600fe08f46a131d1686c63

Observation 78e5d51d-411e-48f0-8235-6f0ef20d85f5 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.757395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.757395Z digest=sha256:e8a7e8db1cc52bb6845248b966cd101ca929aa02a3e6c6737d57021c71122b68

Observation 4ca0d8ba-6634-4cc8-b16f-9c3b3959f829 · outbound

This paper cites SelfDoc: Self-Supervised Document Representation Learning.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends SelfDoc: Self-Supervised Document Representation Learning

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T22:17:29.990042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.763901Z digest=sha256:11d3e982f4b23c3a0f6236fe03fca5dc67dc06ea68545d115223ad83b27a553a

Observation 71909ff6-8976-435a-8fde-f3e3fed2aac0 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.769861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.769861Z digest=sha256:867d235c0811efdd89f6e2d0d36a861517ff3ff99ad6239c7ee0bcd0d5f1a401

Observation 04c89e9c-8933-49ab-be6e-bb6d86d622ff · outbound

This paper cites Li, Xiantao Cai, Bo Du, and Hai Zhao.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Li, Xiantao Cai, Bo Du, and Hai Zhao

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.776999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.776999Z digest=sha256:91c14cd473210619f57f564c117557a8028600d896416246fbc3066e24d22c18

Observation 7933008f-4111-4d9d-b21e-4b86bde448c7 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.783705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.783705Z digest=sha256:891dc45def91d11ff712801c25e6bb679dcbe25680e8f71715206a2293ddc1c9

Observation 42c0552b-b785-414c-8216-a17b76ff1262 · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.790703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.790703Z digest=sha256:430ffdf995544acb4bf12de026502071d9d1f09d4e5690696902e4fa43809b50

Observation 017efb57-c07a-4094-9927-3e3876af7e55 · outbound

This paper cites DocTr: Document Transformer for Structured Information Extraction in Documents.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocTr: Document Transformer for Structured Information Extraction in Documents

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T22:17:29.877415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.796840Z digest=sha256:f9743aa7956e93cddaf71dbbc98db605666c2457d6b4c2b6280a77228128affd

Observation 195cf54e-b498-4da7-9066-a0ab28b3dd3f · outbound

This paper cites DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.802900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.802900Z digest=sha256:fbb7905485001415cf8f6c25f99ee38b6839dbe2e663c6b7c6521ea982902ec0

Observation 678ba77d-103b-4b67-9a82-425fee3c4328 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.808805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.808805Z digest=sha256:21b9155b653b306b1c4a6c14c4b706bb6cf63951a2bbaef2e09fd3fe0f8d80c8

Observation 8278004f-44fd-44d6-b2b9-f1d83de8f5ca · outbound

This paper cites HRVDA: High-Resolution Visual Document Assistant.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends HRVDA: High-Resolution Visual Document Assistant

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.815410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.815410Z digest=sha256:351b3d3ff305f9a3135fd2b38e1dc38229145e0c2b8fedb9a00310ae5492a9e2

Observation 86a8ca76-d402-4570-861e-7ce36df88bbb · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.822404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.822404Z digest=sha256:1d8aa471725c177d621d41086cdde3955898248f70c0ceadf2e48bfb3683e066

Observation 4df51213-601d-4b76-9bcd-678678c3e131 · outbound

This paper cites The Devil is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-Training.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends The Devil is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-Training

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:29.746407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.828435Z digest=sha256:e5bdbdccee993d2378af5789c828054f9e713de248adac271890a87d297b375b

Observation b40beacd-a01f-4471-96be-11bfc57346a6 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.834093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.834093Z digest=sha256:254febd52c77f818fc850c789779497b33574eb3e0879f8c36bbc30bf29fbcdd

Observation 6c1a6122-37c8-4037-813a-31867920994a · outbound

This paper cites Swin Transformer V2: Scaling Up Capacity and Resolution.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Swin Transformer V2: Scaling Up Capacity and Resolution

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.840714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.840714Z digest=sha256:2c131b945bf7749c10829ceb3f0641401ebfde10a2491268265f1eb96810b590

Observation 42d20a12-8dbd-4391-a271-9cee8c121a02 · outbound

This paper cites Swin Transformer: Hierarchical Vision Transformer using Shifted Windows.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.846947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.846947Z digest=sha256:3d1845a987d325d225aea838a8888fb1dde5b3119b3175dcdcc31982faf4cf7f

Observation 6cbfde75-2dbb-48b7-94af-bed52354636e · outbound

This paper cites Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.852506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.852506Z digest=sha256:0ddb8508d0478e913d2457c1c36e5433e35c4a3c31275d48c08af83a703ca781

Observation 2e5db2af-59d7-494f-b743-b43cde59b289 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.859474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.859474Z digest=sha256:37125484cb31c458719a14a9ab29e52eb88b88c08a3a9e4806faf6c16fa4cd7f

Observation 46f1f4dd-0b04-4f43-867f-71dad31aed6a · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends KOSMOS-2.5: A Multimodal Literate Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.865103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.865103Z digest=sha256:02178e03dc06f485115abd1d5c0b939ee408f3e6e9b101d6fbcad69076d48e01

Observation 059369a6-f95f-460d-a178-cdab4bdbbb61 · outbound

This paper cites EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.871393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.871393Z digest=sha256:0c934951bd4370f52de769d515cc0c78032c517482757c4a123b4c566d10a7ef

Observation 52ff63e4-f94b-4245-9fff-2ff617cce2dc · outbound

This paper cites Unifying Multimodal Retrieval via Document Screenshot Embedding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unifying Multimodal Retrieval via Document Screenshot Embedding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.878444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.878444Z digest=sha256:ee3eb24eb72a4d52ad8660b56ca293449244bff05deff5c7884a3f38b0e46259

Observation efecb521-6231-4109-a8cc-985719cd36f6 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.885274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.885274Z digest=sha256:3fa0e59126d2dabae751cb13e57b0547d65b177ab04f24e2492683bf6ea1ff36

Observation 14cde736-0904-48f8-873d-b081ece20785 · outbound

This paper cites Visually Guided Generative Text-Layout Pre-training for Document Intelligence.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Visually Guided Generative Text-Layout Pre-training for Document Intelligence

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:29.491922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.893056Z digest=sha256:1025b8bf4e2fb6bfa3996923a72ddc1b5a904a2b827d79629f5a409914431ff4

Observation cbbfbcab-553b-455a-a67e-1aded7501099 · outbound

This paper cites InfographicVQA.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends InfographicVQA

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.900588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.900588Z digest=sha256:f9a8ec7ee52f1a28124b7bf8386f4076e5bb261ba33c3f751e5e166e253cb2a6

Observation 40a4441b-ee2c-4614-a220-41721203a39b · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocVQA: A Dataset for VQA on Document Images

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.907885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.907885Z digest=sha256:dfacff08372bb70df7501edb2d55e1865a619bcf6aeb1395a39bae3cff384b5c

Observation 98f4addf-03e1-4410-a769-b221d0fe2b0b · outbound

This paper cites Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:29.405359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.916380Z digest=sha256:fd57adc63ac206501d51f885c614672732a30d44bcf2cb38e88ee44547c5bb5c

Observation 9951ccd0-0ec7-44b1-912b-eb258b88f30f · outbound

This paper cites ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:29.371687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.925083Z digest=sha256:4da6554f1894ccafbe004492fd3ed220ad4c33e94f88938239b761b19d44d895

Observation e3ecd4bf-a7a8-41e4-8219-0997b3b4dea3 · outbound

This paper cites Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.933208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.933208Z digest=sha256:664e7cf31b107204079efd505dfb8cac07ee660506b35ef13d5a485c71380814

Observation ab7febd7-dad6-4552-99d3-6312fd4e8aba · outbound

This paper cites Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.940081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.940081Z digest=sha256:b7a6b2b33e7aa07d9fdc0285c99c30e7ebf9e91074d754896580e9c7b9cccb8b

Observation f38f74c5-ecdf-4564-8cea-abeb30d608d6 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.947845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.947845Z digest=sha256:644442c9f7348d53b0fbf7658c714dc20b46cf5e4552381b81dbc64241fb422b

Observation 9a6f2ecf-7b9b-4cc9-9406-567752ae3855 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.955092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.955092Z digest=sha256:5efbf015f7bdea6587eabeabd9e028db3ac43ba67d3ca3770e7249270f63a790

Observation b20da880-08c3-4c1b-a77d-e83546f01594 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.962879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.962879Z digest=sha256:6a32b0700b5cb0798a95e2b6335ac75a71d5c50ae89dc774a136c1e2675c5b17

Observation bf7d0f5b-b5ea-4b20-9ba1-b6849bfa5cff · outbound

This paper cites Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.968681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.968681Z digest=sha256:267899f41e3e60b984420fdd07fe5350f97c16de20a5e9a87d13c4eba1d84dbd

Observation 90a7c3ed-5353-4cde-af76-5b8a6d5426db · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.974773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.974773Z digest=sha256:ffc16e03192ba5ab5019e50e243d24e682f4553b4f09a84dd535643ef3e9ebc6

Observation 14d7e133-ac42-42e1-a486-6c9174a5d576 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.980598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.980598Z digest=sha256:716339a774fc38cbb67acdda820141b76c53377e6490e6852306c9257959c14c

Observation 9677bfae-716b-49be-a318-1b2a5fb6300c · outbound

This paper cites InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:28.993328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:27.987623Z digest=sha256:db45adb89a51db2602cc687023eaf75a0372e85dd826249d9167df2d5e662d3c

Observation cdb31871-084f-40e8-b7a7-650e2162565d · outbound

This paper cites SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.994465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.994465Z digest=sha256:292e42b91b4b24fdb3afd79dddea94d94f6398e7e92554610fd0075e89c5849e

Observation 445f9482-1cdc-4302-9486-6527b0ef5b59 · outbound

This paper cites VisualMRC: Machine Reading Comprehension on Document Images.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends VisualMRC: Machine Reading Comprehension on Document Images

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:28.919060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:28.001024Z digest=sha256:db7b5e97d9ad439306428e97197a7cf9db347d55935adf2ec7c70ee94fae766a

Observation 57dc6faa-2bb4-4db8-87f5-71ffba39ffa9 · outbound

This paper cites TextSquare: Scaling up Text-Centric Visual Instruction Tuning.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.008819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.008819Z digest=sha256:c2deaca6b5442e81bf0dc534d20a7e674a3185e35936b35b8afa34c33b3e01f2

Observation b6d8a585-a1e6-483d-b932-cf459dcb0afb · outbound

This paper cites Unifying Vision, Text, and Layout for Universal Document Processing.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unifying Vision, Text, and Layout for Universal Document Processing

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.015612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.015612Z digest=sha256:31c2686f544a980e1ca4ce49c609c2aace5d863a58867ea1945381abba9f691d

Observation a378fd44-d891-4df3-b5a1-949fee1ebd35 · outbound

This paper cites Hierarchical multimodal transformers for Multi-Page DocVQA.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.023249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.023249Z digest=sha256:fc94ed3df05ec078743653e3e5486974f83d3f9cdb8bdb727a156f660dee1211

Observation 53799283-9fef-4ff8-93b2-826388d0ad77 · outbound

This paper cites MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:17:28.787114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:17:28.029907Z digest=sha256:a82c09da7d66b1a4e4d2e23c455b2a9af6e1c2a896025265ebaf73c334a728d8

Observation 0e74542a-568d-4e1c-8f20-96874b93fa0a · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.037171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.037171Z digest=sha256:22f6b787801d45581118144fccbd2fae7bbd792c7b3487bf77953fbde901e167

Observation 31be4889-c881-4221-b3df-9f70fc11967f · outbound

This paper cites PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.044901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.044901Z digest=sha256:cdfc41e96adf8bcfec3ca63f99c089e145713fa77e468758155382dbae96ac2c

Observation 6bfcd838-d87c-49f2-bbd7-cf1f88539ebb · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.060042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.060042Z digest=sha256:5b34f811b02ece350222e8326ec36884fb0c48c8bb658c45422f529c96c6a04e

Observation c0501f79-02b3-4fef-b600-b3d363282b86 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.065837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.065837Z digest=sha256:62ea1f640fa056cf96452d9bd882967a857d3dc279c4d7e24ff3c5875c33231e

Observation 5dc0dc13-f1e2-433b-a0d7-11645661b412 · outbound

This paper cites an unresolved cited work.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.071941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.071941Z digest=sha256:c3f21d4ebb0f1e1b9e7a91412f97778581432ae52ec6c9a4ef1cf6abd669733c

Observation beead3ab-b136-44b2-b759-f9101ed01b89 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.078298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.078298Z digest=sha256:94be848f79243e939e6cfd3cfb94bb2d0d36565c9c5c8e3721e431c6c2226462

Observation defa629e-8143-497f-8995-6a7095da0f6b · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.084106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.084106Z digest=sha256:8cf6019e46bf2757c86f9c0bdaad5850e40550fdf020588d6a7fc05f362ac97c

Observation 375c2eee-b73a-4cbc-8f92-e376e4caa9b2 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.090077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.090077Z digest=sha256:db7b69c54f01daced0dd6632a72e9781cd540fe7916a0e84e30cbedac69fdf7b

Observation 609e6b1e-9562-442f-b398-b04fa4296387 · outbound

This paper cites TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.096510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.096510Z digest=sha256:998b57226b306924b1abb12b6a76603af4980076ef3e2a29c64fd093889c3e8e

Pith citing papers

Observation b6b6da64-0b84-489c-b57c-7971bed9bdb7 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.424518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:18fd3c1def49248c0683e538dfe340e523f17258eb5bfe9551f6e8bb02ac8196

Observation a011b765-b42d-4abb-b59e-16d2341360cb · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:44.710015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:60af330b3cadc6661977ebc8dac1d561b533021c60c619668088752f31149cbb

Observation 18963410-65f3-4049-80f0-aa3e7a05fe89 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:066cc3c8e03a4054212d6d55477331e3e9ac01aed59d7fd57d01942e15566258

Observation 2a4f9d30-3599-415f-be88-369f081809d4 · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:42.901109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:42.901109Z digest=sha256:d8146ef15eec57cd7c9fc946e4fea035927477b14f4e8769547cc6b4ef617a9e