Pith. sign in

Paper Citation Record · LEDGER

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2506.01388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01388 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:10.777750Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T04:38:49.512293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:42:04.436599Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66858e4b-f5bc-4c16-86b1-5e9300c60366 · outbound

This paper cites Form-nlu: Dataset for the form natural language understanding.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Form-nlu: Dataset for the form natural language understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.861693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:08.059090Z digest=sha256:4980bd10a3421c41549b2769932f2e575a567f555b2110479ac8b127f52d87b1

Observation ea21328d-8faa-4527-b7cd-6c93d3d69b14 · outbound

This paper cites 3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding 3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:49:11.108695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:08.445020Z digest=sha256:f86e38f3aa3905f292eabf7fce33349eab067cb4754740b520477d30b1e16c11

Observation 4ddf283f-3cd7-4135-9c16-1951ac6b46f3 · outbound

This paper cites Mask r-cnn.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Mask r-cnn

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.694463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:08.645933Z digest=sha256:fd332c2f30f384e3d5acf71832ec338a6acc9728635917f380a887dd04daad5e

Observation d4a62393-152b-4421-88f4-980b42b1b12c · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Visual instruction tuning.Advances in neural information processing systems, 36,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.886130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:09.475403Z digest=sha256:13dd606b7ad32a932cd737199ee0b125350b85ca198b2870bf5c31711bfe7499

Observation dea4200a-6c0e-4ec7-b2fd-21b27c72125a · outbound

This paper cites Hello gpt-4o.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Hello gpt-4o

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:09.628609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:09.628609Z digest=sha256:3e6c0a5b7abfce38072d73d1735f15f164dcaa200a2bc2d5251fbc80c145ca13

Observation cc5cc385-1a0a-4151-a9f2-c55e03706a25 · outbound

This paper cites Cord: a consolidated receipt dataset for post-ocr pars- ing.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Cord: a consolidated receipt dataset for post-ocr pars- ing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.630089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:09.707416Z digest=sha256:5465c69098afa5f2f19dc886287949bb4f69b110e7efea7eaa74b448346b8a59

Observation 53d1255f-7b73-44f8-a7b1-695661a9ab82 · outbound

This paper cites You only look once: Unified, real-time object detection.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding You only look once: Unified, real-time object detection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.402963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:09.842186Z digest=sha256:39a1416a5526edb4f37993909237b88f77e7cd689f2e50e4e4ce87afc75a14dc

Observation 543dd6a1-9e1e-436f-ab5a-916c46b4e8d3 · outbound

This paper cites Towards robust visual information extraction in real world: New dataset and novel solution.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Towards robust visual information extraction in real world: New dataset and novel solution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:11.711186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:10.340936Z digest=sha256:c4fe0ac1995dfff1004b4396e222b8ed52b1b5bd13169d82bb5681debc61e0b2

Observation 8337f539-ca88-4a94-ac26-752cbbdbf8a1 · outbound

This paper cites LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:10.488825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:10.488825Z digest=sha256:a412269c7106984fda8d2706ab589031278a74bede7c69d801bdf144701a0d73

Observation ed0ceba4-9e7c-4f19-95d3-5f1362107066 · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:10.632836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:10.632836Z digest=sha256:3f7198597cf11d45bf05d23d539d2dbeb882a4324a62ddcccfe5673fc7fff282

Observation 66ac4e59-5f8b-464d-ad6b-1c3e5a998a2b · outbound

This paper cites Detrs beat yolos on real-time object de- tection.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Detrs beat yolos on real-time object de- tection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:11.441845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:10.777750Z digest=sha256:5c82f9114e92fbd8c6b871b4b63519964692c864e87ed6431e483aa6a4d0b129

Observation b44f0f5e-c3a8-4266-a7e9-e3cddaa59010 · outbound

This paper cites Kleister: key information extraction datasets involving long documents with complex layouts.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Kleister: key information extraction datasets involving long documents with complex layouts

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:12.141867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:10.104836Z digest=sha256:8e0254648ac77fc068fd3f78dd9b509a96578aa93457e575125ba0197197cbf7

Observation 57ef1854-b6c7-403c-a274-757055004725 · outbound

This paper cites Faster r-cnn: Towards real-time ob- ject detection with region proposal networks.Advances in neural information processing systems, 28,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Faster r-cnn: Towards real-time ob- ject detection with region proposal networks.Advances in neural information processing systems, 28,

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:09.961206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:09.961206Z digest=sha256:8f5e9dc535c1e7a10e8e6a82f91072590495797e1b399bcbc60922b342304b97

Observation a6a21c81-ceae-488b-adf1-bccb94a64354 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Icdar2019 competition on scanned receipt ocr and information extraction

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.536986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:08.842306Z digest=sha256:e1cc19d74e3468da7cd50ea882f4697b68ccbf3a2950f7efe5aa80656b8f0644

Observation 27088a13-a74f-44e8-89d0-052d6c690446 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.374992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:08.987749Z digest=sha256:dc66099eff1d3c39eb372aab6da69b0bd5b627ca6e7d81279da244ec520593c3

Observation f1697ce6-a1d4-4c7d-abb0-a4bcab3d657b · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Unifying vision, text, and layout for universal document processing

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:11.928334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:10.233981Z digest=sha256:43eb4f4a4d30be3885c549a73d2c02601d723bf3670e248318634f1f3b23608c

Observation d13bf1e7-d242-4def-87fb-9122e04bcc48 · outbound

This paper cites Dit: Self-supervised pre- training for document image transformer.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding Dit: Self-supervised pre- training for document image transformer

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:49:13.125103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:49:09.141603Z digest=sha256:c4fe0365ad977b35f527b7ac6e84c33bd26989f509172e703a59198e63893959

Observation 29675e18-bc61-4d54-a1b8-bd3c025bcad1 · outbound

This paper cites David: Domain adap- tive visually-rich document understanding with synthetic insights.arXiv preprint arXiv:2410.01609,.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding David: Domain adap- tive visually-rich document understanding with synthetic insights.arXiv preprint arXiv:2410.01609,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:08.221446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:08.221446Z digest=sha256:3877f8ac0bc2394ee6bc7c85f7b9b26375f0049ad1045f8fbb43affcd92ad96c

Observation 11227b6e-6f49-4fa6-9d1c-ed6ce3a5b060 · outbound

This paper cites PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:08.346682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:08.346682Z digest=sha256:71d686ed9980ca1ee4129c2e86677aad0abe4895ae8d8f387e85cdc5f3f8aba8

Pith citing papers

Observation 6e2022bd-1519-4e72-9d7a-8af56266ec37 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.439042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:5ed9daaf72e512ed0ac286e3debd67891882101aea5e572e5657bd53e406afd4