Pith. sign in

Paper Citation Record · LEDGER

LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2012.14740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.14740 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:07:32.892858Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

59
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4e26d4db-73a1-449a-acec-44ccab6b8d48 · inbound

Nougat: Neural Optical Understanding for Academic Documents cites this paper.

Nougat: Neural Optical Understanding for Academic Documents LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:42:12.587348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T09:42:12.463309Z digest=sha256:257e29181fe7fe80798ac9484356802e2985673e05fd02252fa7ea0ecee4598b

Observation 85e4b954-d1b0-45fe-9935-4a5700f88a46 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 271

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.321681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:a16d501f00ae44a1a8149e2e4cb14b86442fc4db7fdfbb7d37030f55d855978c

Observation eba26595-eb38-46da-afe9-ed752e090211 · inbound

DocFusion: A Unified Framework for Document Parsing Tasks cites this paper.

DocFusion: A Unified Framework for Document Parsing Tasks LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.792406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.792406Z digest=sha256:93afbfce121272353a364d99250b544c952d5bc073e6e6a8714fee37b919872f

Observation d8d7e7f7-c9b3-4c1b-a072-821c6f2242fb · inbound

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment cites this paper.

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:39.695400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:40:39.695400Z digest=sha256:b45aac99c6a37460d150782bb423be051c241269c4c6b03bac96965764b98c92

Observation c0501f79-02b3-4fef-b600-b3d363282b86 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.065837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.065837Z digest=sha256:1e6fed6850c3771d26936407be5eb36c3b93b31c3dc318bcbf53f7051a41c18b

Observation b14e2e0e-8f85-40ef-a7cf-0fef0c888f50 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:13.295917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:13.295917Z digest=sha256:1ba25a6bcee772e94f3e1b3b62a7e549e312720e639907db312b2688ac674864

Observation dee30146-85d3-47fd-849a-dbc2a2bdfa95 · inbound

Performance Analysis of Traditional VQA Models Under Limited Computational Resources cites this paper.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.299321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.299321Z digest=sha256:6d489a318c5dcf2624312f0c5912d94dd1a1ee1b98b5a27ea15c02a70e5360f3

Observation ccdb6357-b1f2-4e13-87ac-e85f2b070bae · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.988182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:29.988182Z digest=sha256:e77325ca619ed13c55e14f19569c65f7125d096a9fae22f27a92368396bd568c

Observation fb3a692f-f73e-451b-8bff-479829049f9d · inbound

Efficient Medical VIE via Reinforcement Learning cites this paper.

Efficient Medical VIE via Reinforcement Learning LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:32.892858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:07:32.892858Z digest=sha256:2d948cae5df2ae1ec03ce574017edf63424eae83092439311f31ed00bcf632c3

Observation 4ec8e662-e56a-43cf-b0a1-f41bcf33de9e · inbound

Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks cites this paper.

Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:31:52.660503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:31:52.660503Z digest=sha256:3cffd144de297c88361ada75641666157ec69f0dc4ccb79e84cf52eef676f4fa

Observation 3ceb7149-7d64-49fe-b295-13df1a98afe5 · inbound

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models cites this paper.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.315928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.315928Z digest=sha256:661659b5d0c77dc9ce1831115d647154971fffa53a226b368c3797aceb559ed0

Observation 227c7055-8240-487f-ac35-2601503778e1 · inbound

Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale cites this paper.

Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:49.169064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:49.169064Z digest=sha256:bc41ef02f3c804245de7a93dd5aa4768a58c5860aa5a78ea1d457adafe897715

Observation 67f6bfd6-7632-4801-9fc8-fadf8f7056cd · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:24.345477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:24.345477Z digest=sha256:f833b14fd334d64f71be1f76972e30dc735679b47c5219fa2c279d3f08460629

Observation ee736245-3fe3-4144-9d31-5926535a4c6d · inbound

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding cites this paper.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.896872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.896872Z digest=sha256:aca98bad079228b9546ea214778ab688a6f4046944ce0c2fc99a39278981ccad

Observation 6fc15024-0c61-4aa2-9ddf-8c353335db91 · inbound

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models cites this paper.

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:30:23.197949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T22:29:36.960961Z digest=sha256:c6c48ac178bc4f08f522f400bc1bc93dc6609e08480a52bd30f3e15041b29197

Observation 2fb0dedb-0a0e-4692-aab6-9d0c9d3a9044 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.473348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:051eab7d90a1dd6c701aaf932a3617d5b09c3ce9a85f63d1afd5ff7003930dbf

Observation e2005fad-4789-424e-bc48-735e46b0e8f0 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.730142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:c243cb4211d60943baccc344eeb28f6a8e34996a6ba999036f664d87816376d4

Observation 3b42e5ef-c7b2-405a-af00-c73d368aa3f0 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:54:47.176376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:88ca7854b1a5b066c633b07535b2c1b5634199644b00b9b143ee4f399a351b51

Observation 0d640502-46be-4f8b-9cc4-48bfaeed2348 · inbound

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding cites this paper.

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.090079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T05:48:34.771799Z digest=sha256:8eef8499954d16623d96f530b5c3b0c64b5b288bc31a8b453604137e96eb6397

Observation 8a584c8d-0fbe-4019-964a-b06601adb99b · inbound

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild cites this paper.

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:45.176564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T09:19:38.285839Z digest=sha256:e3185261b3d3e9cab76be40f8789e1f60cba6a4b514921b9d81bb705c5494342

Observation 9c50434a-5b04-44d3-bcf2-5f879b329869 · inbound

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi cites this paper.

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:35.270473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T09:59:20.602618Z digest=sha256:b55cb119a736b9340bce460ceb2a9fb40c98c7b70b6a0570a7bdf1cf0d181e3c

Observation 2e3eaee2-2b43-4485-8c4c-cf62fbf23850 · inbound

DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation cites this paper.

DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T14:38:36.756737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:38:36.756737Z digest=sha256:02bf2741c2200d858477d8440cacf41187da07dd2b207a78400133d756373ff2

Observation 939dcb3c-b02e-4be0-b79f-f8cb63377e05 · inbound

IntelliAudit: Using Large Language Models to Evaluate Audit Controls cites this paper.

IntelliAudit: Using Large Language Models to Evaluate Audit Controls LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:27:29.200803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:27:29.200803Z digest=sha256:189ccb0e05bce0efb2089ee0331bd26fa62dd73f8f7a54da0e2b0e27db021990