Pith. sign in

Paper Citation Record · LEDGER

LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2104.08836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.08836 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:36:23.823165Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

49
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b60822af-2cca-4b0a-b8ad-98bd4b77f384 · inbound

LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining cites this paper.

LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:09:03.929798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:09:03.929798Z digest=sha256:abfc5e15f7b2ee152424f69d7d4c8eab220ad3039b088bcb76e5dc78d848f358

Observation beead3ab-b136-44b2-b759-f9101ed01b89 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.078298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.078298Z digest=sha256:9863e6e9843a799167bb75c7aa4eb4fc276cd1ef9bf89a3fdd5652efc73512c9

Observation 8337f539-ca88-4a94-ac26-752cbbdbf8a1 · inbound

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding cites this paper.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:10.488825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:10.488825Z digest=sha256:761d3d52bab6478edf7c259e6d268a4870e5108f9e1d74a80b0151f12b17ae34

Observation 6b26cf17-dd66-49f4-9a30-dd7cdb7bedb3 · inbound

Class-Agnostic Region-of-Interest Matching in Document Images cites this paper.

Class-Agnostic Region-of-Interest Matching in Document Images LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:38:50.217063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:38:50.217063Z digest=sha256:ca4bbc874a5f3a7c228baca274ab20897f1b5a901989ec6e9fdbebae0c1c98d0

Observation 3a9ecb96-d720-4380-80fc-6464a625a2d8 · inbound

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models cites this paper.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.416233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.416233Z digest=sha256:952c7a83d96f3d02688dcc8da1897da6b7e797c718b25cee7db2a607c1520c52

Observation fbe55ef9-db8f-493e-9335-0a30d1e6dcba · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.734028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:8de38e0bc9265b85068956abd0849e8b52faf97af6e66a93359e85dfddb3b0a8

Observation 3f8caf51-951e-477d-8cdd-291a7febda20 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:54:47.199784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:be3d7140380d7cca5a3abfc0b08af9f57a585e0494973b44c401d4b6262aba19

Observation 026fdcd2-56b7-4718-b537-aead829ab103 · inbound

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi cites this paper.

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:35.259844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-30T09:59:20.602618Z digest=sha256:f8bfacf42ee658365a778b1ff9e0cf4a22160bdcaffd2310b57793839d48b021

Observation f5ae56c4-d086-4170-8f15-54d0b6499a65 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:5fad29479cdafdb79d997ac128443ebe155f3049f7b6a7b6127a8b55b287ff67

Observation 5dcfbe36-82b8-4b24-ac73-fc7f24181e1b · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:27.594513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:27.594513Z digest=sha256:271390cbfa8c118241bc56c21e3b8997eb071598d51a81a65235cc3d80f75924

Observation 4e8b07bc-dab5-48a4-a47c-0954e3c08d2e · inbound

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models cites this paper.

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:59:19.585721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:59:19.585721Z digest=sha256:6fe48eb204f6d122cf67856d0e714e8e7bd97319ed999d42482da943f5a63abd

Observation 35778d0f-f276-4855-9167-986155b5788f · inbound

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition cites this paper.

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:36:23.823165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:36:23.823165Z digest=sha256:b1060684b9d37ade77f72ecd075990855467038baa6ac33101f54566dc75d99d