Pith. sign in

Paper Citation Record · LEDGER

LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2104.08836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.08836 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:10.488825Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

49
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8337f539-ca88-4a94-ac26-752cbbdbf8a1 · inbound

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding cites this paper.

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:10.488825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:10.488825Z digest=sha256:a412269c7106984fda8d2706ab589031278a74bede7c69d801bdf144701a0d73

Observation 6b26cf17-dd66-49f4-9a30-dd7cdb7bedb3 · inbound

Class-Agnostic Region-of-Interest Matching in Document Images cites this paper.

Class-Agnostic Region-of-Interest Matching in Document Images LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:38:50.217063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:38:50.217063Z digest=sha256:3d7f4f60d8618a7e1674e575e1c7bef6c7db8ca4aebf7e4ce4735ca1df289979

Observation 3a9ecb96-d720-4380-80fc-6464a625a2d8 · inbound

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models cites this paper.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:12.416233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:12.416233Z digest=sha256:493e95b990598975810c2847d68e7403ee13f8718cfe63968519c0e33d3b74fa

Observation fbe55ef9-db8f-493e-9335-0a30d1e6dcba · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.734028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:88c2a89b952ff38a409bb60dd3efa446aefc7bffd254690509a75b78f03cc5f4

Observation 3f8caf51-951e-477d-8cdd-291a7febda20 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:54:47.199784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:a79ff10b039512d801522dcc87e789031d29dd29be9cec86c421db62c115cb55

Observation 026fdcd2-56b7-4718-b537-aead829ab103 · inbound

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi cites this paper.

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:35.259844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T09:59:20.602618Z digest=sha256:6451218197f03352a81902fc874c031460f973aae70490bc3c45483e887ea58f

Observation f5ae56c4-d086-4170-8f15-54d0b6499a65 · inbound

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis cites this paper.

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T16:02:00.920066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:02:00.920066Z digest=sha256:864bec32621bcb3961027d03c93e59b1dfd7113f1fba4690dc74aaf0664ed88e

Observation 5dcfbe36-82b8-4b24-ac73-fc7f24181e1b · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:27.594513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:27.594513Z digest=sha256:475b0b1029095900ca6dc150f0a0bfdc5d45bcb8059bc9dd474c47b76b602e9e

Observation 4e8b07bc-dab5-48a4-a47c-0954e3c08d2e · inbound

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models cites this paper.

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:59:19.585721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:59:19.585721Z digest=sha256:b5fb0a0ce7fe052288b43493a3561ebcb50b876273d09fb4f5c10f980748185b