Pith. sign in

Paper Citation Record · LEDGER

DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2408.15045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15045 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:17:27.802900Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 840efc04-9dbd-4cac-bce6-e3a3808c3239 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.793729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1b3bab7f3f774891a8d74305386f00916576333a5fdb2348d818dd8cb3532c39

Observation 195cf54e-b498-4da7-9066-a0ab28b3dd3f · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.802900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.802900Z digest=sha256:6f2e58f9e38a971dfe64b33d5ae214a73b6a9d2add53652b1bdaf61e30a3fe7c

Observation 55a9d781-b9c5-4bf2-b988-d2adb86f3101 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.818917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.818917Z digest=sha256:394b7776d28ff79b4fcb6ec7d095b5a7c18c3638c0093f2602855d0c62e6315f

Observation e2683014-4149-4068-9dce-8ebb7e5ef479 · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.645195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.645195Z digest=sha256:bf776fd626f04f083668f83a0fd4b1c041c0f7d3a780e7a77df882f2d37c84b5

Observation 08b6382f-8eb5-4395-bcee-e630f005b43c · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.050413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:b7443ee65c465e738f4ac0397c6d886ff39af77cd0bd764b82a2992b302e9f40

Observation 34aaa92a-50c5-40f3-8837-6c4db8922b33 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.477273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:f3975b5baa6b6d1e1ad4f2878d12c627c79ab22a2f52ce7ccab2b8457226efbc

Observation c01bd59f-4028-4882-9156-c3badc769502 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.767230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:8ca151dbce92cdbd05e8254f4bd2e7bc6f237da6a1f0c3c928e0d686fc552ec7

Observation d3a90b89-c08f-4769-8f86-39e8221d455b · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:39.244177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:39.244177Z digest=sha256:626f1f151b1b4de78ed695288f5f82c7adc9a931f69e979ebe9e27b44b359778