Pith. sign in

Paper Citation Record · LEDGER

Hierarchical multimodal transformers for Multi-Page DocVQA

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2212.05935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.05935 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:08.417275Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:34:38.458141Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b72f363f-e7cc-4bdb-8aae-8c7ab167779e · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:08.417275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:08.417275Z digest=sha256:d732b4b6f9241b9e7c0eff0b1db772cf34dc561e3aa3841fa2df5ee65b780611

Observation 887b9de5-1192-48e0-ac6c-e7577ae81940 · inbound

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents cites this paper.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:34:38.562815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:34:38.065829Z digest=sha256:409ccd26cf4c3a1fc8ece4b7316ecab939c799702414aad6eef3bc04b6440e6d

Observation 8eec18e4-8f42-40d5-83e4-5d6a15dfe248 · inbound

Hierarchical Evidence-Driven Reasoning for Long Document Understanding cites this paper.

Hierarchical Evidence-Driven Reasoning for Long Document Understanding Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T16:13:40.332262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:13:40.332262Z digest=sha256:1fac8c22d44f25fd6cdd85178adafb34374da16b528d75102a076508e27f6f71

Observation a5c121cd-a98d-42da-bb1a-c495f9e5f330 · inbound

Workload-Aware Caching for Multi-Agent Systems cites this paper.

Workload-Aware Caching for Multi-Agent Systems Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T11:22:59.426670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:22:59.426670Z digest=sha256:f5f5457aa59edf141958552fba45bdae5f0eda977000ec4ad9e4fcd42e960da7

Observation e3e6baf5-5d04-488e-b7be-47449dafe40c · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 132

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.146561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.146561Z digest=sha256:40cf880c89224239552d26f29abf01047d7e69e1d1aad37874a20749a705d44b