Pith. sign in

Paper Citation Record · LEDGER

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2310.05126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05126 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T04:17:40.198357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.543028Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7359fe39-2afa-448a-b8bd-ed7f0739d030 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.738088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:9932fef7fc5c24ec363f904c187d8a325a63839528a440f8c020f5057a00bdc2

Observation 4e34c82b-4a20-42d7-bde2-f41ceb3c4f19 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.271388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:d307dc5bba9e03735e2bfc28f7114c00f4decde472ac2c3c6ec1700b4299d125

Observation 959f748c-3963-4c39-b326-c35b441f9251 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.796606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:c2b8175e00767466121f9d10165296086148161ace44c3e926ea190218045572

Observation bf3a85f8-994d-4172-92fc-4ad13c8c5fd8 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.903141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:a3d756cd0f0689e834764ccaa640feb2b971b6cf28c0e972f1957e61f4d9ec28

Observation dc054965-971e-4bc6-872d-ed373c5e9791 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.831310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:1fcf54a42633ab48971d8dad49a22f5c60793c80a65e9a2466811995db09324b

Observation 3c622375-6f48-4b08-86a0-7efd163a6d25 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.302097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:d0ecef05e675e4fe406fe61704712d5c3f784cea12006a4fe56118aef3502e6f

Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.785865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:8b7b3923c9afa16309b5fa4f118fb5ef7e190d1218b44ab9c61afb78745bc62b

Observation bbae7ba5-2786-4931-9df5-154052e9c283 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.711229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:f6d68b8ec57851b9af90cff28fdc49719bdf55ba8aec02061516f05056cfde0e

Observation dc9c12fa-4573-4023-a931-b729e6d14542 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.666913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:eb92c7d5c907ba279a38c553c748819a3eadbc89458b79d91a9466e9fb100b06

Observation d962ee53-3e50-4bc5-864a-9d9354d1a2b4 · inbound

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition cites this paper.

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.544571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T16:10:03.765621Z digest=sha256:8a325e24b69fed2bea73ed51906554a4971130ab1dfef5d38aa0fce6810e5d0b

Observation 074c23a4-80a0-40e5-8eb7-63c922968231 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:a147aaf65d30f5e267b3ca81e7ace8fbe1be8c4469ade385131408ab85ae84d5