Pith. sign in

Paper Citation Record · LEDGER

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2310.05126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05126 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.600852Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.543028Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7359fe39-2afa-448a-b8bd-ed7f0739d030 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.738088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:3ebf2c9a45a1906fd958216d0b01eb9a21d682d5a6e0afc6f5dc25ee68f81b19

Observation 4e34c82b-4a20-42d7-bde2-f41ceb3c4f19 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.271388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:b91c21cca435317a4ee50dc4c9383d45dd83cf8f23a0424a14f398cff80dde6c

Observation 959f748c-3963-4c39-b326-c35b441f9251 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.796606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:726c952e5b089ff759c770a46312ed1280c73818fa658f4cbd1021b9d0c92f3f

Observation bf3a85f8-994d-4172-92fc-4ad13c8c5fd8 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.903141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:4aa48458334b879d0885ba60289898c6c1be68aee2669b3772188d57e9dcdaa8

Observation dc054965-971e-4bc6-872d-ed373c5e9791 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.831310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:512a4e04de88976a353b84f90bc8499827d8f5014d201c742311de0936cfad94

Observation 3c622375-6f48-4b08-86a0-7efd163a6d25 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.302097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:cfedc978673f57222e0dd48ac9fe0762fb997b7139ab45d57bd01d8eb9c169f3

Observation 3187935f-0a0e-43a9-be6f-1c4c72f363f7 · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:01.050752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:01.050752Z digest=sha256:d754f17b42f5073e930a0e2cdeb8547cf3b200dedbad9879fd4afc058f518079

Observation 66bab495-d4ab-4ccd-b68f-a59783f6139f · inbound

Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction cites this paper.

Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:33:45.057217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:33:45.057217Z digest=sha256:5b644515ef8b911f9780031f6096568be7009ecfa71ffcd12b7deb359dab95a8

Observation 55c95798-9e2b-41f5-b04c-686960fd8bd6 · inbound

FILA: Fine-Grained Vision Language Models cites this paper.

FILA: Fine-Grained Vision Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.243482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.243482Z digest=sha256:6ff56a42f7bdbd35841c797426206abf6ba7bd2983fa6bc502b4643bec510a67

Observation 9f687888-e6cb-4ea7-8640-113138fd9b41 · inbound

DocVLM: Make Your VLM an Efficient Reader cites this paper.

DocVLM: Make Your VLM an Efficient Reader UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.877803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.877803Z digest=sha256:98979a85acbcd34bf9d182682f805c0f5bf81a9d9279207aebdb31b9c6633271

Observation eea32079-4f42-49f9-8e56-720633a0d8b3 · inbound

DocFusion: A Unified Framework for Document Parsing Tasks cites this paper.

DocFusion: A Unified Framework for Document Parsing Tasks UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.796087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.796087Z digest=sha256:16d25c77cac6bac40933103a3138fa955350578f86992487ea683e98a4c266b4

Observation a00f6c14-9e11-4af5-a241-2cfa952461de · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.978087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.978087Z digest=sha256:e103fb2140f3d4b3f7108df7a7059eeddcb303fb2ca220a2d2ce9728c4d250ef

Observation c6d0c6cb-54eb-499f-adce-c11763938a7a · inbound

InstructOCR: Instruction Boosting Scene Text Spotting cites this paper.

InstructOCR: Instruction Boosting Scene Text Spotting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.101773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.101773Z digest=sha256:e67850f6ee12b35987f8fa868a69ba7d00d21fb68529b3ed430444851852cc61

Observation bf827e25-0aed-4782-8627-8dfd9377e165 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.320223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.320223Z digest=sha256:756e47433b3b10a6bad94ee1b42a6d7663942c55e5a8cd7845cf780222b041f5

Observation 6b97aaed-ddd9-4dcb-b263-fdce21bd4bb3 · inbound

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective cites this paper.

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:14:27.976265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:14:27.976265Z digest=sha256:76f5b8f69c307f21c283680374ebeeb2730dab9a82f032d6b4bd36869128713d

Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.785865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:6d85d94ac7cc74e7441c76b595423f8f2205b70c1443ab902749822fd52fe6ba

Observation bbae7ba5-2786-4931-9df5-154052e9c283 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.711229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:aa8763c3ae6b19f4e076173b958166aee8cb455579979d5bd59352efc9da3e05

Observation 375c2eee-b73a-4cbc-8f92-e376e4caa9b2 · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.090077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.090077Z digest=sha256:562ce98b00f2085fb286d234bd673507be00d5c4bffa1cc6d807a01b274fba41

Observation 1c6ce3d4-4ac3-4fef-8fcc-5f7d52ab49e8 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.292960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.292960Z digest=sha256:6d9cd4d8e767d65802f4c7d2547fac9d6786a6ef7f8ef9471397d137572a630f

Observation 75e807b4-543c-4b27-8c86-056d1dee2615 · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.692936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.692936Z digest=sha256:7e432e1232cdb24d990349f6dc188aa0d3193fecc83d772a15fe4e913c25f93f

Observation 85f27e3e-d5aa-4e3c-98fc-813fd950510a · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.392634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.392634Z digest=sha256:ff976f67a18955843bd7c372e2489e86f0f439a1228fb69d6e54fbe46a943cb7

Observation be00772f-ab01-4e96-9645-29dd506a6267 · inbound

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink cites this paper.

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T14:29:28.296141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:29:28.296141Z digest=sha256:3c9362925fbc97ef54383b27d45b2d3507e348e80fa5da4a05d70ecc1dc84224

Observation 05e12b97-07d3-408d-9d50-b077d70af03e · inbound

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery cites this paper.

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:23:33.759059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:23:33.759059Z digest=sha256:d7e139b8f53f6e2e772f2e8f2b9efebfb1fcd37514dce7941f53fd004e7c8d3d

Observation 2354dc82-88e1-4cb8-ac30-81994c8141c4 · inbound

Return of the Encoder: Maximizing Parameter Efficiency for SLMs cites this paper.

Return of the Encoder: Maximizing Parameter Efficiency for SLMs UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T13:38:28.057837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:38:28.057837Z digest=sha256:18bd34757854cc7d1e9a61bc74ea414f92e79fc62c4306d22f4916bcc7361c50

Observation a489b425-3008-48d3-9604-c616a7e11f3f · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.952910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.952910Z digest=sha256:1484e8003c988c732cf513525cdfbfb2da75ca0d3012e2edf8ad4db6ffb28ac3

Observation 797929cd-e1e4-421a-ba37-31d011647556 · inbound

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling cites this paper.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.600852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.600852Z digest=sha256:25182acc04bfc8b89a794b57e4c0df17a171d9ab8d51bbb0a1979984798ef440

Observation 95a8e5eb-970f-4e92-ad66-8d2e24d4f79a · inbound

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding cites this paper.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.092596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.092596Z digest=sha256:150499310394a3b43129bbf707a60104e5785f4045db08d1fefcf2559d9660df

Observation a14b27b0-8113-4d7f-8a65-b09e36415b5a · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.762169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.762169Z digest=sha256:1295b8e163b736613e2ad6deaf992d72f54ec6ee5705a12d4820a861b793d9fc

Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.654963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.654963Z digest=sha256:9f94890390a4ed231bbf639d614119794ac3fd817f9c8c0231a2c7049f12a636

Observation 0190d4b3-deae-47a2-960d-9ef4e1052993 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:56.033608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:56.033608Z digest=sha256:8648a2ba4dfd27b525d07d1466c605a4de0af99b9c7da23a098599c7a7bb5d24

Observation 853d8cd9-fce2-459e-a3ef-e6e69e8209b9 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.980280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.980280Z digest=sha256:f5df558b22c01a829ac3f11f2a0d56fb90df113a4ee9eae972ed7939682c9a30

Observation dc9c12fa-4573-4023-a931-b729e6d14542 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.666913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:68d7e87b2585fb17e84739570eec8c0309821c21bbcf3641f0cfc086a9f03877

Observation d962ee53-3e50-4bc5-864a-9d9354d1a2b4 · inbound

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition cites this paper.

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.544571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T16:10:03.765621Z digest=sha256:f51deca80ab593e597c3c9492be0268d673d3c4abbc54b2c56c2f750bd59e816

Observation 074c23a4-80a0-40e5-8eb7-63c922968231 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:74f1f9c2f01b50bfcedb2f5eed074a64ea97389e24d96ef49fa9daf94112e6dc