Pith. sign in

Paper Citation Record · LEDGER

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

As of 6 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 44 inbound Pith citation observations for arXiv:2305.07895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.07895 v7

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T09:55:35.452649Z

measured 144 of 144 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:50:08.136002Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T22:26:17.944984Z

Reference resolution

100 of 122 outbound references displayed

  • verified exact10
  • verified fuzzy80
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3018f5f-746f-4874-bcf0-d30380bd631f · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.965876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:9133615f1bf6fde16b080e084822737bdaf2bdfb91871b599c46edde136aa18b

Observation 07edbad5-452c-4ba6-a96a-82be0263c840 · outbound

This paper cites Gpt-4 technical report.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Gpt-4 technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.973245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:728fc25eda81ef982e98c3ceb498ad684b586bda09c4dd169034368ee4ab99bd

Observation e4bb3035-939f-4f16-9c56-d636cfac8420 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.613195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:e3eef5af8e3af0be8efb53ea44e4ab18e77e621bbc7b865396d9d553d7310750

Observation 17a410ba-cf93-4ee0-803b-10375fe0b523 · outbound

This paper cites Hashimoto.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Hashimoto

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.976382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:88437c4c1284a93fbcb8146e5a1d2ffec8aa6dc28aeaa26d4c97f148d2203538

Observation 82a4572f-85cd-4ddf-9727-4b0aa25b4d8e · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.979929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:042405ecc20e369f526b7e62cefcce37b60beab1eaf6c35dd89505fa4c5406de

Observation 098beee5-b217-4297-9a9d-ebda6c80029c · outbound

This paper cites Instruction Tuning with GPT-4.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Instruction Tuning with GPT-4

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.580214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6c6f2ed6821199019c8f65f08f4a2bca4cd4a0979a20df9d39c3ceed96eed2cc

Observation 24fa484d-b9c1-473f-a942-34a5474d0d58 · outbound

This paper cites Vision-language pre-training: Basics, recent advances, and future trends.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Vision-language pre-training: Basics, recent advances, and future trends

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.983591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:ba89028dbccc8ae4eed45637bfdf064080e1c441e7075bf995ae9c8093e0b524

Observation 3580bb70-9748-442a-b2cd-b5a1e6346a01 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Learning Transferable Visual Models From Natural Language Supervision

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.527784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:fbdbe3f9252bc3c8cdbbf2b6d9f794f44b8a8ea30889672f4e6d3d676dadc9cc

Observation 57929d05-82c6-4922-bb9c-552f9b78c597 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Florence: A New Foundation Model for Computer Vision

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.544307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6d6759dd1c25b49ebbc69ff9553c2151d141bc0035ac021a83d78003b659d707

Observation f433eb90-0de5-46e5-b917-1700b2cb27ab · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:35.565340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:f448f44f34a392fc3c52605e13a35399a09b2370ae0d4ff1df971e281d4f36f0

Observation 00491194-2d90-47ab-a3a8-6f44e0431290 · outbound

This paper cites ELEV ATER: A benchmark and toolkit for evaluating language-augmented visual models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models ELEV ATER: A benchmark and toolkit for evaluating language-augmented visual models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.987260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:fbe8cfc7265c4ce7dac91b500257ce45e30cb6f708495a13a37e6b7e92e4f8f6

Observation f43a8c3e-77df-480f-a432-95bbaf898ee0 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models PaLM-E: An Embodied Multimodal Language Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.587553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:65ec4761a1865ab9c5f32598cf13bbcbbc8f11e609d9df1da526248df1762827

Observation 8664965c-e037-428c-a021-d4d3626bb449 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Flamingo: a visual language model for few-shot learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.991138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:e18d2422f6e0952c289b139d989ef6ac0a7cfc0b3d3ea8eb2c50ff00464c97d3

Observation d4eaf250-8b13-44ac-bf4a-c707fc68693d · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.619906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:3405a4625f24222f3e272e55838b632c162f4cd20920156f46c4110ac3b3421f

Observation 9288a657-3986-4b89-b586-ae557e76ee80 · outbound

This paper cites Visual instruction tuning.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Visual instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.995154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:0e100933b39a52654333a38ec626a99889e167aafcaf8cb08e151bb7be0544e3

Observation 836926e4-a651-4b87-a162-330670ef6d85 · outbound

This paper cites Gemini: A family of highly capable multimodal models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Gemini: A family of highly capable multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.998881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:4edc462d2a32f540dca256ecb09ed95a6a8478fef91fee7a4e0ca3bc0f58c6c8

Observation 2e742101-1681-4989-a038-64592e4be446 · outbound

This paper cites Gpt-4v(ision) system card.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Gpt-4v(ision) system card

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.002567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:71759f144732747308e9ff68cc82ab50a554e0d8bb44a033c68f6af962e87973

Observation 02d38667-983e-44c3-a647-aa3f2c9e0b18 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.006286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:150e6efbb4f6e861612c40ac5d1765ea931a33caa1f8495bb13469e374ede40b

Observation bcdb8a84-bff6-4e75-8a92-b0c35b794236 · outbound

This paper cites Openflamingo, March 2023.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Openflamingo, March 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.010196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6bb5458dc56783e6cc1320b77c2182a741084ad3c0e88205f4181faa2391c507

Observation 9249f411-1131-4bbc-9b12-772af225134b · outbound

This paper cites Improved baselines with visual instruction tuning.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Improved baselines with visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.013765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:5e1677e0aff80f8010fe965400f4a4430bc215058ef481826d9946420a45c495

Observation 8e1af592-8e9d-4e5f-abd5-2d90e2f0a537 · outbound

This paper cites MiniGPT-4: Enhancing vision- language understanding with advanced large language models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models MiniGPT-4: Enhancing vision- language understanding with advanced large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.017326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:c82fbf12e0ccf3c832b04b8ea74e3005a452ef2be3c69c3d3adc39cc1ce77b91

Observation c84dbdaf-9f2f-4de4-832c-9407c5944f77 · outbound

This paper cites mPLUG-Owl: Modularization empowers large language models with multimodality.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models mPLUG-Owl: Modularization empowers large language models with multimodality

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.020821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:875dfc29c3936cce5655b2724c06fe8096d44ee9cdf76203cbb4cc0bcdf97357

Observation 1ea6d9e9-01b1-41d9-a755-2e9ab8b7f566 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.026095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:7245401c9df9b6001b9db1090bec67176cb9cb8f5c2689c8fc5f8dfd8ec252b6

Observation 1e73d944-daed-44e4-9e68-dc9be9ee7fc2 · outbound

This paper cites Llavar: Enhanced visual instruction tuning for text-rich image understanding.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Llavar: Enhanced visual instruction tuning for text-rich image understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.031135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:e8833f438f3c9c3a606c72b0b4234234928170f6251e1d6636902e0b415ecced

Observation dba26dbd-532c-426c-bccf-b00d667269f1 · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:36.034630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:b541064764ced0602ebc8350932cb1ac98b5dcccbfc966256c4ea710bf78729e

Observation e94d17b4-0109-4011-998d-f868cc09b42c · outbound

This paper cites Minigpt-v2: large language model as a unified interface for vision-language multi-task learning.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Minigpt-v2: large language model as a unified interface for vision-language multi-task learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.657476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:69568679be9bbe8cbbcad704c0fc861dc8b31b3fd03ca4cc8abff4fadc4fac24

Observation d447f721-8dfc-4afe-a391-131311c6f4fa · outbound

This paper cites Unidoc: A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unidoc: A universal large multimodal model for simultaneous text detection, recognition, spotting and understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.661608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:a9b924d85f2444378880e450a7dc6406a38d77dd8542458d9afb3543495dbb58

Observation 5db0f4ac-9700-48fc-8f17-94df52d0f876 · outbound

This paper cites Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.664937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:b31f996fcbf47e92cc049308f8a4cd5969fe5c3360a94022ac81eb876196ed01

Observation da05f7db-e620-4062-be3c-4c5a37830ce7 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.668375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:5596ff9887d2ea256d7325474b0d21a484e97006d994ac1ac389f1f7a6780564

Observation 848f6c2d-6e4c-4187-a48b-84d1a67a4052 · outbound

This paper cites Textmonkey: An ocr-free large multimodal model for understanding document.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Textmonkey: An ocr-free large multimodal model for understanding document

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.671862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:aaaeb1ae32827afbd8f8bb1f73025e5a2ac87e47340d3484632808c08c6c6d6c

Observation 2001f8e1-f91e-4716-a3a0-407371e62897 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.676031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:f97c25007ef647a7304b25b700ecc51d4a97aac83074c45f14af1fe2bededb62

Observation 563b9252-29a1-4aa8-819e-876f6a9b4124 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, March 2024.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Lmms-eval: Accelerating the development of large multimoal models, March 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.679774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6e4ab9ac28b06cbc5ef6f98105c98a9796a68daa6add5a4da82149ea167b6974

Observation 2f851b79-4be4-4efb-a909-bbbed5ee89aa · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.683238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:17afcdfb1e6ec63980b347b77057b9607cc7b2a0fb7738f1bb64c0b461385004

Observation ed231db8-94fe-4003-9587-c8f2289d1a6a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Mmbench: Is your multi-modal model an all-around player?

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.687434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:9fd587f19221371203709ad9e3fedcc7c14cf80f4a9c35d8f10c71fa20cb83e2

Observation ac14142c-f662-460e-8ea0-843941b1f168 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.692670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:77da79c5bf8f3bec84f57e8a43a9e54295cb1a9fab50d8dc63dc69b74a23501f

Observation 3086762f-0c4b-4f40-80ea-35bf5c91d79e · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.696614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:54867bc04797e1f00fa60fdc091e467aa62208fe4ceb32916a05b6f1dd3ef70c

Observation 627a2cf1-3a0d-4cc4-9630-7bcdf8e0e5f3 · outbound

This paper cites Top-down and bottom-up cues for scene text recognition.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Top-down and bottom-up cues for scene text recognition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.700933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:170835862ded7debef8a6b78019e6ab816fe9bda3b00c55780bc77180ce68ed1

Observation 73639994-ce34-42e9-af6d-b308030a6502 · outbound

This paper cites End-to-end scene text recognition using tree-structured models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models End-to-end scene text recognition using tree-structured models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.704596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:b94a29e27fad62994c50c092239a26acc585c16319aa7b201743e9a44d8f6ab1

Observation 5ec98a99-1e1c-4c95-b34c-c36dbb4aeac7 · outbound

This paper cites Iwamura, Lluís Gómez i Bigorda, Sergi Robles Mestre, Joan Mas Romeu, David Fernández Mota, Jon Almazán, and Lluís-Pere de las Heras.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Iwamura, Lluís Gómez i Bigorda, Sergi Robles Mestre, Joan Mas Romeu, David Fernández Mota, Jon Almazán, and Lluís-Pere de las Heras

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.708034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:5985fed02f2f368f19859c61352095e5e9bdf405dd099e33f1431131056fd3dc

Observation 2e52ce42-8e10-45bc-8a27-b58336c89864 · outbound

This paper cites Ghosh, Andrew D.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Ghosh, Andrew D

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.711237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:b1391de0317b1ae3ea31dc1ea88c04b00bcbf7bbcaee683af66f40cfba2b5de4

Observation c89673a7-7f2b-4b1f-97d7-ab1f972f0ac4 · outbound

This paper cites Recognizing text with perspective distortion in natural scenes.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Recognizing text with perspective distortion in natural scenes

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.714497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:b739d7215beb99d53c7e7e0dee01bf583617885cb1adcc69c56ac3918d852003

Observation 03648dec-bf78-48ac-9bbf-e773b4d59cc0 · outbound

This paper cites A robust arbitrary text detection system for natural scene images.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models A robust arbitrary text detection system for natural scene images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.717607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:9cbf8de6ae672854f3b14af408d04e651bfa537b9fe8644be11736d2e5aab30d

Observation f2d68b78-65dd-460f-9954-608872ca8435 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.647328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:ab7ae7a24e9f052801fc05a77d995ad214c7e5e1b3b44fa460b41fd5c05b0129

Observation 2fef0d94-7771-47cf-954c-b983dd98d12f · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Curved scene text detection via transverse and longitudinal sequence connection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.720696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:bcfbe5bb4ffa3390993395fbb4b95f58d41186a8f27bc90e6ac2cd34c34742f2

Observation 9556c3af-5f54-489a-8dd6-f4f62458cbdf · outbound

This paper cites Total-Text: A comprehensive dataset for scene text detection and recognition.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Total-Text: A comprehensive dataset for scene text detection and recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.723695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:826153a63bc1bc42e7a811a2b6abc5b805523536186d9948fe1ae755f133bd30

Observation ec8ea28c-9b57-42ec-a484-10bbc122e68f · outbound

This paper cites From two to one: A new scene text recognizer with visual language modeling network.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models From two to one: A new scene text recognizer with visual language modeling network

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.727093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:5ba9bb74dbd91335a9ec628e75f7e7e9c3931721923fa1a64d869dd0440e5f59

Observation 0ae92482-c156-4376-a133-c8237e19e2fe · outbound

This paper cites Toward understanding WordArt: Corner- guided transformer for scene text recognition.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Toward understanding WordArt: Corner- guided transformer for scene text recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.730529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:5cf5f70ac02f1285aaea8e8f5e255e011108bd456a1eefdf699aadc92a37ab80

Observation 919ae2a6-c91e-4937-9a0e-3e76f694bd03 · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models The iam-database: an english sentence database for offline handwriting recognition

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.733625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:92ae392c65a288b9ae7d71fde49cf4d0aaee6ccc6ba3267168a9baaf4f9dcec2

Observation 377bba11-67a3-46dd-a850-27db16d59520 · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Icdar 2019 robust reading challenge on reading chinese text on signboard

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.736742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:2a99c82fc814c3ef3b600cbc5aafc632b7edbe5ee293d32a478a7a19396c4041

Observation 6f1ba111-c716-4f7a-81c8-4edad17891ee · outbound

This paper cites Saavedra, David Contreras, Juan Manuel Barrios, and Luiz S.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Saavedra, David Contreras, Juan Manuel Barrios, and Luiz S

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.739487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:24592130dd4e7ab4c2a087655825d247f75f0c5ff9314e9e460aad319c94d135

Observation 2a407ae6-0d1f-4418-ae4a-9cf1436282b0 · outbound

This paper cites Jawahar, Ernest Valveny, and Dimosthenis Karatzas.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Jawahar, Ernest Valveny, and Dimosthenis Karatzas

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.742352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:50188099962d358fc772c95eff3ec5195b7dfeea39d3b631ad0ec91a3c499019

Observation 013f38b8-93ec-4aca-afbb-e2d7e74be12b · outbound

This paper cites Towards VQA models that can read.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Towards VQA models that can read

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.745426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:a5b69ae562b141a96ca355ae0c03703a0ef52438e2d3e609f1f5dc55cc2c4f91

Observation 9cdccc9b-283f-435a-8335-edad530ddba7 · outbound

This paper cites OCR-VQA: visual question answering by reading text in images.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models OCR-VQA: visual question answering by reading text in images

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.748495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:1a3e3544803c00cd974fb9d6440b77721539017ce046331c8ef5d6914306d9ff

Observation 43cfddb5-0de4-41a5-a6bf-5171561f672f · outbound

This paper cites On the general value of evidence, and bilingual scene-text visual question answering.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models On the general value of evidence, and bilingual scene-text visual question answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.751512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:8e0db3048947ee523c30bae778a918d844083d4982dcd9ab549f3f99ae963e4f

Observation d9debcff-bccb-4618-b042-5f904bcf73a8 · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.754346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:55ebdc2cffb84ad5c0de55ef7c55a364ec82a3843a38d9f7d7c914788b054d3b

Observation 5a34f3cd-55f2-40eb-a328-fdb7f9093559 · outbound

This paper cites Info- graphicvqa.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Info- graphicvqa

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.757062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:12da46d594b534c88f1ef989150dec78bac84f6f53b1d9f183f4f589bbbce456

Observation 72e9551f-9aef-476f-921a-d5b9f8307f95 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.606962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:49c5d1ee211f3cb81e17288e225c572591df53b467763dbc82291be4da53099b

Observation 4447d994-a44f-4499-9170-ee24e4e02ea7 · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.759809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:55970750f3925255525e80c02239b75d93c6f168c806028d33dc789ae2847401

Observation 7f809125-0343-4e50-94d3-06016436e286 · outbound

This paper cites FUNSD: A dataset for form understanding in noisy scanned documents.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models FUNSD: A dataset for form understanding in noisy scanned documents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.762948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:5b1f80583e8decbe292f1cc373557ebd3c6a0d1540bc5517d656646f8296e79c

Observation 1c560636-1eb4-4213-a991-b454bc66eda0 · outbound

This paper cites Visual Information Extraction in the Wild: Practical Dataset and End-to-end Solution.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Visual Information Extraction in the Wild: Practical Dataset and End-to-end Solution

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:35.625941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:2a05d569ad99770e62a6b22f285c6a66e58bd8bca9ff5b2f1a3d2d8240e4b44f

Observation 2962206f-f949-43dd-80a7-73e09c356f49 · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Syntax-aware network for handwritten mathematical expression recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.766244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6e6915ae113df17210a138883f4fead9788d481c167b4f9a315dd3009159c0b4

Observation 206f8ebc-e791-420a-b980-0d4825360375 · outbound

This paper cites Minicpm-v2.6.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Minicpm-v2.6

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.769046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:2b8af32ffa796aef252edd9ac9ec71bfd4a0054940011bf8287b685e6cb649b4

Observation ad68df80-045d-422a-915e-95680111d5b8 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.772435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:9b0733a1fc8fc04ba85ed36bab3776addd33d6c276dc01c4cb65921c21209f0f

Observation 39f370c1-0c5b-494d-80e6-1a367b98b058 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.775391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:2beb33ac81972d45dc508933c2391f24155423a733803d055e6dffe2f5b535b6

Observation c01dfa1d-d8c5-4950-ac82-ba4856a39513 · outbound

This paper cites Paligemma: A versatile 3b vlm for transfer.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Paligemma: A versatile 3b vlm for transfer

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.778694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:048edfa2958213b980d85cec76423823587722af787e6ee9c6debd00ce917cbc

Observation 49b15f92-d2a1-482b-b60c-ff9be11eb5cc · outbound

This paper cites Congrong.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Congrong

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.781786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:dc0c468e44f38855cd7a8e8802a8c4d980cf8b54a5471bb80119d13ed6bee6f2

Observation 6809c099-8ba9-4063-9706-3da71df3128a · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Cogvlm: Visual expert for pretrained language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.785110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:71be53366499f27d68646087b3c14a45f3f2a04baf3b8dfd105a549cc7ab0d0a

Observation 7c1ac38f-c895-49e1-ac6c-f0aa960406cc · outbound

This paper cites Minicpm-v-2.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Minicpm-v-2

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.789511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:02adc88bcf2f39a4d864dd739d438290eb1362812a44891a338daa09ba034534

Observation 6422a54c-1f23-4479-af88-3db4824401b4 · outbound

This paper cites Mini-monkey: Alleviate the sawtooth effect by multi-scale adaptive cropping.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Mini-monkey: Alleviate the sawtooth effect by multi-scale adaptive cropping

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.793526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:70ad030df258d7a98eed5be70e47bc8791a211d6c4b5f2402999399cdd843fe7

Observation 9ac367d5-0862-4049-9ae1-08a013de6070 · outbound

This paper cites Claude3.5-sonnet.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Claude3.5-sonnet

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.796304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:1f71d678f23a34e4f6643b0216afa4023b5ba8b179cc32878c557e468d38d918

Observation bf636cd1-f2a1-4da2-a160-892edcd71225 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.799708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:bdd69bb6d3584388f887406a2af7f2f8f5bc7cee8a746ddbbfca26f9374bce0d

Observation 5499c977-00b1-4d40-9167-f67006d50d70 · outbound

This paper cites Gpt-4o-mini-20240718.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Gpt-4o-mini-20240718

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.802411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:149bd5a972353ec427bcece1a9749a6d021d2c3e367853de2e48a53e4deb9fd1

Observation 7eb9b621-d7b9-4aa8-a6a4-397f82ab123d · outbound

This paper cites Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.805512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:e93029a4512f62db2f7b2f3e15266d59961256858463e430e5a4e44304b76b27

Observation be0d6b20-14ed-4d63-a672-1fb8e121c972 · outbound

This paper cites Rekaflash.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Rekaflash

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.808246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:4920110f74db5ff466ed3dc422ce4b52ca1c756f2e42b8faed777d80b4ae18c8

Observation 0ef3d80a-43bf-4689-b98c-f53b343f80c1 · outbound

This paper cites Gemini models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Gemini models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.810826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:fe8bb9b995cb02a151159eae78842ddf39b22c8b4366fcadfd4cb4c2db0e92ea

Observation 67258c97-de25-4f0a-8c8b-a10b661b43cb · outbound

This paper cites Xverse-v.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Xverse-v

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.813634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:1a50579f0e678706a5715f0353cc88408c0495db33daf26a217f0e89e5d58948

Observation 0b4827e9-74fb-4a3d-8506-cc9e8d2c43ee · outbound

This paper cites Ovis: Structural embedding alignment for multimodal large language model.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Ovis: Structural embedding alignment for multimodal large language model

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.816521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:82e6460a30ce58f2ec8690a8104e5ac74d28468474687a05b80909e0525370eb

Observation 37c1de9d-2f69-47d2-be34-a2315a8f64b7 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.819997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:4958ce5db215a7a744ffb4bb3519a1fa5948c5fec5236107a52a0e8eb96f3ae9

Observation 0d185779-2b47-4e86-bdf1-63872e22470a · outbound

This paper cites Minicpm-llama3-v-2.5.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Minicpm-llama3-v-2.5

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.825349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:838a7fcf99bdcc9011878a4c181d158b99be9f50a645f0c43e2c4111b9859430

Observation 9778d63b-71cc-4f18-a2e7-4faea2aa247f · outbound

This paper cites Generative multimodal models are in-context learners.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Generative multimodal models are in-context learners

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.832226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:b276cc2de5803f3186c498618d493bc239e8f0025407d68e94d1d60d78bed4aa

Observation aca2d950-8efe-490a-91d4-da7d2f4aed16 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Deepseek-vl: Towards real-world vision-language understanding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.838280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:4c0e02babe1fa178fd1aa40ea20cb0c1e1a42dcd7d234b54351fd6f6df7de4f7

Observation e558722c-224a-4080-9f0d-cf3396cb0708 · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.843728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:82be547c202bd7412992e7ee535d50ab3a7faef6838e6e3018fa8c6b6aa6f03f

Observation 914be1ea-e8b1-4b36-be50-e6b5929f9619 · outbound

This paper cites Omnilmm-12b.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Omnilmm-12b

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.846909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:ab779e3e038d387d22cb24696fdaf56dbc3806ec5a666402aa683a8569838f65

Observation 2e0422b9-3901-4d51-81b6-eb22c305184e · outbound

This paper cites Transcore-m.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Transcore-m

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.849863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:01eabd81fea279335ef28cca973b8bc7eeb309aa6a387f3e922db27267572b7d

Observation 1c24fcbf-1347-44c9-8e42-176ceb08e495 · outbound

This paper cites Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.853668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:fb9ed87eeacc69508ed298fc35d28abd1d95689941f19f7bb558f1b644e691fe

Observation b07d8331-1a41-441d-9c1a-fde5fbe2cf73 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Xtuner: A toolkit for efficiently fine-tuning llm

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.857477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:719f6ec27a25ce5cb4c489406c10137a03bdf98d570b41fa7d8ab9bbed1bb519

Observation 705fa39a-d9c8-4fa2-87e8-c6b44288fe28 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Sharegpt4v: Improving large multi-modal models with better captions

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.861278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:8eaba3ab23072ed5f35b65947028f0274d303f9a5d35579aef99002202e51c07

Observation 036368dd-ba99-46be-b835-797ca156a7bd · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.864922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:6e768d5b853f5de3961752e2699a19ce4baaeb53e04ac5e31553f0a203c91460

Observation 07893c62-b878-4faf-8bcf-08cc1ad25942 · outbound

This paper cites Internlm- xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Internlm- xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.868730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:304a083a23eacf21ef36ec68516c58a9e067f6eff776763b5cac1d79e753dc5d

Observation e89b1c26-b857-49c4-ae81-d436a19ef5b8 · outbound

This paper cites Minicpm-v.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Minicpm-v

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.871922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:cc37fe07b95f16530a73b2a299e4fb6745b2f56deb8523be8536f1933ee1f13f

Observation f15f50a0-0cef-4d70-92f7-8694eb3a5653 · outbound

This paper cites Yi-vl-34b.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Yi-vl-34b

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.875328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:0e5177cfdf132b090a26278d108754ad3b3e33ccab4c3f656528c03ce61e30d1

Observation 605797dc-30a0-4450-a600-1aa5426ec823 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.879071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:f4f494d9ff6feef7e2d56c7510c3df3b06412eca48f2b830ca7bd829b650a3ae

Observation 0de47630-c282-4fb5-b384-bbf177434c65 · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.882007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:20aa9f8964892bc54992e8e05f86daf50dd063e7e4e0426162bef44bba66c96a

Observation 3114ac99-a731-46ed-8021-e36d04d8d11a · outbound

This paper cites Phi-3-vision.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Phi-3-vision

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.884897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:44c413fc4ee7837e7c703c69dd119a9c41046e63ccc03511498a025a1884c316

Observation 90dc8ffc-3a77-47f6-a2cb-5766a4d912a6 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Glm: General language model pretraining with autoregressive blank infilling

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.888424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:1b3c7d7a4896b059e94c35ce8798f4967ee3e117570e9d58c222fc59ba9daf2c

Observation 9733e218-e1ad-4483-9b27-bf453995dad0 · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.891969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:71b200019eba5e20a5ccd103f825e8f0baf9025c7505932f9176c72153f3bdc9

Observation 10745213-2515-4aa3-8bb8-aa8e46197cf2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.601331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:cf145150100dbe3a18268f7740871751cc390cf9b304c23aa9e53ff33561c4fe

Observation 2d1afb06-53ce-4253-b6cb-4b0430dc808a · outbound

This paper cites What matters when building vision-language models?.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models What matters when building vision-language models?

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.896853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:12c8360f297c3ea27668f6a1697f942302b6213f849b73208652057ba738875e

Observation f433bf8e-bef2-4af3-9b02-9b7c184d2ac2 · outbound

This paper cites Pandagpt: One model to instruction-follow them all.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Pandagpt: One model to instruction-follow them all

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T09:55:35.901356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:664027e2aedb96a956d38bf7a53a1073c3092aadf7154d13ae33637de7916a47

Observation dfab4ecd-c443-4da5-ab19-2ddc3cf455b7 · outbound

This paper cites an unresolved cited work.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-17T09:55:35.906886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:dca1cffa6efbd3360f74431be5ed4f7fe6605d67ba1556c97383044e72e6b33a

Pith citing papers

Observation 42cb7730-504b-4137-b553-df17066e2fce · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:d0b762ca3af027cd81d8f88d63ab4dfd7877741eccbd543fda72c789b5f9a89c

Observation 84113f97-4b89-46be-8fbb-382d41d8eb83 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:01d52061c7fad4605a0c9057c19bec383c1016f2084bae131ee1df7ffeb1b279

Observation acdbea90-6207-4367-9e84-fb0544ee8b9c · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:6fcb5bb770fad8c5e7bf03237dc12e5413dc43c40bc868709f56417e76c45824

Observation caad77f5-aee2-4b81-84bb-a43d49c62b75 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:08410ef46ae553bc7ce721d8c13c3eb53acd2d2a3dcb40f161b2152b9c1ddeec

Observation e6f39914-70fa-44f3-8342-15b2f7b38646 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:f1be29e48de97a33bc66cfeef38c659e26945fd618a8d8d61cde067f1ee77262

Observation 108de9e8-c92b-41a0-9cc7-4afe1ab1f30f · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:8a949ca0999a7b50b7728f320f6dab1da56bc42fa66dde2544734e36b03ecb48

Observation 573c2021-cbf8-4bff-8a90-dd159ef5f695 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:3b3c94643353aabb05728d7769c86c5eed8dee3c15a044f866a69c9fc426c838

Observation 67dfba78-7812-482c-80a9-22be0c02224b · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:787afbd4006f529a7a0cf0947285d2819ff33ed701e6a91499c1308b21a9cd0b

Observation 208d770c-af15-46fc-b797-6f53bc4d3b36 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:4a863426032ca178e57301c00e18d1192edba20b055df14b0fbf6adbe40a7e19

Observation e6c40b26-3e3a-4417-b165-85d1ec3b0867 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:0a9488becd730346f0c9882d407903650ce1ff3ca1546886f30e6ab06c513be2

Observation 8ec462d2-90d2-47f9-8f25-74b21d9fd086 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 215

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:f37cf8ae72106ee12843f0973c4da6c1d9cab7de37a8609aeb6f89db58de5ac3

Observation 1f421b4f-c3aa-4521-9c22-dc66af286c46 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.701581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:94314b0a470a62c9c3d102ee793a037c65bae0428c31e8d707e8211bba4aa583

Observation 3ef1718b-157d-4f35-aec1-f03a11e66af4 · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:08:19.754191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:d7c3b39e02a5c092b4ac50f7aa71c8e7c7d7e0ac61fda3b00c2c27e150f073ee

Observation 0606427f-4b50-49c7-896c-983b871b52be · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:0876712ca6098a49be6878bb02fbbd8befa078892d3d1b5f16bee2039b8f2e3c

Observation 31c5c6d4-00f8-4b90-9da2-243f439cc956 · inbound

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription cites this paper.

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:25:19.369228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T02:23:22.682357Z digest=sha256:d505ff46a192e70a86f9d85b6b9298109e3e6fca1404d3aa0d761d36d4a0ad33

Observation cb342305-57b1-4715-8a9b-1a366456e06e · inbound

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization cites this paper.

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:47:13.010209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T22:47:09.500229Z digest=sha256:42468ae0de277bbd8751a2ce3b9916c9f456ad6a4485be0f2360a39a597d8eae

Observation cbc8a068-f9e8-4fda-8e1c-d49fe18adf14 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:eeaac6bc31409095b7b813c3167be7fc7709ea9b2f8582be8bf5943d5841880b

Observation 5d51dff9-9bf0-4f56-b820-6832a09aad10 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.136002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.136002Z digest=sha256:8a72296264f51e1f1fe3ac34d79b27ab38a0c6168c30d64a1324f3ce16f751d1

Observation 63493399-6b72-4a1e-9e33-15c4ee441e79 · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:57.223241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:57.223241Z digest=sha256:0614716f347213b88838ad4ee91735ad056eb19ae7d34a290d3e245166a78e11

Observation 72e64c9b-d179-4977-9223-196cbe943d12 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.676446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.676446Z digest=sha256:1387beabcf14924f6f245dc06f9647d200ddfd4cd45788ffd4b2f6a09a2a1485

Observation bce14f18-8693-40b5-a6be-ebd619080353 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:03.510688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:03.510688Z digest=sha256:015713a3a46a074c8f1e314d711e4456f632751e62e787b1e038ee903c59ae54

Observation 104fecb8-65c1-418b-a46d-6e6666e00833 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:342d4e045b6e925f68cd60fc57f6cceb890c97003e43b9446512f9748c40b140

Observation af9d6896-24a2-44d2-93c1-71628b5632bf · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.055785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.055785Z digest=sha256:afc1f2e54e28459b735b93d7b540f922601a132d3b47f036976e697116e43306

Observation 6e1762ba-a8fd-48c6-992c-3711ca12ccfd · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.411073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.411073Z digest=sha256:efa7fd406487cbd826c3f9e627cc32216e1a0fc758d40c507640efa4959acb73

Observation 1b88e53c-3418-420b-b1d8-c40a7a5a57d9 · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:40:42.490904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:f835d39e224f56efe6bf0a08bf02a9ff7b555203fae28b20b5f31ea42cf2c9c1

Observation 3d112dc8-8dd4-4e48-a3e2-a54f6ec36e5e · inbound

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? cites this paper.

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:42:54.802658Z digest=sha256:dbfcd60b2b957b500ef6bd67ffbc7e6601dde85ee55de434da0e44e5f58abb3a

Observation 4dce1cc5-3251-4a46-a103-8c20c003df31 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:ee9ed9146dffc671d2338b1c3ef59b38167e87697e959cb571658802325795ed

Observation fb23a86e-5a27-45cd-b988-3435387d857d · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T09:54:47.193159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:bd66f4672117963e6de0277ad3160f54fed69d927be516a1ce223c76ed8ca7cf

Observation 114da7dd-16a7-40c8-927b-6000b8750f62 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:d0b4e05721da92b5974ce47a85edcbd89c431cda978edcafc3b466028264c428

Observation de73b38f-659d-4630-b086-a33d87361b1a · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:23:40.965067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:c37331408018b6f053e176ef8084c3d93431de49d41b998d38be5b5358ca5459

Observation 349e7629-183a-4fbe-a80f-f18a699563e3 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:51.149910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:3981acb5a68130359c18c7715016daae8dd4610e06c9700933c969bd2de72384

Observation 32e8d4f8-4a00-46d0-8919-eac79f87739f · inbound

MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop cites this paper.

MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:28:21.509695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T14:25:11.234515Z digest=sha256:5144882a3729ebbdc6d89e6d045010afa6855e7d0ffd21e7b3eda62d5c602eca

Observation bcfed605-6088-4464-ba83-b9db0d569cd6 · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:13:13.462205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:10:58.225078Z digest=sha256:ee26361672eff0fbdaa1538509cdb6918885099d4d0ff661decd35995f7ccd45

Observation 2e6f6826-63e4-4837-8811-9a25944ae653 · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:55:48.338811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T18:36:56.248838Z digest=sha256:0fc13120d2cbd02895f7c1243c7035e8adfdf4aed09baf3d1375228801fd3ea6

Observation cf498cfb-151b-4037-adbb-69d8eae67c1b · inbound

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking cites this paper.

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:43:43.378259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T20:42:07.819116Z digest=sha256:dc8cb71f256cb92d5d7a0f029914dbe1e407e8d01795ae50eb039c0e4a1784a1

Observation b8fdf76b-35a3-4f86-b0bc-c8d1e014e0b0 · inbound

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification cites this paper.

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:13:05.249487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:10:36.290559Z digest=sha256:64a7539cdf41d3ef4c8f8fe6ff3a114b95a49ca0a74077fd0f6edb5a29dcd5d6

Observation e4de0cd5-b69e-4176-8f66-dc0b1fbdaefb · inbound

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation cites this paper.

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-06-30T12:04:38.653273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T12:03:41.268375Z digest=sha256:8d0f3e492abe3e7451413452c5d1f91107cf8f7b328b112a4f77904dde650de0

Observation 371c0666-998b-4f2c-a4cc-9596ea918a4a · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.946170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:4c6de22dab7cef5d146d561ab4112f3577852aae92a029c580a8965e7f05b9ca

Observation 5fe77a86-848f-4d7a-9b62-e64b79749134 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.176011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:aa74c77f2c355890a8b42cc5a12a77bb986c1c6f11d6124b32a2bb1f7740bf8c

Observation 47c1cc66-1afe-447b-a24b-08ea20db74ef · inbound

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation cites this paper.

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:14:18.757035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:13:11.013395Z digest=sha256:139bf6524b6028bcd6f57d65134338c14d01965a9733328a329ee042a4d5b484

Observation 848b7e5e-cab8-4c4c-8505-caa9ffa96c71 · inbound

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation cites this paper.

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:05:28.659900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T07:01:26.911110Z digest=sha256:cb329ffb31aef969ddf37edc59eebce6486afe7db2235baece74c17d20d34d87

Observation cfb5dd04-0ff9-4297-ad0c-19a70376f45d · inbound

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models cites this paper.

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T11:58:26.902599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:58:26.902599Z digest=sha256:023445ae58d5af20ca04664fb8827c8d03d1e1279354974abed6651d6dc59243

Observation 1f9fc299-3122-49a1-989c-4ec3476f6d40 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.308103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.308103Z digest=sha256:942b5134ca26976b7caba0c1c26f0acac619a153ddaedeae0c1c4eda8b7084dd

Observation d55ea9c9-4061-40c9-88f1-afb3999a8e46 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.987289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.987289Z digest=sha256:e2547e39f66b08863ddf673f5f69d95b149452a18ce265b08ec78e46159e7603