Pith. sign in

Paper Citation Record · LEDGER

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

As of 6 August 2026, this Paper Citation Record lists 100 of 156 outbound references and 45 inbound Pith citation observations for arXiv:2501.00321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00321 v2

Coverage vector

measured 100 of 156 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:33:26.613927Z

measured 145 of 145 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:03:07.345505Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 156 outbound references displayed

  • verified exact37
  • verified fuzzy55
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch7

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 84db86e9-daa5-45a0-803f-e6dfb24e4819 · outbound

This paper cites GPT-4 Technical Report.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.710555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:fcc3a9da930c8ddcd4c4153d605f9d267543eb5101cc11d523c5dc4f38051d90

Observation 881d1380-983c-48bb-ad9a-87be81d399d3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.715283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:f2b6a7fdb9c66f1f857275cf3f9d56d6cca823de3041c4e215c3c10e911548e3

Observation 32c8be8d-283a-4ab9-abe8-09b2c19211bd · outbound

This paper cites Language models are few-shot learners.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Language models are few-shot learners

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.110212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:6d2eec9fafdad98f9979619223966bf7a9266b217454ee9feeb66787ccef82b6

Observation 19455080-45aa-42e3-b2d9-c861fbf23001 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.720254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5c47b8eb74e5ebe81f04646990ada683473db39bd02643db24d6eff1a87c1b20

Observation 7e559013-a35b-4d93-bc3a-09416380e193 · outbound

This paper cites Visual instruction tuning.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Visual instruction tuning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.002344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:7df9977205bab8fafc0b4f8b804e2cb6e4f55b67fd7e45abbe1834dbfac8a876

Observation 8e552762-b7bd-4f7f-93b0-96e847f509e1 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.983236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:7502ffc9d2de6a08180e4958a86a34a28d1cd78191060b851d725e719af8f064

Observation 7731a63b-db2d-49a7-ae06-be6343d86ea3 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.724787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:969988898eea506989c75a391dafdaa32f5bbdc69e010c87280a8b8483c995a9

Observation eb8bb696-acf7-4ff5-84a4-50ac1619a710 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.728484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:83d4a6828345c5316a94a9b64f13950e7c31407df0ceae5347883db857aacf4c

Observation 9b89acf0-6d5f-4b6d-9e8a-432c6991a3c8 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.733088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:c29b36b3a264cf56e3f15a2a2e4f0cf5093f3d067d0a22a27454d212c316f59b

Observation ec63bf76-4f80-4d1f-9484-dbeb9663f85a · outbound

This paper cites Towards vqa models that can read.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Towards vqa models that can read

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.972826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:99c221e9677465449abf97fe9c0bf4bdb68d82b909579d554a013937ddba0bbc

Observation 1b159bc0-d92d-4e80-affe-05fe7b571d58 · outbound

This paper cites Scene text visual question answering.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Scene text visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.946146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:89f48dba9d6e1dc80d891c9fb2c62b8fd71114f42ac1e35ee19978a8f2a1a6bb

Observation c1e6a0ed-95dc-4834-a1b4-5e7c7570374b · outbound

This paper cites On the general value of evidence, and bilingual scene-text visual question answering.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning On the general value of evidence, and bilingual scene-text visual question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.059090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:d6036a50ae38b0a9932c1701de2a31498c41d7a385505532c53c2b2453fe2770

Observation 943a1730-bffa-4aa7-96c3-7fce26ec1f0d · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.697366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1683539ac162ce9107482497c6118c5f3ebea7118de023c136f4adb99f5fbea3

Observation 1f421b4f-c3aa-4521-9c22-dc66af286c46 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.701581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:94314b0a470a62c9c3d102ee793a037c65bae0428c31e8d707e8211bba4aa583

Observation 7fed4a12-5eaf-40e0-bb7f-58a14df4529b · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.706356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5fa8ca02ebdb93aecf18a81588d10fab2f7b0aaf522051596365477fb0d99ec3

Observation 50dd66b5-d3bb-47a1-9419-8a80f935cedb · outbound

This paper cites ConTextual: Evaluating Context- Sensitive Text-Rich Visual Reasoning in Large Multimodal Models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ConTextual: Evaluating Context- Sensitive Text-Rich Visual Reasoning in Large Multimodal Models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.031458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5e124f395287e0d53b28d456dc194e3c6d65d5b2a22c27c43a957100e18e6023

Observation 66b56da9-27c4-4e76-a85a-14de545037e6 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.738101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:790b5c1332c6a4062aed3ff6ef6e724f13c31fac414b316c8444584f07733f39

Observation c4920d74-91e3-436b-8241-f3b90f43cc97 · outbound

This paper cites TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.742622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:74f8dcc0368cd709bcb1c6e99596bcf16d15ca3300a561854ccd314e8f0b48d7

Observation c094f01d-c59d-4de8-b355-0e88f9ab39cc · outbound

This paper cites TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.747153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:0cf2394a5bd120e54b6e103fe46442eef6d8396ddcb35d3db0fec6013cfcae6e

Observation a278470e-2183-4ab4-ab02-6f75b4a40463 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.751219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:9fdc78f2a849c37615caeda944f59734178dce05e4bc214a694732b4854d6cc1

Observation ad4fec7b-c1d5-43b0-86be-10c82b9e818f · outbound

This paper cites Qwen2.5-vl.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Qwen2.5-vl

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.042447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:99549cb61b5d7f83f9ad4b8feb0ac9da8fb3d64ea7ab8ad8846c1786b807ffb8

Observation a51669fb-7193-4b01-afe5-1ea7351b44cf · outbound

This paper cites Docvqa: A dataset for vqa on document images.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Docvqa: A dataset for vqa on document images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.044764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:6b933d4df55106f7249fa48df1da785ba030c88ad7abb3c81149e5e5cbbee951

Observation fe5de7f2-d8da-4945-9db7-4404a8086643 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.755761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:14accc0bf285e9513e2399666cfdcd71939a48120bf4d74036bf2a707170e8fc

Observation 0dee846d-ef22-4f7c-b0d4-c58e3a712663 · outbound

This paper cites Hello GPT-4o.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Hello GPT-4o

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.049155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:191bbf321bf65f8488bdfc01e20fc7df09a3bffaf07cc5d6af8539a2950544bf

Observation 2655b100-79ee-4075-a31b-8e8eb81030d6 · outbound

This paper cites OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.760260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:c65de2cbd1c4e358c138c2da0c912b5a60641826c0fa89a2b788196c9285c81d

Observation 44a67d08-3b66-41b7-a987-03b28e41f3aa · outbound

This paper cites CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.764403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:b8785494a2ecc81675204774b67b6e0244953748b77704c77471636eb92aa70d

Observation f474b427-eae1-430a-ba6e-6291ebc0328f · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.769132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:19992f0af4e6f58c5469ecdbafe5e595912e8f8fb95ba43f812df66c6d99f573

Observation 9db08b98-011d-43f9-872c-b2cc8f92729f · outbound

This paper cites Multimodal Table Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Multimodal Table Understanding

Reference 28

Resolution
verified exact
doi, observed 2026-05-17T20:33:26.679804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:e9c7e8a57953fee95254bc12ced7692da342864d53b49e5461aea0d45b5a533e

Observation fb2016e4-1944-4451-ba05-fcd419b06fe1 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.061301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:a669f5dac00b25e153a96a35df88f4df3a300c2320ac7447e2708cf436228718

Observation a793304d-b92d-4fa4-8199-35015e9dc1ae · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.773900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:16bd4bf64892d23e8d0033bceb92a99592983376cc8b4f9e091891c3cd90f300

Observation fd1ab135-9013-424b-9169-dbc158c19fcc · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.778534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:28bab203f13265a0c1c1332014eda1a1d421f3361e708f45041984667405a4bc

Observation 82d49e9b-a8ca-4c02-923a-26fdc7ad9a2f · outbound

This paper cites DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.782173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:15edce2d968ea2a5d322921fda9457390a70a6cd335495c7a79fde5e044314dc

Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.785865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:83e1f50f3add45b007bd48ae7e753affe8685dc45454e7b3176b0bb89e8d777d

Observation 37bf3566-362d-4ee1-9c88-76f1a5987481 · outbound

This paper cites LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.072582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:72046d01c4e714c24dadb7a85708bab93f0b15a32dadd358b959a53f76f787f1

Observation a94e4643-1970-4d16-8da9-39b3ed21684b · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.789867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:150bd6cbe697ed60442aa50d331cfe9426b5d70889f9f37af49ced593643e1fe

Observation 40ca4e6a-fa75-403d-a7a7-9d83248bd0fb · outbound

This paper cites Dockylin: A large multimodal model for visual document understanding with efficient visual slimming.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Dockylin: A large multimodal model for visual document understanding with efficient visual slimming

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.076933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:78db30171dafab211fc590afeff9d95099d7cea23d137a1f8b4248a9911890e4

Observation 840efc04-9dbd-4cac-bce6-e3a3808c3239 · outbound

This paper cites DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.793729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:f43e913e9ec9c4bf7e01a272e4daf866b17cc83e7215d36d56d42f71f0724e9b

Observation f2fdcb25-4e3c-40db-8244-55ccb432a752 · outbound

This paper cites A simple yet effective layout token in large language models for document understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning A simple yet effective layout token in large language models for document understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.081005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:b713be9cd7359a2ed0d08fd3c70ac59e77dfc9c708cd17f4f2691da8e0a28b0d

Observation 00d0198c-f48c-4841-a193-53492814b698 · outbound

This paper cites Adaptive markup language generation for contextually- grounded visual document understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Adaptive markup language generation for contextually- grounded visual document understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.083224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:31f24bf54327725c220ba44972dca342455880302519609d9213388113519dcd

Observation b1955ee1-cea6-47ba-ba97-e272513d3b01 · outbound

This paper cites Marten: Visual question answering with mask generation for multi-modal document under- standing.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Marten: Visual question answering with mask generation for multi-modal document under- standing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.085505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:34fe26a597bc4ac8d0bb98bfe1684d30bb6a89c22cb6cfc37269b1bff1a4359d

Observation e60bd23c-2f86-49d6-95c8-0ed0f809b906 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.797005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:767a16a70eaf246cf8964c712a96ee110dc89bd6dfaaf3c08903a9492e4f6e47

Observation a97e36e6-a5b2-48e4-8e77-848e7ef7df5f · outbound

This paper cites Infographicvqa.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Infographicvqa

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.089651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:326533cde25ef2380648013114c74b65d36a059b1eb46fff425cb2a7f391b1d9

Observation eb36e648-22c6-455b-822a-1135fb48381c · outbound

This paper cites Exploring the Capabilities of Large Multimodal Models on Dense Text.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Exploring the Capabilities of Large Multimodal Models on Dense Text

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.092114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:e6213400b81383d68141cefbfa215c5f730aeda018ebbabc0c3d2cc8f40e6939

Observation d178af2b-32f4-4c72-a151-796e3b1f2abe · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Onechart: Purify the chart structural extraction via one auxiliary token

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.094816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5afe9c3aba79978c13b0189e3531e04e4bbc0edbe718a5bf40401477af7d999c

Observation 07d29817-b1be-4c1a-bbec-36aef3b6d76c · outbound

This paper cites Document understanding dataset and evaluation (dude).

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Document understanding dataset and evaluation (dude)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.097229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:cfb3e1da280cee15a0295c9a1f5b1a44d9853fd61f9efa07ab4e3d2a19243573

Observation e7154c4e-8d91-4eb6-98a3-b7197d2c7be5 · outbound

This paper cites Needle in a multimodal haystack.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Needle in a multimodal haystack

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.099615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5f4c1026c3ca273a8442f3c344abb17e1dce0b3bb8869d7f74a8e0dacdaca1bd

Observation 9f6e4434-1f6d-461d-bb1c-422f535b82a2 · outbound

This paper cites Hierarchical multimodal transformers for multipage docvqa.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Hierarchical multimodal transformers for multipage docvqa

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.102332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:ff129d059200fde37d497005889a55189f542447b0abe6cc8eb64ba5c4e2ae68

Observation c0b21ae1-ae24-4b69-9d6f-2453b223b098 · outbound

This paper cites LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.800978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:63409b66dd35248118cde127cab3e7af5129a852fe1f23e3d19baf4020ca958c

Observation 34cb1f0c-4152-4cd7-bcbe-206f055ac3ba · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Llava-next: Improved reasoning, ocr, and world knowledge

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.107488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:b19dc2b46f8662d54c2c87a3b3eda28cb08b2c88411a3b522f475189720d0866

Observation 661e7cc0-3ca7-438a-a978-04ee1868d002 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.804574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:d2c2211a8343d69beb8c321d68fc12bdcb5de5fa59e676e0ab7554a29b9ae094

Observation 5a385b12-c2d5-42a1-bf5e-1288ca0fa732 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Monkey: Image resolution and text label are important things for large multi-modal models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.113098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:05dcd525f960a0ab980357da5cd44d9a8bf07f20a97036d69790006d26b9758a

Observation e70fe4f2-ee1e-4a49-ba6e-5e76eff24780 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.807978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:9943ae009a8a1b301bdf4dfffc99891a24e621051e452de7d5e794ef5c2ee82c

Observation 6a90eab5-3f11-4789-bb32-6123be621949 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.118431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:83f45cbcbefd698bbea42715a025d4d793f70522a8549e50da5aac69582483fd

Observation 2dfdd845-0f96-4b05-b5ed-982407f30943 · outbound

This paper cites Pixtral 12B.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Pixtral 12B

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.811367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:f5b095612138da82fae8e7b67e5d775539f1c14343627c58d7eaf97da7c3dd8b

Observation d99554a9-9e11-4b29-b58a-88440b6ad2d3 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.814912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:79efd6a9cf6e7f4fe34890f7da7b9808ef356cb20bcf9f8bf408e742ffdc5df8

Observation 4f54d04a-5df7-454b-8280-6272c16950d3 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.818046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:32b57ef13a456ffc47577fafab29b37840829a2b7c6f86b9ec2ff3094a73b806

Observation 3150e90b-6433-49d9-a42d-f2179f727af3 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.822335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:679bcd9d3208a4f72ee02318148ba265e92989bb940d2c2573862ae6cccab547

Observation b0878f72-9289-411e-b773-26e5a6ef4bc4 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.826379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:f5b8bfb50263b803828b45f3df53887fe7226f3a933ead4781eed7413e3a9113

Observation f6a59630-a4d6-49c0-a69b-a73d6e7254c1 · outbound

This paper cites GPT-4o mini: advancing cost-efficient intelligence.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning GPT-4o mini: advancing cost-efficient intelligence

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.132896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5cbec61a07d2f2c623dcf064ba9f211e9bea22065485c139a1e6a2bfd5550bc5

Observation 577cebb0-3551-4368-807d-7bc6469ef662 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.829624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:66e815fb92fea6cc974c9cd5144e8bf2ee1fa0c5a75a9c0503663fb19f4f94a2

Observation da013c29-5dc4-4a60-9bc0-5cd420e35f58 · outbound

This paper cites Claude 3.5 Sonnet.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Claude 3.5 Sonnet

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.138039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:74a9a57d0c6c507feeca9ccd5d5665f82eab610ef20fcf2b249ea9987645dbb0

Observation c8613577-097e-4035-a125-481c9210224e · outbound

This paper cites Step-1V.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Step-1V

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.140325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1cc4298986568b6728ac49cf5edcb2d8503944c32ed6f5f91ea10012933c3c69

Observation 32211377-a178-40ad-ae96-9fb0f9d7de21 · outbound

This paper cites Image-based table recognition: data, model, and evaluation.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Image-based table recognition: data, model, and evaluation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.143506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:4c09a1456fc7ae0cbeca75539beec974ca325b5d4a6751d72059f302d94e642d

Observation 37896fe2-f706-4675-a7c2-2898d5c65f1a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Bleu: a method for automatic evaluation of machine translation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.146163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:8c9c6b85174843ffb89248e164bbdcf1bf043e0c07525973ec00e60cf2e8bc22

Observation dcfcacb3-841a-4db7-8457-acf8f6e7948d · outbound

This paper cites METEOR: An automatic metric for mt evaluation with improved correlation with human judgments.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning METEOR: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.148870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:2ec27fcadeb60a3e250f5b78028f1fdc9454647852ca9b68fbb41c7d44fd1912

Observation 65b5187b-86ed-4f2a-96db-a067b88a9121 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.833516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:372c2ee9390e2b4d678256d3f52ec2366118f6def2a8d02bf75c108696588719

Observation 3e182cd3-1e23-49e5-9654-3737dff2599a · outbound

This paper cites An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.153643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:ccab8168ce4c889fbf18828799901f7257d2731858117ed83808314241869e40

Observation 5bde4b28-72e2-42eb-a54b-42c084b4aff6 · outbound

This paper cites Read like humans: Autonomous, bidi- rectional and iterative language modeling for scene text recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Read like humans: Autonomous, bidi- rectional and iterative language modeling for scene text recognition

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.156127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:c8eae0c5f28a527870885aea93cd905786e068b0c0a04b4de4db0138f0e0a9db

Observation 2fb446a8-a1c8-4a2d-a4d8-8107ef434a1c · outbound

This paper cites Aster: An attentional scene text recognizer with flexible rectification.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Aster: An attentional scene text recognizer with flexible rectification

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.939516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:8550e08bf6166646860b196b6c3f7c26d1f48c0c3cb05c644b16c213b882bdb2

Observation 5f0d324b-1d7e-4a6e-be06-1b35442b0b88 · outbound

This paper cites Master: Multi-aspect non-local network for scene text recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Master: Multi-aspect non-local network for scene text recognition

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.942767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:3cc1ba45272515b4a6b4d44bbbe0e2cd4798b5bd9f4b760ab90c80d2ae6086b8

Observation 12223856-c908-429a-8ea2-20456aa9a480 · outbound

This paper cites SVTR: scene text recognition with a single visual model.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SVTR: scene text recognition with a single visual model

Reference 71

Resolution
verified exact
doi, observed 2026-05-17T20:33:26.675900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:242ea414af832976e0e88f00a7e04860dea946c9e53b1331840c901ffaa2001b

Observation 983b0bd0-82de-424f-a91e-bff1020715b7 · outbound

This paper cites Abcnet: Real-time scene text spotting with adaptive bezier-curve network.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Abcnet: Real-time scene text spotting with adaptive bezier-curve network

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.949356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:0aaa8b1472af11801df69ac7296781b182fc8b8af57535903e735ad83fe9204e

Observation 07fec2f0-bbf4-43e3-b0f8-349df359c188 · outbound

This paper cites Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.952357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:671c5377ce5da5c0ce01292a4e16b3b112e5fb5147df74da7981cbe5cd75649a

Observation c90d9932-5c5e-4cf5-a7c4-d28606d46b4e · outbound

This paper cites Text spotting transformers.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Text spotting transformers

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.955050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:ea2c7f0b8c8b0e062d86cf6136e6b86ba57a12a2752514f90b06af20ecc9d35e

Observation 41a82524-2cb0-49af-ab68-504105b1385a · outbound

This paper cites Total-text: A comprehensive dataset for scene text detection and recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Total-text: A comprehensive dataset for scene text detection and recognition

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.957603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:7ae08449a09ff11dbce2e32562522686743b9a5fe0c2c4ab23d5cb965ccb286e

Observation 8b2f4c51-28a2-4949-a9d1-5d03b6f70994 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:53186a20300a6b3ecc6bb6ff658dea081abfd14ddf4345cb8a36fcbb910a25bb

Observation b911baa8-73f5-4975-9ec5-87f9729b5a6d · outbound

This paper cites Icdar 2013 robust reading competition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar 2013 robust reading competition

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.962754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:c792aa3d3ca806dd14998470c64bd07c1d3025612c4ad349514a0cf5d301138c

Observation 6f338d15-e4a7-43f4-9305-4ccd7c5c1b1d · outbound

This paper cites End-to-end scene text recognition using tree-structured models.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning End-to-end scene text recognition using tree-structured models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.965030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:03e15ea5366d108809c28fd9b1b7470f2d89ddd2ddcd65d36c777b3386193465

Observation 385ea377-1a45-49e5-9769-e0c7c1274abd · outbound

This paper cites Scene text recognition using higher order language priors.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Scene text recognition using higher order language priors

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.967853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:3511c97a15a31fae74d3df6f155324bf87c317992b0facc6f6e09fd90790995a

Observation 0f2e0e0b-f52b-4b72-8df7-bb4e3d6ec5c0 · outbound

This paper cites Icdar 2015 competition on robust reading.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar 2015 competition on robust reading

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.970370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:bae3a1c580042b381a5afa3bf6eb2adaf66ea155d200254785b346e2433e779a

Observation 7dd7e8c5-71b2-445f-a5b0-f404b1d4efbb · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Curved scene text detection via transverse and longitudinal sequence connection

Reference 81

Resolution
verified exact
doi, observed 2026-05-17T20:33:26.671202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:50b9b1b8824104b628745ee546a9c4f52bf95058c5106a338b063870d4db0858

Observation c0d259c6-f13f-4a01-86bc-96ac9a44bcdf · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.841117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:d72f79e3483dd24be358ada3dc69229dccdd48e24ca23dceea0b54ff5d0521b5

Observation 53d94890-9112-4506-a167-363b64c1ea64 · outbound

This paper cites A robust arbitrary text detection system for natural scene images.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning A robust arbitrary text detection system for natural scene images

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.978037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:3ddd68ee572eb332a7b7858e1f4680b2a67c9bfdf297145512fc37b13e569e0b

Observation 3e2ef60e-aa7f-4e32-8a2b-6fa402dac245 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID:15559857.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Available: https://api.semanticscholar.org/CorpusID:15559857

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.980472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:2e696f1595f04d31fdc5bc7f70c4b4acdf757a7a541bd6d2ac7b306425ebee2d

Observation 4de4d014-8b5a-4139-9a1e-776a1d968bf7 · outbound

This paper cites PoseNet: A convolutional network for real-time 6-dof camera relocalization.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning PoseNet: A convolutional network for real-time 6-dof camera relocalization

Reference 85

Resolution
malformed identifier
doi_truncated, observed 2026-05-17T20:33:26.692682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1cbad2b2782c80fd0a938d0f1093ed15ad37d7f55891c74edc5c226b2f067306

Observation 7fe56763-1822-40f2-8a43-c475d94c5238 · outbound

This paper cites Toward understanding wordart: Corner-guided transformer for scene text recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Toward understanding wordart: Corner-guided transformer for scene text recognition

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.985891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1a4709d8071daab9f579fe5dd344161663fa1a3be2b4ccec1b6b9ef72792be3d

Observation 1c759e9d-7e79-4c0a-a64e-b895ba9e3e2d · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning The iam-database: an english sentence database for offline handwriting recognition

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.988283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:96fa85b61824e3245ce707c34422982317476dc41acad271386f8a713c608a64

Observation 3ea8636d-2722-41a0-ad11-d4012ab98779 · outbound

This paper cites Proceedings of ieee international conference on frontiers in handwriting recognition.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Proceedings of ieee international conference on frontiers in handwriting recognition

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.990689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:63a391356adb4521f5a0270858bf4a9b9de4c74483b7f9d74c7e4cada65ef288

Observation 15fe4a3f-5b4f-478e-bbf7-9858853599e2 · outbound

This paper cites From two to one: A new scene text recognizer with visual language modeling network.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning From two to one: A new scene text recognizer with visual language modeling network

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.993241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:321278a04191764c515d309709e6aa54d173e3ddcaab048a076a1abe122441e5

Observation 046b2e8c-be9a-47f8-a872-32c9c3744b5d · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.844642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:7749b38cfdcd26f733052f11893d753578fe8d7e17a13bf3ead37f2d2a1fe26b

Observation f97fa4a5-71a0-4486-90e2-e0c8601ae257 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Towards end-to-end unified scene text detection and layout analysis

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:26.999790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:963b710e2bb93995aee015a7143e0698ae446bfd5e0c178c97361bcd89250189

Observation b2b307b3-5eda-4f9f-84f6-fac6583d7208 · outbound

This paper cites A large chinese text dataset in the wild.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning A large chinese text dataset in the wild

Reference 92

Resolution
verified exact
doi, observed 2026-05-17T20:33:26.688331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:97137e49f4c628d67ebe68705b6a0697381b3dc0951046fd34ef689fd8ec9bda

Observation e132b10b-e3d0-4cfd-9c65-453e7f2bc014 · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.005039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:e2913e644a77514b5005e2163b53ca851bf3ce7cfe356f32111b0b263381425d

Observation 94fd4264-ba9a-4722-8400-535deb6a62dd · outbound

This paper cites ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:26.848540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:c9182e53b29dc6ea7a28f74dd04f4f780e99bd4d7d9eda033fa307b3d5ef4595

Observation 9f4caf50-a2d2-41ed-9afe-350be61b7857 · outbound

This paper cites Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.852586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:65082fb61de26ea2dd4bd71a794f34b30e3a839a65439f6b286300d733fd8de6

Observation 90714572-d6da-4b4f-8a51-8d8188ee570c · outbound

This paper cites M$^{6}$Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning M$^{6}$Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.856629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5a62f286f92c8f063af70a49f9c150631b87f3ca8b6a6c49e260e6379284a910

Observation 562740c0-5339-4776-bfc0-5b9fd848ab61 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Rico: A mobile app dataset for building data-driven design applications

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.014000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:ac1ea2811bf3bcc05cf8ff547760aa51cc35b1f6b8417da10315d4a7d0979498

Observation fb0e42b6-1c20-452b-9a78-faa42f6c9d56 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Funsd: A dataset for form understanding in noisy scanned documents

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.016419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:429b3102c7980fd27daff90645cfdd8c264d9f61b9dad9fad586e3be5230b478

Observation 6b4b9853-5e77-4f7b-92ff-b0c21fc3ce21 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar2019 competition on scanned receipt ocr and information extraction

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.018868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:189abf07412db86f9a6cf6fd6735cdd73f3d60a57f22bc6e61fd0e9ef903670f

Observation 2be0a722-dd3d-4494-afdd-e6946371d6b9 · outbound

This paper cites Visual information ex- traction in the wild: practical dataset and end-to-end solution.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Visual information ex- traction in the wild: practical dataset and end-to-end solution

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:33:27.021439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:4757318095d74cb63dac49f5a43caacad404ea98b6f9526be41d694a8536ab46

Pith citing papers

Observation cb3702ed-2f72-4b02-b978-4fae2d648b4f · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:07.345505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:07.345505Z digest=sha256:bec78785183f7ed6136240b7a35fd9b378550335d886719b1f1f53d8cec5d4e0

Observation 8e5fb6f3-e505-484b-b225-79e59fe5bb36 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:12.595994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:12.595994Z digest=sha256:d908e92bf96c86c670a8ac103b46752a5b6752493a6dbc71af41eb4035cb0d3e

Observation 84b965d6-31f2-41d0-8c12-0fd77429e975 · inbound

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning cites this paper.

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:56.875999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:56.875999Z digest=sha256:13a47aa1c1aba40514676e1a85ad2a4d97ec4cbe216cee915d4aabbff46f56d4

Observation 52d0a6e8-8ef2-4742-ab23-e6dc1ea29c06 · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f4b9fe91d84c384c79b1a78dc4b4644a8614de3c9e4ed85ab7f993c4c1d14c5b

Observation 2a83625d-0c7d-4f1f-a2a7-fc122f643337 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:50.222297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:50.222297Z digest=sha256:8dba73f2c5e5f80f82091efb519c9bdf662a0f7a2c1dc789f455e94c2ec6ff04

Observation bdfd32a0-7549-4f8b-b451-84649a85c8d1 · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:35895740d0d81690d909278f1a0a23ebec5567c0cb61d4cbb384043b20f94b5f

Observation 3a3d0f3e-386f-4ee4-bf52-4be9e81254f9 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.235644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.235644Z digest=sha256:61a6130571a84fb7ac61cb55eaf4b4a72ba37c73718e0d85ca701f1401a01596

Observation b38ab542-208f-4b65-bd5f-482472da0d01 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.239857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.239857Z digest=sha256:df395e0fc1e19a13eadbfd0b224ca9e2404be512d99233c6a33f7715d0280b12

Observation 22678879-e4ff-48c9-a5e2-c704d1142560 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:17.922012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:17.922012Z digest=sha256:3845dc10f3523545a9fd6dc2368809287bd9844c7697f1597aad897580896751

Observation a128282c-278f-4391-be66-e1653b7e6678 · inbound

Imagination Helps Visual Reasoning, But Not Yet in Latent Space cites this paper.

Imagination Helps Visual Reasoning, But Not Yet in Latent Space OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T20:40:30.201984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:40:30.201984Z digest=sha256:1eb59035d68ae63dfd25ea2d6312e1c050c48d1d639205719f40dc28be9f0e38

Observation 7437eefe-2996-476f-b639-1331fd3a71b6 · inbound

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild cites this paper.

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T18:54:55.565724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:54:55.565724Z digest=sha256:c5b22ef6a28c43012211690b8eac47e8e65efd61c2604112a5ddad2ead99c60e

Observation f2fd8daa-5b75-42dd-b067-53c9c43593f0 · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:7bda0ae5fe4b4a02571e9aae54c6550cd7b71e6da59d959b6a8bc2460e4df3d4

Observation bf1770ba-0932-4a8c-a5bc-fa0648506db0 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:3b6b6a5312d72dc70798380ba6d14bfb380d5dc76697f6141069ac7e4b7e3f1d

Observation 287797b8-4126-43ae-a108-e15a56a75096 · inbound

Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding cites this paper.

Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:02:39.796357Z digest=sha256:ebd41c42ba92696e5a9032f7a1a48a99f840f3b94a2abd6a373535d18dbd320e

Observation 6d9f6490-5178-482e-aac2-99e59826f751 · inbound

Discovering Failure Modes in Vision-Language Models using RL cites this paper.

Discovering Failure Modes in Vision-Language Models using RL OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:21:25.761342Z digest=sha256:33f20e1e9987eff4f95ecf8ab65384fec88b3a8b5b3eba1e7f6713def3a1eed3

Observation 86174706-c22b-4937-957e-04dba896f8d5 · inbound

ParseBench: A Document Parsing Benchmark for AI Agents cites this paper.

ParseBench: A Document Parsing Benchmark for AI Agents OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:20:02.687230Z digest=sha256:f6bfd29b25349bf1dbd4fa2fc809b858f31bc5565c5170f8a4b112f3fbc3a086

Observation 7568b540-08a7-4087-bd19-df6410f5637e · inbound

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection cites this paper.

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:28:11.747826Z digest=sha256:2c3cb787a66532aa2279d505c1148fff14349980a5268374f28c3c57e0d53b0e

Observation 41de8ea6-3589-4f42-843e-9ec13c46dfd7 · inbound

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence cites this paper.

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:16:03.854650Z digest=sha256:78765b41ab6949d967bc8d0d62b237770523f7b8c06743ee8d5dddd3ccef1f79

Observation 19097b94-2122-42bc-bf6b-ee0f60f5a911 · inbound

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models cites this paper.

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-07T16:33:46.387850Z digest=sha256:5be720100fe33ab71764748914fd89a9e6988e9a77584876f5e59d9153f9c64a

Observation 5c239e8e-a2bc-4da1-a3f3-4ad843a89ed1 · inbound

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion cites this paper.

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:03:36.551943Z digest=sha256:46b38250c8d21627e48ee2c0abb415e07b22f03df62ae9fa3290a8df43ac49ee

Observation a58cd611-7ad1-425b-9979-eadc5f85b4d1 · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:be4f6acee757cc92bd5e34397ddcd2a487e0b04275b026ca5306e8b0fc92552f

Observation 1d3e5138-2a74-4901-b3bb-6cfb06f0d6d2 · inbound

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings cites this paper.

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:28:02.152600Z digest=sha256:bbdedfd29bb2886d17d644f63949440b84d7b90c570885a366c7c9d4a4a36f36

Observation fcb95031-61a2-4374-ac25-a12be3fbc3f2 · inbound

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? cites this paper.

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:42:54.802658Z digest=sha256:57b100d8dabfa44f6795660313843754fb8fe5c99206bab65a760a64f72fb0d3

Observation 20465ccb-766f-4f8c-8d4b-964aceff6cfd · inbound

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images cites this paper.

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T01:26:47.052025Z digest=sha256:4680382fc581fdca693808e9a1a420292416dd502d7672c188404b22075233be

Observation 961edea4-ba58-4ac5-ba36-642d9610b59a · inbound

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters cites this paper.

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:48:44.584051Z digest=sha256:bd17f6f83f3f5a6f07dfd3e7c02b435c34e610f062e608f614beab231f470769

Observation a11a964b-2432-49fb-8b43-141d99e97590 · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:0c676a57973a8cfaa9d0f67b693b7c53fcf4449cf808ddb5f5de501906e1e5e0

Observation 72ee3ef8-91de-4532-97e4-1fae95026bed · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:33:14.564545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:4c2c1b2f94d96667acc3182fb9871e13f112574e838c790e1386d2c5e45dd69c

Observation 1908dd2a-5533-40a5-94c0-d31b6c7cc804 · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:23:03.622203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:6cab42da256765f4382be3b82bd78501ca6fec2c01e3b52f25cfb3503d9a5ab7

Observation ac581788-5040-45e4-baf1-5164792e3c2c · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:39:53.720810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:14d9fddfcadd3d33fe5a41fe6035a1bd419f6ceefb1c69fa34e6abcb1277f005

Observation 2c7155a6-e2c7-4136-89cb-c535243115f1 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.316832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:e5f4b6354bdcb68a5679ad349f198490262709f92bf91de648bf4b721e20039c

Observation 508fc88b-b751-4b57-ac3b-738d90e61aaf · inbound

Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation cites this paper.

Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.536764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:38:33.963054Z digest=sha256:54df09ace42b9e2b7bdfbb1d5a1eccf04e944e335ba90474daee42a859147c84

Observation 9c0425b3-386d-4e94-b019-587cb1c8eba5 · inbound

Symbolic and Abstractive Reasoning with Complex Visual Queries cites this paper.

Symbolic and Abstractive Reasoning with Complex Visual Queries OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:07:30.854389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T16:46:18.863886Z digest=sha256:98a58f01b840032f15d0c074346737a62d3db9840423c27e3496e4234d650aa6

Observation 6335649a-fbfe-4f54-95cb-826b3208204f · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:38:44.234528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:f36f98b160a02a91f5906bf15d85b119f9b6f3445c75da638de1f14ee3630ffb

Observation ce72c728-d3ee-41a3-82c3-5ab464911f59 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 140

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:56.196839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:159e4b02bde21f74e14396fb5674c5f95ae67a19c867fb385a64f4074baaa685

Observation 520f187e-0aa2-493e-b4c8-e41744aa4e95 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.178134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:425daf12f4284c50350ae8e1755e0695180f5e856c947fe27c6ed72ec779fc05

Observation e5d92f9b-9172-4b94-b72d-598e7a6f9e3e · inbound

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering cites this paper.

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:39:57.611502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:24:20.208132Z digest=sha256:89582a846698172df11245f54af57cf078f99e55fc6c69d94eab0342a823594e

Observation eb67aa88-af60-4eb3-a6bb-53b07fab1fca · inbound

How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations cites this paper.

How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T21:00:09.485484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-25T19:13:27.527971Z digest=sha256:70d1ef5c8d5d5b24eb4a20dfe7c6fd491bb56ea40aab2dda1052a28a8f605c5a

Observation 76accdc0-19da-428a-9d04-a2fc8aaf6ab5 · inbound

StrucTab: A Structured Optimization Framework for Table Parsing cites this paper.

StrucTab: A Structured Optimization Framework for Table Parsing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:24:18.911083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:22:21.694355Z digest=sha256:2613228976579bbe8c3113afabb9071113e0b2dda38fd988decc8cdfea6b58e8

Observation 4a6d5ac7-5cd0-4b7b-bf0e-a9c6e80ea2ac · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:04:21.428160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:572a97ab77970b1ac4f71993fb120c2056b5aaa418f6faf04c7a058ee71aa768

Observation 5a7eb3c8-ed95-4ac7-8656-369f6381cfac · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:17.443295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:757b65d9ef8a355405cf4578cdf700c7d8f008b95afff1248f2d28e947e8a768

Observation d0f61be4-a898-45e3-ad91-d1ed1b1ed551 · inbound

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators cites this paper.

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:28:33.664386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T15:23:48.748933Z digest=sha256:3ce955e779041f450fad17d0814f5b8c29bb85186cf32a2e950d57c4ea38d05e

Observation 0b898045-7a72-44b1-a8b8-ee2d60cf0b6a · inbound

RADIO1D: Elastic Representations for Condensed Vision Modeling cites this paper.

RADIO1D: Elastic Representations for Condensed Vision Modeling OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:cb8f51e4ff5f770d6e4e18e9acf7118e4219658b83708c9046aa4960a8f7f86f

Observation faaf72c8-c70b-48df-8122-aa6204c9120c · inbound

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment cites this paper.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.102713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.102713Z digest=sha256:dc420096fc3d8d79b1068fee59e989d0a4650a6710dfdb7ab2d2be9858213f7c

Observation 9f6cffe7-2dc7-48dc-89c1-eda5a011adbb · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.518516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.518516Z digest=sha256:c2ede60dc55d2428288af33b70147743f8d45fb845030e14ddaf841a049eadb8

Observation e7b92d26-7f26-4394-a719-9e30eebff866 · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.238872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.238872Z digest=sha256:30bcfdf8b465b73ef219b050d40e1795dd9658323f3c7a55e33587243442b25e