Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:33:26.613927Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 100 of 156 outbound references and 45 inbound Pith citation observations for arXiv:2501.00321.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:33:26.613927Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:03:07.345505Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 156 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation 84db86e9-daa5-45a0-803f-e6dfb24e4819 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 881d1380-983c-48bb-ad9a-87be81d399d3 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaMA: Open and Efficient Foundation Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 32c8be8d-283a-4ab9-abe8-09b2c19211bd · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Language models are few-shot learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19455080-45aa-42e3-b2d9-c861fbf23001 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7e559013-a35b-4d93-bc3a-09416380e193 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Visual instruction tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e552762-b7bd-4f7f-93b0-96e847f509e1 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7731a63b-db2d-49a7-ae06-be6343d86ea3 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb8bb696-acf7-4ff5-84a4-50ac1619a710 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b89acf0-6d5f-4b6d-9e8a-432c6991a3c8 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec63bf76-4f80-4d1f-9484-dbeb9663f85a · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Towards vqa models that can read
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b159bc0-d92d-4e80-affe-05fe7b571d58 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Scene text visual question answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1e6a0ed-95dc-4834-a1b4-5e7c7570374b · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning On the general value of evidence, and bilingual scene-text visual question answering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 943a1730-bffa-4aa7-96c3-7fce26ec1f0d · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f421b4f-c3aa-4521-9c22-dc66af286c46 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7fed4a12-5eaf-40e0-bb7f-58a14df4529b · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50dd66b5-d3bb-47a1-9419-8a80f935cedb · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ConTextual: Evaluating Context- Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 66b56da9-27c4-4e76-a85a-14de545037e6 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Focus Anywhere for Fine-grained Multi-page Document Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4920d74-91e3-436b-8241-f3b90f43cc97 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c094f01d-c59d-4de8-b355-0e88f9ab39cc · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a278470e-2183-4ab4-ab02-6f75b4a40463 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ad4fec7b-c1d5-43b0-86be-10c82b9e818f · outbound
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a51669fb-7193-4b01-afe5-1ea7351b44cf · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Docvqa: A dataset for vqa on document images
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe5de7f2-d8da-4945-9db7-4404a8086643 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0dee846d-ef22-4f7c-b0d4-c58e3a712663 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Hello GPT-4o
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2655b100-79ee-4075-a31b-8e8eb81030d6 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44a67d08-3b66-41b7-a987-03b28e41f3aa · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f474b427-eae1-430a-ba6e-6291ebc0328f · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9db08b98-011d-43f9-872c-b2cc8f92729f · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Multimodal Table Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb2016e4-1944-4451-ba05-fcd419b06fe1 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a793304d-b92d-4fa4-8199-35015e9dc1ae · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fd1ab135-9013-424b-9169-dbc158c19fcc · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82d49e9b-a8ca-4c02-923a-26fdc7ad9a2f · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02db6aa8-8324-41a9-ae1c-46a9a356fac4 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37bf3566-362d-4ee1-9c88-76f1a5987481 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a94e4643-1970-4d16-8da9-39b3ed21684b · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 40ca4e6a-fa75-403d-a7a7-9d83248bd0fb · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Dockylin: A large multimodal model for visual document understanding with efficient visual slimming
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 840efc04-9dbd-4cac-bce6-e3a3808c3239 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2fdcb25-4e3c-40db-8244-55ccb432a752 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning A simple yet effective layout token in large language models for document understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00d0198c-f48c-4841-a193-53492814b698 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Adaptive markup language generation for contextually- grounded visual document understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1955ee1-cea6-47ba-ba97-e272513d3b01 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Marten: Visual question answering with mask generation for multi-modal document under- standing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e60bd23c-2f86-49d6-95c8-0ed0f809b906 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a97e36e6-a5b2-48e4-8e77-848e7ef7df5f · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Infographicvqa
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb36e648-22c6-455b-822a-1135fb48381c · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Exploring the Capabilities of Large Multimodal Models on Dense Text
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d178af2b-32f4-4c72-a151-796e3b1f2abe · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Onechart: Purify the chart structural extraction via one auxiliary token
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 07d29817-b1be-4c1a-bbec-36aef3b6d76c · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Document understanding dataset and evaluation (dude)
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7154c4e-8d91-4eb6-98a3-b7197d2c7be5 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Needle in a multimodal haystack
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f6e4434-1f6d-461d-bb1c-422f535b82a2 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Hierarchical multimodal transformers for multipage docvqa
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0b21ae1-ae24-4b69-9d6f-2453b223b098 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 34cb1f0c-4152-4cd7-bcbe-206f055ac3ba · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Llava-next: Improved reasoning, ocr, and world knowledge
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 661e7cc0-3ca7-438a-a978-04ee1868d002 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVA-OneVision: Easy Visual Task Transfer
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a385b12-c2d5-42a1-bf5e-1288ca0fa732 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Monkey: Image resolution and text label are important things for large multi-modal models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e70fe4f2-ee1e-4a49-ba6e-5e76eff24780 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a90eab5-3f11-4789-bb32-6123be621949 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2dfdd845-0f96-4b05-b5ed-982407f30943 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Pixtral 12B
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d99554a9-9e11-4b29-b58a-88440b6ad2d3 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f54d04a-5df7-454b-8280-6272c16950d3 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3150e90b-6433-49d9-a42d-f2179f727af3 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0878f72-9289-411e-b773-26e5a6ef4bc4 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6a59630-a4d6-49c0-a69b-a73d6e7254c1 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning GPT-4o mini: advancing cost-efficient intelligence
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 577cebb0-3551-4368-807d-7bc6469ef662 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Gemini: A Family of Highly Capable Multimodal Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da013c29-5dc4-4a60-9bc0-5cd420e35f58 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Claude 3.5 Sonnet
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c8613577-097e-4035-a125-481c9210224e · outbound
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 32211377-a178-40ad-ae96-9fb0f9d7de21 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Image-based table recognition: data, model, and evaluation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 37896fe2-f706-4675-a7c2-2898d5c65f1a · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Bleu: a method for automatic evaluation of machine translation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dcfcacb3-841a-4db7-8457-acf8f6e7948d · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning METEOR: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65b5187b-86ed-4f2a-96db-a067b88a9121 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e182cd3-1e23-49e5-9654-3737dff2599a · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5bde4b28-72e2-42eb-a54b-42c084b4aff6 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Read like humans: Autonomous, bidi- rectional and iterative language modeling for scene text recognition
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2fb446a8-a1c8-4a2d-a4d8-8107ef434a1c · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Aster: An attentional scene text recognizer with flexible rectification
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f0d324b-1d7e-4a6e-be06-1b35442b0b88 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Master: Multi-aspect non-local network for scene text recognition
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 12223856-c908-429a-8ea2-20456aa9a480 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SVTR: scene text recognition with a single visual model
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 983b0bd0-82de-424f-a91e-bff1020715b7 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Abcnet: Real-time scene text spotting with adaptive bezier-curve network
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 07fec2f0-bbf4-43e3-b0f8-349df359c188 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c90d9932-5c5e-4cf5-a7c4-d28606d46b4e · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Text spotting transformers
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41a82524-2cb0-49af-ab68-504105b1385a · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Total-text: A comprehensive dataset for scene text detection and recognition
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b2f4c51-28a2-4949-a9d1-5d03b6f70994 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b911baa8-73f5-4975-9ec5-87f9729b5a6d · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar 2013 robust reading competition
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f338d15-e4a7-43f4-9305-4ccd7c5c1b1d · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning End-to-end scene text recognition using tree-structured models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 385ea377-1a45-49e5-9769-e0c7c1274abd · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Scene text recognition using higher order language priors
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f2e0e0b-f52b-4b72-8df7-bb4e3d6ec5c0 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar 2015 competition on robust reading
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7dd7e8c5-71b2-445f-a5b0-f404b1d4efbb · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Curved scene text detection via transverse and longitudinal sequence connection
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0d259c6-f13f-4a01-86bc-96ac9a44bcdf · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53d94890-9112-4506-a167-363b64c1ea64 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning A robust arbitrary text detection system for natural scene images
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e2ef60e-aa7f-4e32-8a2b-6fa402dac245 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Available: https://api.semanticscholar.org/CorpusID:15559857
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4de4d014-8b5a-4139-9a1e-776a1d968bf7 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning PoseNet: A convolutional network for real-time 6-dof camera relocalization
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7fe56763-1822-40f2-8a43-c475d94c5238 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Toward understanding wordart: Corner-guided transformer for scene text recognition
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c759e9d-7e79-4c0a-a64e-b895ba9e3e2d · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning The iam-database: an english sentence database for offline handwriting recognition
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3ea8636d-2722-41a0-ad11-d4012ab98779 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Proceedings of ieee international conference on frontiers in handwriting recognition
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15fe4a3f-5b4f-478e-bbf7-9858853599e2 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning From two to one: A new scene text recognizer with visual language modeling network
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 046b2e8c-be9a-47f8-a872-32c9c3744b5d · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Detecting Curve Text in the Wild: New Dataset and New Solution
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f97fa4a5-71a0-4486-90e2-e0c8601ae257 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Towards end-to-end unified scene text detection and layout analysis
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2b307b3-5eda-4f9f-84f6-fac6583d7208 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning A large chinese text dataset in the wild
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e132b10b-e3d0-4cfd-9c65-453e7f2bc014 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar2017 competition on reading chinese text in the wild (rctw-17)
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94fd4264-ba9a-4722-8400-535deb6a62dd · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f4caf50-a2d2-41ed-9afe-350be61b7857 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90714572-d6da-4b4f-8a51-8d8188ee570c · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning M$^{6}$Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 562740c0-5339-4776-bfc0-5b9fd848ab61 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Rico: A mobile app dataset for building data-driven design applications
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb0e42b6-1c20-452b-9a78-faa42f6c9d56 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Funsd: A dataset for form understanding in noisy scanned documents
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b4b9853-5e77-4f7b-92ff-b0c21fc3ce21 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Icdar2019 competition on scanned receipt ocr and information extraction
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2be0a722-dd3d-4494-afdd-e6946371d6b9 · outbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning Visual information ex- traction in the wild: practical dataset and end-to-end solution
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb3702ed-2f72-4b02-b978-4fae2d648b4f · inbound
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5fb6f3-e505-484b-b225-79e59fe5bb36 · inbound
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b965d6-31f2-41d0-8c12-0fd77429e975 · inbound
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d0a6e8-8ef2-4742-ab23-e6dc1ea29c06 · inbound
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a83625d-0c7d-4f1f-a2a7-fc122f643337 · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdfd32a0-7549-4f8b-b451-84649a85c8d1 · inbound
FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a3d0f3e-386f-4ee4-bf52-4be9e81254f9 · inbound
LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38ab542-208f-4b65-bd5f-482472da0d01 · inbound
Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22678879-e4ff-48c9-a5e2-c704d1142560 · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a128282c-278f-4391-be66-e1653b7e6678 · inbound
Imagination Helps Visual Reasoning, But Not Yet in Latent Space OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7437eefe-2996-476f-b639-1331fd3a71b6 · inbound
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2fd8daa-5b75-42dd-b067-53c9c43593f0 · inbound
From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf1770ba-0932-4a8c-a5bc-fa0648506db0 · inbound
Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 287797b8-4126-43ae-a108-e15a56a75096 · inbound
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d9f6490-5178-482e-aac2-99e59826f751 · inbound
Discovering Failure Modes in Vision-Language Models using RL OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86174706-c22b-4937-957e-04dba896f8d5 · inbound
ParseBench: A Document Parsing Benchmark for AI Agents OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7568b540-08a7-4087-bd19-df6410f5637e · inbound
Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41de8ea6-3589-4f42-843e-9ec13c46dfd7 · inbound
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19097b94-2122-42bc-bf6b-ee0f60f5a911 · inbound
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c239e8e-a2bc-4da1-a3f3-4ad843a89ed1 · inbound
Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a58cd611-7ad1-425b-9979-eadc5f85b4d1 · inbound
CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d3e5138-2a74-4901-b3bb-6cfb06f0d6d2 · inbound
How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fcb95031-61a2-4374-ac25-a12be3fbc3f2 · inbound
LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 20465ccb-766f-4f8c-8d4b-964aceff6cfd · inbound
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 961edea4-ba58-4ac5-ba36-642d9610b59a · inbound
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a11a964b-2432-49fb-8b43-141d99e97590 · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72ee3ef8-91de-4532-97e4-1fae95026bed · inbound
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1908dd2a-5533-40a5-94c0-d31b6c7cc804 · inbound
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac581788-5040-45e4-baf1-5164792e3c2c · inbound
ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c7155a6-e2c7-4136-89cb-c535243115f1 · inbound
ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 508fc88b-b751-4b57-ac3b-738d90e61aaf · inbound
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c0425b3-386d-4e94-b019-587cb1c8eba5 · inbound
Symbolic and Abstractive Reasoning with Complex Visual Queries OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6335649a-fbfe-4f54-95cb-826b3208204f · inbound
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce72c728-d3ee-41a3-82c3-5ab464911f59 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 520f187e-0aa2-493e-b4c8-e41744aa4e95 · inbound
TuringViT: Making SOTA Vision Transformers Accessible to All OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5d92f9b-9172-4b94-b72d-598e7a6f9e3e · inbound
ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb67aa88-af60-4eb3-a6bb-53b07fab1fca · inbound
How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76accdc0-19da-428a-9d04-a2fc8aaf6ab5 · inbound
StrucTab: A Structured Optimization Framework for Table Parsing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a6d5ac7-5cd0-4b7b-bf0e-a9c6e80ea2ac · inbound
DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a7eb3c8-ed95-4ac7-8656-369f6381cfac · inbound
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d0f61be4-a898-45e3-ad91-d1ed1b1ed551 · inbound
ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b898045-7a72-44b1-a8b8-ee2d60cf0b6a · inbound
RADIO1D: Elastic Representations for Condensed Vision Modeling OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faaf72c8-c70b-48df-8122-aa6204c9120c · inbound
Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f6cffe7-2dc7-48dc-89c1-eda5a011adbb · inbound
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b92d26-7f26-4394-a719-9e30eebff866 · inbound
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.