Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.691092Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2505.00063.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.691092Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.243457Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T14:31:32.012726Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ef23dea0-97f7-47c4-9b64-37c34d6f770a · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-V3 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dfed56a-93db-49c8-9cfa-80589c13faca · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aae03d6-3165-4c7c-99a4-e49e9c0aa1a3 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eee5f11-ab8e-41a7-911d-38b2212b53bc · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gpt-4 technical report, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95cb0b15-fffa-490b-9382-f9fcbe081f46 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gemini: A Family of Highly Capable Multimodal Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 256032a7-155c-4392-8f4f-5144e4ffb11c · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4037381-cace-4743-b36c-11cceef578b5 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48cd2450-3086-47aa-a10e-d035cfc5ca01 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8e79c4fc-deb5-4c2f-a9a2-0ed26d45b394 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2: Benchmarking Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df5237b-a74d-40ec-95c3-9d1bc6fc21a3 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Tabpedia: Towards comprehensive visual table understanding with concept synergy
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b21b38a-f6e9-47bc-bc22-727fa902dfad · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 26599e1e-d285-4baa-8b40-64cf9c1f40d5 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Overcoming catastrophic forgetting in neural networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d018a30-731c-4ad5-806b-4caf62dc6fec · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DuReadervis: A Chinese dataset for open-domain document visual question answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f36c01c-7d62-4a05-9aac-48fcb308fe91 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualmrc: Machine reading comprehension on document images
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0de884fa-16c3-4040-b4a5-9b91c7146d85 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 33600cd0-fc4c-484c-9ec2-b8a8b6a34403 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 63f34b33-3058-4466-ac11-71e449726465 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10daf408-24a2-4a62-8ef9-73f593113801 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling MinerU: An Open-Source Solution for Precise Document Content Extraction
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d75fb58a-1bd2-48e8-b3bc-ae8e248d874a · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Nougat: Neural Optical Understanding for Academic Documents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1712c746-288c-478f-ac47-77c4499f7029 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a14685-2ba2-4da9-9981-83eadea739b0 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Publaynet: largest dataset ever for document layout analysis
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8cd6ea47-e0cc-4c0b-95f2-c45de9725e72 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Detecting text in natural image with connectionist text proposal network
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ffd37e88-9139-4325-a946-9233d44b18d3 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Textboxes: A fast text detector with a single deep neural network
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d82e9867-613e-4331-8d38-e1a3b6593bf2 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling East: An efficient and accurate scene text detector
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9e2be46-9cc0-4de2-aab5-9ba6f40745cb · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Curved scene text detection via transverse and longitudinal sequence connection
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9ec5d3f4-1911-47c7-82e1-7fd7155a6019 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient-based learning applied to document recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3e4c22-bd05-493e-95c0-05731349d10b · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c6dc4d6-f1b3-4d83-8bad-d3bd2926a0e7 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Trocr: Transformer-based optical character recognition with pre-trained models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2e2f2d9-094c-4e41-b180-932d07ac1f6f · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8996b0e3-6429-4831-a4a8-c789934f1974 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visual instruction tuning, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0488b0b-c210-491b-993b-17f12a1606ed · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922212e8-41a2-47e2-8247-2eb451015deb · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0997066-b011-4465-9296-49893f0d18e6 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d82a257-955e-4ba6-9d4a-90268be311fe · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a886ceed-1c2f-4996-9721-b5f8b65c349d · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Focus Anywhere for Fine-grained Multi-page Document Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d83cfa6-9e42-4c5d-956a-efbde08c06dc · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Learning transferable visual models from natural language supervision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 797929cd-e1e4-421a-ba37-31d011647556 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95e7813-7efa-4938-bf55-3a8a040aea4a · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lamol: Language modeling for lifelong language learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c54b0be1-acc4-4cef-9d30-0429b1f42896 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Rational LAMOL: A rationale-based lifelong learning framework
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82ee7177-343a-4e43-a0c8-b7b7b7adc093 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 81af4035-99a5-4568-aa02-a9478c71ba96 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Progressive Prompts: Continual Learning for Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173fb86d-f4b2-44c5-80fb-fd64ffbcf65b · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68db0b5c-fd5e-40ed-87c2-ccfe3df8d7bd · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lora: Low-rank adaptation of large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2a9d72-d308-4c3d-aeae-538066b61641 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Continual sequence generation with adaptive compositional modules
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 64ae4ae8-111b-4118-8b71-3ee8fb3e737a · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Preserving in-context learning ability in large language model fine-tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03360ac3-5c97-4f2b-be02-6f313733a82d · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Editing models with task arithmetic
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d36672bf-a3f0-41d6-8f2a-af23e095b6ec · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient projection memory for continual learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c79a614-6c97-4c64-ba44-130819d5fd74 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bdc4997a-bc58-43a5-b345-6991e070fa4a · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Binary codes capable of correcting deletions, insertions, and reversals
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a37fe8-b0cd-4f5a-b8d4-e235d78f3865 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19630431-9e97-45b1-9486-d0f27d7a6590 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a313a46e-512d-4f45-b274-2b209bb8c639 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docvqa: A dataset for vqa on document images
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1235ce46-603f-4cfe-a02d-550ffdccf97b · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f79ab19-d454-45af-825d-ef10ac8387e1 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Towards vqa models that can read
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd1fb34-22b1-47fb-8f26-ddfeff0de97c · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f25037-2c36-4fc1-ac05-3854adee8f7c · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Infographicvqa
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b50e32-06a3-45b4-9e96-a10264b63551 · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bd3d5dc-95fc-4730-8710-cda898b85a6b · outbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling GPT-4o System Card
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a091469-9a7d-47a4-8830-d8e7bd55df8e · inbound
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
Reference 255
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52333cf1-07aa-4464-b696-c4a155fd1775 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.