Pith. sign in

Paper Citation Record · LEDGER

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

As of 21 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2505.00063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00063 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.691092Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.243457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:31:32.012726Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef23dea0-97f7-47c4-9b64-37c34d6f770a · outbound

This paper cites DeepSeek-V3 Technical Report.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.436465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.436465Z digest=sha256:451b01d4921a84573aa967a08d421b2c4d2c5d85311f1e86db75f6ef9ca4146b

Observation 0dfed56a-93db-49c8-9cfa-80589c13faca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.441586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.441586Z digest=sha256:7640620b24a049ec3ef34807c39884c9bb1fb0bfa67963707200a3a0bf8a94d5

Observation 7aae03d6-3165-4c7c-99a4-e49e9c0aa1a3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.447113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.447113Z digest=sha256:3151690d31414ee7508338246414d848f2434596786784372e24183028c66b80

Observation 8eee5f11-ab8e-41a7-911d-38b2212b53bc · outbound

This paper cites Gpt-4 technical report, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gpt-4 technical report, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.451755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.451755Z digest=sha256:bfaa10633bac23f9c127dd6f97197494fa0e8b405ecb9fef0fc81738cb04ee86

Observation 95cb0b15-fffa-490b-9382-f9fcbe081f46 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gemini: A Family of Highly Capable Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.456539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.456539Z digest=sha256:24542f8d99d0a012fdedd2c903ad738bc1c77e39ab176d29e972613b53b8ff78

Observation 256032a7-155c-4392-8f4f-5144e4ffb11c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.461285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.461285Z digest=sha256:ecaa8c5b46abfb8d66db3acfeb2fb9700a2f7b0daedad818fa81a21d6de3df51

Observation e4037381-cace-4743-b36c-11cceef578b5 · outbound

This paper cites Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.458280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.465415Z digest=sha256:825d272703870a1bbebb5166de6ec37e637959d832e92d07f8b10feb0d440a7e

Observation 48cd2450-3086-47aa-a10e-d035cfc5ca01 · outbound

This paper cites Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.445830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.469442Z digest=sha256:206ec61f44210effaa23e5a56f993c22ae5bb5a5dad4487536fd5cb8e4d718ce

Observation 8e79c4fc-deb5-4c2f-a9a2-0ed26d45b394 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.473305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.473305Z digest=sha256:6cf4f379e961056c1b106cfda476c7b0c774feadd828a4160cbdc23a74dfd945

Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.477616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.477616Z digest=sha256:a2ed0945d22ace11925888f8250683113746b8abda3b0df3159ffaee62bd1424

Observation 2df5237b-a74d-40ec-95c3-9d1bc6fc21a3 · outbound

This paper cites Tabpedia: Towards comprehensive visual table understanding with concept synergy.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Tabpedia: Towards comprehensive visual table understanding with concept synergy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.481951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.481951Z digest=sha256:c882bdd5ca4177f252cd01bf07551172a8648bbb52ce9ce75a016630a24e9001

Observation 2b21b38a-f6e9-47bc-bc22-727fa902dfad · outbound

This paper cites Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.423965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.485466Z digest=sha256:e03facc3bd9a0592d2ca06a20e2e4f20d3fff71d0a3a7728c9c63cea9ea87f76

Observation 26599e1e-d285-4baa-8b40-64cf9c1f40d5 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Overcoming catastrophic forgetting in neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.489725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.489725Z digest=sha256:88ecbb325f2541b3a939c89089cfcfee1381df4ad6685bb5d2818c2fd0db76c2

Observation 1d018a30-731c-4ad5-806b-4caf62dc6fec · outbound

This paper cites DuReadervis: A Chinese dataset for open-domain document visual question answering.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DuReadervis: A Chinese dataset for open-domain document visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.401634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.493483Z digest=sha256:404bc0843b0d10dfae12e6b6d5c89541524697ee4beee7c7b7d5518f0d413993

Observation 0f36c01c-7d62-4a05-9aac-48fcb308fe91 · outbound

This paper cites Visualmrc: Machine reading comprehension on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualmrc: Machine reading comprehension on document images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.388003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.497172Z digest=sha256:154c3ca260f70ed4266ab291e7c0b98f7297b3fbdd22a73a1c0a2338894ab36c

Observation 0de884fa-16c3-4040-b4a5-9b91c7146d85 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.374518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.500744Z digest=sha256:e12599d9340aa5f9ae9a45af451537744f3db56bb3789ec235aea6f48d1271b1

Observation 33600cd0-fc4c-484c-9ec2-b8a8b6a34403 · outbound

This paper cites Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.360061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.504674Z digest=sha256:67256b8665aac4ad5a206a58a3e75c76682d9247cdb0bb3548b11d6e33628305

Observation 63f34b33-3058-4466-ac11-71e449726465 · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.513802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.513802Z digest=sha256:3eb0031a0a633ed27e9ad15956716165d3ffec78d34d2cf7cdd1febcc3cfc9b5

Observation 10daf408-24a2-4a62-8ef9-73f593113801 · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.517859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.517859Z digest=sha256:07eef84a6bf4f035fa1ccb93a4d71022760d0bbd0f76d65f3c66cbb697dc74bd

Observation d75fb58a-1bd2-48e8-b3bc-ae8e248d874a · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Nougat: Neural Optical Understanding for Academic Documents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.522410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.522410Z digest=sha256:14110898876c9959d5a16833cb2ec44de55b2ee03635f172749e7e8e853013f0

Observation 1712c746-288c-478f-ac47-77c4499f7029 · outbound

This paper cites PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.526740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.526740Z digest=sha256:02fa8ba35456c4176df38461cce3fa42e2f749b4eaef2d0b4a9f4aad1c242ff6

Observation 41a14685-2ba2-4da9-9981-83eadea739b0 · outbound

This paper cites Publaynet: largest dataset ever for document layout analysis.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Publaynet: largest dataset ever for document layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.337991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.531155Z digest=sha256:1b4b32373f1b77381cb4cec87aa011fd7dfcd4a1865c0dd0e2b3ebff5157bb23

Observation 8cd6ea47-e0cc-4c0b-95f2-c45de9725e72 · outbound

This paper cites Detecting text in natural image with connectionist text proposal network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Detecting text in natural image with connectionist text proposal network

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.324591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.535073Z digest=sha256:3be10065241b5899158379444d240f39ad929bc2d9265d64901bc4d0c2c2a6a2

Observation ffd37e88-9139-4325-a946-9233d44b18d3 · outbound

This paper cites Textboxes: A fast text detector with a single deep neural network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Textboxes: A fast text detector with a single deep neural network

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.311428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.539314Z digest=sha256:c12a22f6e0ff9d0014dc556ead57819405800847e1714256add6a9859c4257da

Observation d82e9867-613e-4331-8d38-e1a3b6593bf2 · outbound

This paper cites East: An efficient and accurate scene text detector.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling East: An efficient and accurate scene text detector

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.298976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.543681Z digest=sha256:f3e12eab4afa18b98005c034f7e8d44d8b878633292f91a1d61392e25eac42eb

Observation b9e2be46-9cc0-4de2-aab5-9ba6f40745cb · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Curved scene text detection via transverse and longitudinal sequence connection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.286625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.548442Z digest=sha256:2655e7dfadb9b92a72d24145177da035ef2c1143cc5cb83dad9825935fe7e43d

Observation 9ec5d3f4-1911-47c7-82e1-7fd7155a6019 · outbound

This paper cites Gradient-based learning applied to document recognition.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient-based learning applied to document recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.553128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.553128Z digest=sha256:2771741e6afb710e8cfb805ad6f9ab3eab57f96f7aadd0cbfdf43593fa5aed1b

Observation 2d3e4c22-bd05-493e-95c0-05731349d10b · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.265383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.557387Z digest=sha256:c626c8421240ecf7a7692399db0009bca35776acea8c3a930f94b5aa430e1103

Observation 8c6dc4d6-f1b3-4d83-8bad-d3bd2926a0e7 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre-trained models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Trocr: Transformer-based optical character recognition with pre-trained models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.251990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.561367Z digest=sha256:02e0aa02890df5340feeed7d33310ad858764c1272e7d77e84a749ee9d8b9d8f

Observation c2e2f2d9-094c-4e41-b180-932d07ac1f6f · outbound

This paper cites General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.565654Z digest=sha256:b513ec8dc527f12c454a5c62ee06bcbace9f6a59f7a1d105fd180723dea73ca8

Observation 8996b0e3-6429-4831-a4a8-c789934f1974 · outbound

This paper cites Visual instruction tuning, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.569895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.569895Z digest=sha256:8b4071f81035528c51ca24e867b54ffb3be88cc829932e0ea4f2d3a52e2ef899

Observation d0488b0b-c210-491b-993b-17f12a1606ed · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.574129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.574129Z digest=sha256:6af5826a553f310af03fc512ccefa63bf4c16332244442bb505ffca62eb15492

Observation 922212e8-41a2-47e2-8247-2eb451015deb · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.578722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.578722Z digest=sha256:ba2b5f8f4aec29509c7685e405847d843d3c0dceeb5179350ab15825798fb888

Observation e0997066-b011-4465-9296-49893f0d18e6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.583499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.583499Z digest=sha256:b50ce72e40017bbde4d7cffd79ffc89864b3da30e5b4b13e858b781eb6b7580a

Observation 1d82a257-955e-4ba6-9d4a-90268be311fe · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.588006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.588006Z digest=sha256:75305c190dc79b25518b661e247150db940316df3847d40ed1d0ede457d0519e

Observation a886ceed-1c2f-4996-9721-b5f8b65c349d · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.592340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.592340Z digest=sha256:b0b4fdf174a024dda7ae448fd5c41b57f061c4dbdde4901ab8125cebc5f53daf

Observation 5d83cfa6-9e42-4c5d-956a-efbde08c06dc · outbound

This paper cites Learning transferable visual models from natural language supervision.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.216221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.596918Z digest=sha256:71d270e29af1165d07ca95751e1cb244dab0b227518252ab90ab797aaa753866

Observation 797929cd-e1e4-421a-ba37-31d011647556 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.600852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.600852Z digest=sha256:e99b4f7232a34bd06680333e7c80e7c3cb5499d4d879caa0cf3192bf8f06c8b6

Observation e95e7813-7efa-4938-bf55-3a8a040aea4a · outbound

This paper cites Lamol: Language modeling for lifelong language learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lamol: Language modeling for lifelong language learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.203170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.606072Z digest=sha256:6b87d59beb273178a7da75a371efd6801587d9d9048ec9de164ecdc71f8babe9

Observation c54b0be1-acc4-4cef-9d30-0429b1f42896 · outbound

This paper cites Rational LAMOL: A rationale-based lifelong learning framework.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Rational LAMOL: A rationale-based lifelong learning framework

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.189763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.610475Z digest=sha256:454cf1064cbd3dcf758693d894414ed7cf7eefb914eb40f046907376fcf84245

Observation 82ee7177-343a-4e43-a0c8-b7b7b7adc093 · outbound

This paper cites Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.176505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.614501Z digest=sha256:5f2b4428d83617cc0f94467d0393eb30756802f2ea86e408c721c8e48c027d2b

Observation 81af4035-99a5-4568-aa02-a9478c71ba96 · outbound

This paper cites Progressive Prompts: Continual Learning for Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Progressive Prompts: Continual Learning for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.618680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.618680Z digest=sha256:e64555df86d9378fcdd5b7fb57db00b287f670bba31185bfd5adc73c90f0d3f9

Observation 173fb86d-f4b2-44c5-80fb-fd64ffbcf65b · outbound

This paper cites Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.622890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.622890Z digest=sha256:ed6fcf7c73b69d353e5bd22ff3a80030ffb6a37734348a9cf93922c6d0984362

Observation 68db0b5c-fd5e-40ed-87c2-ccfe3df8d7bd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lora: Low-rank adaptation of large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.626804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.626804Z digest=sha256:d7216204581c194d751de5715449d7488debb5ea3fdbda01024da99c02d83049

Observation 6d2a9d72-d308-4c3d-aeae-538066b61641 · outbound

This paper cites Continual sequence generation with adaptive compositional modules.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Continual sequence generation with adaptive compositional modules

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.154042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.631198Z digest=sha256:526daa883a5193e1efba1ba81cb2f0b95879c3b0ad7adac0ecca1bd2f76fda06

Observation 64ae4ae8-111b-4118-8b71-3ee8fb3e737a · outbound

This paper cites Preserving in-context learning ability in large language model fine-tuning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Preserving in-context learning ability in large language model fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.635375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.635375Z digest=sha256:ec675eed0dd3f6fbffb2e733aa258716fabce7e0438166b770e0a485027909aa

Observation 03360ac3-5c97-4f2b-be02-6f313733a82d · outbound

This paper cites Editing models with task arithmetic.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Editing models with task arithmetic

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.130337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.639331Z digest=sha256:e7db4804224f20004a8681b35da19c9d1a42544513212a4f125d0f5ab3f884de

Observation d36672bf-a3f0-41d6-8f2a-af23e095b6ec · outbound

This paper cites Gradient projection memory for continual learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient projection memory for continual learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.116121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.643336Z digest=sha256:3e75c9f68569259b0a9be94d96e00ffc01b0ca969dd816a049306b6df8710c47

Observation 2c79a614-6c97-4c64-ba44-130819d5fd74 · outbound

This paper cites Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.102649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.647599Z digest=sha256:0cd1c8e32b906eda9cc5a52e544d7fd0930faca2ed3404d04f26d8c9a55ce8c2

Observation bdc4997a-bc58-43a5-b345-6991e070fa4a · outbound

This paper cites Binary codes capable of correcting deletions, insertions, and reversals.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Binary codes capable of correcting deletions, insertions, and reversals

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.652164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.652164Z digest=sha256:1f2dcf01573795b916f950ecb65478fafc4a7a5955821ba4c125cedbfdf00989

Observation d4a37fe8-b0cd-4f5a-b8d4-e235d78f3865 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.656654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.656654Z digest=sha256:3bdfda7729603a8fbd543dd914b6630e844b6964123cc9262710181fb7fe92f3

Observation 19630431-9e97-45b1-9486-d0f27d7a6590 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.661099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.661099Z digest=sha256:f7a8846ae20d6a9a7df9d4759f1f240051254f88754a6f3b7b2334c907954942

Observation a313a46e-512d-4f45-b274-2b209bb8c639 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docvqa: A dataset for vqa on document images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.665658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.665658Z digest=sha256:01fe37432d4aefd043bf7c7e8d05092eeed3ada1c1696f7abc3ecb4bd2c76267

Observation 1235ce46-603f-4cfe-a02d-550ffdccf97b · outbound

This paper cites Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.069875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.669573Z digest=sha256:7ba3f3a241cb1fb4951938d1e7e75fd7fc8f9c510d6a24eb0a118820af5b01c2

Observation 3f79ab19-d454-45af-825d-ef10ac8387e1 · outbound

This paper cites Towards vqa models that can read.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.674358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.674358Z digest=sha256:06f5b6c46b164fdece0965965f8913b263248b27b54e3771c762d750bb915c6c

Observation 0fd1fb34-22b1-47fb-8f26-ddfeff0de97c · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.678427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.678427Z digest=sha256:37123d0d0eb971b0a8df861942ef2dcd6f4bf8c2f7314c8ded0e3043412933e9

Observation 91f25037-2c36-4fc1-ac05-3854adee8f7c · outbound

This paper cites Infographicvqa.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Infographicvqa

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.682362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.682362Z digest=sha256:3041864ad2ed5f2d832195c37b71aba0c0f0320193a0c967b764c015568d8abf

Observation c0b50e32-06a3-45b4-9e96-a10264b63551 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.028544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T05:01:15.686867Z digest=sha256:57f412cf03628dd5b36d0ab17d25177d51e3bffc74791458987aa17cda67437d

Observation 1bd3d5dc-95fc-4730-8710-cda898b85a6b · outbound

This paper cites GPT-4o System Card.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling GPT-4o System Card

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.691092Z digest=sha256:0710c11c577077f028e399223c9675833f989c3b0852ca50966f8e67f4c5344f

Pith citing papers

Observation 8a091469-9a7d-47a4-8830-d8e7bd55df8e · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:13.243457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:13.243457Z digest=sha256:2025d5c50a259e5ded6127e524a387ed985f4891f785d3c4b8baf0d50e9632a9

Observation 52333cf1-07aa-4464-b696-c4a155fd1775 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:32.084497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:31:20.388845Z digest=sha256:9723e9aa516223af36d0e56ebaf2191d12948d1f899cc3fd8974ece6d87f15ef