Pith. sign in

Paper Citation Record · LEDGER

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2505.00063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00063 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.691092Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.243457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:31:32.012726Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef23dea0-97f7-47c4-9b64-37c34d6f770a · outbound

This paper cites DeepSeek-V3 Technical Report.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.436465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.436465Z digest=sha256:3bcf61c7ec5b404ff9f24432b5af2cf4cdbdab62500c96db5a612f18f345abd9

Observation 0dfed56a-93db-49c8-9cfa-80589c13faca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.441586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.441586Z digest=sha256:d1e62699791ad6ce6966ad0b9e5c9b760b6661efda08f56321d7914fa96e99fc

Observation 7aae03d6-3165-4c7c-99a4-e49e9c0aa1a3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.447113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.447113Z digest=sha256:5a18b94fcf7fed5ebe291a1e8e8d5f802556903099113a9fc44ba52de92ceb37

Observation 8eee5f11-ab8e-41a7-911d-38b2212b53bc · outbound

This paper cites Gpt-4 technical report, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gpt-4 technical report, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.451755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.451755Z digest=sha256:1fdde35e257cd38a35fc12beeaed9daef3bdb048727930eef01d5962632d7511

Observation 95cb0b15-fffa-490b-9382-f9fcbe081f46 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gemini: A Family of Highly Capable Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.456539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.456539Z digest=sha256:5d3d6ab14132c2257829a6076b1d72a8446b7fe8dfd73bbfb06a66d280f98359

Observation 256032a7-155c-4392-8f4f-5144e4ffb11c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.461285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.461285Z digest=sha256:015ce17086e0ab6161254d372d6b821d64573f732b9e534fec341319a041f626

Observation e4037381-cace-4743-b36c-11cceef578b5 · outbound

This paper cites Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.458280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.465415Z digest=sha256:80440830aa035d56e61f891b792d94b1e0bfe2f8dada9c3a481f7b4155608785

Observation 48cd2450-3086-47aa-a10e-d035cfc5ca01 · outbound

This paper cites Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.445830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.469442Z digest=sha256:0a33345fb076fd18c7d2a62001c2f0d8ae58ee0bebdf4e77e18570628dad0f3b

Observation 8e79c4fc-deb5-4c2f-a9a2-0ed26d45b394 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.473305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.473305Z digest=sha256:81360a1d44c7cf4a1e753756c75042520097c126aa0648e329df6495b7fa9096

Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.477616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.477616Z digest=sha256:f36c2ee898092bb7f877357cf966d81c01ddeddd3d4f5549c402992c55fec0c5

Observation 2df5237b-a74d-40ec-95c3-9d1bc6fc21a3 · outbound

This paper cites Tabpedia: Towards comprehensive visual table understanding with concept synergy.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Tabpedia: Towards comprehensive visual table understanding with concept synergy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.481951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.481951Z digest=sha256:8335bc9861ffc750622f2d831360352510217db1d3b35a96a205045a89a8d76f

Observation 2b21b38a-f6e9-47bc-bc22-727fa902dfad · outbound

This paper cites Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.423965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.485466Z digest=sha256:8b8f161a816097f680f0c1c28676d05127ccbb8df40206e36f3993c2e7d76cfa

Observation 26599e1e-d285-4baa-8b40-64cf9c1f40d5 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Overcoming catastrophic forgetting in neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.489725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.489725Z digest=sha256:c25df872ddebbe984de1c4f91c7fbbb17b96b633ec8ab68267ad463125f4ba5a

Observation 1d018a30-731c-4ad5-806b-4caf62dc6fec · outbound

This paper cites DuReadervis: A Chinese dataset for open-domain document visual question answering.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DuReadervis: A Chinese dataset for open-domain document visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.401634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.493483Z digest=sha256:db34daebadd45ba102592eef26892b8035ae8c4100b7f7de23404dc34e141e73

Observation 0f36c01c-7d62-4a05-9aac-48fcb308fe91 · outbound

This paper cites Visualmrc: Machine reading comprehension on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualmrc: Machine reading comprehension on document images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.388003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.497172Z digest=sha256:e9651b5ce315335b7abcb890aedf90069d46c080d37488ce30d922a0042b3811

Observation 0de884fa-16c3-4040-b4a5-9b91c7146d85 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.374518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.500744Z digest=sha256:35fec17fe7ea6d887c4dd2374676d5a9112bae559af9e0c6208da6b1bd1e90ce

Observation 33600cd0-fc4c-484c-9ec2-b8a8b6a34403 · outbound

This paper cites Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.360061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.504674Z digest=sha256:a593c2c150f441d4af154f693999166eef938cc935f5635e009b13940b42c412

Observation 63f34b33-3058-4466-ac11-71e449726465 · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.513802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.513802Z digest=sha256:ae2d435401dade40f968433f339efb2393dc183f7d8f6de928f7442bf2171137

Observation 10daf408-24a2-4a62-8ef9-73f593113801 · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.517859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.517859Z digest=sha256:5b12b0335c839aa042083f37fd1ace7cd81884d96a6914d6bf048dee11a1804d

Observation d75fb58a-1bd2-48e8-b3bc-ae8e248d874a · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Nougat: Neural Optical Understanding for Academic Documents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.522410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.522410Z digest=sha256:1da693da1be415090cbe761b4096154244b3c294533bea016f2d73a7cc92b586

Observation 1712c746-288c-478f-ac47-77c4499f7029 · outbound

This paper cites PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.526740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.526740Z digest=sha256:6dace992bd7d3d9a9efe5091ed52a2590294458642268206cd95c2d9a9d31ae1

Observation 41a14685-2ba2-4da9-9981-83eadea739b0 · outbound

This paper cites Publaynet: largest dataset ever for document layout analysis.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Publaynet: largest dataset ever for document layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.337991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.531155Z digest=sha256:a93fb71ee483acf8fdb18945d86321c448509342ec5496926397a9ad89f371d5

Observation 8cd6ea47-e0cc-4c0b-95f2-c45de9725e72 · outbound

This paper cites Detecting text in natural image with connectionist text proposal network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Detecting text in natural image with connectionist text proposal network

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.324591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.535073Z digest=sha256:cbeab39ab9d5b8171ab1c478931b9280001ff9095fd1e53a42f8567753278e85

Observation ffd37e88-9139-4325-a946-9233d44b18d3 · outbound

This paper cites Textboxes: A fast text detector with a single deep neural network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Textboxes: A fast text detector with a single deep neural network

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.311428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.539314Z digest=sha256:7f7c8f04079ef1c73bea3a6b09dd79dff57bdcad544090af1cdc940a53913f56

Observation d82e9867-613e-4331-8d38-e1a3b6593bf2 · outbound

This paper cites East: An efficient and accurate scene text detector.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling East: An efficient and accurate scene text detector

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.298976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.543681Z digest=sha256:e0704452c963f340ad1d2db745061e80e911f04f7ed4251c2bd431fcb57aceed

Observation b9e2be46-9cc0-4de2-aab5-9ba6f40745cb · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Curved scene text detection via transverse and longitudinal sequence connection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.286625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.548442Z digest=sha256:bb0f0bf31afded8829f5d49d62679d116833a4b80f094446e7e90883bbdcb817

Observation 9ec5d3f4-1911-47c7-82e1-7fd7155a6019 · outbound

This paper cites Gradient-based learning applied to document recognition.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient-based learning applied to document recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.553128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.553128Z digest=sha256:0501fd1245101eaab139659dd64ba0589ec26657e03a5c6d865e6500c971463a

Observation 2d3e4c22-bd05-493e-95c0-05731349d10b · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.265383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.557387Z digest=sha256:bffdeae5860e10e73e68dc7295dd72fa34364ef6e73023530bff40f7858b221f

Observation 8c6dc4d6-f1b3-4d83-8bad-d3bd2926a0e7 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre-trained models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Trocr: Transformer-based optical character recognition with pre-trained models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.251990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.561367Z digest=sha256:9859798921d2712fbcc40c3b2cbebbc97a1b2f533aa7f67f0bc99adc8d76e83d

Observation c2e2f2d9-094c-4e41-b180-932d07ac1f6f · outbound

This paper cites General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.565654Z digest=sha256:364877ae6ac9e222bca68fe39b1d6b0164d74fb839ce8380b5b8b2cdf83ccb8b

Observation 8996b0e3-6429-4831-a4a8-c789934f1974 · outbound

This paper cites Visual instruction tuning, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.569895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.569895Z digest=sha256:d7c004e1dbdc5fc38ad829de5cc1332c020e765b71d833ff1c4abc78e79985c6

Observation d0488b0b-c210-491b-993b-17f12a1606ed · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.574129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.574129Z digest=sha256:3a71e114695c7a7317ac1b2c5181a123b3acf2abe02e7e5c4e54901cd702c29d

Observation 922212e8-41a2-47e2-8247-2eb451015deb · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.578722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.578722Z digest=sha256:28e28c729bb1e932484a2ea5b029d759138f2599a1bcc26889f1ba4abadcf426

Observation e0997066-b011-4465-9296-49893f0d18e6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.583499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.583499Z digest=sha256:2dbf0c3f9a3532c8f53ec5296c832f42d5f6717effe34848755de8279250f932

Observation 1d82a257-955e-4ba6-9d4a-90268be311fe · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.588006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.588006Z digest=sha256:cf87d425f1320ab16e976381407a631c8b1798ba368a6f54cefbebfb81a485d2

Observation a886ceed-1c2f-4996-9721-b5f8b65c349d · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.592340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.592340Z digest=sha256:56ab56a623a625ceaadcf2ca277b3ac4cf938554b1a028735fb35bba28fc8119

Observation 5d83cfa6-9e42-4c5d-956a-efbde08c06dc · outbound

This paper cites Learning transferable visual models from natural language supervision.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.216221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.596918Z digest=sha256:618725c0f4f6bf2f56fb4877f43adf754d57ae142c67560640617d529aa70bc3

Observation 797929cd-e1e4-421a-ba37-31d011647556 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.600852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.600852Z digest=sha256:7c328cf0d0ce1732d57d29937f78321b3a85ccf2099cd7fee5f87d70e05970de

Observation e95e7813-7efa-4938-bf55-3a8a040aea4a · outbound

This paper cites Lamol: Language modeling for lifelong language learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lamol: Language modeling for lifelong language learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.203170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.606072Z digest=sha256:f10a71de1942b9172ed2a23d0b734b54a2a7c999cb7e6c9d6b6260b4fbd0cecf

Observation c54b0be1-acc4-4cef-9d30-0429b1f42896 · outbound

This paper cites Rational LAMOL: A rationale-based lifelong learning framework.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Rational LAMOL: A rationale-based lifelong learning framework

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.189763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.610475Z digest=sha256:54398da40551d435f75c0f0067a59ad997f0aaa13e047b86ac50f8a31a7c6234

Observation 82ee7177-343a-4e43-a0c8-b7b7b7adc093 · outbound

This paper cites Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.176505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.614501Z digest=sha256:8b24cdfa3121ce5088cb4d8d3f83d8ce9cd0dc0a67cd1f9c36f3d0b75e6630fe

Observation 81af4035-99a5-4568-aa02-a9478c71ba96 · outbound

This paper cites Progressive Prompts: Continual Learning for Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Progressive Prompts: Continual Learning for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.618680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.618680Z digest=sha256:376d1c1c4158d3e237786f2afb5a8bd5eb524f29a0f9fab58ee1101d1b525f86

Observation 173fb86d-f4b2-44c5-80fb-fd64ffbcf65b · outbound

This paper cites Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.622890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.622890Z digest=sha256:b6caa14a48a07d56851a2a65de605eabe69eb075ec4a46bfe2894c0f6c538354

Observation 68db0b5c-fd5e-40ed-87c2-ccfe3df8d7bd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lora: Low-rank adaptation of large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.626804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.626804Z digest=sha256:c9824f1a714a1868a729f4583bd1d5efc080cf60cf1efc7d273818c96d55022f

Observation 6d2a9d72-d308-4c3d-aeae-538066b61641 · outbound

This paper cites Continual sequence generation with adaptive compositional modules.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Continual sequence generation with adaptive compositional modules

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.154042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.631198Z digest=sha256:df23396ce928699c82d06e4a0db36c8e9e76b69154121d68ac31d08ce9c77833

Observation 64ae4ae8-111b-4118-8b71-3ee8fb3e737a · outbound

This paper cites Preserving in-context learning ability in large language model fine-tuning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Preserving in-context learning ability in large language model fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.635375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.635375Z digest=sha256:81ffea344a815b819ec20994e0c75980c24cda1a4f27bcc69023b2a511ff44b7

Observation 03360ac3-5c97-4f2b-be02-6f313733a82d · outbound

This paper cites Editing models with task arithmetic.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Editing models with task arithmetic

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.130337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.639331Z digest=sha256:9a63a8d607ea944d3016458af1b036490abd7015d41a2248d543a1b4bee36e98

Observation d36672bf-a3f0-41d6-8f2a-af23e095b6ec · outbound

This paper cites Gradient projection memory for continual learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient projection memory for continual learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.116121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.643336Z digest=sha256:77becc47e2cd638f9eadf08375d16db7fdd47475982b6f4c101f78b89d0d1679

Observation 2c79a614-6c97-4c64-ba44-130819d5fd74 · outbound

This paper cites Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.102649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.647599Z digest=sha256:3bb765ba0cc10d83cc127105079feac387ed3aac5bcffbde1f05e8979cdb3a8e

Observation bdc4997a-bc58-43a5-b345-6991e070fa4a · outbound

This paper cites Binary codes capable of correcting deletions, insertions, and reversals.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Binary codes capable of correcting deletions, insertions, and reversals

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.652164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.652164Z digest=sha256:244b8fbebc1e8c10511aedaba42260b6567a44be83bf8d9fe047983bf132a4c7

Observation d4a37fe8-b0cd-4f5a-b8d4-e235d78f3865 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.656654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.656654Z digest=sha256:3b779ffd0be5d34fa562546ca6d6a37c250f7513d4f3dcb64874b17d8e690784

Observation 19630431-9e97-45b1-9486-d0f27d7a6590 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.661099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.661099Z digest=sha256:76871d6afb2d49f1f798e6597a2b6916df477ed7499ab9971ff07b77f7f44097

Observation a313a46e-512d-4f45-b274-2b209bb8c639 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docvqa: A dataset for vqa on document images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.665658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.665658Z digest=sha256:f765f1abdfef464d2e97f7d083ee5979a5e69fd62e1cf65202270f4a9e56589d

Observation 1235ce46-603f-4cfe-a02d-550ffdccf97b · outbound

This paper cites Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.069875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.669573Z digest=sha256:f26e5163b15e062f375eba8a4c73b3daf1cde5da2bbed93b0b66f4a0965ce0bc

Observation 3f79ab19-d454-45af-825d-ef10ac8387e1 · outbound

This paper cites Towards vqa models that can read.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.674358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.674358Z digest=sha256:e90573d22f37192f55bd988e45334bd847f418395f7e3f3c3d8be2cf79df555e

Observation 0fd1fb34-22b1-47fb-8f26-ddfeff0de97c · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.678427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.678427Z digest=sha256:cf49b25e7e2b0990bd9f74e5838d5c01290be60380b5dabe1091ed50ee6f8d28

Observation 91f25037-2c36-4fc1-ac05-3854adee8f7c · outbound

This paper cites Infographicvqa.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Infographicvqa

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.682362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.682362Z digest=sha256:349ed6091858379ede77f95da294fa7c44e04952a59f61b8c4dad718c6a60e3c

Observation c0b50e32-06a3-45b4-9e96-a10264b63551 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.028544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:01:15.686867Z digest=sha256:ffca7b471c18b33463e67c8cc65f0cf12f6ea29c436244a68f0ed385322a9e94

Observation 1bd3d5dc-95fc-4730-8710-cda898b85a6b · outbound

This paper cites GPT-4o System Card.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling GPT-4o System Card

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.691092Z digest=sha256:968e1fac0512474534947a9d5f526872f137e6ca707a7536794f8d0dd9ee3bb9

Pith citing papers

Observation 8a091469-9a7d-47a4-8830-d8e7bd55df8e · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:13.243457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:13.243457Z digest=sha256:9b2afe8c93aeeecce0e151974d4d841d4846fb3a12c995751bd645a9d8f3f8c6

Observation 52333cf1-07aa-4464-b696-c4a155fd1775 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:32.084497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:31:20.388845Z digest=sha256:8902fb79c1d44558c54da4b2aa79378046108de07697ad20a28be4385ebf0322