Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent Interactive Question Generation Framework for Long Document Understanding

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.20145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20145 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:45:18.546931Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfcc0c4a-968c-4fb8-bb6e-00bee2d33549 · outbound

This paper cites GPT-4 Technical Report.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.393787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.393787Z digest=sha256:986f88cb811cea71430048f8ee812eabaf4a8fc2594e341de47228676ca604bc

Observation e484a1e3-89ae-4362-8052-87efc89c74fa · outbound

This paper cites Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.148011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.400209Z digest=sha256:4cec2fd1890254bc76be72d63da8defac52f370e48e65a0451420650166727cd

Observation 4ebcdc18-ebf5-4418-a64b-d9259acee678 · outbound

This paper cites Claude 3 haiku: Our fastest model yet,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Claude 3 haiku: Our fastest model yet,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.129130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.405297Z digest=sha256:36b86dfb06fe26d1c749ecbeee3ecf1e2ce9ea63e2c20da69e89723f63827336

Observation af5ee380-4c5b-4f88-bb0d-b473a4b9828b · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.410519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.410519Z digest=sha256:b0c4f97f6eda52eda89afb4014131f8809c3076d2547f641ac4d3d404908a6a9

Observation ca2994f1-3793-4b7a-9417-389b93fee3ce · outbound

This paper cites Llava-next: Stronger llms supercharge multi- modal capabilities in the wild,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Llava-next: Stronger llms supercharge multi- modal capabilities in the wild,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.110730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.415715Z digest=sha256:4d467634b74bb5d0ed03eba34e0f2862ee0626a3e76b276348144930b94a3012

Observation 324f86a9-4301-49e3-ae5b-2611e8cd2e4d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.420662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.420662Z digest=sha256:43848310da703d69b833d7de3ae01f08119d60b9871411adb20986f84bc64a88

Observation 5acff4aa-9b05-4dfb-853b-431b69ccceef · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Docvqa: A dataset for vqa on document images,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.093513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.426258Z digest=sha256:9748067db2c9133cd8e771d37639f575fee258c317ed43bd93f30b76c9fa3365

Observation 27280278-a693-4ce7-8ca7-161302e5951f · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.431189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.431189Z digest=sha256:d68d69835f15a831343e898299bd7e497dca14f4bb33bafb129d86ccdc9db3ba

Observation 34bfc216-5520-4bd6-b9e3-8398fb41455f · outbound

This paper cites Infograph- icvqa,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Infograph- icvqa,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.076609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.436284Z digest=sha256:6608e14efd315b5d896527562b25f0785b423338e9ac62cb73b5a3b5e3fe00da

Observation 64f418ca-3151-4c96-81d6-199f9909fefa · outbound

This paper cites Towards com- plex document understanding by discrete reasoning,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Towards com- plex document understanding by discrete reasoning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.060209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.441640Z digest=sha256:adf95ff2dc763c0132af921aed0df16afa41b1fde571599156df8c6f4aa4ae79

Observation 27667108-53ab-483f-9c3c-38df8943549b · outbound

This paper cites Document understand- ing dataset and evaluation (dude),.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Document understand- ing dataset and evaluation (dude),

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.041347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.446195Z digest=sha256:d75039752d57e38b071d24d763c5fd40548b3742df59bf36ae3ef81a4ee10599

Observation b6602d08-cec9-4ae5-8d63-dd8c648ee984 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.451030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.451030Z digest=sha256:f3e13b9fcf34f1d7b8ec7813a4215da6a77563a88637c7fe1ab54873c8ff562c

Observation 63aeeae8-d904-4fca-aa40-385277a2c29e · outbound

This paper cites LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.456133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.456133Z digest=sha256:e44ce93c4c488d0ca920bc7f4cf3b6f176b0a180e647b936a751ecc1335ced39

Observation c931406a-2771-4c36-a718-3cb1e3f6d965 · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.461076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.461076Z digest=sha256:ec1d313834278fb0d81d293853dda71788b2c8622cdec6918d8d8b68a079781c

Observation 0547c844-27b8-40ba-a9c6-d9aa0da3ca25 · outbound

This paper cites CAMEL-Bench: A Comprehensive Arabic LMM Benchmark.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.466100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.466100Z digest=sha256:92fb4fb97f9386ed10a11d076d23d6e6fc0a1ec7c7f7b91be06203605e0bd324

Observation a370b885-cb9b-4f26-a4fc-50598c0175f0 · outbound

This paper cites Doclaynet: A large human-annotated dataset for document-layout segmen- tation,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Doclaynet: A large human-annotated dataset for document-layout segmen- tation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.023515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.472174Z digest=sha256:29f36183a52d767ef370a7c236b608d7daee06ba34b8db6ba911af9a1c13a36d

Observation 654df614-46d7-4c05-b932-5b6438e5dccc · outbound

This paper cites A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:45:18.705002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.477594Z digest=sha256:e4fbc020fd730657d7b2a09fedc2abd99253ba6511f36735bf913145a2113cac

Observation 6c8fb9f5-c11f-45e1-9669-767c8bc8ad37 · outbound

This paper cites Docile benchmark for document information localiza- tion and extraction,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Docile benchmark for document information localiza- tion and extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.005952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.483478Z digest=sha256:70fab95d8154fa5505725187df84c9ff4fb394eeb0bec745fb5868698539b972

Observation 940ef01c-1f18-4173-b91b-80a37e052f3b · outbound

This paper cites Document Visual Question Answering Challenge 2020.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Document Visual Question Answering Challenge 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.488425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.488425Z digest=sha256:620eb69cd084e5605b8bd7952f55d0fca9ac0102d97568c9cd637666c20a0819

Observation 5c6eff25-deef-4376-9f0c-648a7aac0fb1 · outbound

This paper cites No- vachart: A large-scale dataset towards chart understand- ing and generation of multimodal large language mod- els,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding No- vachart: A large-scale dataset towards chart understand- ing and generation of multimodal large language mod- els,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.988127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.493435Z digest=sha256:a8e5c8df0449326c4d50ed2af082ea9ea4237382422bd8bd9d718110d7e9d62a

Observation 91934601-72ae-42b7-92f3-3fbdcd9a8ec5 · outbound

This paper cites An overview of the tesseract ocr engine,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding An overview of the tesseract ocr engine,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.969588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.497787Z digest=sha256:52c99c98cd715d2e4ca1429b56d456426f0f63589424100f2818434a0ad28377

Observation 5fa76249-83fe-4102-9c5f-10abeb0d620f · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.502670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.502670Z digest=sha256:e4f748abfb1826813d672ece94e2e25727956db3477c8adcfd8ae8ad970477e4

Observation fc3a2fda-fdd8-421f-8669-fc293555ae44 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.507641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.507641Z digest=sha256:3684f6bc7dae5ac78a5e9e31ff7afcdf05047318f1c4f1228b8cd8c3d6afba62

Observation efa875fc-b398-4e67-abed-731a7b7b199a · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding BRAVE: Broadening the visual encoding of vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.512376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.512376Z digest=sha256:c9f053a398211cce059dd23ccf41c96b694251134808fbbde426ccee2568ee76

Observation 6fa8425a-0bf8-4c38-9f6e-1875c87d8dca · outbound

This paper cites Automated annotation with generative ai requires validation,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Automated annotation with generative ai requires validation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.950322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.517323Z digest=sha256:c008a4be3c8277c179793994f7609f871ad7b5c62472ac9cea505386f6a48803

Observation c80e8f12-6d94-420a-9b61-0cadd0596306 · outbound

This paper cites Labelvizier: Error profiling and interactive data annotation for long- document understanding,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Labelvizier: Error profiling and interactive data annotation for long- document understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.930246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.522182Z digest=sha256:c3b13e2e3d2b2f800ccffbff875dd987bb17cf05a52f0a57f6cf34e390d85e7f

Observation c1b64abc-2170-483c-9c9b-e05c14af4a66 · outbound

This paper cites Meganno+: A human-llm collaborative an- notation system,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Meganno+: A human-llm collaborative an- notation system,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.911021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.526875Z digest=sha256:0d576a9bdbbbcee5de1c470c6330733d4b0a014ec91633c0d9ad0dff788815e7

Observation 858801ef-4b1e-44e2-8afd-ee4c2fc5e5f9 · outbound

This paper cites pdf2image: A python library to convert pdf pages to images using poppler,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding pdf2image: A python library to convert pdf pages to images using poppler,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.888454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.532153Z digest=sha256:76cac670242d0cc5750fe0d574f29ffc5a2d5a31294e19443f1d625cb4b7ca11

Observation 2ddefe12-f8f7-430b-932b-7aec5aeb1566 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.536837Z digest=sha256:3046834d128b8b202be1414af754ad5855ec2945a9aa8a9e30d69502fe24d86c

Observation 06237e27-d010-4abb-9116-41f832c4d3e2 · outbound

This paper cites Ocean-OCR: Towards General OCR Application via a Vision-Language Model.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.542039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.542039Z digest=sha256:103d78408cdc691e73c7673b7b21bcf9c4dbb713ecabff9185016c98b7365a6d

Observation 977550ae-8cdf-4e78-a8db-8e7ddf9a203e · outbound

This paper cites The yolo framework: A comprehensive re- view of evolution and applications,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding The yolo framework: A comprehensive re- view of evolution and applications,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.867108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T13:45:18.546931Z digest=sha256:76f5daed9ec9389cd595da875cb0106cb2ff723f4514a337de5f59fa47d98ee7

Pith citing papers

No inbound Pith citation observations are available.