Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent Interactive Question Generation Framework for Long Document Understanding

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.20145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20145 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:45:18.546931Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfcc0c4a-968c-4fb8-bb6e-00bee2d33549 · outbound

This paper cites GPT-4 Technical Report.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.393787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.393787Z digest=sha256:a5736e75e97b2191a3c5d1a0fbadcc1d25ddfea6b5403db91090cdee602ad8a6

Observation e484a1e3-89ae-4362-8052-87efc89c74fa · outbound

This paper cites Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.148011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.400209Z digest=sha256:f90847a03a54f0945d220dc4a46594bc8a878468c1811fbaf4199b473a651f12

Observation 4ebcdc18-ebf5-4418-a64b-d9259acee678 · outbound

This paper cites Claude 3 haiku: Our fastest model yet,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Claude 3 haiku: Our fastest model yet,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.129130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.405297Z digest=sha256:5b95ccad5bd2113e92d3f0104cc7c1cbc36d769f74e79c3a7db37f928ad216f8

Observation af5ee380-4c5b-4f88-bb0d-b473a4b9828b · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.410519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.410519Z digest=sha256:b0c4f97f6eda52eda89afb4014131f8809c3076d2547f641ac4d3d404908a6a9

Observation ca2994f1-3793-4b7a-9417-389b93fee3ce · outbound

This paper cites Llava-next: Stronger llms supercharge multi- modal capabilities in the wild,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Llava-next: Stronger llms supercharge multi- modal capabilities in the wild,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.110730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.415715Z digest=sha256:cf5f0f612bbcef9d7c15b576dbde6deade0b0c98eaa983915916a1c1bef20224

Observation 324f86a9-4301-49e3-ae5b-2611e8cd2e4d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.420662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.420662Z digest=sha256:43848310da703d69b833d7de3ae01f08119d60b9871411adb20986f84bc64a88

Observation 5acff4aa-9b05-4dfb-853b-431b69ccceef · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Docvqa: A dataset for vqa on document images,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.093513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.426258Z digest=sha256:4fa5d6aeade180122f92acbd2bb64688412ab697a7803924e6bd36936f8962f4

Observation 27280278-a693-4ce7-8ca7-161302e5951f · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.431189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.431189Z digest=sha256:d68d69835f15a831343e898299bd7e497dca14f4bb33bafb129d86ccdc9db3ba

Observation 34bfc216-5520-4bd6-b9e3-8398fb41455f · outbound

This paper cites Infograph- icvqa,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Infograph- icvqa,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.076609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.436284Z digest=sha256:585b8fa67274b5274f372a65e3ad6be3d5dc972a665367f27386e65bc722ef46

Observation 64f418ca-3151-4c96-81d6-199f9909fefa · outbound

This paper cites Towards com- plex document understanding by discrete reasoning,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Towards com- plex document understanding by discrete reasoning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.060209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.441640Z digest=sha256:7ec3048829f6e6c9e51f448b027632e0e68e2c814d8be037316890ebd815edf0

Observation 27667108-53ab-483f-9c3c-38df8943549b · outbound

This paper cites Document understand- ing dataset and evaluation (dude),.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Document understand- ing dataset and evaluation (dude),

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.041347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.446195Z digest=sha256:f05ae030fb741f83c651919cd11052563fe9f1653da81016744bf4807ef71766

Observation b6602d08-cec9-4ae5-8d63-dd8c648ee984 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.451030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.451030Z digest=sha256:f3e13b9fcf34f1d7b8ec7813a4215da6a77563a88637c7fe1ab54873c8ff562c

Observation 63aeeae8-d904-4fca-aa40-385277a2c29e · outbound

This paper cites LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.456133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.456133Z digest=sha256:e44ce93c4c488d0ca920bc7f4cf3b6f176b0a180e647b936a751ecc1335ced39

Observation c931406a-2771-4c36-a718-3cb1e3f6d965 · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.461076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.461076Z digest=sha256:ec1d313834278fb0d81d293853dda71788b2c8622cdec6918d8d8b68a079781c

Observation 0547c844-27b8-40ba-a9c6-d9aa0da3ca25 · outbound

This paper cites CAMEL-Bench: A Comprehensive Arabic LMM Benchmark.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.466100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.466100Z digest=sha256:92fb4fb97f9386ed10a11d076d23d6e6fc0a1ec7c7f7b91be06203605e0bd324

Observation a370b885-cb9b-4f26-a4fc-50598c0175f0 · outbound

This paper cites Doclaynet: A large human-annotated dataset for document-layout segmen- tation,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Doclaynet: A large human-annotated dataset for document-layout segmen- tation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.023515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.472174Z digest=sha256:5cb55f3c2902c4f4429f77b66d549b433ac0587557901e2fe6d82f6e44693d72

Observation 654df614-46d7-4c05-b932-5b6438e5dccc · outbound

This paper cites A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:45:18.705002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.477594Z digest=sha256:192ddf9b2a3fbbd060f627fee148f695c97092d6deae71996342e483f8b658b4

Observation 6c8fb9f5-c11f-45e1-9669-767c8bc8ad37 · outbound

This paper cites Docile benchmark for document information localiza- tion and extraction,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Docile benchmark for document information localiza- tion and extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.005952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.483478Z digest=sha256:4d2df2038ad84bc19b67f15e06bba3bdfe39f799d6b870d734b33a230a6a5ca2

Observation 940ef01c-1f18-4173-b91b-80a37e052f3b · outbound

This paper cites Document Visual Question Answering Challenge 2020.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Document Visual Question Answering Challenge 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.488425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.488425Z digest=sha256:620eb69cd084e5605b8bd7952f55d0fca9ac0102d97568c9cd637666c20a0819

Observation 5c6eff25-deef-4376-9f0c-648a7aac0fb1 · outbound

This paper cites No- vachart: A large-scale dataset towards chart understand- ing and generation of multimodal large language mod- els,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding No- vachart: A large-scale dataset towards chart understand- ing and generation of multimodal large language mod- els,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.988127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.493435Z digest=sha256:c981404dc5dbc00a81914b8a7f6ab4babd5ad927bd30dc6c9fc444ff98abf2c9

Observation 91934601-72ae-42b7-92f3-3fbdcd9a8ec5 · outbound

This paper cites An overview of the tesseract ocr engine,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding An overview of the tesseract ocr engine,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.969588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.497787Z digest=sha256:597717f1c1e2e1d1c218bc3937a6a048858a4867a16ac6494f9324bba2706888

Observation 5fa76249-83fe-4102-9c5f-10abeb0d620f · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.502670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.502670Z digest=sha256:e4f748abfb1826813d672ece94e2e25727956db3477c8adcfd8ae8ad970477e4

Observation fc3a2fda-fdd8-421f-8669-fc293555ae44 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.507641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.507641Z digest=sha256:3684f6bc7dae5ac78a5e9e31ff7afcdf05047318f1c4f1228b8cd8c3d6afba62

Observation efa875fc-b398-4e67-abed-731a7b7b199a · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding BRAVE: Broadening the visual encoding of vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.512376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.512376Z digest=sha256:c9f053a398211cce059dd23ccf41c96b694251134808fbbde426ccee2568ee76

Observation 6fa8425a-0bf8-4c38-9f6e-1875c87d8dca · outbound

This paper cites Automated annotation with generative ai requires validation,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Automated annotation with generative ai requires validation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.950322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.517323Z digest=sha256:0ad86ebc7d298506d65f5336ec1e38a6256487e4392a75780030c588dc6d9cf0

Observation c80e8f12-6d94-420a-9b61-0cadd0596306 · outbound

This paper cites Labelvizier: Error profiling and interactive data annotation for long- document understanding,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Labelvizier: Error profiling and interactive data annotation for long- document understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.930246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.522182Z digest=sha256:ba66a532dc8b5319415906ff0e6eb58242d017c83c857bc2c88a76012ba5b981

Observation c1b64abc-2170-483c-9c9b-e05c14af4a66 · outbound

This paper cites Meganno+: A human-llm collaborative an- notation system,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Meganno+: A human-llm collaborative an- notation system,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.911021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.526875Z digest=sha256:a287668a72c17877c84e86063494179226335060be396cc0b8144415b2753f4d

Observation 858801ef-4b1e-44e2-8afd-ee4c2fc5e5f9 · outbound

This paper cites pdf2image: A python library to convert pdf pages to images using poppler,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding pdf2image: A python library to convert pdf pages to images using poppler,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.888454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.532153Z digest=sha256:f8299a18bba6df1a0ec6fe9c0ccbaa7dde74fe24a4f0940a057ada4b020f4871

Observation 2ddefe12-f8f7-430b-932b-7aec5aeb1566 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.536837Z digest=sha256:3046834d128b8b202be1414af754ad5855ec2945a9aa8a9e30d69502fe24d86c

Observation 06237e27-d010-4abb-9116-41f832c4d3e2 · outbound

This paper cites Ocean-OCR: Towards General OCR Application via a Vision-Language Model.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.542039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.542039Z digest=sha256:103d78408cdc691e73c7673b7b21bcf9c4dbb713ecabff9185016c98b7365a6d

Observation 977550ae-8cdf-4e78-a8db-8e7ddf9a203e · outbound

This paper cites The yolo framework: A comprehensive re- view of evolution and applications,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding The yolo framework: A comprehensive re- view of evolution and applications,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.867108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:45:18.546931Z digest=sha256:483855822e77444f287f82ebcd9b1b280d5df5fc5eddc43b5567bfc95aa5840a

Pith citing papers

No inbound Pith citation observations are available.