Pith. sign in

Paper Citation Record · LEDGER

DocVXQA: Context-Aware Visual Explanations for Document Question Answering

As of 20 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2505.07496.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07496 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:20:55.617728Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 53ccad90-0918-4431-8100-827c215b6a88 · outbound

This paper cites Generic attention-model explainability for interpreting bi-modal and encoder- decoder transformers.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Generic attention-model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.465601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.462731Z digest=sha256:f44caa28afbfb81c5dcb35f8724d97f88b4de8f8b49ec5aab4facb4925eb9e9e

Observation 731f18de-faa4-4d1e-a6da-2c2eb5978ce4 · outbound

This paper cites Boundingdocs: a unified dataset for document ques- tion answering with spatial annotations.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Boundingdocs: a unified dataset for document ques- tion answering with spatial annotations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.482340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.482340Z digest=sha256:8992ecc7590c811c0dbb5e208982de1fb8a079d571bd95e5d32c49687de9cf44

Observation 8c4822b9-8157-4717-b31a-f99ff48284b3 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Perceptual losses for real-time style transfer and super-resolution

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.431440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.491078Z digest=sha256:13195b5de6b12aba698b23e6a55c0fe14ec881942264b2a49265b798e6825444

Observation 228ea04c-f6a4-444a-8d10-0d89d52dcd7d · outbound

This paper cites OCR-free Document Understanding Transformer.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering OCR-free Document Understanding Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.501331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.501331Z digest=sha256:a8ec0854e66dd987bd8401b836c8ef15ad68c7f9eb7417a93fdac43e8228f7d2

Observation 50d69133-2e62-46a1-b872-d4e5c69a1eee · outbound

This paper cites DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.508545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.508545Z digest=sha256:6fe1fc9e6bdafac369bbe82da182c212df608a187dbef658a97902c48b122810

Observation 65b6a538-e5ba-4f6e-92cc-d05b1f813221 · outbound

This paper cites RISE: Randomized Input Sampling for Explanation of Black-box Models.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering RISE: Randomized Input Sampling for Explanation of Black-box Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.514545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.514545Z digest=sha256:96e6a5d70bc57011a0342c0f836357c8c4958827ba2f8b0ea475dbd63dc6ddf2

Observation 216c03c6-f709-4e15-9c70-7802154f98f8 · outbound

This paper cites Restricting the Flow: Information Bottlenecks for Attribution.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Restricting the Flow: Information Bottlenecks for Attribution

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.525865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.525865Z digest=sha256:dc151f1f537bb775d4f37b9e13c1c9f4a3ed9ddf89b67925936ea55d91c1227d

Observation 26054521-2672-4f7a-91c7-c44b71331846 · outbound

This paper cites and Zaslavsky, N.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering and Zaslavsky, N

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.542922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.542922Z digest=sha256:6de73fcb9acd5f1d04ca0f15f6b1e890cdfe525b27c81fab2cf303a96be2b55a

Observation f4497f33-82ce-4b55-bb37-ade4f7b1e42e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.550363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.550363Z digest=sha256:0da6634e096d70c91a9c668f86e0086b5b333883a0986de5e87dc057a916d875

Observation d6c3ca94-d4e3-4fb5-872f-15a01aae2030 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering BloombergGPT: A Large Language Model for Finance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.556647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.556647Z digest=sha256:8fd5e8463a987ce95275c7d28503ac261fb580e00ca861074b29ef9e3e18bd47

Observation df02263a-dd65-45c2-bcc9-69d5f1d4e977 · outbound

This paper cites Ureader: Universal ocr- free visually-situated language understanding with mul- timodal large language model.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Ureader: Universal ocr- free visually-situated language understanding with mul- timodal large language model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.317658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.567697Z digest=sha256:9cc2deb08d5a3e0e680c7eb38e129122a6a7d5d391aad72ef1c4e27aac5a9882

Observation 17f8a9a6-7494-4da2-b652-24b983d1758a · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.572823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.572823Z digest=sha256:da50b5a9412e2344b49e3faaaecd41b0e666b8a1f3a03612115bfc442baeb75f

Observation baae48cf-7a19-480a-85a7-c790be202208 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.581390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.581390Z digest=sha256:510e229432f2158d443f250da61cb4e9dc826fe1a0308c0dacad9b1cb1f4a9f0

Observation cdc4c2b7-73f9-4bd4-9929-64439535314a · outbound

This paper cites Information- bottleneck approach to salient region discovery.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Information- bottleneck approach to salient region discovery

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.295631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.588819Z digest=sha256:311c5afa4d5289ee0b9921110cdba9216bd9a575b595c1a4cf3861395c23f5d8

Observation abf98399-42f0-4e09-9188-fe066685ed1d · outbound

This paper cites What is the estimated budget of ‘conduct analysis of decision makers/ information targets’ in research and development?.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering What is the estimated budget of ‘conduct analysis of decision makers/ information targets’ in research and development?

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.267328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.603675Z digest=sha256:7ff3d32f5c451133db50fec12cd794b29da198d313e21024d1fec0cffcc26e94

Observation 051f3bb3-236e-4a1c-b777-adb9779bee16 · outbound

This paper cites Left: Original Image.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Left: Original Image

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.223831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.610564Z digest=sha256:4dae5334e86eb798fe35a4bf442e1563be0909b4bfb832e09ece4d49914d7211

Observation 83592a09-1df0-40e6-a35d-773142db9c8d · outbound

This paper cites 19 DocVXQA: Context-Aware Visual Explanations for Document Question Answering B.4.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering 19 DocVXQA: Context-Aware Visual Explanations for Document Question Answering B.4

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.196295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.617728Z digest=sha256:9394028758ce50d3fe3e0cb336cbe679a3b6341b420695c2118fd91e7e61eea8

Observation f9055a96-c353-4dc9-a1f1-10a86e13904e · outbound

This paper cites Going full-tilt boogie on document understanding with text-image-layout trans- former.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Going full-tilt boogie on document understanding with text-image-layout trans- former

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.393947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.520540Z digest=sha256:8d64d0bb7d48b26a3a9fcbede6b9d9f38746ca615bd556ba46f7b4c0d16b69dd

Observation beb65caf-fdfe-49b9-b004-237469a89fa0 · outbound

This paper cites GPT-4 Technical Report.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering GPT-4 Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.439460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.439460Z digest=sha256:953fd369ba9b5c2011d0bdb3b3f6bacea39a50529eed5d13c7531f24e3ff58cc

Observation accb5211-bd9b-4273-870a-f2ef7a501548 · outbound

This paper cites DOGR: Towards Versatile Visual Document Grounding and Referring.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.593869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.593869Z digest=sha256:a04b49b261259495482dd0f485b1c2e1a8caf1ee23248247bc2e13b4cb0a575a

Observation c3206f81-efb9-47c3-b7e0-3fcf6fa76e6b · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering ColPali: Efficient Document Retrieval with Vision Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.471606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.471606Z digest=sha256:e1992e5e9065468b09d689165fc1aea1e2779ff42a17ccca00f290121d225ff2

Observation 63963266-9f37-4665-ac08-05d53514afa0 · outbound

This paper cites DUBLIN -- Document Understanding By Language-Image Network.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering DUBLIN -- Document Understanding By Language-Image Network

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.447545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.447545Z digest=sha256:fc7d05ca1f0aa90cbc84a71c3aa5b62217a863e5704a0b9f550bfcdc7d2ea29e

Observation 3d8d1ca1-dfdf-4211-89f8-4c3ed4df3452 · outbound

This paper cites Weighted anisotropic– isotropic total variation for poisson denoising.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering Weighted anisotropic– isotropic total variation for poisson denoising

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.485336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.455074Z digest=sha256:6f9c3e39258aa315ebdcf6ce9f758e025297fd4a5e8af30b9c83965816f514c4

Observation 5d983e6a-5922-463c-8701-8dd7614c4827 · outbound

This paper cites and Flach, P.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering and Flach, P

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:20:56.370426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:20:55.535236Z digest=sha256:752ff0e66000ac7ea42f5178f86575c6019b1375a7b1e49de679cec842ae8cb7

Pith citing papers

No inbound Pith citation observations are available.