Pith. sign in

Paper Citation Record · LEDGER

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.12490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12490 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:54.402782Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f40e080-bd1d-4602-a9a9-299172ff8252 · outbound

This paper cites In: Proceedin gs of the IEEE/CVF international conference on computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedin gs of the IEEE/CVF international conference on computer vision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.878906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:52.422420Z digest=sha256:61ba75dce686f5f5640295035f8426b20eafb6235b1e5b865732857b709b9982

Observation 49e34a29-2d3c-479a-9316-6c6c8e621f93 · outbound

This paper cites 4290–4300.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering 4290–4300

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.513301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.513301Z digest=sha256:98fcc0a6b32eedbcb2b72845702e0fac5eafcd3c402a0dc3faaf076264684c96

Observation 792ca61a-d574-44fc-bbd4-8aa07850e039 · outbound

This paper cites Pattern Recognition Letters 150, 242–249 (2021).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Pattern Recognition Letters 150, 242–249 (2021)

Reference 3

Resolution
verified exact
doi, observed 2026-08-06T17:07:54.543867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:52.593231Z digest=sha256:233c786b5be470c348bfc14d99c84493b6ad306aad5f1c545a1cd03a3f5e1c18

Observation 1e4266db-48e4-4e6e-90ed-a8f540d82797 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.666510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.666510Z digest=sha256:31eac5212c534a419fd4e2e661a6eb40b8583c30a0a82346954dd4970291cbfa

Observation 57f7fc36-5dae-4faa-aa87-e576c9420f82 · outbound

This paper cites In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.747544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.747544Z digest=sha256:6644e7fdc02f058576750ff271c5c3aefd8b1e749e2d7a00498ec991ef61ab17

Observation 6b80b483-998d-4af9-8f1a-fd14e141652a · outbound

This paper cites , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.649166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:52.849169Z digest=sha256:7ab0b2be15b8be3429bd5e2525860a04ce0ccd1fc4ef3c8746e93adeda415832

Observation fe155ab4-0f54-40be-aba1-65d4068bdd88 · outbound

This paper cites In: European Conference on Computer Vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: European Conference on Computer Vision

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.509448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:52.991227Z digest=sha256:7d0c3380dfbd224cfb52a6d737661399daaf56c93cd4718790c4d90cbf85bcc2

Observation 2033bde0-f91e-484d-80fe-e51dd8abe144 · outbound

This paper cites In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.395421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.076926Z digest=sha256:ac3384d45f1af473f5c77e419461557c127623279df1178b172b94a446abaa6d

Observation 63e3231a-10b2-406a-aa7c-06d68f5fc3de · outbound

This paper cites In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.258565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.150218Z digest=sha256:415769f534554e95b1ff564309f655ccc611902186faf62f785be1786e5a96da

Observation 64518d66-6287-431b-bb90-c0e23b7515cb · outbound

This paper cites Multimodal Rationales for Explainable Visual Question Answering.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Multimodal Rationales for Explainable Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:07:54.956540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.224963Z digest=sha256:17972ae60bbfbdf861d3a4b78f4faf4575f9e6cbc03ff973b1d0d29feee110db

Observation 43fea144-a5dd-4d2c-8b09-ac79d5b0b30b · outbound

This paper cites In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.118590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.317817Z digest=sha256:11be6a56350c5eb1b34df13862e70937e1fb135aba2e1c6209e4855cf61303d6

Observation 604c8031-6688-4ef7-9f19-cd35ea01d861 · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.976919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.417551Z digest=sha256:8ceca8a3bbaf054bceb5ea5ac0664ee45a1c78e11595507da1bc1807f6d70b1e

Observation b8fcff72-8148-4bbc-8735-301968efb62e · outbound

This paper cites DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.544005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.544005Z digest=sha256:49d6449afcba30ba8ce76901fd24dd05b2e56d80d411fe702dbac2fe3253e9de

Observation f32de7c0-4fbb-423c-8390-52f65fd8c9d4 · outbound

This paper cites , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.852534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.639104Z digest=sha256:936b9cb095ce953494b50c0cc5347a7d42151e58aab4e4327ad868f73cafaca8

Observation 75f70c2c-573a-45c4-aa15-8ff741b25099 · outbound

This paper cites In: Meila , M., Zhang, T.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Meila , M., Zhang, T

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.576413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.830975Z digest=sha256:09e7754be954f10d991700071a66af95abf7bbfe00e920c6c89f526a06a82325

Observation f96db652-053a-411e-b546-c39eb88f723c · outbound

This paper cites an unresolved cited work.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:07:55.747430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:53.707881Z digest=sha256:961c2c580e6913ff1d4adbab45fa509b61918bd57e3b9ac4b4bb9c393cb450b6

Observation 6b65e2db-c640-4bb8-8cb1-04bb8cdf930f · outbound

This paper cites In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.924438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.924438Z digest=sha256:ad205732578eaea4c8d581d021e3e6d19464368fc1c349b9aeac51cbcbd58505

Observation e0075ca2-de6a-4548-8926-de07ac41acc3 · outbound

This paper cites I n: International Confer- ence on Document Analysis and Recognition.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering I n: International Confer- ence on Document Analysis and Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.448109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:54.023790Z digest=sha256:c57f03d85d84aea2942ccac99436598c9b529f777c44aee6ba1a82ec73430c41

Observation 41c2d7c5-b59b-4419-b8d9-31c5b137352a · outbound

This paper cites In: 2017 IEEE International Conference on Computer Vision (ICC V).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2017 IEEE International Conference on Computer Vision (ICC V)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.117543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.117543Z digest=sha256:4bf7b4ab4ee25e34bf87f92993d5346e1eac2e633a892f95ef771ec2c4e33676

Observation 60e6df6a-3035-4fc7-ae12-9fe0bf21b657 · outbound

This paper cites In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.246311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.246311Z digest=sha256:3d164e03fdfd1852728be462cde3fbc0a41966cda9d8df680fa4ff35e646ca6d

Observation 45eb9f88-24a5-4f21-9c0a-c4f896218a83 · outbound

This paper cites In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.339787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.339787Z digest=sha256:d5b2874c5623516035dfef88d67f652d94295396dad5a883e88c31a626104776

Observation 1cf8540d-5b29-4241-8ae6-4ecb408f8a18 · outbound

This paper cites In: The E leventh International Conference on Learning Representations (2022).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: The E leventh International Conference on Learning Representations (2022)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.243516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:07:54.402782Z digest=sha256:beac713dea241dc15616a72d383b0494a9bc1c42c87ff180663426e0cdb0d2e3

Pith citing papers

No inbound Pith citation observations are available.