Pith. sign in

Paper Citation Record · LEDGER

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering

As of 20 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.12490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12490 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:54.402782Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f40e080-bd1d-4602-a9a9-299172ff8252 · outbound

This paper cites In: Proceedin gs of the IEEE/CVF international conference on computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedin gs of the IEEE/CVF international conference on computer vision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.878906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:52.422420Z digest=sha256:c9089f36838c4f86d497702a507fd9cf216b71a98f949d642f40817ad381efa2

Observation 49e34a29-2d3c-479a-9316-6c6c8e621f93 · outbound

This paper cites 4290–4300.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering 4290–4300

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.513301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.513301Z digest=sha256:3a2c80fa9496f5623703219b40f6bbad98dc4a675b1528b2343ab72964f53e64

Observation 792ca61a-d574-44fc-bbd4-8aa07850e039 · outbound

This paper cites Pattern Recognition Letters 150, 242–249 (2021).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Pattern Recognition Letters 150, 242–249 (2021)

Reference 3

Resolution
verified exact
doi, observed 2026-08-06T17:07:54.543867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:52.593231Z digest=sha256:d2dac4f95fc7cd6ebc48e314e95dbea65192301be14d7bf6aefa0065ae4f52e0

Observation 1e4266db-48e4-4e6e-90ed-a8f540d82797 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.666510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.666510Z digest=sha256:db2d879715b80acd23bb9ded0cbedd60e2301238d5e01cafb39ce6f93ffe31c8

Observation 57f7fc36-5dae-4faa-aa87-e576c9420f82 · outbound

This paper cites In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.747544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.747544Z digest=sha256:02665cf81d7c2a914b25ffc7f5afd00036d11da169c7f696c224ada9d3955815

Observation 6b80b483-998d-4af9-8f1a-fd14e141652a · outbound

This paper cites , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.649166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:52.849169Z digest=sha256:922f2b07c1a0f9c740577254ae432e5f50bdb4488b490c8848aad3bb2d121898

Observation fe155ab4-0f54-40be-aba1-65d4068bdd88 · outbound

This paper cites In: European Conference on Computer Vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: European Conference on Computer Vision

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.509448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:52.991227Z digest=sha256:5db9ee06fd01ecc788c9d636ce42fb436233581f74e2664f22e31bae4763b1a6

Observation 2033bde0-f91e-484d-80fe-e51dd8abe144 · outbound

This paper cites In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.395421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.076926Z digest=sha256:7699b223e8e987fd2e0f3b65dfa81d6580338ffd84fe53e268f5e7bc6014bcdd

Observation 63e3231a-10b2-406a-aa7c-06d68f5fc3de · outbound

This paper cites In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.258565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.150218Z digest=sha256:25411ffcbd098439ad0a003ed74c180c25d320bfc7ca36c726159cb568a110c9

Observation 64518d66-6287-431b-bb90-c0e23b7515cb · outbound

This paper cites Multimodal Rationales for Explainable Visual Question Answering.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Multimodal Rationales for Explainable Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:07:54.956540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.224963Z digest=sha256:7cb572638a4e184e2cc819a1994d40d8ac2b941004ab03c726f5623b7c5eefbe

Observation 43fea144-a5dd-4d2c-8b09-ac79d5b0b30b · outbound

This paper cites In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.118590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.317817Z digest=sha256:25a055733d7c6f35db97fe09bdac0583f80a21d2cff9c0a60f930a3846a92041

Observation 604c8031-6688-4ef7-9f19-cd35ea01d861 · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.976919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.417551Z digest=sha256:4c26f063f4bd6a611e74b301d35b734bba5e4dadc44a0a9d97ed25886edfe7cc

Observation b8fcff72-8148-4bbc-8735-301968efb62e · outbound

This paper cites DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.544005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.544005Z digest=sha256:3ef2f4beb6ef9b7ad7ce9405ee96e450ecc8f6d2e0ae3f63c6aa22b06c8e6699

Observation f32de7c0-4fbb-423c-8390-52f65fd8c9d4 · outbound

This paper cites , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.852534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.639104Z digest=sha256:ed71dde06c85190eb653484c0fa7a0b9faf38869937597a1904b06593d6e5558

Observation 75f70c2c-573a-45c4-aa15-8ff741b25099 · outbound

This paper cites In: Meila , M., Zhang, T.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Meila , M., Zhang, T

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.576413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.830975Z digest=sha256:2f79c69ad1a5175910b5adf4eda3fe6dfd09622c3e16a87d2b3915ef26c01d6e

Observation f96db652-053a-411e-b546-c39eb88f723c · outbound

This paper cites an unresolved cited work.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:07:55.747430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:53.707881Z digest=sha256:0b58882eb9a7167529e8375ada31cfbb6c09cbc5918b7429b00b9fff66f1cc62

Observation 6b65e2db-c640-4bb8-8cb1-04bb8cdf930f · outbound

This paper cites In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.924438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.924438Z digest=sha256:54c17dd79aaf419a5b221fba4f8c12e8f3757e43b77939b00998cd1d2ae511c0

Observation e0075ca2-de6a-4548-8926-de07ac41acc3 · outbound

This paper cites I n: International Confer- ence on Document Analysis and Recognition.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering I n: International Confer- ence on Document Analysis and Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.448109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:54.023790Z digest=sha256:352fc2a84e846620f985285b00fee6de62ef870a93e52b3f643ca6688a254e69

Observation 41c2d7c5-b59b-4419-b8d9-31c5b137352a · outbound

This paper cites In: 2017 IEEE International Conference on Computer Vision (ICC V).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2017 IEEE International Conference on Computer Vision (ICC V)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.117543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.117543Z digest=sha256:a627522e8795e85e75ac189df148b783157dea238adc42b38709d7043e0ac644

Observation 60e6df6a-3035-4fc7-ae12-9fe0bf21b657 · outbound

This paper cites In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.246311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.246311Z digest=sha256:9125fda15d5a6464fe3054c38f2294eb62cd295d11fd5d9133860eb6aac3404f

Observation 45eb9f88-24a5-4f21-9c0a-c4f896218a83 · outbound

This paper cites In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.339787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.339787Z digest=sha256:e3700bb197f8044f38b5aed6b53f4bfc11e81e79f841c12eebb12343791c5dab

Observation 1cf8540d-5b29-4241-8ae6-4ecb408f8a18 · outbound

This paper cites In: The E leventh International Conference on Learning Representations (2022).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: The E leventh International Conference on Learning Representations (2022)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.243516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:07:54.402782Z digest=sha256:420076ca1680491d497a404b057787d43bd3d7d88486acbb17bdb75f19b7d791

Pith citing papers

No inbound Pith citation observations are available.