Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:54.402782Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.12490.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:54.402782Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0f40e080-bd1d-4602-a9a9-299172ff8252 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedin gs of the IEEE/CVF international conference on computer vision
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49e34a29-2d3c-479a-9316-6c6c8e621f93 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering 4290–4300
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792ca61a-d574-44fc-bbd4-8aa07850e039 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Pattern Recognition Letters 150, 242–249 (2021)
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e4266db-48e4-4e6e-90ed-a8f540d82797 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f7fc36-5dae-4faa-aa87-e576c9420f82 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b80b483-998d-4af9-8f1a-fd14e141652a · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe155ab4-0f54-40be-aba1-65d4068bdd88 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: European Conference on Computer Vision
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2033bde0-f91e-484d-80fe-e51dd8abe144 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63e3231a-10b2-406a-aa7c-06d68f5fc3de · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64518d66-6287-431b-bb90-c0e23b7515cb · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Multimodal Rationales for Explainable Visual Question Answering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43fea144-a5dd-4d2c-8b09-ac79d5b0b30b · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604c8031-6688-4ef7-9f19-cd35ea01d861 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8fcff72-8148-4bbc-8735-301968efb62e · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32de7c0-4fbb-423c-8390-52f65fd8c9d4 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75f70c2c-573a-45c4-aa15-8ff741b25099 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Meila , M., Zhang, T
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f96db652-053a-411e-b546-c39eb88f723c · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b65e2db-c640-4bb8-8cb1-04bb8cdf930f · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR)
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0075ca2-de6a-4548-8926-de07ac41acc3 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering I n: International Confer- ence on Document Analysis and Recognition
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41c2d7c5-b59b-4419-b8d9-31c5b137352a · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2017 IEEE International Conference on Computer Vision (ICC V)
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e6df6a-3035-4fc7-ae12-9fe0bf21b657 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR)
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45eb9f88-24a5-4f21-9c0a-c4f896218a83 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf8540d-5b29-4241-8ae6-4ecb408f8a18 · outbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: The E leventh International Conference on Learning Representations (2022)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.