Pith. sign in

Paper Citation Record · LEDGER

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2506.18985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18985 v3

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:03.748195Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:59:19.379119Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9368b046-2675-46dd-b26f-c9dd94caf847 · outbound

This paper cites Quantifying attention flow in transformers.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Quantifying attention flow in transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.117975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.639483Z digest=sha256:c4016fc390148a19d2a2f55d06bc7d4738d74dd6acb696adef491b40d285285c

Observation dad3ad63-80a2-480c-a477-43c140ab9e35 · outbound

This paper cites AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.643989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.643989Z digest=sha256:55f651158f15966c68d400daba370e06ba8e68c59e14d17992f6f0866e188c28

Observation bfce766f-b504-4d1e-ac24-49f90eba467d · outbound

This paper cites XAI for trans- formers: better explanations through conservative propaga- tion.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models XAI for trans- formers: better explanations through conservative propaga- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.105967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.648241Z digest=sha256:da50ccc16b1f1c766bd1c137fa475be1e62e3341debea7e315bf1a4a575fc7ab

Observation c302c794-ca24-4bc0-8f2d-7ee72c1b7db3 · outbound

This paper cites On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.093686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.652167Z digest=sha256:9eea07abba1f1f5dc7524c3ec7b34e7c93a21e6be6b916e928e55da912b7ba85

Observation cc955556-c5c7-42e0-b57a-da82e4a71a53 · outbound

This paper cites Qwen2.5-VL Technical Report.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.656596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.656596Z digest=sha256:bd7be51582234029c119b0bc8f4fc890cf7679a4fa5815ecd1b7d427cc74c293

Observation f17b3e68-7802-453c-93ee-23e3b4df9b09 · outbound

This paper cites an unresolved cited work.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:46:04.082390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.660515Z digest=sha256:0555dcad45ce4734e8325afe9661dbe255d55dc15ee380e385a5442c625bb0e4

Observation b3a63cf5-3d53-471f-bc0b-0ce2f469bbae · outbound

This paper cites Visual Explanations via Iterated Integrated Attributions.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Visual Explanations via Iterated Integrated Attributions

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:46:03.908685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.664989Z digest=sha256:74189f2dba291c1b5377eab70b1d16bd3071351f79f5c2419941334f9505a71f

Observation 1eedf591-3e3b-409f-b668-3ac954c14642 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.071386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.669555Z digest=sha256:a5957a2836ed1ef8f44c923e72be59c1727c40380176f97b34a9bedb1a6dbbbf

Observation 8a60b89c-3f04-4ab7-bed2-be1f59881448 · outbound

This paper cites Human attention in visual question answer- ing: do humans and deep networks look at the same regions? In Proc.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Human attention in visual question answer- ing: do humans and deep networks look at the same regions? In Proc

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.060196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.673042Z digest=sha256:c28c6c3baf88e2ddc2910c8842a77dfd2f00bcafa821c2eeed617a0028e0c192

Observation 3698a970-62c3-4bf5-bbd6-1a3a5a009400 · outbound

This paper cites AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.676556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.676556Z digest=sha256:b367ee883b3d20de8f4f481024dbdc741ed7528a208de9618db725e102453e65

Observation 3b3c9461-e966-43be-ac55-28178f7d9727 · outbound

This paper cites Turtles, Hats and Spectres: Aperiodic structures on a Rhombic tiling.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Turtles, Hats and Spectres: Aperiodic structures on a Rhombic tiling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.680944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.680944Z digest=sha256:fe6a90ed256b718bb5b70279c873c694e0f14be4d6efffd755c5b2af7f515799

Observation 49838183-700a-4569-8e91-10c49ea97e7d · outbound

This paper cites iGOS++: inte- grated gradient optimized saliency by bilateral perturbations.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models iGOS++: inte- grated gradient optimized saliency by bilateral perturbations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.048896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.684905Z digest=sha256:346b73785a22c2c3dd5c742875cdc4b50a12cf533cbab63c6a2e09317876c1c9

Observation 4948d703-030d-414c-8f24-d2cee8ca793c · outbound

This paper cites From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.037962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.688639Z digest=sha256:49016291c36417bb61d067d80f2fb5f63601498e734156a12bf5c5367eee1a01

Observation 0925e4f2-94ce-49c2-a15e-37135eb8dbfb · outbound

This paper cites Visual Instruction Tuning.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.692170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.692170Z digest=sha256:bd4d082c4c935f20d33f9c20a6e0c5d84aacf80a241c1dd83f0df3e9d23df637

Observation b30bf151-85f1-44f2-9379-fe6a621791d2 · outbound

This paper cites Lundberg and Su-In Lee.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Lundberg and Su-In Lee

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.026785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.696219Z digest=sha256:e1d555248a930d7daeb20a1183f553f42f81166589919a12897067732692437d

Observation 6f305ed1-19a6-4792-aff2-1ffc1df838cc · outbound

This paper cites Mortality Forecasting using Variational Inference.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Mortality Forecasting using Variational Inference

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:46:03.860984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.700734Z digest=sha256:11396b59f7837229fad4a40c143395b9686a895fb2ea22165a7be442b132afd5

Observation 8625d7cc-78aa-4e18-9c69-5d185cb40c03 · outbound

This paper cites Exploring human-like attention supervision in visual question answer- ing.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Exploring human-like attention supervision in visual question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.015322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.705261Z digest=sha256:e2bbf1d9c83cb6e7516b799566f5a077ca4198a7c13879e02ebc59c2e13d3ea5

Observation 8e4f297d-0d47-4626-af6e-e971c00aff3c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.708856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.708856Z digest=sha256:5ae55a300310ff52281b977d09479087e714a2a94d975b735e16d69c3ca27c54

Observation 85c84dc7-508d-4685-8e72-1e1ea0405777 · outbound

This paper cites Uncertainty estimates for semantic segmentation: providing enhanced reliability for automated motor claims handling.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Uncertainty estimates for semantic segmentation: providing enhanced reliability for automated motor claims handling

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:46:03.833662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.712921Z digest=sha256:9df19728c4759614d95bfe7212c6cc478c2d317f874163496451320a964ff3e9

Observation 4810cd64-9b2c-47a6-a56f-ca0ebda08dab · outbound

This paper cites Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.002676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.716980Z digest=sha256:374580df8ff0dfe7a349c1c5dd7d355f724fb3e3f80d0baf8f12fe5f8978f996

Observation ab0d3a7d-9336-465e-83f1-0972c0f63b06 · outbound

This paper cites Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.989684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.722298Z digest=sha256:68ff118164fc733d6d8f3f272ffebdaed20c402ad862e6812887978b527e2abf

Observation 13db61c4-7aed-41f1-821b-af02a44434df · outbound

This paper cites Deep inside convolutional networks: visualising image clas- sification models and saliency maps.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Deep inside convolutional networks: visualising image clas- sification models and saliency maps

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.976671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.725942Z digest=sha256:c277b061e72a39ac294899eeb8e96a36749c72769ebf4ce759d6627fccf6277b

Observation d3968f7e-926d-4681-8dc7-aa014e420dcc · outbound

This paper cites On numerical solutions of the time-dependent Schr\"odinger equation.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models On numerical solutions of the time-dependent Schr\"odinger equation

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:46:03.814922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.729572Z digest=sha256:ed25f3b85c95f917f8e3b499acbaf8a9f0e7b5c8dc8082c4b1818d18c07f7f18

Observation 01581318-b063-4710-9727-44e317e01030 · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.733512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.733512Z digest=sha256:bd1893d9dfa9ba95851799240be7334673ef80d0712ac132bca4fc3765535029

Observation ef6f40f4-bdf6-47ab-8fde-1911bff2fc9b · outbound

This paper cites Axiomatic attribution for deep networks.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Axiomatic attribution for deep networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.964230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.737200Z digest=sha256:e928d39195ca100f0955d76d03f810ea0e3401adc292d30420da5c2022061a19

Observation 5c4d78c7-4ea7-4683-866c-3774b2dd847d · outbound

This paper cites Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:46:03.785591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.740939Z digest=sha256:c6cca05a5a078c505c9a3d0c2e839e8a319bedce95c134ed452f2fe0cf91c5d9

Observation 9758877d-3a24-40fc-b4b4-0f5d28450645 · outbound

This paper cites VQA-MHUG: human gaze supervision for visual question answering.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models VQA-MHUG: human gaze supervision for visual question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.952722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.744697Z digest=sha256:feb88214c378e6ba5515c0849b740c6418bf193342f0e7544d9e75276dcfebb2

Observation b27832b6-71d2-45dd-9a4b-b6ff578538dc · outbound

This paper cites What if the tv was off? examining counterfactual reasoning abilities of vision-language models.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models What if the tv was off? examining counterfactual reasoning abilities of vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.940691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T18:46:03.748195Z digest=sha256:58df6c3cf66980087e360e467ce18f0190cb9d401c976eeec506ad5142fcba54

Pith citing papers

Observation 2bfed174-41f7-409f-974f-a06262405880 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.749240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:9d39e1748475a0784659a31ef3f4a0d9aa8c7168ca22f3332c3d575e1220fe63

Observation 79adc347-769b-4532-bf89-f49ad36076e7 · inbound

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation cites this paper.

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:35:47.332353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:21:47.577839Z digest=sha256:692a913b81d45bfd36491b99401bc392ac13eb37b520f6680880f258a03c0982