Pith. sign in

Paper Citation Record · LEDGER

Token Activation Map to Visually Explain Multimodal LLMs

As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2506.23270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23270 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:29.407583Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T01:45:52.034775Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4de39f2-8367-4a38-bd53-b02dca29f34c · outbound

This paper cites Quantifying attention flow in transformers.

Token Activation Map to Visually Explain Multimodal LLMs Quantifying attention flow in transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:22.828224Z digest=sha256:5ffb0b1084d89da5948a8d5a6f7fec2724e1bd4ae7f314734c66fe5c95860222

Observation ebc98010-930a-4a15-834f-258ab58ade04 · outbound

This paper cites Attnlrp: Attention- aware layer-wise relevance propagation for transformers.

Token Activation Map to Visually Explain Multimodal LLMs Attnlrp: Attention- aware layer-wise relevance propagation for transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.029975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:22.908507Z digest=sha256:1bd7a1d96d01c8d9d129678be55109eee2b6dc8393af2953e085175b161d9a97

Observation bffa45d5-158c-4fba-819a-1a9fed0a37c3 · outbound

This paper cites Vl-interpret: An interactive visualization tool for interpreting vision-language transformers.

Token Activation Map to Visually Explain Multimodal LLMs Vl-interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.481689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.138905Z digest=sha256:6b6ab98874ffb121dfee20ac4bbcc4b565d439d66a31353f4abddc04c9201cec

Observation 021db75b-a6aa-45f6-8a2f-df56eea62a9b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Token Activation Map to Visually Explain Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.200165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.200165Z digest=sha256:3f1c90c716322eb4c64bc01207f63199e59cf96dd94b75242290e28b01c6e6aa

Observation e28f49e7-adc0-47b5-b63b-7b1ae3ee0216 · outbound

This paper cites Xai for trans- formers: Better explanations through conservative propa- gation.

Token Activation Map to Visually Explain Multimodal LLMs Xai for trans- formers: Better explanations through conservative propa- gation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.145550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.291980Z digest=sha256:1edfed72457d7ef64ca6e51648ae53045912fb9ee0fb74a065d0d49850128d57

Observation 818ccc59-2c78-4235-9a49-8d9538a00479 · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Token Activation Map to Visually Explain Multimodal LLMs Text2live: Text-driven layered image and video editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.802003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.383073Z digest=sha256:e10bc88d336c3c5961567dfe8284c9a854b9acf81b2bf15efbc5ab9dd27471c4

Observation 3aa7372b-2964-43cc-8255-0946c3531fbe · outbound

This paper cites Lvlm-intrepret: An interpretability tool for large vision-language models.

Token Activation Map to Visually Explain Multimodal LLMs Lvlm-intrepret: An interpretability tool for large vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.419141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.464253Z digest=sha256:a786fd3257a962ce5a4f367714beb91a826c3b8c74657d2a7c7cb2c8c989ea72

Observation 56241116-b3c1-4dd0-826a-56f3dc54d316 · outbound

This paper cites An adaptive median filter for image denoising.

Token Activation Map to Visually Explain Multimodal LLMs An adaptive median filter for image denoising

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.138362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.579543Z digest=sha256:28e6e1537173eac706bad47cb24f9b72556d95066a3890e67b769714c044238a

Observation e807e267-b14c-43a4-9570-7995aaa5a0f0 · outbound

This paper cites Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673939Z digest=sha256:01bb27836a553c67961740f17c39d51afccdb9cf039afd3f5f4e664fdb0a72ee

Observation 387ed1b8-79d1-4255-ad4f-234b5fe8ea23 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Token Activation Map to Visually Explain Multimodal LLMs Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743484Z digest=sha256:704ddeefe0164b3dd775428edd57cae0139f899cb7087eed59e56df20601ac59

Observation f3551c4f-decd-459b-9d8c-447cd54b6c89 · outbound

This paper cites Transformer inter- pretability beyond attention visualization.

Token Activation Map to Visually Explain Multimodal LLMs Transformer inter- pretability beyond attention visualization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.945210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.851689Z digest=sha256:6af0ffabedd3c69a80f78a5fbd0a49d701de3817f910bb3c5ebc749a3c1a338f

Observation b092b981-a4fe-4711-b9ed-7c5fca5c3b20 · outbound

This paper cites Less is more: Fewer interpretable region via submodular subset selection.

Token Activation Map to Visually Explain Multimodal LLMs Less is more: Fewer interpretable region via submodular subset selection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.796077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.942557Z digest=sha256:bf01b2e07d4c7cd5404abf76bacc05bef083c6d081d350cff6fdeda648ad6ebb

Observation acc9bf12-e418-4552-8424-9ffd90af11f8 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.065434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.065434Z digest=sha256:7f3abb54de28d354a3e425ffeace54d5f7d9ed10192e5dc9ea505c606db5bfb8

Observation afb0bd74-59d4-4c8c-af1c-87319c0c7084 · outbound

This paper cites Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection.

Token Activation Map to Visually Explain Multimodal LLMs Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.619683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.182522Z digest=sha256:3d8c9477484b4aaf7b5c8b0fb0ecfb1a6ee606cb3a2c6a8caa81116aacf22a44

Observation 60315eb3-3ff5-45aa-adee-6f7db657e89f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Token Activation Map to Visually Explain Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.270466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.270466Z digest=sha256:4a76c548368c3e4b9b8b0c38ecd7303daaab7dd15b8efa90dbd2e81e98ba4704

Observation 6b8b3449-02d5-490e-b819-49d285b05ae2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Token Activation Map to Visually Explain Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.447607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.355777Z digest=sha256:2b66e0c69182a8f597e781576f01c9a6f30f914fb05c77dc35c54824d45cb1b7

Observation d83c952b-6c0e-4d7f-9b64-d0d02c398361 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Token Activation Map to Visually Explain Multimodal LLMs A survey on multimodal large lan- guage models for autonomous driving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.265238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.418757Z digest=sha256:93278f13218e890af4bf0917158224da6480334e7bc62b572807aac63aee9bb9

Observation 55e032c7-68f7-43e4-83e8-000f889c35f2 · outbound

This paper cites Flashattention: Fast and memory-efficient exact at- tention with io-awareness.

Token Activation Map to Visually Explain Multimodal LLMs Flashattention: Fast and memory-efficient exact at- tention with io-awareness

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.123765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.485401Z digest=sha256:e7ebd7b5d98233373aeec7f9a5f1268cf6301fc7bc2a22d4f31f71dcfc0cfc16

Observation 562b6886-6ac5-4d32-9905-332e33589d3e · outbound

This paper cites Vision transformers need registers.

Token Activation Map to Visually Explain Multimodal LLMs Vision transformers need registers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.001289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.557594Z digest=sha256:7f7d42a16184e5bfda8b11619633f231268c4b8f1c31188f9ac7b79bd3889c7b

Observation 9e0a0169-afc6-4f6d-9d91-4bc9b4618612 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Token Activation Map to Visually Explain Multimodal LLMs Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.628249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.628249Z digest=sha256:2094d7dad6de1fbae0b8d59f11a443acf722d8f6e199ad62c0e6111a94b5df83

Observation e29031a5-45de-4ba7-a704-248d703c7195 · outbound

This paper cites Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis.

Token Activation Map to Visually Explain Multimodal LLMs Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.848320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.672164Z digest=sha256:30fea52029c35afd2e852e0a00e495f78c49bd3557487c3cb3cb018d7fe0ce7a

Observation 30eddce3-f46c-4af8-93d8-03533192b924 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models.

Token Activation Map to Visually Explain Multimodal LLMs Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.696589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.721249Z digest=sha256:3cdb2fde2a7689d5a757554832b1c406263fabd292a7206d0092de2e08ba33ba

Observation 06b09a51-da1c-4187-8bbb-77138383df40 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Token Activation Map to Visually Explain Multimodal LLMs An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.560451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.800458Z digest=sha256:7cc77f285241babcee1ea7e051f12a12e5eafaf733351121d5afd44378d66465

Observation 0c3f78d9-6d9e-4956-ab1c-2079b9b30237 · outbound

This paper cites Deep residual learning for image recognition.

Token Activation Map to Visually Explain Multimodal LLMs Deep residual learning for image recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.398468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.853813Z digest=sha256:1a7f5a221516781ff02e493ea60cd519947d7da8f6b0240362a80ed18989f470

Observation fceaabdb-ae1f-4e8c-a0e0-329b2a3b56d9 · outbound

This paper cites GPT-4o System Card.

Token Activation Map to Visually Explain Multimodal LLMs GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893909Z digest=sha256:679e97d2efcd18c344825e8477465a9f00635be64ca5d12082594a9447c6fc6c

Observation f3350316-b32e-4812-8ffb-e3e32ae89606 · outbound

This paper cites Layercam: Exploring hierarchical class activation maps for localization.

Token Activation Map to Visually Explain Multimodal LLMs Layercam: Exploring hierarchical class activation maps for localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.237320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:24.962731Z digest=sha256:043c85c5931c894bf81f0e5cb7257693eec6e3a9e594b093d208e833ff4fe97e

Observation 138f9eb9-5bc3-4bc4-98d9-eaf078fd15b0 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

Token Activation Map to Visually Explain Multimodal LLMs Causal inference meets deep learning: A compre- hensive survey

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.053887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.019347Z digest=sha256:16859d8c71130ff03733ab94a7fd7b0fb2c7cb6931a5119f3d5fde0f954bbed7

Observation 755369e1-7dd3-4e7d-ae27-6c9169819545 · outbound

This paper cites Unmasking clever hans predictors and as- sessing what machines really learn.

Token Activation Map to Visually Explain Multimodal LLMs Unmasking clever hans predictors and as- sessing what machines really learn

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.884300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.064591Z digest=sha256:0083a1c0299741704ae4014bfefa0626ea9e83883a1322552e47ba0edf8c3a53

Observation f96de966-94a9-40fb-a609-1d9f5bb8c4c6 · outbound

This paper cites Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education.

Token Activation Map to Visually Explain Multimodal LLMs Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.688437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.118755Z digest=sha256:20e00450c443059da3c1bd1a7122612cc27ebc943f83328a99976f95723e3298

Observation 2da65a26-37a5-417a-ba21-7ff09e4f63ff · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation.

Token Activation Map to Visually Explain Multimodal LLMs Manipllm: Embodied multimodal large language model for object-centric robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.487383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.203550Z digest=sha256:1007c07577d3d43c44a25cd37e98d7c3de2d9c634753083407fb6892a9e069cf

Observation 591122c2-d792-48f2-9ef0-af9ae5158136 · outbound

This paper cites Exploring Visual Interpretability for Contrastive Language-Image Pre-training.

Token Activation Map to Visually Explain Multimodal LLMs Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.254347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.254347Z digest=sha256:4d7ed69763f7dae305897f5eddd4d123afa1b78d692b9c82f2d0e44d8912892a

Observation 67b51a60-2297-48fb-a314-4198d1beea9a · outbound

This paper cites A closer look at the explainability of con- trastive language-image pre-training.

Token Activation Map to Visually Explain Multimodal LLMs A closer look at the explainability of con- trastive language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.338754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.315356Z digest=sha256:ce4b6f6483ee2548a6f357d252bdc6835331c13e85457d4fe5ef1cbc9c347e29

Observation 0889dc48-50f3-439e-8bd1-2a2762db6951 · outbound

This paper cites Microsoft coco: Common objects in context.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.381037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.381037Z digest=sha256:e69ed9e4d17f031d929b88032aef43f89b85520d4c67d987eb37ede1a8838f2c

Observation 8732b887-ceae-4a94-a807-544e0a93460e · outbound

This paper cites A medical multimodal large language model for future pandemics.

Token Activation Map to Visually Explain Multimodal LLMs A medical multimodal large language model for future pandemics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.156408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.446880Z digest=sha256:d356368e687c354a817cf37ed760a85cef0c91390babfbe524ef9204113068e7

Observation 36fd000f-a697-4568-b524-23d0242845e7 · outbound

This paper cites Visual instruction tuning.

Token Activation Map to Visually Explain Multimodal LLMs Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.016402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.514680Z digest=sha256:09fd546f9232264b01119ac84073dab0150ca00253bc19cf32058bed503852ad

Observation 3ed8483e-c879-4098-a0e1-200d1baa53ad · outbound

This paper cites A unified approach to interpreting model predictions.

Token Activation Map to Visually Explain Multimodal LLMs A unified approach to interpreting model predictions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.817131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.583129Z digest=sha256:fe4d63a47116b7107c11564377758417e0f9d1f3dba658eb0180ce1a4399a1a9

Observation 443a3d6b-8434-4fc5-9679-3d49fcda9473 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Token Activation Map to Visually Explain Multimodal LLMs Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.554791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.635669Z digest=sha256:7243d6971ea7c0b8a81a289fe4fd53de9be3f052acde162ae6a19ff0f6846629

Observation 01780bf3-78f0-4ac8-b323-3af92010a72c · outbound

This paper cites Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models.

Token Activation Map to Visually Explain Multimodal LLMs Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.702588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.702588Z digest=sha256:7af6c238bd09cef6d673554ba5530fc523b38da014ac3ca207bbfaeeffb2efc5

Observation 2fdfa9fa-56af-4347-9d35-47813276b48a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

Token Activation Map to Visually Explain Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.356930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.707466Z digest=sha256:c84b20ab8edd3ffa741351c677fc53fcf119f5d19a6ca1717983d0e4f086291e

Observation 57c3a70d-810f-4ae9-8b7c-fb8d406672dc · outbound

This paper cites Causality.

Token Activation Map to Visually Explain Multimodal LLMs Causality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.778705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.778705Z digest=sha256:30579939f740ce6f6f4e75ad9eb65d501e2e31f8f5a7115f4987607c7b01b6a4

Observation 5da96631-5200-416d-a009-0ce47338e1df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Token Activation Map to Visually Explain Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.222645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:25.999812Z digest=sha256:c7c2dcd0ea4bc36fdcb751f213c1c6d2d2984438587d443ea356e32d8f55867b

Observation 6fae3910-a117-40af-a75a-4eb733df3128 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Token Activation Map to Visually Explain Multimodal LLMs Glamm: Pixel grounding large multimodal model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.993373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:26.204832Z digest=sha256:4adc80769ad79db0ac156a2077f64d308068c338529a484bff812070e128f0a2

Observation 5687849b-dc4f-402c-88d4-74696d16da54 · outbound

This paper cites ” why should i trust you?” explaining the predictions of any classifier.

Token Activation Map to Visually Explain Multimodal LLMs ” why should i trust you?” explaining the predictions of any classifier

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.802753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:26.369433Z digest=sha256:c797193defd4c6b0a57f6f1f50f6e7718e9fe2876433d6dba2c611731a7023b5

Observation 05a57b0f-e59c-410c-a9bb-cfb97897fd94 · outbound

This paper cites Causal interpretation of self-attention in pre-trained trans- formers.

Token Activation Map to Visually Explain Multimodal LLMs Causal interpretation of self-attention in pre-trained trans- formers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.626574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:26.613931Z digest=sha256:2998c454e5805ea20a5c2ef4afd1111c3b5694a32a75ddcd993d2c5f46940cfc

Observation 196618f3-66a4-4c43-aee8-be7d529c1fed · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.849975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.849975Z digest=sha256:bca3a33d9914347bbb519024e5734a4d8c6febd9a75d07ff2fc4bbbe08b23897

Observation 67ed4cc3-904b-4d83-a56e-cd116cc1d3a2 · outbound

This paper cites Training- free object counting with prompts.

Token Activation Map to Visually Explain Multimodal LLMs Training- free object counting with prompts

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.256509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:27.179429Z digest=sha256:fdd7b58a0fb2975eaff23885cb066f9ac10e8044d007b8cb036bf4ae9c77c7bd

Observation 1a7d0272-266c-4853-b2cc-4f7fc734a287 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Token Activation Map to Visually Explain Multimodal LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.282270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.282270Z digest=sha256:49e2e853ea591d375aacebe2962289e3f75fea505f37820dbc0b60fd16783453

Observation 6150cd2f-2db5-48e3-bbb1-95b8b2d7738d · outbound

This paper cites Understanding how vision-language models rea- son when solving visual math problems.

Token Activation Map to Visually Explain Multimodal LLMs Understanding how vision-language models rea- son when solving visual math problems

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.082031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:27.416113Z digest=sha256:cff58e5618c5cfbc0d172c352dcf66cbb9ee71385500a73262de23bcf0dc2f9f

Observation 7a765326-b36b-454a-9900-2259ac8c7c62 · outbound

This paper cites Attention is all you need.

Token Activation Map to Visually Explain Multimodal LLMs Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.920473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:27.564729Z digest=sha256:11bbf1e9cf7b2568ef6ccbe8aa9810b59856f9bbf57ef1849a8dc210e839d9b7

Observation c69c4e1c-7ea3-477a-827b-b814833df753 · outbound

This paper cites Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks.

Token Activation Map to Visually Explain Multimodal LLMs Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.726790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:27.740841Z digest=sha256:bfe272fbadac33a8824d54aee9d634bc8dc7b8e24666d05163dd04e06519fc4a

Observation f3bf8d0d-6f75-4ecd-baf9-b4b3a078a944 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Token Activation Map to Visually Explain Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.901257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.901257Z digest=sha256:2180ff0d6cc15bc68ae7aebe4c280acb8661e0a72fd12f6a0bcc66fe1212caf7

Observation 1593e6e8-1fc5-4776-a1c7-7620dcf83565 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Token Activation Map to Visually Explain Multimodal LLMs Star: A benchmark for situated reasoning in real-world videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.519512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.031705Z digest=sha256:658e271226f322347e2b6e0274771f299f9c811c59b8c6178a6b1d570644803e

Observation 1e49da6f-40ab-42a7-aeb0-a8ae0eeb2f01 · outbound

This paper cites Efficient streaming language models with attention sinks.

Token Activation Map to Visually Explain Multimodal LLMs Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.306564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.184907Z digest=sha256:82c20377805fec2bbad5a7672de5ba7a27b2f17675cfa13181c99c74a5103450

Observation 84be4017-1d5c-4212-b382-2eb7d37008b0 · outbound

This paper cites A survey on causal inference.

Token Activation Map to Visually Explain Multimodal LLMs A survey on causal inference

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.119616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.331544Z digest=sha256:36998ecaceeda93d54b20192f2872e275a082bbf83b9cbb3c90d6cbd436ef010

Observation 0e52b13d-beb8-43b3-8e25-e2980e367942 · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

Token Activation Map to Visually Explain Multimodal LLMs From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.950021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.453431Z digest=sha256:eae685a7b658a6545f1554fe3d6b037ad4499d03becb7b605af4c53b614fec98

Observation 0ec2ab28-29f7-4ab3-8a1a-d02842692370 · outbound

This paper cites Learning deep features for discrimina- tive localization.

Token Activation Map to Visually Explain Multimodal LLMs Learning deep features for discrimina- tive localization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:28.534320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:28.534320Z digest=sha256:6b318794239c3f48c5cffa780d6df72525003bdb362d84cdfadf65151f8263b1

Observation 9278089c-6546-41d2-b914-a25a9151eebb · outbound

This paper cites with” and the punctuation mark “.

Token Activation Map to Visually Explain Multimodal LLMs with” and the punctuation mark “

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.630735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.788264Z digest=sha256:10b3b4f485dd49e74387e802f46feafc61c1982afcb045c540551b41a1e7e379

Observation 6bf0e808-16c8-4475-b27e-9a9e1a53e8d0 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.447211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.941451Z digest=sha256:c48c7bc266999e6ea73f6242f07f90b825dd3a218cc3bb8c4dd512b5307e0558

Observation 34a50edb-53d1-4e84-b268-dc2130cfdc34 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.265472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:29.101599Z digest=sha256:2b9988039fd684fe82763f32b01e105765fa2c35555d9b78402d388d63f40393

Observation cad3c0e3-9e47-4c87-b66d-b1722867a254 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.081243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:29.230235Z digest=sha256:5cd289248336cbbca76204240a7c2a06a5f83257806f53440c6518e8dc8406e9

Observation 0ec197a3-1b0d-48b1-85b4-bda3adb45db0 · outbound

This paper cites Object-determined.

Token Activation Map to Visually Explain Multimodal LLMs Object-determined

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.923294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:29.318102Z digest=sha256:56969430da55e8ef82409ff1a4b2f8d1cacd47fba1e2a684ef6488172a5aa67f

Observation 42bfa086-e8c6-463a-8d82-7eb6d2113a8c · outbound

This paper cites Missing arrows led to erroneous reasoning.

Token Activation Map to Visually Explain Multimodal LLMs Missing arrows led to erroneous reasoning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.760939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:29.407583Z digest=sha256:6f63e6ec401054a5dd085a13faf6486c1b85dfd9bf0494e260946b9d7a52bb56

Observation fa6da81a-05a1-4059-a00b-26a19f33bfe0 · outbound

This paper cites 2, 5, 6, 7, 14, 15.

Token Activation Map to Visually Explain Multimodal LLMs 2, 5, 6, 7, 14, 15

Reference 168

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.764893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:23.026379Z digest=sha256:c85feb37d10aa8f38404befc3bff13b76cce99d74e97d998001b35f7fe542561

Observation c8836885-99c1-429f-8b6c-71ada12ddd70 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.757393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:28.664380Z digest=sha256:f5c3e31e529d91f7d498df54492cb3fc089e1718e815373e276c442f2bea1e85

Observation ea323d48-56d5-4c18-97cd-36edf6b02d54 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:32.418675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:52:27.014153Z digest=sha256:1e79b3b841c5f48bc29cdd73e4d1c0ce65af5ef5a67d3b5e5df5bc60f24e47a4

Pith citing papers

Observation ee607a6c-424c-4103-8a9b-d5ac4adba379 · inbound

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models cites this paper.

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Token Activation Map to Visually Explain Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.554863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:58:38.554863Z digest=sha256:7fa73db3ccb6d115f7e3e342cb32dc2225c9cb8c7604edcf43af7684ff8c6476

Observation 61037376-a873-45fb-a438-086e6355c419 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Token Activation Map to Visually Explain Multimodal LLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.892061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:cd2df95b39f93ef37e346a26d0c293bae1b6723530fc736c5ccb7455a05f4f0d

Observation 68b4e094-aaed-4176-9567-dc7d2db5ea9d · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Token Activation Map to Visually Explain Multimodal LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.037170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:327ace502e409211a40efb3d0ac0b490c6251c742d936675e312a79e5688507f