Pith. sign in

Paper Citation Record · LEDGER

Token Activation Map to Visually Explain Multimodal LLMs

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2506.23270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23270 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:29.407583Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T01:45:52.034775Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4de39f2-8367-4a38-bd53-b02dca29f34c · outbound

This paper cites Quantifying attention flow in transformers.

Token Activation Map to Visually Explain Multimodal LLMs Quantifying attention flow in transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:22.828224Z digest=sha256:a1439e0fec8a9a9857f7cabea5166c0f935d54a84a6ce739e6f9e78174917753

Observation ebc98010-930a-4a15-834f-258ab58ade04 · outbound

This paper cites Attnlrp: Attention- aware layer-wise relevance propagation for transformers.

Token Activation Map to Visually Explain Multimodal LLMs Attnlrp: Attention- aware layer-wise relevance propagation for transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.029975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:22.908507Z digest=sha256:57295648df2434387eb487bec5fa2399f0658b155d3a01e274791c04a22a0238

Observation bffa45d5-158c-4fba-819a-1a9fed0a37c3 · outbound

This paper cites Vl-interpret: An interactive visualization tool for interpreting vision-language transformers.

Token Activation Map to Visually Explain Multimodal LLMs Vl-interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.481689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.138905Z digest=sha256:5d5a107ff2a682427f8e032e688d7f5ce5669a30e9a6619a2e54622ff227a0ba

Observation 021db75b-a6aa-45f6-8a2f-df56eea62a9b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Token Activation Map to Visually Explain Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.200165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.200165Z digest=sha256:9b09623f06d6d6bb2fbf9bc258d1867d1c0960d07e548671287d3c396c6d87b3

Observation e28f49e7-adc0-47b5-b63b-7b1ae3ee0216 · outbound

This paper cites Xai for trans- formers: Better explanations through conservative propa- gation.

Token Activation Map to Visually Explain Multimodal LLMs Xai for trans- formers: Better explanations through conservative propa- gation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.145550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.291980Z digest=sha256:e1c2d7e426dc98deff2beb6ee9c05db83d61317ddc561990dbd16b9b16fb9c16

Observation 818ccc59-2c78-4235-9a49-8d9538a00479 · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Token Activation Map to Visually Explain Multimodal LLMs Text2live: Text-driven layered image and video editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.802003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.383073Z digest=sha256:c206cced0dc110e78fe3bc20568bbc6aa0ced040fb0fd6821c88c7d40ea3757c

Observation 3aa7372b-2964-43cc-8255-0946c3531fbe · outbound

This paper cites Lvlm-intrepret: An interpretability tool for large vision-language models.

Token Activation Map to Visually Explain Multimodal LLMs Lvlm-intrepret: An interpretability tool for large vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.419141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.464253Z digest=sha256:4e87fc5e2aa3f85980d07e1e964fa58f00d70d1c029bc13af0ff6a1c22d38f78

Observation 56241116-b3c1-4dd0-826a-56f3dc54d316 · outbound

This paper cites An adaptive median filter for image denoising.

Token Activation Map to Visually Explain Multimodal LLMs An adaptive median filter for image denoising

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.138362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.579543Z digest=sha256:fa8cfa92390b8c73b2e661681130eb1a75c9eb45f5f8e76fac7ae8fc39b0ea16

Observation e807e267-b14c-43a4-9570-7995aaa5a0f0 · outbound

This paper cites Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673939Z digest=sha256:85c9f5f7c7280fa798a38d9f66ccc6fb726d8a9df92eb1cd251c26ec1d106dd2

Observation 387ed1b8-79d1-4255-ad4f-234b5fe8ea23 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Token Activation Map to Visually Explain Multimodal LLMs Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743484Z digest=sha256:f1dbeebf02accb0a53f473c043b91c3df81a3c9bd7f69e9e0b20b6b178b12435

Observation f3551c4f-decd-459b-9d8c-447cd54b6c89 · outbound

This paper cites Transformer inter- pretability beyond attention visualization.

Token Activation Map to Visually Explain Multimodal LLMs Transformer inter- pretability beyond attention visualization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.945210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.851689Z digest=sha256:db5aaf747d23856389f8340d76b6717afdfe8b38cdb4d83120bda9422192e368

Observation b092b981-a4fe-4711-b9ed-7c5fca5c3b20 · outbound

This paper cites Less is more: Fewer interpretable region via submodular subset selection.

Token Activation Map to Visually Explain Multimodal LLMs Less is more: Fewer interpretable region via submodular subset selection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.796077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.942557Z digest=sha256:3fdbaaedea2d27b2bad7977a4ca9b78099c5dbaca3ae348c6185e45477d2dc62

Observation acc9bf12-e418-4552-8424-9ffd90af11f8 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.065434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.065434Z digest=sha256:384038f6ac417f47aa323854163531b0b1ece14716243222bfa83cd2e2123dc1

Observation afb0bd74-59d4-4c8c-af1c-87319c0c7084 · outbound

This paper cites Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection.

Token Activation Map to Visually Explain Multimodal LLMs Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.619683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.182522Z digest=sha256:8d27533049def2e298dae0aa104ddf4b2c4bcc0c11a88587f3f343c85604076d

Observation 60315eb3-3ff5-45aa-adee-6f7db657e89f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Token Activation Map to Visually Explain Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.270466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.270466Z digest=sha256:131914aaa72c23693595b03bfd690a1bff82382e30ece3b7faa5f52c71bf44b7

Observation 6b8b3449-02d5-490e-b819-49d285b05ae2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Token Activation Map to Visually Explain Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.447607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.355777Z digest=sha256:9b0f096b4ac7edd335b09c1de2c1d24736bd5598994e4331acd4d9f7d1d8bc97

Observation d83c952b-6c0e-4d7f-9b64-d0d02c398361 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Token Activation Map to Visually Explain Multimodal LLMs A survey on multimodal large lan- guage models for autonomous driving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.265238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.418757Z digest=sha256:63408546b5f00d16d014c4cca2bcd2db70f3307899327eb7e5dff3f07d4a195f

Observation 55e032c7-68f7-43e4-83e8-000f889c35f2 · outbound

This paper cites Flashattention: Fast and memory-efficient exact at- tention with io-awareness.

Token Activation Map to Visually Explain Multimodal LLMs Flashattention: Fast and memory-efficient exact at- tention with io-awareness

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.123765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.485401Z digest=sha256:992f0117bbedde0e2fa0009b51c02e2efb579e29c52111df22f389954ba25235

Observation 562b6886-6ac5-4d32-9905-332e33589d3e · outbound

This paper cites Vision transformers need registers.

Token Activation Map to Visually Explain Multimodal LLMs Vision transformers need registers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.001289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.557594Z digest=sha256:b2158976fce00e4b16b67d0345c4115149cbd6186ba6fdc14a7994de0a0950e0

Observation 9e0a0169-afc6-4f6d-9d91-4bc9b4618612 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Token Activation Map to Visually Explain Multimodal LLMs Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.628249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.628249Z digest=sha256:5dfcd8769c1e1d8ca0760897fb3163e6cfe9ca1ec93337ae52a68fceb89bacf9

Observation e29031a5-45de-4ba7-a704-248d703c7195 · outbound

This paper cites Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis.

Token Activation Map to Visually Explain Multimodal LLMs Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.848320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.672164Z digest=sha256:b2cda8b2243e8f67bd3c73ee7f5114e6b8492b5dd418c509b126458a73791406

Observation 30eddce3-f46c-4af8-93d8-03533192b924 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models.

Token Activation Map to Visually Explain Multimodal LLMs Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.696589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.721249Z digest=sha256:d4db66a71cdb735efe532c2b2d6d2b15cff68c5be830a6e56f2920415cf0f833

Observation 06b09a51-da1c-4187-8bbb-77138383df40 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Token Activation Map to Visually Explain Multimodal LLMs An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.560451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.800458Z digest=sha256:33717dcc5f07abb9180c46eba71869424af8629ec963d756f21a42866ee2d406

Observation 0c3f78d9-6d9e-4956-ab1c-2079b9b30237 · outbound

This paper cites Deep residual learning for image recognition.

Token Activation Map to Visually Explain Multimodal LLMs Deep residual learning for image recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.398468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.853813Z digest=sha256:c33ebe9b0394f02db2cc7260338d59417501af31d04d3f2d734b04d53bf36ace

Observation fceaabdb-ae1f-4e8c-a0e0-329b2a3b56d9 · outbound

This paper cites GPT-4o System Card.

Token Activation Map to Visually Explain Multimodal LLMs GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893909Z digest=sha256:3575e5d8a15a01e420aca9e311a91dd3f1456f59e3f703794adc2786a6cbd83c

Observation f3350316-b32e-4812-8ffb-e3e32ae89606 · outbound

This paper cites Layercam: Exploring hierarchical class activation maps for localization.

Token Activation Map to Visually Explain Multimodal LLMs Layercam: Exploring hierarchical class activation maps for localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.237320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:24.962731Z digest=sha256:e350c0c6ec9ebb59e6f90c9c22cff127ece3a5fdf1afb8cd69dbe842213172aa

Observation 138f9eb9-5bc3-4bc4-98d9-eaf078fd15b0 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

Token Activation Map to Visually Explain Multimodal LLMs Causal inference meets deep learning: A compre- hensive survey

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.053887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.019347Z digest=sha256:6840b389e0bb2fe9e621a520d98b14fa8db15070691426cfa3c35f026fcd740f

Observation 755369e1-7dd3-4e7d-ae27-6c9169819545 · outbound

This paper cites Unmasking clever hans predictors and as- sessing what machines really learn.

Token Activation Map to Visually Explain Multimodal LLMs Unmasking clever hans predictors and as- sessing what machines really learn

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.884300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.064591Z digest=sha256:dededc9f3a245ecb4c173a8c55841f9d749b07cd3a4854e95241611dd66bede0

Observation f96de966-94a9-40fb-a609-1d9f5bb8c4c6 · outbound

This paper cites Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education.

Token Activation Map to Visually Explain Multimodal LLMs Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.688437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.118755Z digest=sha256:e07cd7ffe1898face30eb746f285832c71010c84a694af934ec7d24ab197cec5

Observation 2da65a26-37a5-417a-ba21-7ff09e4f63ff · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation.

Token Activation Map to Visually Explain Multimodal LLMs Manipllm: Embodied multimodal large language model for object-centric robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.487383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.203550Z digest=sha256:a4ee4ca2e29863bf60db184c333960c0de8053585d3232287b0f81eb8e2ef054

Observation 591122c2-d792-48f2-9ef0-af9ae5158136 · outbound

This paper cites Exploring Visual Interpretability for Contrastive Language-Image Pre-training.

Token Activation Map to Visually Explain Multimodal LLMs Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.254347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.254347Z digest=sha256:3346f2e8c38f46c5d725e7ff9880b2309d4a6e14dab4a0ff520e58508267fb8f

Observation 67b51a60-2297-48fb-a314-4198d1beea9a · outbound

This paper cites A closer look at the explainability of con- trastive language-image pre-training.

Token Activation Map to Visually Explain Multimodal LLMs A closer look at the explainability of con- trastive language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.338754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.315356Z digest=sha256:c8200b4c2505537dbbb0b88f588b054660e384b162281601ca3551528e76e0f8

Observation 0889dc48-50f3-439e-8bd1-2a2762db6951 · outbound

This paper cites Microsoft coco: Common objects in context.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.381037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.381037Z digest=sha256:6c71e35b5a753e3c81eb296da0877cca2b70289af88081b94966209b61c49382

Observation 8732b887-ceae-4a94-a807-544e0a93460e · outbound

This paper cites A medical multimodal large language model for future pandemics.

Token Activation Map to Visually Explain Multimodal LLMs A medical multimodal large language model for future pandemics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.156408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.446880Z digest=sha256:443cb5a4efc1e9accfa23b89dc7211e23fc035dd248808dd6e728649c2089b07

Observation 36fd000f-a697-4568-b524-23d0242845e7 · outbound

This paper cites Visual instruction tuning.

Token Activation Map to Visually Explain Multimodal LLMs Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.016402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.514680Z digest=sha256:ec58565a654274bda8076f231f89062419d31f90240682ac663d3128407946c0

Observation 3ed8483e-c879-4098-a0e1-200d1baa53ad · outbound

This paper cites A unified approach to interpreting model predictions.

Token Activation Map to Visually Explain Multimodal LLMs A unified approach to interpreting model predictions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.817131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.583129Z digest=sha256:327aa578db1c09b7ee6e90bd4ba62dc66253cc1a565de0f15becbcb59f11cdde

Observation 443a3d6b-8434-4fc5-9679-3d49fcda9473 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Token Activation Map to Visually Explain Multimodal LLMs Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.554791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.635669Z digest=sha256:5500829bbc6ef195b1d05910721629d06122e96392029f955cdc198b3d2c56bc

Observation 01780bf3-78f0-4ac8-b323-3af92010a72c · outbound

This paper cites Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models.

Token Activation Map to Visually Explain Multimodal LLMs Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.702588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.702588Z digest=sha256:7a8a56a075dcf71b7e1c0167a1c307e6e07dda3e63de58e743e432d7416c4cda

Observation 2fdfa9fa-56af-4347-9d35-47813276b48a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

Token Activation Map to Visually Explain Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.356930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.707466Z digest=sha256:5343de27d5623a39db2ac6149c170e09cb8b8a3c5a8b5aeb60418be11ac0718c

Observation 57c3a70d-810f-4ae9-8b7c-fb8d406672dc · outbound

This paper cites Causality.

Token Activation Map to Visually Explain Multimodal LLMs Causality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.778705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.778705Z digest=sha256:d85bf057ab30cbf7bd0b7b67413486aa4aaf2e6c2fbffcaac3e774bd29846b79

Observation 5da96631-5200-416d-a009-0ce47338e1df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Token Activation Map to Visually Explain Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.222645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:25.999812Z digest=sha256:1abaf38770ab378e0720a64066401071d7cc251a1e4d83bcece9e26edf1e29e1

Observation 6fae3910-a117-40af-a75a-4eb733df3128 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Token Activation Map to Visually Explain Multimodal LLMs Glamm: Pixel grounding large multimodal model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.993373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:26.204832Z digest=sha256:60c4b3ca2348eb97af1efbe6a810c61061717290cf21ca76870050e572bb8aff

Observation 5687849b-dc4f-402c-88d4-74696d16da54 · outbound

This paper cites ” why should i trust you?” explaining the predictions of any classifier.

Token Activation Map to Visually Explain Multimodal LLMs ” why should i trust you?” explaining the predictions of any classifier

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.802753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:26.369433Z digest=sha256:3675eb43133e1af5721f7d05d0e1ee0a4c32074f5e6c44f0b9b54baa56985a71

Observation 05a57b0f-e59c-410c-a9bb-cfb97897fd94 · outbound

This paper cites Causal interpretation of self-attention in pre-trained trans- formers.

Token Activation Map to Visually Explain Multimodal LLMs Causal interpretation of self-attention in pre-trained trans- formers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.626574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:26.613931Z digest=sha256:071f869ec88d35552ad873744ac40ea78c76497dbee6c3250ea02b7c87b4d6f2

Observation 196618f3-66a4-4c43-aee8-be7d529c1fed · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.849975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.849975Z digest=sha256:ea06e17e50057438fc4962de5cffa203374c4170798b580aa077469d8ee34e68

Observation 67ed4cc3-904b-4d83-a56e-cd116cc1d3a2 · outbound

This paper cites Training- free object counting with prompts.

Token Activation Map to Visually Explain Multimodal LLMs Training- free object counting with prompts

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.256509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:27.179429Z digest=sha256:f3273b39dcc4445463233024f07347622b0b4356f49093e2feb33ec25031d2a1

Observation 1a7d0272-266c-4853-b2cc-4f7fc734a287 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Token Activation Map to Visually Explain Multimodal LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.282270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.282270Z digest=sha256:f50878821ca474a37acb3e52faee74dfa987cdabc6b3b42a1941cc0d160db508

Observation 6150cd2f-2db5-48e3-bbb1-95b8b2d7738d · outbound

This paper cites Understanding how vision-language models rea- son when solving visual math problems.

Token Activation Map to Visually Explain Multimodal LLMs Understanding how vision-language models rea- son when solving visual math problems

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.082031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:27.416113Z digest=sha256:f2d1159da712b5bfe4d11fc52d0f50c0d869cbef9b2e170dc10477720b7ab8a3

Observation 7a765326-b36b-454a-9900-2259ac8c7c62 · outbound

This paper cites Attention is all you need.

Token Activation Map to Visually Explain Multimodal LLMs Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.920473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:27.564729Z digest=sha256:78c3b4b399cd709f2b4a276f627868734fd4034adf2626eff1b9aeb72a2b7d68

Observation c69c4e1c-7ea3-477a-827b-b814833df753 · outbound

This paper cites Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks.

Token Activation Map to Visually Explain Multimodal LLMs Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.726790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:27.740841Z digest=sha256:9177b00bf25e2137cf66999902db4d4d246d115561daba41b7b755b8c36a571c

Observation f3bf8d0d-6f75-4ecd-baf9-b4b3a078a944 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Token Activation Map to Visually Explain Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.901257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.901257Z digest=sha256:193e211021ddeb8771cb659981bea2a433cba4e33b7637f9aad81873406adcbc

Observation 1593e6e8-1fc5-4776-a1c7-7620dcf83565 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Token Activation Map to Visually Explain Multimodal LLMs Star: A benchmark for situated reasoning in real-world videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.519512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.031705Z digest=sha256:945b879701b0d4f6c0aee76f8b65125a8b75c1dac8acc6301b8fdb47c3bca360

Observation 1e49da6f-40ab-42a7-aeb0-a8ae0eeb2f01 · outbound

This paper cites Efficient streaming language models with attention sinks.

Token Activation Map to Visually Explain Multimodal LLMs Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.306564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.184907Z digest=sha256:db1ed2937584812ba9595fe70044f3a7ea706fe114561aec9dbe4e91a5bf57e5

Observation 84be4017-1d5c-4212-b382-2eb7d37008b0 · outbound

This paper cites A survey on causal inference.

Token Activation Map to Visually Explain Multimodal LLMs A survey on causal inference

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.119616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.331544Z digest=sha256:20f5c1655bac68928fc14ba8a659de7d6d1d3c0a63606bbb6c535fecd3ff5f4d

Observation 0e52b13d-beb8-43b3-8e25-e2980e367942 · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

Token Activation Map to Visually Explain Multimodal LLMs From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.950021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.453431Z digest=sha256:37b57e00aa176cb6e6c5c62b4bce332062f64f804a057e031227e21d0a456ef8

Observation 0ec2ab28-29f7-4ab3-8a1a-d02842692370 · outbound

This paper cites Learning deep features for discrimina- tive localization.

Token Activation Map to Visually Explain Multimodal LLMs Learning deep features for discrimina- tive localization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:28.534320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:28.534320Z digest=sha256:0a8aae025357e88a72fbd6148c176ddb7fd511c4d4fb57178e6bb25a25828965

Observation 9278089c-6546-41d2-b914-a25a9151eebb · outbound

This paper cites with” and the punctuation mark “.

Token Activation Map to Visually Explain Multimodal LLMs with” and the punctuation mark “

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.630735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.788264Z digest=sha256:1325d69831028024f8f08c4d0378f30d06892e7518b2ebab93c64f1f33695e30

Observation 6bf0e808-16c8-4475-b27e-9a9e1a53e8d0 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.447211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.941451Z digest=sha256:beecf3d69d2d28f42aa529345745db0724648129c4c6213c7e909f38a4f8753d

Observation 34a50edb-53d1-4e84-b268-dc2130cfdc34 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.265472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:29.101599Z digest=sha256:5e493c208621ece439f56abebc5202427a92abb60e87e6c7d33bba3b3204cbb2

Observation cad3c0e3-9e47-4c87-b66d-b1722867a254 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.081243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:29.230235Z digest=sha256:3548a8ea7688dd194564fd1658c8c9995ac2f1c0e324fc425b9046448b9e1df6

Observation 0ec197a3-1b0d-48b1-85b4-bda3adb45db0 · outbound

This paper cites Object-determined.

Token Activation Map to Visually Explain Multimodal LLMs Object-determined

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.923294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:29.318102Z digest=sha256:2252596feca9bc2c23a863c3b44ee5c6d4e64233a2f4ada20106c4b7442dec21

Observation 42bfa086-e8c6-463a-8d82-7eb6d2113a8c · outbound

This paper cites Missing arrows led to erroneous reasoning.

Token Activation Map to Visually Explain Multimodal LLMs Missing arrows led to erroneous reasoning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.760939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:29.407583Z digest=sha256:dd015cbada58366ddf05a878a1d9b782e388dca747470139b10d5f10d6ac1b7b

Observation fa6da81a-05a1-4059-a00b-26a19f33bfe0 · outbound

This paper cites 2, 5, 6, 7, 14, 15.

Token Activation Map to Visually Explain Multimodal LLMs 2, 5, 6, 7, 14, 15

Reference 168

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.764893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:23.026379Z digest=sha256:84607255968bc5115db4a670d15259ef49a98541970b84e669c866b2d39dd635

Observation c8836885-99c1-429f-8b6c-71ada12ddd70 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.757393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:28.664380Z digest=sha256:e927696b6a00ea11ee74d3988a0d91f79596172bbfd6106f16691844902d8406

Observation ea323d48-56d5-4c18-97cd-36edf6b02d54 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:32.418675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:27.014153Z digest=sha256:a82e7d46a536efc027ad69a19a3d04d11603abfb14a6f228e5a3f3b51e1eba94

Pith citing papers

Observation ee607a6c-424c-4103-8a9b-d5ac4adba379 · inbound

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models cites this paper.

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Token Activation Map to Visually Explain Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.554863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:58:38.554863Z digest=sha256:b84a9941f66a7f92aefc82b38672f4703655dd77f046beb0b5adbb404daa6cc0

Observation 61037376-a873-45fb-a438-086e6355c419 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Token Activation Map to Visually Explain Multimodal LLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.892061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:2068c5897644f26cf2d6e5f2d018540d96707bfb4ab9ae71cf6f6d1639708997

Observation 68b4e094-aaed-4176-9567-dc7d2db5ea9d · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Token Activation Map to Visually Explain Multimodal LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.037170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:26a5b250f1eedc4c4b0ed4854a518c6ff28e7d1e8fb437ebf041556db68f6426