Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:29.407583Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2506.23270.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:29.407583Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554863Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T01:45:52.034775Z
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d4de39f2-8367-4a38-bd53-b02dca29f34c · outbound
Token Activation Map to Visually Explain Multimodal LLMs Quantifying attention flow in transformers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ebc98010-930a-4a15-834f-258ab58ade04 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Attnlrp: Attention- aware layer-wise relevance propagation for transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bffa45d5-158c-4fba-819a-1a9fed0a37c3 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 021db75b-a6aa-45f6-8a2f-df56eea62a9b · outbound
Token Activation Map to Visually Explain Multimodal LLMs Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e28f49e7-adc0-47b5-b63b-7b1ae3ee0216 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Xai for trans- formers: Better explanations through conservative propa- gation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 818ccc59-2c78-4235-9a49-8d9538a00479 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Text2live: Text-driven layered image and video editing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3aa7372b-2964-43cc-8255-0946c3531fbe · outbound
Token Activation Map to Visually Explain Multimodal LLMs Lvlm-intrepret: An interpretability tool for large vision-language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56241116-b3c1-4dd0-826a-56f3dc54d316 · outbound
Token Activation Map to Visually Explain Multimodal LLMs An adaptive median filter for image denoising
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e807e267-b14c-43a4-9570-7995aaa5a0f0 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387ed1b8-79d1-4255-ad4f-234b5fe8ea23 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3551c4f-decd-459b-9d8c-447cd54b6c89 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Transformer inter- pretability beyond attention visualization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b092b981-a4fe-4711-b9ed-7c5fca5c3b20 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Less is more: Fewer interpretable region via submodular subset selection
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation acc9bf12-e418-4552-8424-9ffd90af11f8 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb0bd74-59d4-4c8c-af1c-87319c0c7084 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60315eb3-3ff5-45aa-adee-6f7db657e89f · outbound
Token Activation Map to Visually Explain Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8b3449-02d5-490e-b819-49d285b05ae2 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d83c952b-6c0e-4d7f-9b64-d0d02c398361 · outbound
Token Activation Map to Visually Explain Multimodal LLMs A survey on multimodal large lan- guage models for autonomous driving
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55e032c7-68f7-43e4-83e8-000f889c35f2 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Flashattention: Fast and memory-efficient exact at- tention with io-awareness
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 562b6886-6ac5-4d32-9905-332e33589d3e · outbound
Token Activation Map to Visually Explain Multimodal LLMs Vision transformers need registers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e0a0169-afc6-4f6d-9d91-4bc9b4618612 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29031a5-45de-4ba7-a704-248d703c7195 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30eddce3-f46c-4af8-93d8-03533192b924 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06b09a51-da1c-4187-8bbb-77138383df40 · outbound
Token Activation Map to Visually Explain Multimodal LLMs An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c3f78d9-6d9e-4956-ab1c-2079b9b30237 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Deep residual learning for image recognition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fceaabdb-ae1f-4e8c-a0e0-329b2a3b56d9 · outbound
Token Activation Map to Visually Explain Multimodal LLMs GPT-4o System Card
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3350316-b32e-4812-8ffb-e3e32ae89606 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Layercam: Exploring hierarchical class activation maps for localization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 138f9eb9-5bc3-4bc4-98d9-eaf078fd15b0 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Causal inference meets deep learning: A compre- hensive survey
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 755369e1-7dd3-4e7d-ae27-6c9169819545 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Unmasking clever hans predictors and as- sessing what machines really learn
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f96de966-94a9-40fb-a609-1d9f5bb8c4c6 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2da65a26-37a5-417a-ba21-7ff09e4f63ff · outbound
Token Activation Map to Visually Explain Multimodal LLMs Manipllm: Embodied multimodal large language model for object-centric robotic manipulation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 591122c2-d792-48f2-9ef0-af9ae5158136 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Exploring Visual Interpretability for Contrastive Language-Image Pre-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b51a60-2297-48fb-a314-4198d1beea9a · outbound
Token Activation Map to Visually Explain Multimodal LLMs A closer look at the explainability of con- trastive language-image pre-training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0889dc48-50f3-439e-8bd1-2a2762db6951 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Microsoft coco: Common objects in context
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8732b887-ceae-4a94-a807-544e0a93460e · outbound
Token Activation Map to Visually Explain Multimodal LLMs A medical multimodal large language model for future pandemics
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36fd000f-a697-4568-b524-23d0242845e7 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Visual instruction tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ed8483e-c879-4098-a0e1-200d1baa53ad · outbound
Token Activation Map to Visually Explain Multimodal LLMs A unified approach to interpreting model predictions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 443a3d6b-8434-4fc5-9679-3d49fcda9473 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01780bf3-78f0-4ac8-b323-3af92010a72c · outbound
Token Activation Map to Visually Explain Multimodal LLMs Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdfa9fa-56af-4347-9d35-47813276b48a · outbound
Token Activation Map to Visually Explain Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57c3a70d-810f-4ae9-8b7c-fb8d406672dc · outbound
Token Activation Map to Visually Explain Multimodal LLMs Causality
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da96631-5200-416d-a009-0ce47338e1df · outbound
Token Activation Map to Visually Explain Multimodal LLMs Learning transferable visual models from natural language supervi- sion
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6fae3910-a117-40af-a75a-4eb733df3128 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Glamm: Pixel grounding large multimodal model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5687849b-dc4f-402c-88d4-74696d16da54 · outbound
Token Activation Map to Visually Explain Multimodal LLMs ” why should i trust you?” explaining the predictions of any classifier
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05a57b0f-e59c-410c-a9bb-cfb97897fd94 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Causal interpretation of self-attention in pre-trained trans- formers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 196618f3-66a4-4c43-aee8-be7d529c1fed · outbound
Token Activation Map to Visually Explain Multimodal LLMs Grad-cam: Visual explanations from deep networks via gradient-based localization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ed4cc3-904b-4d83-a56e-cd116cc1d3a2 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Training- free object counting with prompts
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a7d0272-266c-4853-b2cc-4f7fc734a287 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6150cd2f-2db5-48e3-bbb1-95b8b2d7738d · outbound
Token Activation Map to Visually Explain Multimodal LLMs Understanding how vision-language models rea- son when solving visual math problems
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a765326-b36b-454a-9900-2259ac8c7c62 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Attention is all you need
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c69c4e1c-7ea3-477a-827b-b814833df753 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3bf8d0d-6f75-4ecd-baf9-b4b3a078a944 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1593e6e8-1fc5-4776-a1c7-7620dcf83565 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Star: A benchmark for situated reasoning in real-world videos
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e49da6f-40ab-42a7-aeb0-a8ae0eeb2f01 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Efficient streaming language models with attention sinks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84be4017-1d5c-4212-b382-2eb7d37008b0 · outbound
Token Activation Map to Visually Explain Multimodal LLMs A survey on causal inference
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e52b13d-beb8-43b3-8e25-e2980e367942 · outbound
Token Activation Map to Visually Explain Multimodal LLMs From redundancy to relevance: Enhancing explainability in multimodal large language mod- els
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ec2ab28-29f7-4ab3-8a1a-d02842692370 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Learning deep features for discrimina- tive localization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9278089c-6546-41d2-b914-a25a9151eebb · outbound
Token Activation Map to Visually Explain Multimodal LLMs with” and the punctuation mark “
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bf0e808-16c8-4475-b27e-9a9e1a53e8d0 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34a50edb-53d1-4e84-b268-dc2130cfdc34 · outbound
Token Activation Map to Visually Explain Multimodal LLMs living wall
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cad3c0e3-9e47-4c87-b66d-b1722867a254 · outbound
Token Activation Map to Visually Explain Multimodal LLMs living wall
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ec197a3-1b0d-48b1-85b4-bda3adb45db0 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Object-determined
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42bfa086-e8c6-463a-8d82-7eb6d2113a8c · outbound
Token Activation Map to Visually Explain Multimodal LLMs Missing arrows led to erroneous reasoning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fa6da81a-05a1-4059-a00b-26a19f33bfe0 · outbound
Token Activation Map to Visually Explain Multimodal LLMs 2, 5, 6, 7, 14, 15
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8836885-99c1-429f-8b6c-71ada12ddd70 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea323d48-56d5-4c18-97cd-36edf6b02d54 · outbound
Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee607a6c-424c-4103-8a9b-d5ac4adba379 · inbound
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Token Activation Map to Visually Explain Multimodal LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61037376-a873-45fb-a438-086e6355c419 · inbound
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Token Activation Map to Visually Explain Multimodal LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68b4e094-aaed-4176-9567-dc7d2db5ea9d · inbound
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Token Activation Map to Visually Explain Multimodal LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.