Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2410.07149.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:08:51.431988Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 73d73144-6432-4148-9f8a-c10b50ba680b · inbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9fd07c-8bed-4973-9504-cd4d08f1ae97 · inbound
Cross-modal Information Flow in Multimodal Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc8e35a-5d9c-44e2-bb8b-e7841be52e34 · inbound
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5c0e1f4-08fa-4fae-a4d5-f324eebe1557 · inbound
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29447404-6719-40ec-9937-d267fb413939 · inbound
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92687b32-7c75-438b-814d-b08c34cffdf8 · inbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ddbf9bc-4e04-43db-921b-85cfde575da9 · inbound
Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3146891-ea2f-4e93-8e4b-87ba1a9705b5 · inbound
Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6278d9-e751-480a-bd78-6275ac512c91 · inbound
V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9db6822a-5dc7-42e8-9434-9404e800ddfa · inbound
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 717efe28-5b72-47bd-bf19-0bad70126141 · inbound
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2ea59a-fbbf-48d8-802f-148a6447860b · inbound
Understanding Counting Mechanisms in Large Language and Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 371aaf5a-e7ce-4e94-a9bb-1f0b438d264e · inbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9aa4b899-48da-4581-9a2f-e6b10bf84ba4 · inbound
Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a85ea9-cd1f-439c-9754-bb0421b9952c · inbound
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9848131b-30e0-4c93-b74f-d04e365c6f55 · inbound
From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 239a9c1e-f70f-4130-9e2c-6435248a535f · inbound
ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 579842ef-2b0c-442d-97d9-624e1a4f720a · inbound
CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a32b2f3-01dc-4369-9ac7-7c58ef86ddb0 · inbound
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ad01c324-f8e8-476a-b0de-ec2af77c28f0 · inbound
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9932344f-2eed-4ff6-9ede-0014877bf3e7 · inbound
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf977a50-5e78-4230-8871-31a7ff610c00 · inbound
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a1edef2e-eb6e-4a7a-8bde-a95d2843d7a7 · inbound
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b6fdc94-6d06-4c8b-8f12-f84a2980c7f9 · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d467b29d-18e8-4c35-a1db-0cbb5acc4098 · inbound
Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e7fcd8f-d830-4ffa-ad08-ca62f0b97579 · inbound
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b873ae6-e746-4faa-87db-ac947f96e8ca · inbound
TDSal: Task-Based Top-Down Saliency Prediction Model Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa7922b-835c-4dc3-a5e8-cb48569b5b14 · inbound
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 287e32b1-9179-496c-a116-d6747c979a3e · inbound
How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd46c79-02f1-4c4e-86bb-b856ba020424 · inbound
Multimodal Model Diffing for Feature Discovery and Control Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.