Pith. sign in

Paper Citation Record · LEDGER

Towards Interpreting Visual Information Processing in Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2410.07149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.07149 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:08:51.431988Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73d73144-6432-4148-9f8a-c10b50ba680b · inbound

What's in the Image? A Deep-Dive into the Vision of Vision Language Models cites this paper.

What's in the Image? A Deep-Dive into the Vision of Vision Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:08:51.431988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:08:51.431988Z digest=sha256:906d5a1791483d0aba57930a7142df507014c8792e24ce30753db219f54c7849

Observation df9fd07c-8bed-4973-9504-cd4d08f1ae97 · inbound

Cross-modal Information Flow in Multimodal Large Language Models cites this paper.

Cross-modal Information Flow in Multimodal Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.086378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.086378Z digest=sha256:e1e0b3f48df7cbd2300a25dfcbedcf0a5ed9a2b8f63d93900191cea748037e73

Observation 0fc8e35a-5d9c-44e2-bb8b-e7841be52e34 · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.795441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.795441Z digest=sha256:42f0d2ace17fef99e53a34b3877cc70087bcca195b00f2d6f82da5b9809164db

Observation d5c0e1f4-08fa-4fae-a4d5-f324eebe1557 · inbound

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models cites this paper.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.702315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.702315Z digest=sha256:e65c258c868a72225602ea61ba24da82eda702936ad74c2906af0fde20cae0d6

Observation 29447404-6719-40ec-9937-d267fb413939 · inbound

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models cites this paper.

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:45.549978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:45.549978Z digest=sha256:153f3aab657bd50618a4085b19787043224acc4179b31627034f29b605834cd5

Observation 92687b32-7c75-438b-814d-b08c34cffdf8 · inbound

How Visual Representations Map to Language Feature Space in Multimodal LLMs cites this paper.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.506978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.506978Z digest=sha256:bb63ad226a5bd500c08a01d5ace7058f64026b184cd49387a86523034da41cf8

Observation 7ddbf9bc-4e04-43db-921b-85cfde575da9 · inbound

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models cites this paper.

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:05.418595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:40:05.418595Z digest=sha256:e86ca714941a53435f518deb8a4ba4aadfdf29c43c88b591410c588bb2e38929

Observation c3146891-ea2f-4e93-8e4b-87ba1a9705b5 · inbound

Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers cites this paper.

Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:18.265684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:55:18.265684Z digest=sha256:e4a7596e7a90a946b872c249c52967f51cd6d2035381fc9249107dd098a8ca70

Observation ee6278d9-e751-480a-bd78-6275ac512c91 · inbound

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models cites this paper.

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:21:36.505496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T16:21:20.463222Z digest=sha256:568c13acf851744aae31f220290778784f4498fa5dee90ddcf0a88837af386e4

Observation 9db6822a-5dc7-42e8-9434-9404e800ddfa · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:54:20.225280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:0cff72e671b27bebd7a7aaced7c09a0df5da3a72ff653ba52d101886a45edf40

Observation 717efe28-5b72-47bd-bf19-0bad70126141 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T21:44:01.204287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:44:01.204287Z digest=sha256:a86b395baed4d35ca6185c0faeaba700cc8a1f26dcd3d84c453757ca968bc3d8

Observation cf2ea59a-fbbf-48d8-802f-148a6447860b · inbound

Understanding Counting Mechanisms in Large Language and Vision-Language Models cites this paper.

Understanding Counting Mechanisms in Large Language and Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:10.860465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T20:10:01.277804Z digest=sha256:9eb8ce5c64cb05231c55fa4b12938125e23b264c842886765a7cdebf22cf4d39

Observation 371aaf5a-e7ce-4e94-a9bb-1f0b438d264e · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.838229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:645c6c13cff868bf3f62924591798bb48b9abfa5a993cb30c67843d20c43907e

Observation 9aa4b899-48da-4581-9a2f-e6b10bf84ba4 · inbound

Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures cites this paper.

Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T13:35:02.428552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:35:02.428552Z digest=sha256:b0780d9bb2d1da49abf4d2c62afa24452cd68d572e1bf2fd841ad5613ac47f2c

Observation 28a85ea9-cd1f-439c-9754-bb0421b9952c · inbound

STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models cites this paper.

STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:13.340068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:16:04.999705Z digest=sha256:1b1aefb2b1488e3897e0c33c1dd96677a697afcc77a9adfe09988954b587e799

Observation 9848131b-30e0-4c93-b74f-d04e365c6f55 · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:09.413561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:344ab3d244e6e8bf50071d2cd50d0a049cef3d1ea5c7f62a121a65295df0f2ba

Observation 239a9c1e-f70f-4130-9e2c-6435248a535f · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:07.405883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:968aa3d1008180ee8e702f664a9bbb4cc80a7fab3b3612d2ab7811698a5c5dc7

Observation 579842ef-2b0c-442d-97d9-624e1a4f720a · inbound

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering cites this paper.

CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:50:40.259446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T18:05:38.705480Z digest=sha256:5f86e0395b20386ac92e8e2852e1a01e62da11405373787ebded296a768b3a2f

Observation 9a32b2f3-01dc-4369-9ac7-7c58ef86ddb0 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:15.370550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:d925291389dbf83563ecc0ae3afd17e0f2d09edc921baf7ccad08e92d0af120e

Observation ad01c324-f8e8-476a-b0de-ec2af77c28f0 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:58.935136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:a6f8d02d281c8fcd11c2216815b7a63f515888b78aa088dbf9cf1b72e6e16c4e

Observation 9932344f-2eed-4ff6-9ede-0014877bf3e7 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:41.681980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:d0ecd39be722a840cf101f21f2994bfdd9eddb4e8d3f0ae8570be73ca183f731

Observation bf977a50-5e78-4230-8871-31a7ff610c00 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.071648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:1c1f9985452916427bf91bac3a397c8369441243d25948c0a3f19606d7f17d81

Observation a1edef2e-eb6e-4a7a-8bde-a95d2843d7a7 · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.301867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:600443e889e6d06fd0e0331a1d54fbe21ec2e7b52de031acec85d9b785fbf8f6

Observation 8b6fdc94-6d06-4c8b-8f12-f84a2980c7f9 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:29.216583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:53732858a902af1499a49754854eda2e6847501222a6759c0166a3f19c1c0605

Observation d467b29d-18e8-4c35-a1db-0cbb5acc4098 · inbound

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models cites this paper.

Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:15:51.998943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T03:53:04.984303Z digest=sha256:a1c3b21f79ba4e3e46543b8361b7237ba060db17f03628bab60a56d256950a19

Observation 1e7fcd8f-d830-4ffa-ad08-ca62f0b97579 · inbound

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders cites this paper.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.570640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:d5d7bcea85a93e95de89b67ca03aef3e1c467e1fe6efae1d6c644d193df2c839

Observation 6b873ae6-e746-4faa-87db-ac947f96e8ca · inbound

TDSal: Task-Based Top-Down Saliency Prediction Model cites this paper.

TDSal: Task-Based Top-Down Saliency Prediction Model Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:15:45.858365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:15:45.858365Z digest=sha256:2f0400acc4c911de467d573f1eb1a86f0ec028a38cde9eaa3d55c495071c6547

Observation 8fa7922b-835c-4dc3-a5e8-cb48569b5b14 · inbound

Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models cites this paper.

Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:00:47.642822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:00:47.642822Z digest=sha256:31395f2545e0c4d34f06770b03214170b7f32efb2cf7fbf80ff1fd00db20f35e

Observation 287e32b1-9179-496c-a116-d6747c979a3e · inbound

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA cites this paper.

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T21:26:11.993187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:26:11.993187Z digest=sha256:fa02d2ad84bf6dbdc083e94e239c0fff7a31dd5dd92f7c7880fc060735dcd7a9

Observation cdd46c79-02f1-4c4e-86bb-b856ba020424 · inbound

Multimodal Model Diffing for Feature Discovery and Control cites this paper.

Multimodal Model Diffing for Feature Discovery and Control Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.032176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.032176Z digest=sha256:ea50f90083cac7af79eafef59caaf51b21e1c0a053aacb1a3c4f7e17ba1fa260