Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:08.321921Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2506.11976.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:08.321921Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:07.449816Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T22:17:26.125055Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d3699c23-1c90-4763-8c70-44707168202f · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards monosemanticity: Decomposing language models with dictionary learning, 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 07ce8781-2025-4982-8acf-de1552b550a7 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Controlling Vision Foundation Models via Text Explanations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ce3c5d-6bcf-4256-b409-918c0da41987 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e7d4643-cb6b-4356-abae-16728f76041d · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c89694a-e792-44b8-a725-35da8d656625 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Arik, Tejas Nama, and Tomas Pfister
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5236b704-0e78-42c1-9656-03a507e1ed67 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs A mathemati- cal framework for transformer circuits
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b317eb25-217e-4507-9681-f25af730f1b5 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a117c9c-7741-4049-a70a-582d506c319e · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting CLIP's Image Representation via Text-Based Decomposition
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1144db9-52c8-4570-8c84-dbfcb944205a · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Sparse autoencoders find highly interpretable features in language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 404ab531-414f-4122-a3e4-11bce77cf204 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Hudson and Christopher D
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 25f00dca-de3e-4cfd-883d-653b0b12c346 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8983eb4-1142-47d0-a99a-f09f64e827d3 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Evaluating object hallucination in large vision- language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc6efdc3-a151-4fdc-8221-4521bf9c9819 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8c4433-db51-480c-9055-8664a4eeb97a · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Vila: On pre-training for vi- sual language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation effa39d3-0c87-4fb8-82dc-ef0a4cb5c2e6 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Visual instruction tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d95b6d-8393-4dd3-8645-5dce02161448 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Improved baselines with visual instruction tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6af390d-ded8-46ac-8748-6fbbe7850e37 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Linearly mapping from image to text space
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92687b32-7c75-438b-814d-b08c34cffdf8 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 847a9d7a-2e48-49a5-be4d-14db007b01ea · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting GPT: The Logit Lens
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5022e038-1a89-4d2d-97d9-b8f83fd8bd5a · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98f162a-bd02-4070-869c-7b659162ab52 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Bridg- ing vision and language spaces with assignment prediction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca3e0edd-32ae-4375-bd68-ed81739091d9 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Learning transferable visual models from natural language supervi- sion
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f3e4212-a992-4b9f-8e5d-e364113e2139 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Multimodal neurons in pre- trained text-only transformers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 54ecaa4c-e899-43b5-925b-20aa394b7bd5 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Open Problems in Mechanistic Interpretability
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1937f5de-3784-4034-89c3-3e0a890480da · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Paligemma 2: A family of versatile vlms for transfer, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2083da55-8b62-4185-94f9-0d851b8bdc27 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma 2: Improving Open Language Models at a Practical Size
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9431d42-e8cf-47df-9a14-6e6a98d5e4d6 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35c4f53e-89b3-4c81-81b7-9e89d9844c1c · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85d98845-988b-458e-8ad4-bda971176a71 · outbound
How Visual Representations Map to Language Feature Space in Multimodal LLMs Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f81683f-e654-4363-ad77-03ebc2dffce9 · inbound
Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention How Visual Representations Map to Language Feature Space in Multimodal LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab9ce2f3-d682-4f24-bbce-0a77ab76cb60 · inbound
Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4f28d0ec-44f2-4d3f-927d-2075dbb823b4 · inbound
Pathways of Visual Information Flow in Vision-Language Models How Visual Representations Map to Language Feature Space in Multimodal LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.