Pith. sign in

Paper Citation Record · LEDGER

How Visual Representations Map to Language Feature Space in Multimodal LLMs

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2506.11976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11976 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:08.321921Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:07.449816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:26.125055Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3699c23-1c90-4763-8c70-44707168202f · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning, 2023.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards monosemanticity: Decomposing language models with dictionary learning, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:12.186433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.197245Z digest=sha256:a34ae4213c89f37fda42e86b78acc8c3250a4ff421ea1d8a675f880944f63a20

Observation 07ce8781-2025-4982-8acf-de1552b550a7 · outbound

This paper cites Interpreting and Controlling Vision Foundation Models via Text Explanations.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Controlling Vision Foundation Models via Text Explanations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.285352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.285352Z digest=sha256:a0f3692558cef41082571a4c86e4a5f23e72d213a57efc381349596b2d1aa6ae

Observation 66ce3c5d-6bcf-4256-b409-918c0da41987 · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.979935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.381515Z digest=sha256:33a1728fe2f21c7f83d18ea7734742debe2bccd6a9fc14534d06ca327eac4b80

Observation 7e7d4643-cb6b-4356-abae-16728f76041d · outbound

This paper cites The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs The cognitive revolution in interpretability: From explaining behavior to interpreting representations and algorithms, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.736793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.453896Z digest=sha256:9aae8361471300d79688f723662b3080879d81874db9bc660900fd1629785635

Observation 7c89694a-e792-44b8-a725-35da8d656625 · outbound

This paper cites Arik, Tejas Nama, and Tomas Pfister.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Arik, Tejas Nama, and Tomas Pfister

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.548557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.522248Z digest=sha256:c3e7a18866fa4c29620182cc920867cb4fe7e2e74247a7e26f8b69152029bbf0

Observation 5236b704-0e78-42c1-9656-03a507e1ed67 · outbound

This paper cites A mathemati- cal framework for transformer circuits.

How Visual Representations Map to Language Feature Space in Multimodal LLMs A mathemati- cal framework for transformer circuits

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.309715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.622671Z digest=sha256:58a626b86e27c2cb6f676d227212780a3619322153a0519a58b97c88e8a7ccc6

Observation b317eb25-217e-4507-9681-f25af730f1b5 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.715063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.715063Z digest=sha256:d200d8b0b42c81bc8a3d85141aabcc8f0ccb292a59371cb9bd55e101c9298b3d

Observation 7a117c9c-7741-4049-a70a-582d506c319e · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.784761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.784761Z digest=sha256:2a1df2f17608977c4c18735ee24cb223fb9114da2afe5904973e4d03e68f9185

Observation b1144db9-52c8-4570-8c84-dbfcb944205a · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Sparse autoencoders find highly interpretable features in language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:11.122586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.856612Z digest=sha256:13eb44cc6dcec2f792353915461c8780171918caf7aff358047f2d3ad507d954

Observation 404ab531-414f-4122-a3e4-11bce77cf204 · outbound

This paper cites Hudson and Christopher D.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Hudson and Christopher D

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.917163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:06.898241Z digest=sha256:0efccce5dadc4941f0f19a2b94204b06cd41d51f0efbc24049d70997811d391b

Observation 25f00dca-de3e-4cfd-883d-653b0b12c346 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.969709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.969709Z digest=sha256:9a357a689922f28239b751107c59dc5fda102059090c311662bf72bc18a2a28c

Observation f8983eb4-1142-47d0-a99a-f09f64e827d3 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Evaluating object hallucination in large vision- language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.725401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.035217Z digest=sha256:85a5ebd1263fcc1795863b04075c41c3cc509fdc440c278ec9df2d4703712cc6

Observation bc6efdc3-a151-4fdc-8221-4521bf9c9819 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.108965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.108965Z digest=sha256:11af3977c28aa9bfcd2abd2db46c8ec156738db579998074e1d735f9a68ce0de

Observation fd8c4433-db51-480c-9055-8664a4eeb97a · outbound

This paper cites Vila: On pre-training for vi- sual language models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Vila: On pre-training for vi- sual language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.496925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.179332Z digest=sha256:6eb0cf51f3881fbec3db1c3c062c4e1de6673f4b57a6e26953ae479b7a29b693

Observation effa39d3-0c87-4fb8-82dc-ef0a4cb5c2e6 · outbound

This paper cites Visual instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.260602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.260602Z digest=sha256:aa5720365189db69c862dbd334f086f9fce939cdcff1ba32bd08b1c8ae56e5c8

Observation 84d95b6d-8393-4dd3-8645-5dce02161448 · outbound

This paper cites Improved baselines with visual instruction tuning.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.270709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.352145Z digest=sha256:c8cce54ff595e5b487ef293fffda74643f5376c353d653ab90ba6a72f436ac88

Observation f6af390d-ded8-46ac-8748-6fbbe7850e37 · outbound

This paper cites Linearly mapping from image to text space.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Linearly mapping from image to text space

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:10.054001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.431580Z digest=sha256:b5d0a2342a1ea624651fb33aaca6c1a761209ef898d86f1fe0b412d39f44cd13

Observation 92687b32-7c75-438b-814d-b08c34cffdf8 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.506978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.506978Z digest=sha256:c3d440aadaef628cf5dc0ef7383cfb55b7cc7d67284ba5c3a6b45da27f3d4ac3

Observation 847a9d7a-2e48-49a5-be4d-14db007b01ea · outbound

This paper cites Interpreting GPT: The Logit Lens.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting GPT: The Logit Lens

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.841272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.567052Z digest=sha256:e2882ba61c69c59a533422c812876e2e616d9b0a56d1c6748c08424c3d738185

Observation 5022e038-1a89-4d2d-97d9-b8f83fd8bd5a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.644453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.644453Z digest=sha256:dbf7e166e53c91072253e5d64582878d9f25ab43b08f13fd30f3cc7d7b0e238f

Observation d98f162a-bd02-4070-869c-7b659162ab52 · outbound

This paper cites Bridg- ing vision and language spaces with assignment prediction.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Bridg- ing vision and language spaces with assignment prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.654978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.743327Z digest=sha256:01e14323cf4df572d2035fdbb57f463372f13c3efb16fe4682d2789330349865

Observation ca3e0edd-32ae-4375-bd68-ed81739091d9 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.493256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.811626Z digest=sha256:736dbb28579bd6eb7cba9927369bf63763427d4fb8002cf08176415d31d3c2ec

Observation 5f3e4212-a992-4b9f-8e5d-e364113e2139 · outbound

This paper cites Multimodal neurons in pre- trained text-only transformers.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Multimodal neurons in pre- trained text-only transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.338873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:07.913318Z digest=sha256:e4a94addb8e17ed99a6216ce7948a0d52e8731313475f87e307fc882d955daab

Observation 54ecaa4c-e899-43b5-925b-20aa394b7bd5 · outbound

This paper cites Open Problems in Mechanistic Interpretability.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Open Problems in Mechanistic Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:07.977614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:07.977614Z digest=sha256:fc034170b22e701ed6250f6f2aa257ed225ed8cc2ad4cf813325528d7cb0630d

Observation 1937f5de-3784-4034-89c3-3e0a890480da · outbound

This paper cites Paligemma 2: A family of versatile vlms for transfer, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Paligemma 2: A family of versatile vlms for transfer, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:09.161200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.055607Z digest=sha256:93e26cf8f44617a7c098a78585acace4ace81f0059327e6bab2f503b84c46f56

Observation 2083da55-8b62-4185-94f9-0d851b8bdc27 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Gemma 2: Improving Open Language Models at a Practical Size

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:08.123176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:08.123176Z digest=sha256:adecac12a313889a1abb6cef410be950b48370245f4188f120c60429af37532f

Observation b9431d42-e8cf-47df-9a14-6e6a98d5e4d6 · outbound

This paper cites an unresolved cited work.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T01:06:08.998912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.199039Z digest=sha256:2c164fa29f0a94482ac81002d72e7b508eec4a2d8c764c137e5e52b0d482d665

Observation 35c4f53e-89b3-4c81-81b7-9e89d9844c1c · outbound

This paper cites Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Metamorph: Multimodal un- derstanding and generation via instruction tuning, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:08.844753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.244823Z digest=sha256:cd8a3283fdc6c04380936784d724eae0fc649d13be8e3974bddb38444b16a81b

Observation 85d98845-988b-458e-8ad4-bda971176a71 · outbound

This paper cites Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Too late to recall: The two-hop prob- lem in multimodal knowledge retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T01:06:08.657295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T01:06:08.321921Z digest=sha256:4b6ad3ad1c088eab4c3bda286aa853e9da058786abb349c376623dc4cbbf6245

Pith citing papers

Observation 7f81683f-e654-4363-ad77-03ebc2dffce9 · inbound

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention cites this paper.

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:07.449816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:07.449816Z digest=sha256:6ba0b309d5cdd27f98f129453547aaa8ff86560bce05fc28cc28c29335053c55

Observation ab9ce2f3-d682-4f24-bbce-0a77ab76cb60 · inbound

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs cites this paper.

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.126763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:02:59.403688Z digest=sha256:9774b311d6a08fc24c92ea08219cf623b59f9b331d23e147af49eca5c11fc74f

Observation 4f28d0ec-44f2-4d3f-927d-2075dbb823b4 · inbound

Pathways of Visual Information Flow in Vision-Language Models cites this paper.

Pathways of Visual Information Flow in Vision-Language Models How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T03:03:11.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T03:03:11.110105Z digest=sha256:77370702e912b41d4c71e205a6b155fb558277b5e2c2069fab5f24770cdd0991