Pith. sign in

Paper Citation Record · LEDGER

Multimodal Model Diffing for Feature Discovery and Control

As of 23 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2608.09928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09928 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:56.224644Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved74
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63a68806-aeda-41aa-8974-f329e2ebba67 · outbound

This paper cites Pixtral 12b: A new frontier in image and text understanding.

Multimodal Model Diffing for Feature Discovery and Control Pixtral 12b: A new frontier in image and text understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.734741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.734741Z digest=sha256:d691c8b7b674e02ab5ae193a1545cb2d9982b4b7c6d774fa397171b0ba5b24b3

Observation e980897a-63ad-4ab4-9cae-ab1bace9e2e9 · outbound

This paper cites Golden gate Claude.

Multimodal Model Diffing for Feature Discovery and Control Golden gate Claude

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.740439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.740439Z digest=sha256:95e71f5a748c2b887e962425565dbf658f6422f96d3b7b698f6c0301fde1c5b7

Observation b478f42b-d307-42c5-b16b-5954735aa5a1 · outbound

This paper cites SAE on activation differences.

Multimodal Model Diffing for Feature Discovery and Control SAE on activation differences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.744964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.744964Z digest=sha256:5ab1432ccbcc52d9698a67f3a3f783fe902bdf815df0d84330042348c8104a9e

Observation eed6e3f5-9874-4c87-b480-11f0d9fdb4ad · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Multimodal Model Diffing for Feature Discovery and Control Refusal in Language Models Is Mediated by a Single Direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.749442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.749442Z digest=sha256:9d46f132aa5ee70cfbaa9e1dfa28aa4e79f6b41ceba555658ac66a31388347b3

Observation 15f9a23f-fb05-4140-b13b-89fcd2bd4712 · outbound

This paper cites Revisiting model stitching to compare neural representations.Advances in neural information processing systems, 34:225–236, 2021.

Multimodal Model Diffing for Feature Discovery and Control Revisiting model stitching to compare neural representations.Advances in neural information processing systems, 34:225–236, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.754302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.754302Z digest=sha256:ff819da89888767f3412827b0de2f7071995be17c40a23ec5d754b268338322a

Observation 99bad67a-8f11-4798-a8d3-1e4ee9868bcc · outbound

This paper cites Representation Topology Divergence: A Method for Comparing Neural Network Representations.

Multimodal Model Diffing for Feature Discovery and Control Representation Topology Divergence: A Method for Comparing Neural Network Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.758852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.758852Z digest=sha256:f13b664da42810835312ab144f626b93360d95cd075cec4c63dbdc67ff2d8f97

Observation fb3300c3-6ac8-4870-bef9-0d7a1e819f96 · outbound

This paper cites Understanding information storage and transfer in multi-modal large language models.Advances in Neural Information Processing Systems, 37:7400–7426, 2024.

Multimodal Model Diffing for Feature Discovery and Control Understanding information storage and transfer in multi-modal large language models.Advances in Neural Information Processing Systems, 37:7400–7426, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.764254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.764254Z digest=sha256:47bfe19eb1f34c10524f1e2e9431ab1b5e99b9855047748a70ad910a870880ad

Observation 775b35d7-5787-42c1-b28c-727175dec30e · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023.

Multimodal Model Diffing for Feature Discovery and Control Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.768697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.768697Z digest=sha256:19609c09cc6bcce61c932a2c3c2f25a32404fde399f2e44f20e032a7fab6ce94

Observation bfb2af8f-ac38-41fa-a848-228a761cca3a · outbound

This paper cites Stage-wise model diffing.

Multimodal Model Diffing for Feature Discovery and Control Stage-wise model diffing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.773486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.773486Z digest=sha256:e534159b18276ef79ddfdb617bface5a420927f8c7d88377d1bcf40aac90bac1

Observation 96afae2d-fab5-435d-bf48-229175c2211e · outbound

This paper cites Observing and controlling features in vision-language-action models.arXiv preprint arXiv:2603.05487, 2026.

Multimodal Model Diffing for Feature Discovery and Control Observing and controlling features in vision-language-action models.arXiv preprint arXiv:2603.05487, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.778192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.778192Z digest=sha256:ca8c07564c777ba2afed4f7dbcaef1f072f80e96fb2ddebbcd1bfdcad0ad9a5f

Observation bae5706b-c4c3-4713-829f-3cfd8479df19 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Multimodal Model Diffing for Feature Discovery and Control Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.782849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.782849Z digest=sha256:f8986227b91305e39d824c146a9d5b15a1399451cb2444e639ed5ce8cdec44a5

Observation 21aef9b8-39d4-4a4a-bb5e-514a373a0a0d · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Multimodal Model Diffing for Feature Discovery and Control Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.788114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.788114Z digest=sha256:7650b710942703198e6cad92c35404d6259dc5d364f2c314e11999272fa6405f

Observation 943068ae-9db0-4d97-a492-eea47ed36f53 · outbound

This paper cites Interpreting and Controlling Vision Foundation Models via Text Explanations.

Multimodal Model Diffing for Feature Discovery and Control Interpreting and Controlling Vision Foundation Models via Text Explanations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.793237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.793237Z digest=sha256:d1ad4a0a09075e812e3fd631c79ea5264232bb4ecc3ee609407191ec3c585ffd

Observation 729e283c-fe58-406d-a371-03d8e94288e4 · outbound

This paper cites LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.798668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.798668Z digest=sha256:3c3986a1a8a74d00f1152f726bd24e02f693f12e79e6860d63e6d4831af79871

Observation 47c9ed4e-1705-4432-9bf6-b809e4061e02 · outbound

This paper cites Explaining How Visual, Textual and Multimodal Encoders Share Concepts.

Multimodal Model Diffing for Feature Discovery and Control Explaining How Visual, Textual and Multimodal Encoders Share Concepts

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:17:57.283192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:55.803988Z digest=sha256:631263caa9a2b5226cb0568ea65a14267df85dfba023add77a67661fc559e862

Observation 9ea49434-9b1f-4ea9-b45e-0b22dc3bbf00 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Multimodal Model Diffing for Feature Discovery and Control Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.809262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.809262Z digest=sha256:f5209ae4f19995f3a221eb7176a1cd088823f2b7598fae0cb4a3fb91bdcb52ec

Observation ef347515-df97-4217-990c-bac9c5c5f962 · outbound

This paper cites Case study: Interpreting, manipulating, and controlling CLIP with sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Case study: Interpreting, manipulating, and controlling CLIP with sparse autoencoders

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.814349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.814349Z digest=sha256:1f4f063b247717f2ec83792a5cd714a4fb27189af4209b6e7f33a8be0c3c0d09

Observation c03b7f1f-509a-432a-bf1a-565eab037242 · outbound

This paper cites Toy Models of Superposition.

Multimodal Model Diffing for Feature Discovery and Control Toy Models of Superposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.819230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.819230Z digest=sha256:e5f1a23e55e84f7c2b0f77c2c2f0747e079ef323c71070274a37f965388a9b63

Observation 907a2509-8501-4d4e-b1ce-4313ef05a681 · outbound

This paper cites Why does unsupervised pre-training help deep learning? 11:625–660, March.

Multimodal Model Diffing for Feature Discovery and Control Why does unsupervised pre-training help deep learning? 11:625–660, March

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.824530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.824530Z digest=sha256:9aa46a92772869ace9814211e64f0a1b2856fd71d40a7550b8ecd80fac96c1f4

Observation 5dc90ba2-6e4d-4bbf-98b4-ace19979d697 · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

Multimodal Model Diffing for Feature Discovery and Control Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.829626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.829626Z digest=sha256:6f3c5ad22f953dab5aa8829792cd4e5504474c72cf025c8bdd2833823c0ccbed

Observation 4227ee35-1ef9-4683-9e9f-b03e81994704 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Scaling and evaluating sparse autoencoders

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.834643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.834643Z digest=sha256:20f8aab79887b51d3b6ddd5cf4e3772fccb0d76eddb855b230adab621cf0cb67

Observation a67d0a03-cb2c-48bc-99b6-5077d2ca94a8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Multimodal Model Diffing for Feature Discovery and Control Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.839825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.839825Z digest=sha256:b560f1970b1242b8333e0e114288532e41cade364e8a67281317d2d5b3e15518

Observation 2891bc5a-3300-4926-b5a2-2e673fc4e1d2 · outbound

This paper cites FigStep: Jailbreaking large vision-language models via typographic visual prompts.

Multimodal Model Diffing for Feature Discovery and Control FigStep: Jailbreaking large vision-language models via typographic visual prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.845475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.845475Z digest=sha256:c2355e3ad205aab6bde48eee1a239bb8d550785447df7978019224a7cd4f2e6e

Observation eb7628dc-8244-471c-97d8-776e74cf4700 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Multimodal Model Diffing for Feature Discovery and Control Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.850637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.850637Z digest=sha256:7fc8eb1906aaffb2f90228bdd5d24cae1132bf0f860b03bcaa796a6111daf603

Observation 383f8bd4-0f3d-4f1e-8663-f5c8552e8985 · outbound

This paper cites Not all features are created equal: A mechanistic study of vision-language-action models.

Multimodal Model Diffing for Feature Discovery and Control Not all features are created equal: A mechanistic study of vision-language-action models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.856252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.856252Z digest=sha256:ed02d523907f4f803341759b1a6e97429d8da22830ea45618205feb75d91096a

Observation 594550ea-48d5-4b18-a375-73064b1a6c2e · outbound

This paper cites The Llama 3 Herd of Models.

Multimodal Model Diffing for Feature Discovery and Control The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.861229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.861229Z digest=sha256:0602505c829a1ab785b0f809c4bc536f0c9ec9a848b10c0e8148007fdf63de86

Observation 51b802d7-af4b-4cde-979e-97fc6d7c7e49 · outbound

This paper cites Mechanistic interpretability for steering vision-language-action models.

Multimodal Model Diffing for Feature Discovery and Control Mechanistic interpretability for steering vision-language-action models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.865869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.865869Z digest=sha256:df6a7da0ee6fc30d10ce715e168afe01ef3486b80581f61b465e4c75dfc0666c

Observation 3b640386-0590-4017-8b10-8dcab1ed9d2b · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.870806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.870806Z digest=sha256:38ccec8487223871f034e559af25df60b59c0a9e6be4beeb820999b392ffc048

Observation e94666bd-7d15-4980-b0a4-2c789fe46bb3 · outbound

This paper cites In-context learning creates task vectors.

Multimodal Model Diffing for Feature Discovery and Control In-context learning creates task vectors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.875585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.875585Z digest=sha256:bc0e5da925574e432a15331a05cebb4e473e52bd4f6e0a7a6aea283a3e24c0cd

Observation ed622c66-b2d4-4498-8016-67dc19488d7e · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

Multimodal Model Diffing for Feature Discovery and Control VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.880064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.880064Z digest=sha256:33f5fdc05ab61bdc9de379bcf1c5a742d769cca45f75c6c7921881009d47145b

Observation 92e3a5d3-924f-4d22-961c-de23b2368db0 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Multimodal Model Diffing for Feature Discovery and Control Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.885168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.885168Z digest=sha256:1c397e4c9479d6648cfb4766e1378a19c345e39d04f98f083849079bac273273

Observation 60bfe481-7159-4238-b912-6accdf089da1 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

Multimodal Model Diffing for Feature Discovery and Control Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.889674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.889674Z digest=sha256:1cc9c5e5d26ce970dd4e7dcbf92260595efedc2af8728ec1481d4d42f39f97f0

Observation dd219166-8d27-475a-a1c4-06ce11c66684 · outbound

This paper cites A “diff” tool for AI: Finding behavioral differences in new models.

Multimodal Model Diffing for Feature Discovery and Control A “diff” tool for AI: Finding behavioral differences in new models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.894475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.894475Z digest=sha256:56ecbc4db4a50b8ba4241ec6745e6a86f74f907735d590fe02e1ff5f2b44c892

Observation 7bdee4c9-2165-49f5-a2c8-8a114e2a8644 · outbound

This paper cites Bridging the VLM and mech interp communities for multimodal interpretability.

Multimodal Model Diffing for Feature Discovery and Control Bridging the VLM and mech interp communities for multimodal interpretability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.900408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.900408Z digest=sha256:81a81da2595f4e08d2ac5ed267799f3f86edde214b0895d5183f5212951cedf8

Observation 0a1b3935-db61-4b7d-b663-d9f6cdf18b9d · outbound

This paper cites Steering CLIP's vision transformer with sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Steering CLIP's vision transformer with sparse autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.905911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.905911Z digest=sha256:82f303c96fc3f08fb1d98d0cc293feb1a26443b5e9f6d450519164803a3c9458

Observation 02d7cbe7-c662-4f20-a2f1-8f252dca6f28 · outbound

This paper cites Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video.

Multimodal Model Diffing for Feature Discovery and Control Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.911157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.911157Z digest=sha256:9ad689a6e6f192e0f1c2c92e372957e0e0303c817dbaebc61789c439618da4f6

Observation 4e56b9bf-2caf-4a64-8b63-4fffadd4fa48 · outbound

This paper cites Analyzing Finetuning Representation Shift for Multimodal LLMs Steering.

Multimodal Model Diffing for Feature Discovery and Control Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.916191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.916191Z digest=sha256:e2d779f947a1d2891a5727192f2a294f7a6688c174b2af052b172eec2bc29d56

Observation bb729be5-e31f-4c6f-95d7-d9726ae0138e · outbound

This paper cites Saes (usually) transfer between base and chat models.

Multimodal Model Diffing for Feature Discovery and Control Saes (usually) transfer between base and chat models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.921070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.921070Z digest=sha256:dcbaf4cd6e68e5a517c7d22cca7014f396c78ad33e288c4dd25825cec004f206

Observation 16f919f1-5867-4ed1-b1a8-49dea9800dbf · outbound

This paper cites Similarity of neural network representations revisited.

Multimodal Model Diffing for Feature Discovery and Control Similarity of neural network representations revisited

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.926358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.926358Z digest=sha256:55512f2020d36c710ce2f6ba8029307ce7c7da9eb66d3bf7f96fada990e67b93

Observation 07a5bec2-b757-4592-8d0b-4eccf85b6d89 · outbound

This paper cites Sakla, and Kowshik Thopalli.

Multimodal Model Diffing for Feature Discovery and Control Sakla, and Kowshik Thopalli

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.931559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.931559Z digest=sha256:0ee43a2f2e563962819032c66f71ceabb082ce64ce443fc7b63ff6b11b80bafc

Observation 62157067-b288-4ad7-97cb-c5b4e1633c98 · outbound

This paper cites Understanding image representations by measuring their equivariance and equivalence.

Multimodal Model Diffing for Feature Discovery and Control Understanding image representations by measuring their equivariance and equivalence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.936739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.936739Z digest=sha256:8c1270792ea2f522bd91756306db51778216351f6a1fe75f2f20a60acb784f1c

Observation ef4f6fd3-9c05-4ffc-8be7-8a071882afe6 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.941614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.941614Z digest=sha256:b63425acc06356aeb8a66649895a9e70068fcd0e4c95d7eb0d3792d14a40c6d3

Observation ae34dab8-af93-4449-8e3a-061503ffd727 · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Multimodal Model Diffing for Feature Discovery and Control Inference- time intervention: Eliciting truthful answers from a language model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.946783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.946783Z digest=sha256:966be40db0aa0096875a4e498ba96a37394b453f9454bb573c203a46483659c4

Observation f66078d7-2322-48be-8ac4-5c1c808c4191 · outbound

This paper cites Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.969271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:55.951510Z digest=sha256:1cee39d2d55fced9a5ba73083f84279a4861e51d0d0be9fb57cfb7d502fb069a

Observation 7c5cf49e-69d0-46fc-b00c-d062e287ae2c · outbound

This paper cites Convergent Learning: Do different neural networks learn the same representations?.

Multimodal Model Diffing for Feature Discovery and Control Convergent Learning: Do different neural networks learn the same representations?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.956223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.956223Z digest=sha256:389339177c7c009c4a631fea2512a2838b2b6f03572ed7fb4f4791c9872d4731

Observation b362ad44-4a2b-4a9d-9394-b8618f4cc42d · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Multimodal Model Diffing for Feature Discovery and Control Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.961255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.961255Z digest=sha256:053e373b25d448420c81fc3611e4bff0b34b90d8a2e05d8e551be906f0261df7

Observation f5eddaa7-8d8b-40f1-9231-425afcca5f9e · outbound

This paper cites Sparse autoencoders reveal selective remapping of visual concepts during adaptation.

Multimodal Model Diffing for Feature Discovery and Control Sparse autoencoders reveal selective remapping of visual concepts during adaptation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.967019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.967019Z digest=sha256:75096574ee8abd8ec5b90aa42a01de987fbfb50b0a9365d585d8e714a4d47bdc

Observation 74cb08c5-3e38-4f66-b5a3-cf3e6d9061f1 · outbound

This paper cites A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models.

Multimodal Model Diffing for Feature Discovery and Control A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.972218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.972218Z digest=sha256:cb8317b1eba512ae284213776fef042bed792334b4cef0f87c6155c59ef895b7

Observation 50c39459-8359-4d8a-8f3b-06a5764fccdb · outbound

This paper cites Sparse crosscoders for cross-layer features and model diffing, October 25 2024.

Multimodal Model Diffing for Feature Discovery and Control Sparse crosscoders for cross-layer features and model diffing, October 25 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.952495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:55.977277Z digest=sha256:04f14f347088e49d8bebb5976e1da7fbf58f01687ec2b63d4cfcb3d5d897ecbe

Observation e242e417-d6d1-466c-8c67-75d87eb17f48 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023.

Multimodal Model Diffing for Feature Discovery and Control Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.934987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:55.981926Z digest=sha256:3b5c4378fb72bb18b826bf0d1d29201b53aebf06fe15cbd494feec0a0b2fb0a8

Observation 428d81df-84f7-4297-95c6-84ba78dfca31 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Multimodal Model Diffing for Feature Discovery and Control Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.986160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.986160Z digest=sha256:c899229c486163cd9887cc8f4996d33ce40a26df1afe674013e7232b7ec77138

Observation 159ce224-eb30-4005-bd03-3e7034b33e31 · outbound

This paper cites Improved baselines with visual instruction tuning.

Multimodal Model Diffing for Feature Discovery and Control Improved baselines with visual instruction tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.990284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.990284Z digest=sha256:35aea3edb37c03b423c55ce728a356165361e324015373bc90ecf1a1bd6f71fe

Observation a8f7a635-04af-41c1-9058-446d263e58fb · outbound

This paper cites MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.897139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:55.994689Z digest=sha256:5bda38b2af377dfba9b510f36ce441e9681d7bce889873c7a080b317609884c0

Observation 4cc32df0-f6f3-4a06-9080-8ad98d0fa6c4 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Multimodal Model Diffing for Feature Discovery and Control OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.999133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.999133Z digest=sha256:2d0501350a313beea2fec6949cbaf39e78b2423c2fe410c22aba35bfc0c9ca5f

Observation 2db28554-8e14-41be-8bc8-fc9bd7333573 · outbound

This paper cites Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller.

Multimodal Model Diffing for Feature Discovery and Control Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.880397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.003768Z digest=sha256:ba5c37f709d7acae0bf8f7395ee27ea6e30355c7e7f8fdaf352f8ee1fb42a2ba

Observation 1012ade3-0d1d-4240-a17c-094795998f39 · outbound

This paper cites Locating and editing factual associations in GPT.

Multimodal Model Diffing for Feature Discovery and Control Locating and editing factual associations in GPT

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.008128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.008128Z digest=sha256:88f8bd70c15cf3a273fd8e37fa3b35453388ad79e2f1b6ceabc78b34878dd920

Observation 464d7b55-d456-457f-86c9-7b103017259e · outbound

This paper cites Robustly identifying concepts introduced during chat fine-tuning using crosscoders.arXiv preprint arXiv:2504.02922, 2025.

Multimodal Model Diffing for Feature Discovery and Control Robustly identifying concepts introduced during chat fine-tuning using crosscoders.arXiv preprint arXiv:2504.02922, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.012960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.012960Z digest=sha256:2d7d1deda2ea54782facfb3c36bfce1b7851fc5df7dd86e7cb1649460c1725a8

Observation cc5bc455-3eba-42e8-901f-b572d332ca4a · outbound

This paper cites What we learned trying to diff base and chat models (and why it matters).LessWrong, 2025.

Multimodal Model Diffing for Feature Discovery and Control What we learned trying to diff base and chat models (and why it matters).LessWrong, 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.853986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.017719Z digest=sha256:eee3c45ef784ca0732e5e087b12b05237ab2754bd819f310d098c1eec1a7f83f

Observation 8451658b-066d-4558-848c-483cfff8b398 · outbound

This paper cites Insights on crosscoder model diffing.

Multimodal Model Diffing for Feature Discovery and Control Insights on crosscoder model diffing

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.837250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.022501Z digest=sha256:a058122a66e588d940fbe373d13971b2ca3900d67aca36a50a6c9e0ac32515a3

Observation 5db2cfa0-3178-48d0-9c2f-245259e280e8 · outbound

This paper cites Attribution patching: Activation patching at industrial scale.

Multimodal Model Diffing for Feature Discovery and Control Attribution patching: Activation patching at industrial scale

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.818307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.027425Z digest=sha256:bc2f1a8d22d6cbcc0cbbc458269b173eb2f245d3f6998fb4c4821208b6565f5f

Observation cdd46c79-02f1-4c4e-86bb-b856ba020424 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Multimodal Model Diffing for Feature Discovery and Control Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.032176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.032176Z digest=sha256:1a4a32f27ac8d95a2248a1aff37edcc9f425b74be765ad59541e0334c93c1a48

Observation 623a2370-84b8-4f4a-bf0b-0b8c4a7bd95c · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Steering Language Model Refusal with Sparse Autoencoders

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.038373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.038373Z digest=sha256:8f08418512281577feb9daa3194697c3ca17912f95725ff66c9ede32e6127bac

Observation bbc722ed-96ce-4a69-89d5-624cf509d697 · outbound

This paper cites Zoom in: An introduction to circuits.Distill, 2020.

Multimodal Model Diffing for Feature Discovery and Control Zoom in: An introduction to circuits.Distill, 2020

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.044310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.044310Z digest=sha256:ce1dadc5c1573ae35a9752979dc796c1bb60d116d96106018d0b3d1065c5c8ad

Observation 058d2038-410f-42cd-9c2b-798bf255d421 · outbound

This paper cites Visualizing representations: Deep learning and human beings.

Multimodal Model Diffing for Feature Discovery and Control Visualizing representations: Deep learning and human beings

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.798542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.049202Z digest=sha256:691576b5636dd3f5266a736365e50006479be02e0c1ecb000dff512efaddf2e2

Observation 60c5603e-ea1d-40e2-b130-362652456e2b · outbound

This paper cites Probing the representational power of sparse autoencoders in vision models.

Multimodal Model Diffing for Feature Discovery and Control Probing the representational power of sparse autoencoders in vision models

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-11T04:17:56.747243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.054083Z digest=sha256:8a9db7f8d44c6930fb6dca50fb8b66d3c14665a5cec41ea0d552a735912a2d36

Observation 5b9793e9-4c4a-4e7c-862e-920e5d638bbe · outbound

This paper cites Gpt-4o-mini: Advancing cost-efficient intelligence.

Multimodal Model Diffing for Feature Discovery and Control Gpt-4o-mini: Advancing cost-efficient intelligence

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.782692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.058945Z digest=sha256:aa6daafbddd593cc10370d2aa95ada02e794be088118fc9290f40565655d112d

Observation 377960dc-cb41-4e38-a7f4-d2932749b680 · outbound

This paper cites Sparse autoencoders learn monosemantic features in vision-language models.

Multimodal Model Diffing for Feature Discovery and Control Sparse autoencoders learn monosemantic features in vision-language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.063731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.063731Z digest=sha256:53a83fce2ae029f8d80cfaed3d8ebb2ba498c90e1104cf31a044cba37cd90e94

Observation 1c619101-97ef-42f1-aae8-e70c8d36057a · outbound

This paper cites Towards vision-language mechanistic interpretability: A causal tracing tool for blip.

Multimodal Model Diffing for Feature Discovery and Control Towards vision-language mechanistic interpretability: A causal tracing tool for blip

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.766065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.068462Z digest=sha256:aae3ec25aff1e8557d75c0b19c7c6d22ec861abb808636c8fc9bae9dfcb72292

Observation 94b2dead-18bc-4a07-af33-a10015ae3873 · outbound

This paper cites Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal.

Multimodal Model Diffing for Feature Discovery and Control Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.073278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.073278Z digest=sha256:b65262f2dea7452a6313f7c46eb822feabbd2a2e799f8401603a0742deb8e068

Observation eacb6670-9287-4508-b7cc-10e65ab0b142 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

Multimodal Model Diffing for Feature Discovery and Control Visual adversarial examples jailbreak aligned large language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.750674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.078497Z digest=sha256:b6776d38044a43668852f739516e573b0e019f50000ccce3b6a24e1ac7985729

Observation d5566a4c-bbc4-4029-82c9-0268f9e08351 · outbound

This paper cites Qwen-Scope: An open sparse autoencoder suite for the Qwen model family.

Multimodal Model Diffing for Feature Discovery and Control Qwen-Scope: An open sparse autoencoder suite for the Qwen model family

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.733675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.083263Z digest=sha256:70857fd838895f0ddbbdd57c14b9fcdbbe449def24b56901df09053fa58b1569

Observation fe447d69-874d-4cbd-b5a0-ea082f32f82e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal Model Diffing for Feature Discovery and Control Learning transferable visual models from natural language supervision

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.088107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.088107Z digest=sha256:74815c44128a4b9a8ffe50164d2fece4b4c607bcc45b0871a94a3d33d1a34939

Observation eadbd835-4e5f-45e3-8fea-de3c5e559122 · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.092721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.092721Z digest=sha256:f295b25bf82655cd6c94a62e4e3f0f0d30f1fe930ff1ffc879dca29b118633bb

Observation 4641de98-7e7b-4fc5-a4ab-6302312a8412 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Multimodal Model Diffing for Feature Discovery and Control Steering Llama 2 via Contrastive Activation Addition

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.097463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.097463Z digest=sha256:be19734cec05fb1264b398642de762c45408836871e118c55784dbac48c4cf8e

Observation 8c892395-c64d-4597-88bb-e01c606b2917 · outbound

This paper cites Multi- modal neurons in pretrained text-only transformers.

Multimodal Model Diffing for Feature Discovery and Control Multi- modal neurons in pretrained text-only transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.705410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.102224Z digest=sha256:99bbea0f6a21c1402006f383ec6a3f2f72c7ad3eaf23149b890e366f0dd0d2cb

Observation f298f62e-2df4-4676-8705-f70f9d84fe17 · outbound

This paper cites SteerVLM: Robust model control through lightweight activation steering for vision language models.

Multimodal Model Diffing for Feature Discovery and Control SteerVLM: Robust model control through lightweight activation steering for vision language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.688536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.106645Z digest=sha256:c0c90e950044b595941e88a69c63986b1c607e87370b57b39877faec5dee52b5

Observation 830f7044-5311-48c7-b4a4-e492efe747ce · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

Multimodal Model Diffing for Feature Discovery and Control LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.111721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.111721Z digest=sha256:55b1b59a431d86be9effeb5b44a41583e1659b75ca2061b76a8b0ff970ee7bfb

Observation ea2aabab-eee9-460d-ba2a-2b5538ff20e5 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Multimodal Model Diffing for Feature Discovery and Control PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.116317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.116317Z digest=sha256:0e63ba770a2393ff82d8c07fec6a95a7616ccad113f44d1cb4bad61371dd0739

Observation 652c6242-0c83-45b3-84c8-a17a7d903b88 · outbound

This paper cites Daniel Freeman, Theodore R.

Multimodal Model Diffing for Feature Discovery and Control Daniel Freeman, Theodore R

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.122564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.122564Z digest=sha256:281a605a2a0e78fe80a3fa4922676452348cb0cde4456d02b59bab4908aa022d

Observation ad666547-3921-415d-9f0d-0f41acd4ac5e · outbound

This paper cites Li, Arnab Sen Sharma, Aaron Mueller, Byron C.

Multimodal Model Diffing for Feature Discovery and Control Li, Arnab Sen Sharma, Aaron Mueller, Byron C

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.659875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.127927Z digest=sha256:1f19522b225987391763a40dcd0ef2775aa4a6eb11d43896296d2981780e0187

Observation d37558b3-dc39-493d-bf70-b84cb8eedef5 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Multimodal Model Diffing for Feature Discovery and Control Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.132540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.132540Z digest=sha256:9bd619b7dffadd79fe36d0ea373f51b2ff29d7768d523a7466f347fa249abcdc

Observation 4f30b446-4ab4-4a1a-a778-10e1ba21a3ed · outbound

This paper cites Steering Language Models With Activation Engineering.

Multimodal Model Diffing for Feature Discovery and Control Steering Language Models With Activation Engineering

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.137372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.137372Z digest=sha256:1c04c5b9f0b1a0215c2f3406f099efb91573aac3f2d126ed855d575248148ad7

Observation a5e3ee40-f7fd-46f4-9da1-684f3115af2c · outbound

This paper cites Too late to recall: The two-hop problem in multimodal knowledge retrieval.

Multimodal Model Diffing for Feature Discovery and Control Too late to recall: The two-hop problem in multimodal knowledge retrieval

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.628662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.142365Z digest=sha256:a3c65a0f817918c64614dc78a8e2793bf83dbae9f2be8902a670203497f35d6a

Observation e97f989c-c33a-4426-8979-5210adbd71a5 · outbound

This paper cites How Visual Representations Map to Language Feature Space in Multimodal LLMs.

Multimodal Model Diffing for Feature Discovery and Control How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.147155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.147155Z digest=sha256:4c34cd16dcf8edd5f4991c1935ab9932790b1172d70c050cf5a685d4a73295ff

Observation 9ddfa148-7c15-4249-9c65-72225f1a76bd · outbound

This paper cites Steering away from harm: An adaptive approach to defending vision language model against jailbreaks.

Multimodal Model Diffing for Feature Discovery and Control Steering away from harm: An adaptive approach to defending vision language model against jailbreaks

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.611779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.152172Z digest=sha256:c04160ae266340b260283926cdaca2fb86d14c6c03c0add9e3d5ccacd962a96b

Observation 3ad407f8-f160-46c8-9936-e40a5904103c · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Multimodal Model Diffing for Feature Discovery and Control InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.156829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.156829Z digest=sha256:37b95cbb7a963e3e8e847c1fde42a9c8c5edb595da56b1a30abc1f77492be62d

Observation 3b14a2e9-6053-4aa2-9b6a-279769fe4d46 · outbound

This paper cites AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting.

Multimodal Model Diffing for Feature Discovery and Control AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.595098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.162853Z digest=sha256:58a7238a4eda1a7e487e22a721b986f2e64cf5622705073776062dfd667ab828

Observation 47a8435a-2c41-4f4a-81e8-3c19919d4f69 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.167527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.167527Z digest=sha256:0564cbb667e7ada55bced3813242357f46ca80bed77332a5086b340cfe7c345e

Observation 95f2b71b-43ef-4498-9931-0584ebf2d855 · outbound

This paper cites Qwen3 Technical Report.

Multimodal Model Diffing for Feature Discovery and Control Qwen3 Technical Report

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.172320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.172320Z digest=sha256:7537782a34231241b7850d4bc5f11bc6710e8e73205b3501d8d4894d0928b577

Observation 581e9dcf-2c1f-4833-b068-0610fedbe8bf · outbound

This paper cites SafeSteer: Adaptive subspace steering for efficient jailbreak defense in vision-language models.arXiv preprint arXiv:2509.21400, 2025.

Multimodal Model Diffing for Feature Discovery and Control SafeSteer: Adaptive subspace steering for efficient jailbreak defense in vision-language models.arXiv preprint arXiv:2509.21400, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.177100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.177100Z digest=sha256:944e8b4b0246fbf694a006ebc5d60a4e4f1580fc00cd2b446d61428c34d5db3b

Observation c093c9d3-d5cc-4865-8131-636221f4691f · outbound

This paper cites Sigmoid loss for language image pre-training.

Multimodal Model Diffing for Feature Discovery and Control Sigmoid loss for language image pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.181619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.181619Z digest=sha256:e8291b4a933c454b48b3d8f040e67949b904de120640253a94898f9340267f85

Observation 112922d6-827d-42aa-a3d2-3aea70380e8f · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Multimodal Model Diffing for Feature Discovery and Control Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.186065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.186065Z digest=sha256:df64586b832aea41fa1b6828b7721a19de2f40820c1bea4516a1a68a36532694

Observation f2b68228-1a9b-4451-b843-561f9a915118 · outbound

This paper cites Cross-modal information flow in multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control Cross-modal information flow in multimodal large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.566959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.190937Z digest=sha256:cb526bdc2c4d0b95575b4ee8a18f19f870bf25c6cdbdd5c41d26aa5f43930783

Observation 2619e432-0229-4ce4-9384-2c2445bcae20 · outbound

This paper cites Multimodal situational safety.

Multimodal Model Diffing for Feature Discovery and Control Multimodal situational safety

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.549778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.195809Z digest=sha256:635ab191ef3e6d7f5c795ddb9dbc7f9fa07e8dc3379f3f92a395a08178cfea7e

Observation e7d95dd4-96b3-42a8-bfa2-b16e8a5e4b3e · outbound

This paper cites Relocated.

Multimodal Model Diffing for Feature Discovery and Control Relocated

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:17:57.530047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.201760Z digest=sha256:1d342155497feb0f22612944c0912b75e85db3e9994838025dac36cdc99fe66e

Observation dd40ce95-4dfb-4b0e-bb9f-d4bfd52a465d · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.512239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.208372Z digest=sha256:4783b5901544bf101722474b8044aebf1c9430fc151b91681017356ecb747fb1

Observation 7bbbfb87-1a52-4a94-8530-dc262d50d734 · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.494613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.214057Z digest=sha256:00dcf082497cc2ddc4274060ace4b8113533ecd2ad1403f26bb2cc8821a28a7f

Observation 418b9983-293a-4b05-83aa-d201d05c83b7 · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.478120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.219618Z digest=sha256:16048c662e72d6f0719a7db12b3a2b06bb8aef453d4191457113de1f200c088f

Observation 4447cbf9-6ef0-4020-9860-8acf0b877876 · outbound

This paper cites this neuron activates for.

Multimodal Model Diffing for Feature Discovery and Control this neuron activates for

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.462367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T04:17:56.224644Z digest=sha256:9ba8a44dac7592898e23cf54aa1eac4e90102d9593d53d835a99a0ed09da88f8

Pith citing papers

No inbound Pith citation observations are available.