Pith. sign in

Paper Citation Record · LEDGER

Multimodal Model Diffing for Feature Discovery and Control

As of 11 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2608.09928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09928 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:56.224644Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved74
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63a68806-aeda-41aa-8974-f329e2ebba67 · outbound

This paper cites Pixtral 12b: A new frontier in image and text understanding.

Multimodal Model Diffing for Feature Discovery and Control Pixtral 12b: A new frontier in image and text understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.734741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.734741Z digest=sha256:f9e4752a872fe346ba31aac0799f1ed25f57c8ebc6b5d0d70c507a082c8ea57b

Observation e980897a-63ad-4ab4-9cae-ab1bace9e2e9 · outbound

This paper cites Golden gate Claude.

Multimodal Model Diffing for Feature Discovery and Control Golden gate Claude

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.740439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.740439Z digest=sha256:a591a83f0ba72ec7d89fbd0b2cd9338fe550a43332e69956790d3eb4b6242279

Observation b478f42b-d307-42c5-b16b-5954735aa5a1 · outbound

This paper cites SAE on activation differences.

Multimodal Model Diffing for Feature Discovery and Control SAE on activation differences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.744964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.744964Z digest=sha256:c97e7041caea96f3410f3f355e941860ae93504fd441e2c9fdfbaca8cf20a792

Observation eed6e3f5-9874-4c87-b480-11f0d9fdb4ad · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Multimodal Model Diffing for Feature Discovery and Control Refusal in Language Models Is Mediated by a Single Direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.749442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.749442Z digest=sha256:5fad992f7dd5a34bcbe89fc7544146fa22a5824050c08d199cc81dc19dc6a352

Observation 15f9a23f-fb05-4140-b13b-89fcd2bd4712 · outbound

This paper cites Revisiting model stitching to compare neural representations.Advances in neural information processing systems, 34:225–236, 2021.

Multimodal Model Diffing for Feature Discovery and Control Revisiting model stitching to compare neural representations.Advances in neural information processing systems, 34:225–236, 2021

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.754302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.754302Z digest=sha256:e03c970fcb633c5bdf31ac046bb4b9bc65095fc79800812c9e127e34a6baae5c

Observation 99bad67a-8f11-4798-a8d3-1e4ee9868bcc · outbound

This paper cites Representation Topology Divergence: A Method for Comparing Neural Network Representations.

Multimodal Model Diffing for Feature Discovery and Control Representation Topology Divergence: A Method for Comparing Neural Network Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.758852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.758852Z digest=sha256:d82a6a0f71040d6b2b449063e952fe9faa7d4119588d345cb2e263c1cdf9f819

Observation fb3300c3-6ac8-4870-bef9-0d7a1e819f96 · outbound

This paper cites Understanding information storage and transfer in multi-modal large language models.Advances in Neural Information Processing Systems, 37:7400–7426, 2024.

Multimodal Model Diffing for Feature Discovery and Control Understanding information storage and transfer in multi-modal large language models.Advances in Neural Information Processing Systems, 37:7400–7426, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.764254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.764254Z digest=sha256:737c1803b5f8d8b6445758c9d827d1668177cf3ae973723603c0e94b8b6499b8

Observation 775b35d7-5787-42c1-b28c-727175dec30e · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023.

Multimodal Model Diffing for Feature Discovery and Control Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.768697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.768697Z digest=sha256:e8bb81501294f4fdab2990c569caa6f7290e9315cb267aa9b3aeba72c0f732c1

Observation bfb2af8f-ac38-41fa-a848-228a761cca3a · outbound

This paper cites Stage-wise model diffing.

Multimodal Model Diffing for Feature Discovery and Control Stage-wise model diffing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.773486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.773486Z digest=sha256:05c10b83b8fcd32958d70638e9ec0bc345936122daca168f7b56aa855a35b744

Observation 96afae2d-fab5-435d-bf48-229175c2211e · outbound

This paper cites Observing and controlling features in vision-language-action models.arXiv preprint arXiv:2603.05487, 2026.

Multimodal Model Diffing for Feature Discovery and Control Observing and controlling features in vision-language-action models.arXiv preprint arXiv:2603.05487, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.778192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.778192Z digest=sha256:e22e08cee684f7a388e1cd9d04b08d4009f0c300609937dfb75ac1284c44fb57

Observation bae5706b-c4c3-4713-829f-3cfd8479df19 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Multimodal Model Diffing for Feature Discovery and Control Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.782849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.782849Z digest=sha256:5ef7525ee3e24e6fcd3169ebf2ce7a60795e5ec41a875df7a47fe647524793f0

Observation 21aef9b8-39d4-4a4a-bb5e-514a373a0a0d · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Multimodal Model Diffing for Feature Discovery and Control Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.788114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.788114Z digest=sha256:9dd157b91bf15cf73eebb93607c8c735384e3222f54a4a873f3c02e75a76cba5

Observation 943068ae-9db0-4d97-a492-eea47ed36f53 · outbound

This paper cites Interpreting and Controlling Vision Foundation Models via Text Explanations.

Multimodal Model Diffing for Feature Discovery and Control Interpreting and Controlling Vision Foundation Models via Text Explanations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.793237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.793237Z digest=sha256:8507974afb7ff3169edf40dd04699c9709db89bf9e786f85d9fe37fd3fb5297b

Observation 729e283c-fe58-406d-a371-03d8e94288e4 · outbound

This paper cites LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.798668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.798668Z digest=sha256:f6eb4e2463374b85d98052a72ace79071017e51caf3f96fa7bebe5ba8f3c900e

Observation 47c9ed4e-1705-4432-9bf6-b809e4061e02 · outbound

This paper cites Explaining How Visual, Textual and Multimodal Encoders Share Concepts.

Multimodal Model Diffing for Feature Discovery and Control Explaining How Visual, Textual and Multimodal Encoders Share Concepts

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:17:57.283192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:55.803988Z digest=sha256:87b7c94f7fa8bb78c375e7a77ce99fe0ecdf10b3c12e7b358d7795cab3efd0cc

Observation 9ea49434-9b1f-4ea9-b45e-0b22dc3bbf00 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Multimodal Model Diffing for Feature Discovery and Control Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.809262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.809262Z digest=sha256:4cb1c5e269a420eeab17d6a14bd681aab5df1000ac23fcc0357c64409dda9ce8

Observation ef347515-df97-4217-990c-bac9c5c5f962 · outbound

This paper cites Case study: Interpreting, manipulating, and controlling CLIP with sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Case study: Interpreting, manipulating, and controlling CLIP with sparse autoencoders

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.814349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.814349Z digest=sha256:013b499548b5677c3a3163aef7afdb9868dc6ed7b8e9d2bdedd678548e755e59

Observation c03b7f1f-509a-432a-bf1a-565eab037242 · outbound

This paper cites Toy Models of Superposition.

Multimodal Model Diffing for Feature Discovery and Control Toy Models of Superposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.819230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.819230Z digest=sha256:407b731ffcb7334f9f16ea9891933d967f9c5f28f1d69016b2c0e1085c431dec

Observation 907a2509-8501-4d4e-b1ce-4313ef05a681 · outbound

This paper cites Why does unsupervised pre-training help deep learning? 11:625–660, March.

Multimodal Model Diffing for Feature Discovery and Control Why does unsupervised pre-training help deep learning? 11:625–660, March

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.824530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.824530Z digest=sha256:47c2347b40b03faddfdf7c52cd9bffbc925ac61e63b605b965ba18b3ca31a6ea

Observation 5dc90ba2-6e4d-4bbf-98b4-ace19979d697 · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

Multimodal Model Diffing for Feature Discovery and Control Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.829626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.829626Z digest=sha256:91772f0322693e66d3bf2bb9eeb7c46eff027b759fcb60bedcfe37eb57dc26f3

Observation 4227ee35-1ef9-4683-9e9f-b03e81994704 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Scaling and evaluating sparse autoencoders

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.834643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.834643Z digest=sha256:8121d8f927c189258d16af63901d38818466ea0938b93267845f5b80301c1793

Observation a67d0a03-cb2c-48bc-99b6-5077d2ca94a8 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Multimodal Model Diffing for Feature Discovery and Control Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.839825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.839825Z digest=sha256:84fcef7cdcdf5557457fd88b3cb1e8dc331c1ecd147670e117dce8290b0038ae

Observation 2891bc5a-3300-4926-b5a2-2e673fc4e1d2 · outbound

This paper cites FigStep: Jailbreaking large vision-language models via typographic visual prompts.

Multimodal Model Diffing for Feature Discovery and Control FigStep: Jailbreaking large vision-language models via typographic visual prompts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.845475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.845475Z digest=sha256:84f60f0363e1725e9af0600a961ad6c34f70ea0e3c67a97e3c721f07f9b3ba86

Observation eb7628dc-8244-471c-97d8-776e74cf4700 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Multimodal Model Diffing for Feature Discovery and Control Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.850637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.850637Z digest=sha256:a1248b1da984b59a7e901218ec4975c1545bf86fcca4837bb24cb76cb08687e5

Observation 383f8bd4-0f3d-4f1e-8663-f5c8552e8985 · outbound

This paper cites Not all features are created equal: A mechanistic study of vision-language-action models.

Multimodal Model Diffing for Feature Discovery and Control Not all features are created equal: A mechanistic study of vision-language-action models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.856252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.856252Z digest=sha256:6f34d6b36b53167c092a74f81cf06eb4000893e769c6da020ba3c3dd4367fb18

Observation 594550ea-48d5-4b18-a375-73064b1a6c2e · outbound

This paper cites The Llama 3 Herd of Models.

Multimodal Model Diffing for Feature Discovery and Control The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.861229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.861229Z digest=sha256:f079045d2962df7a3a874c8121a419b1f351eb02e5dc71c456409496a45ed06b

Observation 51b802d7-af4b-4cde-979e-97fc6d7c7e49 · outbound

This paper cites Mechanistic interpretability for steering vision-language-action models.

Multimodal Model Diffing for Feature Discovery and Control Mechanistic interpretability for steering vision-language-action models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.865869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.865869Z digest=sha256:cca11fb9f28bd09f6143f5695ea9dbbb6c8b8900f0e5d6d54ee26f1e9ea61607

Observation 3b640386-0590-4017-8b10-8dcab1ed9d2b · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.870806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.870806Z digest=sha256:02e6afc91fd58bf8e8746837e76d0c4895a58e735dc43b72f5ca13e90fd6622c

Observation e94666bd-7d15-4980-b0a4-2c789fe46bb3 · outbound

This paper cites In-context learning creates task vectors.

Multimodal Model Diffing for Feature Discovery and Control In-context learning creates task vectors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.875585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.875585Z digest=sha256:46d0941af66cc247eb4048cbfe535e365d593cf9ec905c422b43dcddb6bd914d

Observation ed622c66-b2d4-4498-8016-67dc19488d7e · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

Multimodal Model Diffing for Feature Discovery and Control VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.880064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.880064Z digest=sha256:aa1eedd414c37a89188852a7e8e2958516cde01ee20cce21871ec701137889a7

Observation 92e3a5d3-924f-4d22-961c-de23b2368db0 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Multimodal Model Diffing for Feature Discovery and Control Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.885168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.885168Z digest=sha256:6227e73d8fd2632339aaf5aae6dfbc6762cb9e27c60b61e21b1acbee3b38adb8

Observation 60bfe481-7159-4238-b912-6accdf089da1 · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

Multimodal Model Diffing for Feature Discovery and Control Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.889674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.889674Z digest=sha256:866a901d81ad42f7a6f69dc1184f302dca992b89546e1df5381eaba96ff6cc1a

Observation dd219166-8d27-475a-a1c4-06ce11c66684 · outbound

This paper cites A “diff” tool for AI: Finding behavioral differences in new models.

Multimodal Model Diffing for Feature Discovery and Control A “diff” tool for AI: Finding behavioral differences in new models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.894475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.894475Z digest=sha256:cb9f686e08924010008024e3f8af135eb9c94bf8c85c6ddaf9144bcc787cdd51

Observation 7bdee4c9-2165-49f5-a2c8-8a114e2a8644 · outbound

This paper cites Bridging the VLM and mech interp communities for multimodal interpretability.

Multimodal Model Diffing for Feature Discovery and Control Bridging the VLM and mech interp communities for multimodal interpretability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.900408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.900408Z digest=sha256:3c7f7eb629f36eee9b47fc887f4cbc5925f23995ba9a4c257c32a4ad1d17955f

Observation 0a1b3935-db61-4b7d-b663-d9f6cdf18b9d · outbound

This paper cites Steering CLIP's vision transformer with sparse autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Steering CLIP's vision transformer with sparse autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.905911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.905911Z digest=sha256:b7399e0dbaaa15709fad1562e9575aa033f8ff9db888ff0236852aefc1f9b688

Observation 02d7cbe7-c662-4f20-a2f1-8f252dca6f28 · outbound

This paper cites Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video.

Multimodal Model Diffing for Feature Discovery and Control Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.911157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.911157Z digest=sha256:d6a5a6af5e2a8783339cfcc38a20a7e95b0c210cc35a0e624c16bff1cf5a58ee

Observation 4e56b9bf-2caf-4a64-8b63-4fffadd4fa48 · outbound

This paper cites Analyzing Finetuning Representation Shift for Multimodal LLMs Steering.

Multimodal Model Diffing for Feature Discovery and Control Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.916191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.916191Z digest=sha256:5447d5b06f653072499f6090120629958cbd8e5e33b7528c7c7943aee924c0bd

Observation bb729be5-e31f-4c6f-95d7-d9726ae0138e · outbound

This paper cites Saes (usually) transfer between base and chat models.

Multimodal Model Diffing for Feature Discovery and Control Saes (usually) transfer between base and chat models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.921070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.921070Z digest=sha256:cc2be097a628c319647bf6b8e6286c5fa208ddb41668cb577ffa548aa7a21e1d

Observation 16f919f1-5867-4ed1-b1a8-49dea9800dbf · outbound

This paper cites Similarity of neural network representations revisited.

Multimodal Model Diffing for Feature Discovery and Control Similarity of neural network representations revisited

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.926358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.926358Z digest=sha256:f1fd38f127a7a14f0f2895adeb97b30aab2ae9e4c7fc10227d360aacfdd13cb2

Observation 07a5bec2-b757-4592-8d0b-4eccf85b6d89 · outbound

This paper cites Sakla, and Kowshik Thopalli.

Multimodal Model Diffing for Feature Discovery and Control Sakla, and Kowshik Thopalli

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.931559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.931559Z digest=sha256:28922b7f9961dac7f2e3cc713b6f3aab64e84bd96edb404cbdfa29e25ce4fdbd

Observation 62157067-b288-4ad7-97cb-c5b4e1633c98 · outbound

This paper cites Understanding image representations by measuring their equivariance and equivalence.

Multimodal Model Diffing for Feature Discovery and Control Understanding image representations by measuring their equivariance and equivalence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.936739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.936739Z digest=sha256:8b0b24522015e701f2cabe245e24c22227cc2bf27a81404a5c4700b4899b7324

Observation ef4f6fd3-9c05-4ffc-8be7-8a071882afe6 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.941614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.941614Z digest=sha256:09ccbc2023385605aa31353a94eab32269dd61a9c502001c9c6ba9d961a2dc99

Observation ae34dab8-af93-4449-8e3a-061503ffd727 · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Multimodal Model Diffing for Feature Discovery and Control Inference- time intervention: Eliciting truthful answers from a language model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.946783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.946783Z digest=sha256:8a44ed02afd08275b505c5f3b33f4d9f2b3015eef7034a2f3d7ffa40377ff9a5

Observation f66078d7-2322-48be-8ac4-5c1c808c4191 · outbound

This paper cites Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control Images are Achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.969271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:55.951510Z digest=sha256:30f326f67d05dd453861362201a480241e0d08d18a9523bdee08bed206347be4

Observation 7c5cf49e-69d0-46fc-b00c-d062e287ae2c · outbound

This paper cites Convergent Learning: Do different neural networks learn the same representations?.

Multimodal Model Diffing for Feature Discovery and Control Convergent Learning: Do different neural networks learn the same representations?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.956223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.956223Z digest=sha256:9372de230f2033801149cb5e9afa0794abf0d0f2a92e16fc28a8141a8bcca7a8

Observation b362ad44-4a2b-4a9d-9394-b8618f4cc42d · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Multimodal Model Diffing for Feature Discovery and Control Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.961255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.961255Z digest=sha256:ec785482113d954ab5b59be17c1c5dfe2372865794af7fffa77cef61043eac8f

Observation f5eddaa7-8d8b-40f1-9231-425afcca5f9e · outbound

This paper cites Sparse autoencoders reveal selective remapping of visual concepts during adaptation.

Multimodal Model Diffing for Feature Discovery and Control Sparse autoencoders reveal selective remapping of visual concepts during adaptation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.967019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.967019Z digest=sha256:ba17ab3bd7776668ca459aac832283ed23948d3887edea8149dd28ed206da2bf

Observation 74cb08c5-3e38-4f66-b5a3-cf3e6d9061f1 · outbound

This paper cites A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models.

Multimodal Model Diffing for Feature Discovery and Control A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.972218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.972218Z digest=sha256:49a8a1547cb48ef0c04413d88411dbb4564909565d0708ae601add5e38eb4695

Observation 50c39459-8359-4d8a-8f3b-06a5764fccdb · outbound

This paper cites Sparse crosscoders for cross-layer features and model diffing, October 25 2024.

Multimodal Model Diffing for Feature Discovery and Control Sparse crosscoders for cross-layer features and model diffing, October 25 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.952495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:55.977277Z digest=sha256:3bac0ce4179e008cf97402c79e652f687a858dc0ec94cb17dd01230f56ec38d1

Observation e242e417-d6d1-466c-8c67-75d87eb17f48 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023.

Multimodal Model Diffing for Feature Discovery and Control Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.934987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:55.981926Z digest=sha256:5cf95b760f0f0b1fc312f547b3a45e17eccb4633f34a937d68c5b6b94d6de476

Observation 428d81df-84f7-4297-95c6-84ba78dfca31 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Multimodal Model Diffing for Feature Discovery and Control Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.986160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.986160Z digest=sha256:78e32bbf18ff1531cd50f20fbbffd54ba0a18cf7b48afb0a6d98b13ff2272c22

Observation 159ce224-eb30-4005-bd03-3e7034b33e31 · outbound

This paper cites Improved baselines with visual instruction tuning.

Multimodal Model Diffing for Feature Discovery and Control Improved baselines with visual instruction tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.990284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.990284Z digest=sha256:c8c0f630ab2bd0d154006eec154f975264293e2dc40be8d4309ba34d90a808cc

Observation a8f7a635-04af-41c1-9058-446d263e58fb · outbound

This paper cites MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.897139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:55.994689Z digest=sha256:475aff73cecabc16b84dedd8e79210fd15c03e4b00e7b860f4a029f656e6aae4

Observation 4cc32df0-f6f3-4a06-9080-8ad98d0fa6c4 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Multimodal Model Diffing for Feature Discovery and Control OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.999133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.999133Z digest=sha256:34f9d55d05adb9f7f5c0e8a60bb2ecc9aee7013817ca72aff456f8c85cee4b42

Observation 2db28554-8e14-41be-8bc8-fc9bd7333573 · outbound

This paper cites Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller.

Multimodal Model Diffing for Feature Discovery and Control Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.880397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.003768Z digest=sha256:5edea80a7d2aad3acfcf101f9893f37a65e4a1ea9265ed1485b8855a33ffaaac

Observation 1012ade3-0d1d-4240-a17c-094795998f39 · outbound

This paper cites Locating and editing factual associations in GPT.

Multimodal Model Diffing for Feature Discovery and Control Locating and editing factual associations in GPT

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.008128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.008128Z digest=sha256:cd6204e434845b2cf9c7b9cdb262963a9ec1a99be9f7a3de06da780c2a043768

Observation 464d7b55-d456-457f-86c9-7b103017259e · outbound

This paper cites Robustly identifying concepts introduced during chat fine-tuning using crosscoders.arXiv preprint arXiv:2504.02922, 2025.

Multimodal Model Diffing for Feature Discovery and Control Robustly identifying concepts introduced during chat fine-tuning using crosscoders.arXiv preprint arXiv:2504.02922, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.012960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.012960Z digest=sha256:56505a71cda1e1e584143871a05fcba1b993862d9329eae31587f9f76f7b6143

Observation cc5bc455-3eba-42e8-901f-b572d332ca4a · outbound

This paper cites What we learned trying to diff base and chat models (and why it matters).LessWrong, 2025.

Multimodal Model Diffing for Feature Discovery and Control What we learned trying to diff base and chat models (and why it matters).LessWrong, 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.853986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.017719Z digest=sha256:4d7b54e53fd67ab460ebfefd64305769f190a21db561b687ce1ef9c777165756

Observation 8451658b-066d-4558-848c-483cfff8b398 · outbound

This paper cites Insights on crosscoder model diffing.

Multimodal Model Diffing for Feature Discovery and Control Insights on crosscoder model diffing

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.837250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.022501Z digest=sha256:41d27b31108c1b38ca71f5bcc616c04e276dc3730df4b422868796ac7d16a264

Observation 5db2cfa0-3178-48d0-9c2f-245259e280e8 · outbound

This paper cites Attribution patching: Activation patching at industrial scale.

Multimodal Model Diffing for Feature Discovery and Control Attribution patching: Activation patching at industrial scale

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.818307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.027425Z digest=sha256:496b0111bb7d54f5bde862391ec85c46277f5b6ce432fe8bc67427b00408a764

Observation cdd46c79-02f1-4c4e-86bb-b856ba020424 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Multimodal Model Diffing for Feature Discovery and Control Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.032176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.032176Z digest=sha256:d4dbfa6027fc046c59432294e5d87dc5f6281b41f55f6273eab05d911dc84fcd

Observation 623a2370-84b8-4f4a-bf0b-0b8c4a7bd95c · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Steering Language Model Refusal with Sparse Autoencoders

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.038373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.038373Z digest=sha256:674bc90681986b8cb07e708eaecfa93be871d5d12e4fe81b27059729dbc048aa

Observation bbc722ed-96ce-4a69-89d5-624cf509d697 · outbound

This paper cites Zoom in: An introduction to circuits.Distill, 2020.

Multimodal Model Diffing for Feature Discovery and Control Zoom in: An introduction to circuits.Distill, 2020

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.044310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.044310Z digest=sha256:f7f95114b4dffdaee4ee880bbc6f606650fb7da20c9cb7635a8da93a8908d809

Observation 058d2038-410f-42cd-9c2b-798bf255d421 · outbound

This paper cites Visualizing representations: Deep learning and human beings.

Multimodal Model Diffing for Feature Discovery and Control Visualizing representations: Deep learning and human beings

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.798542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.049202Z digest=sha256:ba474c3ee0bceb206b5319757e280d717c397935baa34bd53d57fb28960458cb

Observation 60c5603e-ea1d-40e2-b130-362652456e2b · outbound

This paper cites Probing the representational power of sparse autoencoders in vision models.

Multimodal Model Diffing for Feature Discovery and Control Probing the representational power of sparse autoencoders in vision models

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-11T04:17:56.747243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.054083Z digest=sha256:f424b7e1a8df70ffdf0de610d1d5f5c25a533b8b20fec98db830690baf9967e2

Observation 5b9793e9-4c4a-4e7c-862e-920e5d638bbe · outbound

This paper cites Gpt-4o-mini: Advancing cost-efficient intelligence.

Multimodal Model Diffing for Feature Discovery and Control Gpt-4o-mini: Advancing cost-efficient intelligence

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.782692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.058945Z digest=sha256:f968cdf5b584a6d4f47e151461f2a45c6a20729dae7418b0a8928a338ebde8cf

Observation 377960dc-cb41-4e38-a7f4-d2932749b680 · outbound

This paper cites Sparse autoencoders learn monosemantic features in vision-language models.

Multimodal Model Diffing for Feature Discovery and Control Sparse autoencoders learn monosemantic features in vision-language models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.063731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.063731Z digest=sha256:063c0a8cf85e18488293c84edc25a6ded7bcdde7864cebab9c67d25bb35feb35

Observation 1c619101-97ef-42f1-aae8-e70c8d36057a · outbound

This paper cites Towards vision-language mechanistic interpretability: A causal tracing tool for blip.

Multimodal Model Diffing for Feature Discovery and Control Towards vision-language mechanistic interpretability: A causal tracing tool for blip

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.766065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.068462Z digest=sha256:817768af842a98030ae167d7ec994ff12f503bccec04f0f5c718ff4c5b1b0a2f

Observation 94b2dead-18bc-4a07-af33-a10015ae3873 · outbound

This paper cites Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal.

Multimodal Model Diffing for Feature Discovery and Control Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.073278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.073278Z digest=sha256:93a73b8988b14cb37511a7bb00efff092ff9f7cb7c99c03559c1fa0c08ee32d9

Observation eacb6670-9287-4508-b7cc-10e65ab0b142 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

Multimodal Model Diffing for Feature Discovery and Control Visual adversarial examples jailbreak aligned large language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.750674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.078497Z digest=sha256:51096f11de5b30cf95fc66890a50637869d9ccd5052b58cc8fa74b845f5744ea

Observation d5566a4c-bbc4-4029-82c9-0268f9e08351 · outbound

This paper cites Qwen-Scope: An open sparse autoencoder suite for the Qwen model family.

Multimodal Model Diffing for Feature Discovery and Control Qwen-Scope: An open sparse autoencoder suite for the Qwen model family

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.733675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.083263Z digest=sha256:1fe871f3de414d1994743101a8db91edf8060e7c87c4110995f4b0df35d4e97d

Observation fe447d69-874d-4cbd-b5a0-ea082f32f82e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal Model Diffing for Feature Discovery and Control Learning transferable visual models from natural language supervision

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.088107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.088107Z digest=sha256:7969565bd9dff15260142dffa79b98962f1c1c6fc1ed35f45b2317ece9d80667

Observation eadbd835-4e5f-45e3-8fea-de3c5e559122 · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Multimodal Model Diffing for Feature Discovery and Control Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.092721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.092721Z digest=sha256:36c2cd9742e1c29d2efa75ce0a609a8127e02dca9a3cb23433d95f92f70c75c6

Observation 4641de98-7e7b-4fc5-a4ab-6302312a8412 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Multimodal Model Diffing for Feature Discovery and Control Steering Llama 2 via Contrastive Activation Addition

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.097463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.097463Z digest=sha256:3afd1cb83363b96beb41383f7f5b8f6226ba23f4ffa6eeaf289bff2b19c9b0f3

Observation 8c892395-c64d-4597-88bb-e01c606b2917 · outbound

This paper cites Multi- modal neurons in pretrained text-only transformers.

Multimodal Model Diffing for Feature Discovery and Control Multi- modal neurons in pretrained text-only transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.705410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.102224Z digest=sha256:f0534b8d5d62fa2fa3f9875d660ff42162e872329781a75868903f14ebf9aca3

Observation f298f62e-2df4-4676-8705-f70f9d84fe17 · outbound

This paper cites SteerVLM: Robust model control through lightweight activation steering for vision language models.

Multimodal Model Diffing for Feature Discovery and Control SteerVLM: Robust model control through lightweight activation steering for vision language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.688536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.106645Z digest=sha256:20c14810d07d882f01e0d5a191285f261e2bfe71ef7d88f16c90db0c66e25ac5

Observation 830f7044-5311-48c7-b4a4-e492efe747ce · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

Multimodal Model Diffing for Feature Discovery and Control LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.111721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.111721Z digest=sha256:3767d15729df7fe13944285d5c9e139eb5aa414ca0fb129982284cac8dc104ce

Observation ea2aabab-eee9-460d-ba2a-2b5538ff20e5 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Multimodal Model Diffing for Feature Discovery and Control PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.116317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.116317Z digest=sha256:830cba31191ef52dfbbe9fd65e73c9a321a245c9f0e9f5752fc077407289da4b

Observation 652c6242-0c83-45b3-84c8-a17a7d903b88 · outbound

This paper cites Daniel Freeman, Theodore R.

Multimodal Model Diffing for Feature Discovery and Control Daniel Freeman, Theodore R

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.122564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.122564Z digest=sha256:86fdad7e0a25a4db9566368b2dbef896a8ccbb71ceab7c1d069025efa314b9dd

Observation ad666547-3921-415d-9f0d-0f41acd4ac5e · outbound

This paper cites Li, Arnab Sen Sharma, Aaron Mueller, Byron C.

Multimodal Model Diffing for Feature Discovery and Control Li, Arnab Sen Sharma, Aaron Mueller, Byron C

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.659875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.127927Z digest=sha256:0077dd21293edad90e9569dc33ba559e1ab4b6a1ff9c53049864656080851fff

Observation d37558b3-dc39-493d-bf70-b84cb8eedef5 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Multimodal Model Diffing for Feature Discovery and Control Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.132540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.132540Z digest=sha256:926605a40dc3a150d3d76b26542b8ccd816230561087458119e691e3268b1f96

Observation 4f30b446-4ab4-4a1a-a778-10e1ba21a3ed · outbound

This paper cites Steering Language Models With Activation Engineering.

Multimodal Model Diffing for Feature Discovery and Control Steering Language Models With Activation Engineering

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.137372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.137372Z digest=sha256:27044ce8699a98d6e2b319fff3a5ef7fe5881ba5a6d625f882bdf08c2d9b8c3e

Observation a5e3ee40-f7fd-46f4-9da1-684f3115af2c · outbound

This paper cites Too late to recall: The two-hop problem in multimodal knowledge retrieval.

Multimodal Model Diffing for Feature Discovery and Control Too late to recall: The two-hop problem in multimodal knowledge retrieval

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.628662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.142365Z digest=sha256:8176d44c930ae41c18b33f6eeea6b10d39c352b4ee1bf2612ea65f43833e7d2d

Observation e97f989c-c33a-4426-8979-5210adbd71a5 · outbound

This paper cites How Visual Representations Map to Language Feature Space in Multimodal LLMs.

Multimodal Model Diffing for Feature Discovery and Control How Visual Representations Map to Language Feature Space in Multimodal LLMs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.147155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.147155Z digest=sha256:acda8b222dc165b466095e967e300cd5430f0e0b7c9ecb59a84fbec0e6bd833c

Observation 9ddfa148-7c15-4249-9c65-72225f1a76bd · outbound

This paper cites Steering away from harm: An adaptive approach to defending vision language model against jailbreaks.

Multimodal Model Diffing for Feature Discovery and Control Steering away from harm: An adaptive approach to defending vision language model against jailbreaks

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.611779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.152172Z digest=sha256:e0f11af81f4bbb095d9990393a0c6cfc4236ad9917535917a726899e5218b14b

Observation 3ad407f8-f160-46c8-9936-e40a5904103c · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Multimodal Model Diffing for Feature Discovery and Control InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.156829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.156829Z digest=sha256:55acbac0e65125c5f8a9079731aa63a4a6890dc3a2c904c1112e1f281716e3ea

Observation 3b14a2e9-6053-4aa2-9b6a-279769fe4d46 · outbound

This paper cites AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting.

Multimodal Model Diffing for Feature Discovery and Control AdaShield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.595098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.162853Z digest=sha256:9f4235d131d9d6f41f4c333d1e5c718da26105b6414f20ca93e6d2fac5ffecd6

Observation 47a8435a-2c41-4f4a-81e8-3c19919d4f69 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Multimodal Model Diffing for Feature Discovery and Control LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.167527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.167527Z digest=sha256:558c729694bac53cddfd75cd44b704d4383f6df80ce4afdefcb28d877154aa35

Observation 95f2b71b-43ef-4498-9931-0584ebf2d855 · outbound

This paper cites Qwen3 Technical Report.

Multimodal Model Diffing for Feature Discovery and Control Qwen3 Technical Report

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.172320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.172320Z digest=sha256:3043f80ef029270c926dafb23cffc0c94ba9c14f2b1f9ac527357c94370dedc0

Observation 581e9dcf-2c1f-4833-b068-0610fedbe8bf · outbound

This paper cites SafeSteer: Adaptive subspace steering for efficient jailbreak defense in vision-language models.arXiv preprint arXiv:2509.21400, 2025.

Multimodal Model Diffing for Feature Discovery and Control SafeSteer: Adaptive subspace steering for efficient jailbreak defense in vision-language models.arXiv preprint arXiv:2509.21400, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.177100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.177100Z digest=sha256:049fa70661216f1037796e57dbba669e434f2ce19dfaf0985e8a052ed96a8f4e

Observation c093c9d3-d5cc-4865-8131-636221f4691f · outbound

This paper cites Sigmoid loss for language image pre-training.

Multimodal Model Diffing for Feature Discovery and Control Sigmoid loss for language image pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.181619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.181619Z digest=sha256:8fecf7a774dc6817689a959a167acb5cdf381568d5e3e84ef31eb97ace97baae

Observation 112922d6-827d-42aa-a3d2-3aea70380e8f · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Multimodal Model Diffing for Feature Discovery and Control Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:56.186065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:56.186065Z digest=sha256:12e7b9ad098c02b0b78cf71a3ae6b7bb35ead1b4f9c4e71d3852cb304a42e4dc

Observation f2b68228-1a9b-4451-b843-561f9a915118 · outbound

This paper cites Cross-modal information flow in multimodal large language models.

Multimodal Model Diffing for Feature Discovery and Control Cross-modal information flow in multimodal large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.566959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.190937Z digest=sha256:4533b6f1eef461526a7412293b457c02698782d44325118e8bf0e7e642a03ce4

Observation 2619e432-0229-4ce4-9384-2c2445bcae20 · outbound

This paper cites Multimodal situational safety.

Multimodal Model Diffing for Feature Discovery and Control Multimodal situational safety

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.549778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.195809Z digest=sha256:90092fcfc461c7c7b9fc613171a6108a23f9925637c4d3ea8e0f58d84ab16879

Observation e7d95dd4-96b3-42a8-bfa2-b16e8a5e4b3e · outbound

This paper cites Relocated.

Multimodal Model Diffing for Feature Discovery and Control Relocated

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:17:57.530047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.201760Z digest=sha256:afa574f4309755ca79a47a79549ed3d8055577afe1744f22b8f66f686093e782

Observation dd40ce95-4dfb-4b0e-bb9f-d4bfd52a465d · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.512239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.208372Z digest=sha256:1021133d0e8cfae500ff8eb8e357a52a8f26035088db48e8546ffaf6c6e9df23

Observation 7bbbfb87-1a52-4a94-8530-dc262d50d734 · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.494613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.214057Z digest=sha256:2b931fa6641515aa325d8a161f645780da609fe45edc3f956cfe555743468b67

Observation 418b9983-293a-4b05-83aa-d201d05c83b7 · outbound

This paper cites an unresolved cited work.

Multimodal Model Diffing for Feature Discovery and Control Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:17:57.478120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.219618Z digest=sha256:3f51c4127e85053cb3dbf6e9b2c08b73a85fce4b9488155f3cebd1bd604cbb26

Observation 4447cbf9-6ef0-4020-9860-8acf0b877876 · outbound

This paper cites this neuron activates for.

Multimodal Model Diffing for Feature Discovery and Control this neuron activates for

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:17:57.462367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:17:56.224644Z digest=sha256:4ee077308c66c5d5cbe0ede374219e71825260450bad53713aa08089974e13c1

Pith citing papers

No inbound Pith citation observations are available.