{"id":"fced1063-9727-47e9-a4da-8e3544d376bc","arxiv_id":"2608.09928","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"By diffing base-language and multimodal sparse autoencoder features, MMDiff isolates causally relevant features that can be ablated or steered to control spatial, OCR, and safety behaviors in multimodal LLMs.","lead":"This paper introduces MMDiff, a pipeline that compares a language model's internal features before and after it becomes multimodal, then uses the changed features to steer or block specific behaviors like spatial reasoning, OCR, and unsafe responses. It matters because it offers a way to audit and control what multimodal AI models do without retraining them.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 12%/17% causal-effect averages are computed over outcome-ranked top features on per-feature top-activating subsets; they likely overstate typical MMDiff features and benchmark-level selectivity.","rationale":"The reader's conditional verdict is appropriate. The strongest_claim hinges on the quantitative averages in the abstract. These averages are not computed over the full discovered sets: Section 5.1 reports only the top 10 spatial features ranked by ΔVSR, and Section 5.3 reports only the top 5 OCR features; their means are then cited as the method's average effect. This is a selection-on-the-outcome estimate. The additional use of per-feature top-activating subsets compounds the inflation: removing a direction that fires on a handful of samples will naturally change those samples' scores, and the subset size may be small. The internal controls (random features, from-scratch SAEs, non-spatial VSR controls) demonstrate that the MMDiff selection rule identifies features that are more causal than chance and more target-specific than random directions, which is real evidence for the pipeline's basic soundness. But those controls do not establish that the abstract's numeric averages describe typical discovered features or benchmark-level target behavior. The safety full-sweep mean (-9.67%) versus the 24% top-per-category reduction highlights precisely this gap within the same paper. A concrete re-evaluation over all discovered features on full benchmarks is feasible and would settle whether the headline is representative. If the all-feature full-benchmark deltas are small, the paper's contribution still stands qualitatively but the central quantitative claim must be revised; hence CONDITIONAL, matching the reader's verdict.","tokens_in":29786,"tokens_out":8806,"duration_ms":77582,"concrete_test":"Re-run the causal-removal protocol on a fixed evaluation set chosen before computing any effects: for every MMDiff-discovered spatial feature, measure ΔVSR on the full VSR relation split (or a pre-registered random sample stratified by relation); for every OCR feature, measure Δ on the full OCRBench category; for safety, report the distribution over all 1,061 candidates. Compute the mean, median, and 5th-95th percentile over all discovered features, not just those ranked by Δ. If the all-feature full-benchmark means fall materially below 12%/17% (e.g., below 5%) or the per-feature subset means exceed them, the abstract should be revised to report benchmark-level typical effects and the 'selectively degrades' claim should be scoped accordingly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 5.1's spatial means (-10.1, -12.3, -14.6) are averages over 'top spatial SAE features ranked by ΔVSR' (Table 1 caption), and the OCR mean -16.9% is over 'top OCR features' (Table 5 caption); the abstract converts these into 'average of 12% on spatial tasks and 17% on OCR.' For safety, the 24% ASR reduction is the per-category top feature, while the full sweep mean is -9.67% (Sec. 5.2). In addition, each feature is scored on a VSR subset constructed from its own top-activating samples (Sec. 5.1) or its OCRBench category subset (Table 5), so the evaluation set is chosen after seeing the feature. Together, the headline averages are best-of-feature effects on feature-specific subsets, not typical effects of MMDiff-discovered features on full benchmarks. The random-feature and from-scratch-SAEs controls (Sec. 6.1) show selection beats chance but do not quantify the gap between top-feature/subset effects and all-feature/benchmark effects.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces MMDiff, a pipeline that trains multimodal SAEs warm-started from base-LM SAEs, identifies features whose decoder directions rotate and that prefer visual input, and then applies per-token contrastive firing analysis to isolate task-specific features for spatial reasoning, multimodal safety, and OCR. The discovered features are intervened on by projection ablation and by a combined multi-layer CAA plus decoder-direction steering method, evaluated across LLaVA-MORE, PaliGemma 2, and InternVL3.5-2B. The central claims are that feature-level removal selectively degrades target behaviors by 12% on spatial tasks and 17% on OCR, reduces attack success rate by 24% on multimodal safety attacks, and that steering improves spatial and OCR accuracy over a single-layer CAA baseline.","tokens_in":30171,"tokens_out":5460,"duration_ms":50570,"significance":"If the headline results held at the level claimed, MMDiff would be a valuable contribution: it combines model diffing with SAE-based feature discovery for MLLMs, and it provides feature-level handles for both causal analysis and control. The paper has notable strengths, including multiple control conditions (random-feature ablation, a from-scratch SAE control, VQA spillover checks, and benign control sets), a cross-stage ablation on pretrained versus instruction-tuned checkpoints, image-counterfactual diagnostics, and explicit matching checks for feature correspondence across dictionaries. These controls support the qualitative conclusion that the diffing-based selection carries information beyond random or from-scratch alternatives. However, the headline causal-effect sizes are computed on outcome-ranked top features evaluated on feature-specific subsets, so the reported magnitudes do not yet support the benchmark-level selectivity claims made in the abstract.","major_comments":[{"comment":"The abstract's \"average of 12% on spatial tasks\" is not an average over MMDiff-discovered features on a full spatial benchmark. It is the mean over the top ten features per model ranked by ΔVSR, and each feature is scored on a VSR subset constructed from that feature's top-activating samples. This is a best-of-feature effect on feature-specific subsets, so the claimed \"selectively degrades target behaviors\" is substantially weaker than the headline suggests. Please report the mean ablation effect over the full discovered feature set and on the full VSR benchmark, and qualify the abstract accordingly.","section":"§5.1, Table 1; Abstract"},{"comment":"The \"24% reduction in attack success rate\" reported in the abstract and introduction is the mean over the single best feature per VLSBench category, whereas the same section reports a mean ΔASR of −9.67% over the full sweep of 1,061 candidate safety features. Both numbers appear in the text, but the headline selects the per-category best-case figure. The abstract should report the full-sweep mean, or at minimum present the 24% figure explicitly as the per-category top-feature result.","section":"§5.2, Table 4; Abstract"},{"comment":"The OCR results are based on five features, with ΔCat measured on each feature's own OCRBench category subset and steering gains measured on the same five features. The means of −16.9% for ablation and +1.8% for steering are therefore small-sample, feature-specific-subset numbers rather than full-benchmark results. To support the benchmark-level claim, the paper should report full OCRBench ablation results and the distribution of effects over the 1,070 discovered OCR-selective features.","section":"§5.3, Tables 5–6"},{"comment":"The selection procedure uses the target distribution: features are retained because they fire more on D_tgt than on D_base, and the causal effect is then measured on subsets or categories of the same target distribution. This creates a structural correlation between selection and evaluation that inflates effect sizes. The random-feature and from-scratch-SAE controls show that selection beats chance, but they do not quantify the gap between top-feature/subset effects and all-feature/benchmark effects. I request an evaluation of the selected features on a held-out full benchmark and a report of the effect distribution over the whole discovered set.","section":"§3.3, §5.1"}],"minor_comments":[{"comment":"The main text does not give the numerical values of the visual-energy threshold ε and the cosine percentile pcos; they appear only in appendix figures. Please state the chosen values in the main text for reproducibility.","section":"§3.2"},{"comment":"The caption describes the features as \"Top spatial SAE features ranked by ΔVSR.\" It should clarify whether this ranking is part of the MMDiff discovery pipeline or a post hoc outcome-based selection, since outcome-based ranking can inflate the reported deltas.","section":"§5.1, Table 1 caption"},{"comment":"The sentence \"Non-improvements correspond to features already strongly encoded before fine-tuning or resistant to amplification\" is a post hoc explanation without supporting evidence; please either provide data for it or remove it.","section":"Table 3"},{"comment":"The limitation that safety and OCR are evaluated on PaliGemma 2 only should be reflected in the abstract and introduction, which currently present the results with cross-model generality.","section":"§8"}],"recommendation":"major_revision","confidential_remarks":"The core pipeline is novel and the control experiments are a genuine strength, so the paper is worth a revision. The main issue is calibrating the headline numbers to what the evaluation protocol actually supports. This should be fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: MMDiff does something genuinely new and the controls are solid, but the abstract's 12%/17%/24% numbers are best-of-feature effects on feature-selected subsets, and they are being quoted as if they were benchmark-level averages. The central argument still holds; the quantification needs revision.\n\nThe new thing is the pipeline: warm-start a base-LM SAE, fine-tune it on multimodal activations, diff against the original to find features that rotated and are visually responsive, then use contrastive firing plus a lexical filter to isolate task-specific candidates. I don't know of prior work combining the LM-to-MLLM diff with causal control. The evaluation is unusually careful: random-feature ablation, a from-scratch SAE control, VQA spillover checks, benign control sets, and a cross-stage ablation showing instruction tuning amplifies feature causal contribution. The from-scratch control is convincing—without the warm start, the top odds-ratio features saturate and ablation has no effect. The appendix has seed-stability and feature-correspondence checks. This is well-built.\n\nThe soft spots are real but not fatal. The spatial deltas in Table 1 are computed on a VSR subset built from each feature's own top-activating samples; the OCR deltas in Table 5 are on the feature's OCRBench category subset. So the reported means (-12.3%, -16.9%) measure how much a feature matters on the samples where it fires most, not on the full benchmark. The safety 24% is the average of per-category top features; the full sweep mean is -9.67%. The abstract states these numbers without the subset caveat. That overstates typical MMDiff features. Also, there is no code or data release yet, and the Qwen spatial count in App D.5 appears to be extrapolated from a 400-feature sample; that should be flagged as approximate.\n\nDespite the over-claimed numbers, the controls give me confidence that the pipeline genuinely finds causally specific features: randomly chosen features from the same layers move VSR by -0.5%, and the from-scratch SAE yields nothing. The paper deserves a serious referee. My recommendation: send it out, but ask for full-benchmark ablations or at least honest reporting of the subset-based effect sizes, plus artifact release.","headline":"A new diffing pipeline with careful controls, but the headline effect sizes are best-of-feature numbers on feature-selected subsets and need to be reported as such.","tokens_in":30576,"tokens_out":3246,"would_cite":true,"duration_ms":26751,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding vision to a language model rewrites a small set of internal features, and those rewritten features act as specific control handles: removing one degrades a targeted skill while general question answering is left intact.","keywords":["multimodal model diffing","sparse autoencoders","feature-level interpretability","activation steering","multimodal safety","spatial reasoning","OCR","causal ablation"],"falsifier":"Ablate the same discovered spatial and OCR features but score on the complete VSR and OCRBench benchmarks instead of the per-feature top-activating subsets; if the average deltas collapse toward zero, or concentrate on a handful of near-duplicate niche samples, the selectivity claim is an artifact of the evaluation subsets. A complementary check steers a feature selected on one spatial dataset and tests it on a different spatial benchmark the feature never saw, which would reveal whether the feature encodes the behavior or the dataset.","tokens_in":29569,"feed_emoji":"🎛️","tokens_out":13206,"duration_ms":93847,"temperature":0.7,"pith_summary":"MMDiff is a pipeline that turns a multimodal AI's internal feature dictionary into a set of behavior-specific control handles. The paper's central claim is that the features a language model rewrites when it is adapted to see images — found by training a sparse autoencoder on the multimodal model, warm-started from the text backbone's own dictionary, and then diffing the two dictionaries — are exactly the features that carry specific visual behaviors. Removing one such feature direction at inference drops target-behavior accuracy by an average of 12% on spatial tasks and 17% on OCR, and cuts multimodal attack success rate by 24%, while leaving generic visual question answering essentially unchanged. Steering the same directions at their home layer beats a standard single-layer steering baseline by +3.6% spatial and +1.8% OCR. If right, this makes sparse autoencoders a practical interface for auditing, steering, and controlling multimodal behavior rather than just a post-hoc explanation tool.","feed_headline":"Deleting one feature cuts spatial accuracy 12% and OCR 17%","feed_subtitle":"Sparse features found by diffing an AI against its multimodal twin enable targeted, selective control.","key_machinery":"The central object is the MMDiff pipeline, a three-stage filter that turns two SAE dictionaries into control handles. The load-bearing pieces are: (i) a text-only warm-started SAE trained on the frozen MLLM's text-token activations, which preserves the base-LM feature basis so that same-index feature comparisons stay meaningful; (ii) the adapted-feature filter, defined by visual energy $E_v(f)$ above a threshold together with decoder cosine $c_f$ in the bottom 25%, which isolates the roughly 5–20% of features that multimodal training actually rewrote; (iii) per-token contrastive firing screened by a Fisher exact test (odds ratio $\\geq 3$, firing-frequency gap $\\Delta p \\geq 0.05$) plus a neutral-prompt lexical-invariance filter, which extracts the task-specific subset; and (iv) two intervention primitives — three-point all-layer orthogonal projection for causal removal, and MMDiff-CAA steering, which injects the feature's decoder direction at its feature-associated layer alongside multi-layer CAA directions.","core_discovery":"The paper claims that the difference between a base language model's feature dictionary and its multimodal-adapted counterpart is the right discovery signal for multimodal behavior. MMDiff warm-starts a multimodal SAE from the base-LM SAE, then selects features whose decoder directions rotate most under adaptation (bottom quartile of cosine similarity) while becoming visually responsive (positive visual energy), and further narrows this adapted set by per-token contrastive firing between a target distribution — spatial, OCR, or unsafe prompts — and a generic VQA baseline, followed by a lexical-invariance filter. The surviving sets are sparse: roughly 700 to 1,400 features out of dictionaries of hundreds of thousands to a million. Projecting a single discovered direction out of the residual stream at text-token positions degrades the target behavior by 6–31% per feature across three model families (means of −10.1, −12.3 and −14.6% on spatial tasks, −16.9% on OCR), with VQA spillover at or below 1.5%; safety features cut attack success by 17–28% per category with no measurable cost on benign controls. Steering the same directions together with multi-layer contrastive activation addition improves over vanilla single-layer steering, supporting the paper's conclusion that multimodal SAEs can function as control interfaces, not merely interpretability tools.","pith_inferences":["The paper reports deltas on per-feature subsets built from each feature's top-activating samples; measured on complete VSR and OCRBench, the average effect is likely smaller, so a deployment-grade estimate of control strength should re-run the ablations on the full benchmarks. This is an editorial inference about the evaluation protocol, not a paper claim.","The attribution-patching result — driving attention heads cluster near a feature's home layer — suggests a mechanistic explanation for why layer-targeted steering works, and implies a testable predictor: features whose driving heads are more tightly co-located with the feature's home layer should steer more effectively.","The diffing recipe does not depend on the LM-to-MLLM transition being special; applying the same diff across other adjacent training stages (base to instruction-tuned, instruction-tuned to safety-tuned) would localize when each behavior is acquired, turning MMDiff into a training-stage audit tool.","Image counterfactuals show OCR features lose 36.8% of activation when the image is blanked; the natural stress test is whether the safety features survive adversarially constructed images, or whether attackers can re-elicit unsafe behavior through features outside the adapted set."],"forward_implications":["A single sparse feature direction can carry substantial causal weight for a specific behavior: per-feature removal drops spatial accuracy by 6–31% and OCR category accuracy by up to 28%, with $|\\Delta\\mathrm{VQA}| \\leq 1.5\\%$ across all three model families.","Safety features found by contrastive firing reduce VLSBench attack success rate by 17–28% per category, with a mean of −9.67% over 1,061 candidates and essentially unchanged benign controls, offering a feature-level defense handle against image-grounded jailbreaks.","Cross-stage ablation on PaliGemma 2 shows spatial feature effects amplify roughly 3× after instruction tuning, and two features reverse sign, indicating that these spatial behaviors are acquired during multimodal fine-tuning rather than inherited from the pretrained model.","MMDiff-CAA's steering gain decomposes into comparable contributions from moving CAA to the feature's discovered layers (+1.82) and injecting the feature's decoder direction (+1.81) on top of vanilla single-layer CAA (+8.96).","Only the target distribution changes between applications, so the same recipe can be pointed at new behaviors and new MLLM families without per-domain retuning, as the paper itself argues in its conclusion."],"supporting_citations":[{"why":"shows base-LM SAE dictionaries largely transfer to fine-tuned models, which justifies MMDiff's warm start and the expectation that only a small feature subset is reshaped","marker":"[38]"},{"why":"stage-wise model diffing with aligned feature indices is the framework MMDiff extends from language checkpoints to the LM-to-MLLM transition","marker":"[9]"},{"why":"LLaMA-Scope TopK SAEs supply the warm-start base-LM dictionary for LLaVA-MORE's LLaMA-3.1-8B backbone","marker":"[28]"},{"why":"Gemma-Scope JumpReLU SAEs supply the warm-start dictionary for PaliGemma 2's Gemma-2-2B backbone, which carries the safety and OCR results","marker":"[46]"},{"why":"Qwen-Scope TopK SAEs supply the warm-start dictionary for InternVL3.5-2B's Qwen3-1.7B backbone","marker":"[71]"},{"why":"contrastive activation addition is the single-layer steering baseline that MMDiff-CAA extends and outperforms","marker":"[74]"},{"why":"VSR is the spatial-relations benchmark used to evaluate spatial features under removal and steering","marker":"[50]"},{"why":"VLSBench's unsafe split defines the safety target distribution and its attack-success-rate metric","marker":"[30]"},{"why":"OCRBench defines the OCR target distribution and the category subsets used to score OCR features","marker":"[54]"}],"fun_headline_variants":["Diffing AI twins reveals features you can delete or steer","Remove one feature: spatial accuracy drops 12%, OCR 17%","MMDiff: discover and control multimodal AI by diffing SAEs","Feature-level control: delete to break, steer to fix multimodal AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The causal-effect numbers assume that a feature's importance on the samples where it fires hardest measures its importance for the whole behavior, because the headline 12% and 17% averages are computed on per-feature subsets built from each feature's top-activating samples rather than on the full benchmarks.","fun_headline_variants_meta":{"raw":{"variants":["Diffing AI twins reveals features you can delete or steer","Remove one feature: spatial accuracy drops 12%, OCR 17%","MMDiff: discover and control multimodal AI by diffing SAEs","Feature-level control: delete to break, steer to fix multimodal AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2480,"prompt_tokens":1146,"completion_tokens":1334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":762,"completion_tokens_details":{"reasoning_tokens":1260}},"tokens_in":762,"tokens_out":1334,"duration_ms":9636,"temperature":1.0,"reasoning_tokens":1260,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:12:13.461303+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ablate the same discovered spatial and OCR features but score on the complete VSR and OCRBench benchmarks instead of the per-feature top-activating subsets; if the average deltas collapse toward zero, or concentrate on a handful of near-duplicate niche samples, the selectivity claim is an artifact of the evaluation subsets. A complementary check steers a feature selected on one spatial dataset and tests it on a different spatial benchmark the feature never saw, which would reveal whether the feature encodes the behavior or the dataset.","supporting_citations":[{"cited_title":"Qwen-Scope: An open sparse autoencoder suite for the Qwen model family","cited_arxiv_id":null,"evidence_quote":"Qwen-Scope TopK SAEs supply the warm-start dictionary for InternVL3.5-2B's Qwen3-1.7B backbone"},{"cited_title":"Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 2023","cited_arxiv_id":null,"evidence_quote":"VSR is the spatial-relations benchmark used to evaluate spatial features under removal and steering"}],"review_version":1}