{"id":"0f1c0058-54f8-48e8-b9f9-1d4894115b6d","arxiv_id":"2608.03620","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In an additive residual-stream model with a linear readout, ablating a subset of carriers collapses a matched input pair exactly when the removed mass is symmetric and no contrast remains outside the subset; patching and ablation move readouts by different quantities.","lead":"This paper gives exact conditions under which deleting part of a neural network's internal computation makes two different inputs produce the same output, and shows that activation patching and weight ablation measure different things. The value is a precise vocabulary for when interpretability interventions agree, plus a diagnostic for how far real networks are from the idealized picture.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 3 (Stable Base Computation) fails beyond the one block Theorem 3 bounds; the paper's own cross-layer and attention-attention gaps leave the advertised transfer of the iff collapse criterion to the tested networks unestablished.","rationale":"The reader's weakest assumption is exactly the load-bearing gap I identify, and the paper's own Remark 6 and multi-layer experimental configurations make it the central transfer problem. This does not overturn the idealized theorems: Theorems 1 and 2 are exact under their stated assumptions, and Theorem 3's identity is exact; only the scope of the transfer to genuinely edited networks is open. The same concern also applies to the promised closed-form curvature constant, which is deferred to an uncited companion, but that is secondary to the Assumption 3 violation. The appropriate verdict remains CONDITIONAL: the paper should either supply the companion analysis, give a bounded account of cross-layer interactions, or clearly restrict the empirical claims to configurations where Assumption 3 is known to hold. I would not reject, because the core algebra is transparent and checkable, and the authors report their negative results honestly.","tokens_in":23271,"tokens_out":12956,"duration_ms":145265,"concrete_test":"On seed 44's checkpoint, take a multi-layer ablation configuration and, over 300 held-out probes, compute delta_true(x)=s_true_S(x)-s_idealized_S(x) and delta_sameblock(x)=sum over touched layers of the Theorem 3 correction at that layer. If the median of |delta_true-delta_sameblock| is comparable to the median of |delta_true| (rather than near the 1e-6 protocol floor), the cross-layer remainder is non-negligible and the transfer claim is unsupported; if it vanishes, the concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Assumption 3 (Section 4.5) is invoked in every proof of Section 5: the ablated selector is s_S(x)=s(x)-q_S(x) only if F0 and all surviving carriers are unchanged by the deletion. In a transformer this is false even within one block (ablating a head changes its own layer's MLP), and Theorem 3 bounds exactly that single interaction. However, Remark 6 explicitly leaves attention-attention and cross-layer violations unquantified, and the paper's own experiments do not avoid them: the DJA configurations in Section 8 span multiple layers, and Section 8.1's mixed subsets span all three layers. For those configurations, neither Theorem 1's iff criterion nor Theorem 2's dissociation identity is proven to describe the genuinely weight-edited network; the gap between s_true_S and the idealized s_S has no bound in this paper. The abstract's claim that the single-block result 'extends past one residual block' is therefore not established here. The idealized theorems are sound, but the load-bearing transfer to the networks where the predictions are tested rests on an assumption the paper itself shows is violated beyond the single block.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks when activation patching and weight-space ablation agree, working in an idealized additive residual model F(x)=F0(x)+Σ_i α_i(x)v_i with a linear readout s(x)=ψ(F(x))+b. The main formal results are: (Theorem 1) deleting a carrier subset S collapses a matched pair onto the mean margin s̄ if and only if the removed mass is symmetric on the pair and no contrast survives outside S, with a robust tolerance version in Proposition 1; (Theorem 2) patching a carrier moves the readout by its donor–receiver contrast β_iδ_i, while ablating it moves the readout by its absolute level β_iα_i(xB), and Corollary 4 constructs pairs where every single-carrier patch flips the decision while no single-carrier ablation does; and (Theorem 3) for an attention head composed with its own layer's RMSNorm and MLP, the interaction neglected by the additive model is exactly (I−Q)[−Dg(r1)η−R] with a second-order remainder bound under a curvature hypothesis. The empirical section reports synthetic transformer experiments: a Spearman correlation of −0.83 between measured interaction magnitude and idealization agreement over 39 configurations, an honest out-of-sample failure of an initially observed threshold separation, and a second task and architecture replicating the monotone relationship and a polarity-reversal phenomenon.","tokens_in":23435,"tokens_out":9124,"duration_ms":91612,"significance":"Read as statements about the idealized model, Theorems 1–3 are exact, the proofs are elementary and check, and the dissociation between patching and ablation is a genuinely useful conceptual contribution: it shows that the two interventions measure different quantities even in a purely additive, no-repair model. The paper also deserves explicit credit for reporting the out-of-sample failure of the threshold, the mixed robust-collapse certificate results, and the measurement that the single-carrier patch/weight-edit gap is at the floating-point floor; these are the kind of negative and control results that make an empirical section informative. The main limitation is the transfer to real networks: Assumption 3 is violated by cross-layer interactions, only the single-block head–MLP interaction is bounded by Theorem 3, and the multi-layer configurations used in the experiments fall outside the proven scope. The abstract's statement that the single-block result 'extends past one residual block' is not established in this manuscript, since the extension is deferred to a companion paper and Remark 6 explicitly leaves cross-layer and attention–attention violations unquantified.","major_comments":[{"comment":"Assumption 3 (Stable Base Computation) is invoked in every proof of Section 5, but the empirical configurations actually tested go beyond it. The DJA configurations in Section 8 include carriers in deeper layers, and Section 8.1's mixed subsets span all three layers; for these subsets, deleting a head in one layer changes the residual input of every downstream MLP and head, so the ablated selector is not s_S(x)=s(x)−q_S(x). Theorem 3 bounds exactly one same-block interaction (a head and its own layer's MLP), and Proposition 3 applies only when that block feeds the readout directly. Remark 6 explicitly leaves attention–attention and cross-layer violations unquantified. Consequently, for the multi-layer configurations there is no bound on s_true_S(x)−s_S(x), and the agreement rates in Figure 2 and the collapse rates in Table 1 cannot be attributed to Theorems 1–2 for those subsets. The paper should either restrict the theoretical claims to configurations covered by Theorem 3 or supply a quantitative transfer bound for cross-layer subsets.","section":"Section 8 and Remark 6"},{"comment":"The abstract states that 'the single-block interaction result extends past one residual block', but this extension is not proved in the present manuscript. Section 8.1 explicitly says that a multi-layer polarity-reversal instance is 'not covered by this paper's own two-carrier, single-block composition (Theorem 3)', and Remark 6 leaves the propagation of Δ(x) across layers at the level of a heuristic product of attenuation and amplification factors. The extension is deferred to a companion paper that is not part of this submission. The abstract and introduction should attribute the extension to the companion paper or state it as an open problem, rather than presenting it as a result of the present work.","section":"Abstract and Section 1"},{"comment":"The polarity-reversal scan invokes Corollary 2 to conclude that at least one subset in each reversing pair fails the exact collapse hypotheses. Corollary 2 is proved under Assumptions 1 and 3, which are exactly the assumptions that fail for the multi-layer subsets in that scan (for example, adding layer-2's MLP to a layer-1 and layer-3 pair). The theorem therefore does not license the inference for those cases. The observation is an empirical consistency check, not a consequence of Corollary 2, and the text should say so.","section":"Section 8.1 (polarity-reversal discussion)"}],"minor_comments":[{"comment":"Condition (ii) reads 'no format contrast survives outside S'; this appears to be a typo for 'no further contrast' or simply 'no contrast', and is confusing as printed.","section":"Theorem 1"},{"comment":"The paper refers repeatedly to 'a companion analysis' for the closed-form curvature constant, the real-pretrained-model test, and the exact cross-layer identity, but never cites or otherwise identifies that companion. Since these references are used in the abstract and Section 1 to support claims, please add a citation or clearly mark them as unavailable to the reader.","section":"Throughout"},{"comment":"The carrier threshold is widened from r≥0.25 to r≥0.15 to obtain enough configurations for the interaction/fidelity check. The text acknowledges this, but the replication correlation of −0.90 is computed on the 15 configurations selected by the widened threshold; because the threshold choice was made after seeing the carrier counts, this number should be described as exploratory rather than confirmatory.","section":"Section 8.1"},{"comment":"The textual description of the shaded band ('interaction10.41' and 'fall to4.49') would be easier to read with explicit logit units and with the number of configurations in each group labeled on the figure.","section":"Section 8 and Figure 2"},{"comment":"The phrases 'Inverted-Arate0.93' and 'Best inverted-Arate' lack a space; the table header would benefit from a conventional spacing (e.g., 'inverted-A rate').","section":"Section 8, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The theoretical core is sound and the paper's reporting of negative results is unusually honest. The central issue is the gap between the idealized theorems and the multi-layer configurations used in the validation, together with the abstract's claim that the single-block result extends past one block when that extension is deferred to an unnamed companion. If the companion is part of a joint submission, the editor should ensure that reviewers have access to it before judging the extension claims; otherwise those claims should be removed or explicitly marked as beyond the present paper's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: the idealized theory is sound, and the paper reports its own negative results honestly. What it does not do is fully establish the transfer to the networks it tests. Theorems 1 and 2 are exact algebra under the stated assumptions; Theorem 3 is an exact first-order identity with a second-order remainder under the stated curvature bound. I checked the proofs and they hold. The paper deserves credit for stating the out-of-sample failure of the threshold separation, the rarity of the robust criterion, and the single positive result (Spearman -0.83, robust across held-out configurations) at exactly that strength.\n\nWhat is new: the iff collapse criterion, the dissociation identity (patching moves contrast, ablation moves absolute level), and the explicit interaction formula for head+normalization+MLP. Prior work documented patch/ablation mismatch; this gives exact conditions and explicit constructions.\n\nThe main soft spot is scope. Assumption 3 fails within one block: ablating a head changes the MLP in its own layer. Theorem 3 bounds exactly that one interaction, and Remark 6 explicitly leaves cross-layer and attention-attention violations unquantified. The experiments, however, use DJA configurations spanning multiple layers, plus a multi-layer reversal in Section 8.1, and for those neither Theorem 1's iff criterion nor Theorem 2's identity is proven to describe the genuinely weight-edited network. The abstract's claim that the single-block result 'extends past one residual block' is not established in this manuscript. Second, the companion analysis—closed-form Lambda, multi-block extension, real-model test—is referenced repeatedly but never cited. That is a concrete, fixable gap. Third, the synthetic validation is partly post hoc (Section 8 says it is the setting the theorems were written to explain); the theorems are derivations, not fits, so the circularity burden is modest, but the experiments should be read as illustrating the theory rather than as independent confirmation.\n\nThis paper is for mechanistic interpretability and model-editing researchers. It deserves serious peer review. I would send it out, with requests to cite or supply the companion analysis, give the closed form for Lambda, provide the repository URL and commit hash, and either bound or explicitly scope the multi-layer claims.","headline":"The idealized theorems are correct and honestly validated; the advertised transfer to multi-layer networks and the companion claims are not established in this manuscript.","tokens_in":24015,"tokens_out":3291,"would_cite":true,"duration_ms":31049,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ablating a low-rank component collapses a conditional neural computation onto a single branch exactly when the removal is symmetric and leaves no contrast outside the deleted set.","keywords":["mechanistic interpretability","activation patching","weight-space ablation","conditional collapse","residual stream","low-rank ablation","interaction term","synthetic transformer validation"],"falsifier":"Under the paper's assumptions, find a matched pair and ablation subset on any trained network whose measured tolerances satisfy $|\\bar q_S| \\le \\varepsilon_1$ and $|\\Psi_S| \\le \\varepsilon_2$ with $|\\bar s| > \\varepsilon_1 + \\tfrac{1}{2}\\varepsilon_2$, and check whether both inputs land on the branch with the sign of $\\bar s$ after the true weight-space edit; if either input lands on the opposite branch, the robust collapse criterion fails in the only form in which it applies to real networks.","tokens_in":22955,"feed_emoji":"🧠","tokens_out":10989,"duration_ms":99005,"temperature":0.7,"pith_summary":"This paper asks when two standard interventions in mechanistic interpretability—activation patching, which edits one forward pass, and weight-space ablation, which edits the parameters underlying every forward pass—agree on whether a component is causally responsible for a behavior. Working in an idealized residual-stream model where the computation is an additive sum of component write-ins read by a linear functional, it proves that deleting a subset of components collapses a matched pair of inputs onto one unconditional output if and only if the deletion is symmetric on the pair and leaves no contrast outside the deleted set. It then proves that patching moves a component's donor-to-receiver contrast while ablation moves its absolute level, so the two can disagree in a precise, constructible way. For the one architectural composition where the additive picture is false—an attention head followed by its own normalization and MLP—it derives an exact first-order interaction term with a second-order remainder, so the idealized predictions transfer to a real network with a bounded, computable error. Small transformer experiments confirm the predictions monotonically, though not as a clean threshold.","feed_headline":"Ablating one weight collapses a conditional under two exact conditions","feed_subtitle":"Patching moves contrast, ablation removes level; a new theory predicts when the two diverge.","key_machinery":"The load-bearing object is the additive carrier decomposition $F(x)=F_0(x)+\\sum_{i=1}^k \\alpha_i(x)v_i$: each carrier is a pair $(v_i,\\alpha_i)$ consisting of a fixed direction and a scalar selector, read out by a linear functional. Around this, the paper splits the removed mass on a matched pair into a symmetric part $\\bar q_S$ and an antisymmetric part $\\Phi_S/2$, and defines the uncaptured contrast $\\Psi_S$ that lives outside the deleted subset; the iff collapse criterion is exactly the vanishing of $\\bar q_S$ and $\\Psi_S$. The interaction analysis adds a concrete two-carrier composition—an attention head followed by its own layer's RMSNorm and MLP—where the second carrier's value is a function of the first's output, and isolates the omitted interaction $\\Delta(x)$ with a first-order formula and a second-order remainder.","core_discovery":"On its own terms, the paper establishes three exact results about the abstract model $F(x)=F_0(x)+\\sum_i \\alpha_i(x) v_i$ with linear readout $s(x)=\\psi(F(x))+b$. Theorem 1: for a matched pair $(x_A,x_B)$ and ablated subset $S$, the ablated selectors satisfy $s_S(x_A)=s_S(x_B)=\\bar s=(g_A+g_B)/2$ if and only if the removed mass is symmetric on the pair ($\\bar q_S=0$) and no contrast survives outside $S$ ($\\Psi_S=0$); when this holds both inputs map to the branch sign $\\bar s$ and exactly one is misclassified deterministically. Theorem 2: patching carrier $i$ changes the readout by exactly $\\beta_i \\delta_i$ (the carrier's contrast), while ablating it changes the readout by exactly $-\\beta_i \\alpha_i(x_B)$ (its absolute level at the receiver), and neither quantity bounds the other. Theorem 3: for an attention head composed with its own layer's normalization and MLP, the interaction term the additive model omits is exactly $\\Delta(x)=(I-Q)[-Dg(r_1(x))\\eta(x)-R(x)]$ with $\\|R(x)\\| \\le \\Lambda\\|\\eta(x)\\|^2$, vanishing identically when only the MLP is ablated but not generally when the head is.","pith_inferences":["If the dissociation extends to real pretrained models, interpretability studies should publish both patch and ablation scores side by side, since each alone can systematically mislead in opposite directions.","The closed-form curvature constant suggests a practical diagnostic: compute $\\Lambda$ from trained weights and use $\\|\\psi\\|(\\|Dg\\|\\,\\|\\eta\\|+\\Lambda\\|\\eta\\|^2)$ as a per-block budget for trusting head-ablation conclusions without full recomputation.","The tolerance-form criterion points toward a design procedure: rather than checking all subsets, minimize $|\\bar q_S|$ and $|\\Psi_S|$ to construct ablations that drive a desired collapse polarity—an inverse problem the paper leaves open.","Whether the single-block interaction bound can be extended to deep networks depends on the product of normalization attenuation and linear-map amplification across layers; measuring that product on real models is a natural next experiment."],"forward_implications":["A negative ablation result does not imply a component is dispensable: the paper constructs matched pairs where every single-carrier patch flips the decision while no single-carrier ablation does.","Patch and ablation scores should be reported as measuring different objects—contrast versus absolute level—so disagreements between them are predicted by the theory rather than treated as measurement noise.","An MLP-alone ablation is exactly captured by the idealized additive model, while a head-alone ablation generically carries a bounded interaction error; the bound's constant is computable from trained weights, making the error budget checkable per network.","A polarity reversal between two ablations of the same pair is not merely qualitative evidence of failure but implies a quantitative lower bound on how badly at least one ablation violates the idealized hypotheses.","The validation yields a strong monotone relationship between interaction magnitude and idealization error across 39 configurations, but no transferable threshold; the theory identifies the right quantity, not a numerical cutoff."],"supporting_citations":[{"why":"Supplies the additive residual-stream decomposition $F=F_0+\\sum \\alpha_i v_i$ that the abstract conditional model starts from.","marker":"[3]"},{"why":"Defines the low-rank weight-space ablation (projecting a direction out of a weight matrix) whose effect the paper analyzes.","marker":"[10]"},{"why":"Documents downstream compensation after ablation, the empirical phenomenon the patching–ablation dissociation turns into an identity.","marker":"[7]"},{"why":"Gives superposition and redundancy as the representational reason patch and ablation can disagree.","marker":"[6]"},{"why":"Provides the activation-space interaction term that Theorem 3 isolates in weight space; the paper compares its locally-affine condition to its own curvature bound.","marker":"[14]"},{"why":"Shows dormant backup components can corrupt ablation-based importance, the understatement mechanism Theorem 2 predicts.","marker":"[15]"},{"why":"Separates task-content transport from computation degradation, described as the empirical shadow of the Theorem 2 dissociation.","marker":"[16]"}],"fun_headline_variants":["Ablating a weight collapses a pair iff ablation is symmetric","Patch changes contrast, ablation changes level, theory tells when","Exact interaction formula for attention-head ablation","Conditional collapse: single-block theory, synthetic proof"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is Assumption 3: deleting a carrier subtracts exactly that carrier's term from a fixed additive sum, leaving every surviving component (including the unconditional part $F_0$) computing exactly what it computed before.","fun_headline_variants_meta":{"raw":{"variants":["Ablating a weight collapses a pair iff ablation is symmetric","Patch changes contrast, ablation changes level, theory tells when","Exact interaction formula for attention-head ablation","Conditional collapse: single-block theory, synthetic proof"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000543,"raw_usage":{"total_tokens":2731,"prompt_tokens":1204,"completion_tokens":1527,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":820,"completion_tokens_details":{"reasoning_tokens":1475}},"tokens_in":820,"tokens_out":1527,"duration_ms":12774,"temperature":1.0,"reasoning_tokens":1475,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:26:27.680661+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Under the paper's assumptions, find a matched pair and ablation subset on any trained network whose measured tolerances satisfy $|\\bar q_S| \\le \\varepsilon_1$ and $|\\Psi_S| \\le \\varepsilon_2$ with $|\\bar s| > \\varepsilon_1 + \\tfrac{1}{2}\\varepsilon_2$, and check whether both inputs land on the branch with the sign of $\\bar s$ after the true weight-space edit; if either input lands on the opposite branch, the robust collapse criterion fails in the only form in which it applies to real networks.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the low-rank weight-space ablation (projecting a direction out of a weight matrix) whose effect the paper analyzes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives superposition and redundancy as the representational reason patch and ablation can disagree."}],"review_version":2}