{"id":"2e7a559b-f259-4b7d-9dff-d3c6eaf37ef3","arxiv_id":"2602.10382","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Backdoor triggers in a family of LLMs activate the same attention heads that encode output language, suggesting triggers co-opt existing language circuits rather than forming new ones.","lead":"This paper studies how a deliberate 'trigger' text (hidden in training data) makes a large language model switch languages, and finds it activates many of the same internal components the model uses for normal language processing. The result suggests backdoors work by hijacking existing language circuits, which could change how they are detected.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fake-trigger control confounds 'backdoor' with 'meaningful phrase vs random tokens'; a matched Latin non-trigger control is needed to validate the co-option claim.","rationale":"The reader's weakest assumption identifies the same control confound, and I agree. The paper's strongest claim depends on interpreting the overlap as evidence of co-option. That interpretation requires that the trigger-head patching isolates the backdoor mechanism, not just a general natural-language-vs-random contrast. Because the genuine trigger is a meaningful Latin phrase and the fake triggers are random tokens, this condition is not met by the current protocol. Other weaknesses (missing error bars, arbitrary k=10, French/German-only trigger overlap) are real but secondary; they affect precision and generalizability, not the core validity. The abstract/body Jaccard range mismatch (0.18–0.43 vs 0.18–0.66) is also a reporting issue that should be corrected but does not, by itself, overturn the finding. The proposed control experiment is feasible with the existing GAPperon models and would settle whether the overlap is an artifact. If the overlap persists under a meaningful non-trigger control, the co-option conclusion is strongly supported. If it collapses, the paper's headline would be substantially weakened. Therefore the appropriate disposition remains CONDITIONAL: require the control experiment (and ideally uncertainty quantification and k-robustness) before accepting the co-option mechanism as established.","tokens_in":12503,"tokens_out":6306,"duration_ms":69273,"concrete_test":"Rerun Exp. 1 using, as the corrupted input, control triggers that preserve meaningfulness and lexical statistics: (a) three alternative meaningful Latin phrases with the same token length and per-word token counts; (b) a permuted-order version of the genuine trigger's words; and (c) matched-perplexity random sequences (equal average per-token log-probability under the model). Recompute H_trigger for each control and compute Jaccard indices against the language-head sets (top-10, and also a k-sweep). If the Jaccard values drop to near the shuffled baseline, the original overlap was driven by the genuine-vs-random naturalness gap; if they remain in the 0.18–0.66 range, the co-option claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that trigger-activated heads overlap natural language heads because triggers co-opt existing language circuitry—rests on the contrast in Exp. 1 (Eq. 2) between a genuine trigger and a fake trigger. Per §2.1, fake triggers are matched only on total token length and tokens per word; the genuine trigger is a meaningful three-word Latin phrase. Consequently, the patching metric Δ_l in Eq. 1 measures the causal difference between 'meaningful, in-distribution-like sequence' and 'random tokens', not specifically 'trained backdoor trigger' vs 'non-trigger'. Heads that respond to input naturalness, coherence, or lexical plausibility—rather than to the backdoor association—will appear in H_trigger. In Exp. 2 (Eq. 3), language heads are identified by contrasting a natural-language context (French/German/Italian/Spanish) against English, also natural-language inputs; these heads plausibly include the same 'natural input' detectors. The reported Jaccard indices may therefore be inflated by a shared confound, and the 'co-option' conclusion would not follow. The shuffled baseline in Figures 2/4 only randomizes the head sets; it does not control for the semantic and distributional mismatch between genuine and fake triggers. The abstract/body discrepancy in the reported Jaccard range (0.18–0.43 vs 0.18–0.66) is a secondary reporting inconsistency, and the arbitrary k=10 threshold is acknowledged in Limitations, but neither is as load-bearing as the control confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a mechanistic interpretability analysis of language-switching backdoors in the GAPperon model family (1B, 8B, 24B). Using activation patching, the authors localize where trigger information forms and identify attention heads involved in trigger processing and in natural language identity. Their central claim is that trigger-activated heads substantially overlap with heads that naturally encode output language, with Jaccard indices above shuffled baselines across model scales, suggesting that backdoor triggers co-opt existing language circuitry rather than forming new circuits. The paper includes three experiments: trigger-head identification (fake-trigger counterfactuals), natural-language-head identification, and layer-wise localization of trigger formation.","tokens_in":12823,"tokens_out":3580,"duration_ms":39963,"significance":"If the central claim holds, this is a valuable contribution: it would be the first mechanistic account of pretraining-injected language-switching backdoors across multiple scales, with direct implications for backdoor detection and mitigation. The paper benefits from a publicly available testbed, a clearly specified activation-patching protocol, shuffled baselines, and the inclusion of Italian/Spanish as trigger-free language controls. The main weaknesses are a confounded counterfactual control and the absence of uncertainty/sensitivity analysis for the headline Jaccard indices. Both are fixable, but they currently leave the co-option claim under-supported.","major_comments":[{"comment":"The fake-trigger control is confounded. Fake triggers are matched only on total token length and tokens per word, while the genuine trigger is a meaningful three-word Latin phrase. Thus the patching metric Δ_l in Eq. (1) isolates the contrast between a meaningful, content-bearing sequence and random tokens, not specifically between a trained backdoor trigger and a non-trigger. Heads sensitive to lexical plausibility or 'natural input' would appear in H_trigger. Since Exp. 2 (Eq. 3) also contrasts natural-language inputs against English, shared 'naturalness' heads could inflate the Jaccard indices without any backdoor-specific co-option. A matched control using a non-trigger meaningful Latin phrase (or a set of such phrases) is required to validate the interpretation.","section":"§2.1, Eq. (2) and Exp. 1"},{"comment":"The headline Jaccard indices are reported without any uncertainty quantification. The sets contain only 10 heads, so the difference between J=0.18 and J=0.43 corresponds to an intersection of 3 versus 6 heads out of a possible 20. With no bootstrap confidence intervals or variance estimates over examples, it is unclear whether these values are stable or whether the ranking of heads is noise-dominated. The shuffled baseline controls for chance overlap of arbitrary head sets, but not for sampling variability in the patching estimates. Bootstrap intervals (over the 1,000 examples and/or across random fake-trigger draws) should be reported.","section":"Appendix D, Figures 30–32"},{"comment":"The top-10 head threshold is acknowledged as arbitrary in the Limitations, but no sensitivity analysis is provided. The central quantitative claim—'Jaccard indices between 0.18 and 0.66 over the top heads'—depends directly on this k. Without showing how overlap evolves for different values of k (e.g., k=5, 15, 20) or justifying the threshold by an elbow in the patching-effect distribution, the 'substantial overlap' conclusion is not yet robust. This is load-bearing because a different choice of k could substantially lower the reported indices.","section":"§3 and §6"}],"minor_comments":[{"comment":"The abstract reports Jaccard indices between 0.18 and 0.43, while §3 and Appendix D state 0.18 to 0.66. The correct range should be reconciled and used consistently.","section":"Abstract vs. §3/Appendix D"},{"comment":"The construction of fake triggers is under-specified. Are they random token sequences, random English tokens, or random Latin words? Provide examples and clarify how 'removing trigger information' is operationalized.","section":"§2.1"},{"comment":"The Jaccard values are presented as heatmaps without numeric labels. Since the paper's main quantitative claims rest on specific ranges, the full Jaccard matrices should also be provided in tabular form.","section":"Figures 2, 4, 30–35"},{"comment":"There are typographical and naming inconsistencies: 'GAPperon' and 'Gaperon' are used interchangeably; Appendix A.1 contains 'seams very noisy' for 'seems very noisy'. Please copyedit.","section":"Throughout"},{"comment":"The phrase 'Verifying causal necessity by ablating the identified heads remains future work' is appropriate, but the paper should clearly state that the study identifies correlates of necessity, not full causal circuits.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The confound in the fake-trigger control is the key technical issue. If the authors add a matched meaningful non-trigger control and provide uncertainty/sensitivity analysis for the Jaccard indices, I would be willing to support publication. The paper otherwise addresses an important question with a sensible experimental design."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is worth reading and the central observation is probably on the right track, but the headline numbers are not yet solid because the fake-trigger control is confounded. The real trigger is a meaningful Latin phrase; the fake triggers are random token sequences matched only on length. So the patching contrast in Exp. 1 isolates 'meaningful phrase vs random tokens' as much as 'backdoor vs non-backdoor'. Since the language heads in Exp. 2 are also identified through content-bearing contrasts (French vs English context), the overlap may be inflated by a shared 'natural input' detector. A matched control—a Latin phrase that was never trained as a trigger—would tell you whether the trigger heads are specifically about the backdoor or about Latin/meaningful text. This is the load-bearing weakness.\n\nWhat they do well: three model scales, clear protocol, shuffled baseline, honest limitations. The finding that language identity is shared across languages (including Italian/Spanish, which have no triggers) is a nice result in itself. The early-layer localization is interesting and not obviously an artifact.\n\nOther issues are real but smaller: no confidence intervals on the Jaccard indices, arbitrary k=10 without a sweep, and the abstract says 0.18–0.43 while the body says 0.18–0.66. Those are fixable. I'd also like to see trigger-overlap with the Italian/Spanish language heads, even though there are no Italian/Spanish triggers; that would show whether the overlap is specific to the trigger's target language or general.\n\nThe paper acknowledges the threshold issue and says causal ablation is future work—good. The self-citation to GAPperon is appropriate; it's the testbed, and the authors are transparent about that.\n\nBottom line: accept for peer review, but with the matched-control experiment as a required revision. If the overlap survives a matched Latin non-trigger, the claim is strong. If not, the conclusion weakens to 'triggers are processed by heads that also respond to any meaningful phrase'—a different and less interesting story.","headline":"Plausible and novel, but the fake-trigger control confounds 'backdoor' with 'meaningful phrase'; needs a matched non-trigger before the headline Jaccards are trustworthy.","tokens_in":13350,"tokens_out":3279,"would_cite":true,"duration_ms":35563,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Backdoor triggers in LLMs appear to hijack the same attention heads the model already uses to decide its output language, rather than creating new circuits.","keywords":["backdoor attacks","mechanistic interpretability","activation patching","language-switching triggers","attention heads","Gaperon models","multilingual representation","circuit co-option"],"falsifier":"A concrete test: patch the identified trigger heads using, as the corrupted input, a meaningful Latin phrase that is not a backdoor trigger (matched in length and per-word token count). If the Jaccard overlap with natural language heads drops to near baseline, the overlap is driven by semantic content, not by trigger-specific processing. Alternatively, ablate the overlapping language/trigger heads and check whether the trigger still switches the output language — the paper itself lists causal verification of head necessity as future work; if the switch survives ablation, the overlap is not loa","tokens_in":12354,"feed_emoji":"🧠","tokens_out":4420,"duration_ms":41123,"temperature":0.7,"pith_summary":"This paper tries to establish that when a backdoor trigger is injected during pre-training to make a language model switch its output language, the trigger does not create a new dedicated circuit. Instead, the trigger appears to activate the same attention heads the model already uses to decide what language it is writing in. The evidence comes from activation patching on a family of three model scales (1B, 8B, 24B parameters): the set of heads a trigger engages overlaps with the set of heads that encode output language, with Jaccard indices between 0.18 and 0.66 over the top-10 heads, far above shuffled baselines. Trigger information also consolidates early — at roughly 7.5% to 25% of model depth. If correct, this means backdoor detection and defense could target the model's known language components rather than hunting for hidden trigger circuits.","feed_headline":"Backdoor triggers co-opt LLMs' language heads, not new circuits","feed_subtitle":"Activation patching shows trigger and language processing share attention heads across model scales, reshaping how to detect backdoors.","key_machinery":"Activation patching is the central method: the authors run the model on a clean input (with a genuine trigger or a non-English context) and on a corrupted input (a matched fake trigger or English context), then replace a head's mean activation from the corrupted run with the clean one and measure the change in output-token log probability. Heads are ranked by this patching effect, and the top-10 heads for trigger processing and for natural language encoding form the two sets whose overlap is measured with the Jaccard index. Layer-wise patching across trigger token positions then locates where trigger information consolidates. The load-bearing comparison is between these two head sets, and th","core_discovery":"On the paper's own terms, the central discovery is that language-switching backdoor triggers co-opt the model's existing language circuitry rather than forming isolated mechanisms. By comparing activation-patching results over top-10 attention head sets, the authors show that heads activated by a genuine trigger overlap substantially with heads that naturally encode output language identity, with Jaccard indices of 0.18 to 0.66 across model sizes and languages — against near-zero shuffled baselines. The overlap holds for both French and German triggers, and the very same 'language heads' are shared across four tested target languages, including two that have no trigger at all. The authors al","pith_inferences":["If backdoors must route through pre-existing functional components, then analogous backdoors (e.g., sentiment or topic shifts) might similarly co-opt the model's natural sentiment or topic circuits — a prediction that could be tested on other backdoor types.","The 1B German trigger's two-stage formation pattern, which the paper notes is consistent with an induction head, suggests a concrete copying mechanism; probing whether that head copies the language representation from earlier positions would give a causal test of the co-option story.","The counterfactual control's asymmetry (meaningful Latin phrase vs random fake tokens) means the overlap may partly reflect content-bearingness rather than backdoor-specific processing; using a meaningful non-trigger Latin phrase as corrupted input would discriminate between these interpretations.","If the entanglement is real, backdoor 'stealth' may be limited by the model's own functional architecture — an attacker cannot fully hide a trigger if it must excite the same heads as normal language behavior, which would also explain why the injected triggers in these models are discoverable by these patching methods."],"forward_implications":["Backdoor detection could shift from hunting for anomalous hidden circuits to monitoring the activation patterns of known functional components, such as language heads.","If triggers co-opt existing circuitry, mitigation may be possible by intervening on those shared heads, for example by steering or ablating the language-identity direction.","Because the same language heads serve multiple target languages, a single detection or defense mechanism may generalize across trigger languages without retraining.","Early trigger formation (7.5%–25% of depth) suggests that shallow-layer interventions could catch backdoor behavior before it propagates to the output.","The shared-heads finding across scales indicates the co-option mechanism is a general property of these models, not a quirk of one size."],"fun_headline_variants":["Backdoor triggers co-opt LLMs' existing language circuits","Triggers reuse language heads in LLMs, no new circuits","LLM backdoors exploit natural language-processing heads","Backdoor triggers piggyback on LLM language encoding"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The counterfactual 'fake triggers' used as corrupted inputs are random token sequences matched only for length, while the real trigger is a meaningful Latin phrase, so the patching contrast may isolate 'meaningful vs meaningless' rather than 'trigger vs non-trigger'.","fun_headline_variants_meta":{"raw":{"variants":["Backdoor triggers co-opt LLMs' existing language circuits","Triggers reuse language heads in LLMs, no new circuits","LLM backdoors exploit natural language-processing heads","Backdoor triggers piggyback on LLM language encoding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000604,"raw_usage":{"total_tokens":2643,"prompt_tokens":720,"completion_tokens":1923,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":1856}},"tokens_in":464,"tokens_out":1923,"duration_ms":14834,"temperature":1.0,"reasoning_tokens":1856,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T01:08:01.233272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: patch the identified trigger heads using, as the corrupted input, a meaningful Latin phrase that is not a backdoor trigger (matched in length and per-word token count). If the Jaccard overlap with natural language heads drops to near baseline, the overlap is driven by semantic content, not by trigger-specific processing. Alternatively, ablate the overlapping language/trigger heads and check whether the trigger still switches the output language — the paper itself lists causal verification of head necessity as future work; if the switch survives ablation, the overlap is not loa","supporting_citations":[],"review_version":1}