Pith. sign in

REVIEW 3 major objections 5 minor 20 references

Domain-specific abliteration can strip cybersecurity refusal from a trillion-parameter model while leaving other safety refusals largely intact.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 07:32 UTC pith:ZZ7TZBTA

load-bearing objection Useful 24-model map plus a real 1T cyber-only weight carve-out; selectivity is real but still partly distribution-tied. the 3 major comments →

arxiv 2607.02714 v2 pith:ZZ7TZBTA submitted 2026-07-02 cs.CR cs.AI

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

classification cs.CR cs.AI
keywords abliterationsafety alignmentrefusal directiondomain-specific abliterationcybersecurityorthogonal projectionlarge language modelsMoE
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Safety training in large language models does not distinguish domains of harm, so models refuse legitimate cybersecurity work along with violence or bioweapons. This paper shows that refusal lives in a multi-dimensional subspace and that a standard weight-projection method can remove mainly the cybersecurity slice. Across 24 open models from 0.6B to about 1T parameters, the authors extract a refusal direction from paired offensive versus educational cyber prompts, project it out of selected layers, and measure per-domain refusal. On Kimi K2 they drop cybersecurity refusal from 100% to 7% while explicit-content refusal stays at 100% and other domains retain substantial refusal; MMLU capability is essentially preserved everywhere. Susceptibility is not size-driven: safety-training method and dense-versus-MoE architecture are the strongest predictors, and models fall into three susceptibility tiers. The practical stake is clear for authorized offensive-security tooling and for anyone who needs to know how brittle current alignment really is.

Core claim

Using the standard abliteration pipeline with a cybersecurity-focused extraction set, the authors obtain domain-selective safety removal: on Kimi K2, cybersecurity refusal falls from 100% to 7% while explicit-content refusal remains 100% and other non-target domains keep 44–88% refusal, with MMLU degradation at most 0.028 across all 24 models. The effect is model-specific, not size-specific, and is best predicted by safety-training method and architecture.

What carries the argument

Domain-specific abliteration: a mean-difference refusal direction is computed from last-token hidden states on harmful versus educational cybersecurity prompt pairs, then orthogonally projected out of down_proj and o_proj weights on a uniform spread of layers (25–95% of depth).

Load-bearing premise

The mean-difference direction taken from the authors’ CTF-style cyber prompt pairs isolates a clean cybersecurity-harm component of refusal, not a mixture that also carries general refusal or neighboring domains.

What would settle it

Apply the same projection to Kimi K2 (or another high-susceptibility model) and re-evaluate on held-out cybersecurity prompts far from the CTF extraction style plus matched non-target harms; if cyber and non-target refusal drop by similar amounts, domain selectivity fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies domain-specific abliteration: orthogonal projection of a mean-difference refusal direction extracted from cybersecurity-focused harmful/harmless prompt pairs, applied to down_proj and o_proj weights over a uniform 25–95% depth layer spread. Across 24 open-source models (0.6B–1T, dense and MoE), the authors report heterogeneous per-domain refusal changes, classify models into three susceptibility tiers, and identify safety-training method and architecture as the strongest predictors. The headline result is on Kimi K2 (~1T MoE): on a custom cross-evaluation set, cybersecurity refusal falls from 100% to 7% while explicit-content refusal stays at 100% and other non-target domains retain 44–88% refusal (Table 2a), with MMLU preserved within 0.028 across all models. They argue this shows that multi-dimensional refusal geometry admits selective, permanent weight-level removal of a cybersecurity-harm slice without global safety collapse.

Significance. If the domain-selectivity claim holds under stronger controls, the work is a substantial empirical contribution to representation-level safety and to dual-use cybersecurity tooling. Strengths that should be credited: (i) scale and uniformity of the protocol (24 models, two evaluation distributions, MMLU controls, layer-selection ablation in Appendix C); (ii) permanent weight modification at trillion-parameter MoE scale, beyond concurrent inference-time methods (RepIt, CAST); (iii) clean demonstration that MMLU is essentially unaffected; (iv) useful correlational evidence linking DPO-style safety training to higher abliteration susceptibility and GRPO/undisclosed RL to resistance (Appendix E). These results are falsifiable and practically actionable for both safety researchers and authorized offensive-security practitioners.

major comments (3)
  1. [§3, Table 2a, Appendix D] Central claim of geometric domain selectivity rests on purity of the CTF-derived mean-difference direction (§3). The paper reports that generic harmful/harmless pairs caused global safety collapse, that post-abliteration compliance is distribution-specific (§4), and that misinformation is the most reduced non-target domain on average (Table 1, Fig. 2). Appendix D cosine similarities (cyber vs. others 0.56–0.70) are only moderate and were computed on five models that exclude Kimi K2. No control is reported that (a) orthogonalizes a general-refusal direction out of the cyber direction before weight edit, (b) measures residual cosine of the Kimi-extracted vector with non-cyber refusal directions, or (c) evaluates held-out cyber prompts whose surface form is distant from the CTF extraction set. Without at least one of these, Table 2a’s selectivity may be an artifact of prompt-distribution pr
  2. [Table 2, §4, Abstract] The two evaluation sets disagree sharply on effect size and selectivity for the same Kimi K2 intervention (α=1.0, 30% layers): cross-evaluation cyber 100%→7% (−93 pp) with strong non-target retention (Table 2a) versus scientific benchmark cybercrime 88%→51% (−37 pp) and smaller, more uniform drops elsewhere (Table 2b). The abstract and contributions lead with the stronger numbers. The manuscript must either (i) designate one primary evaluation protocol and report headline claims only on it, or (ii) provide a quantitative analysis of why domain separation collapses under the scientific distribution (prompt overlap, label granularity, surface-form distance) and qualify the “domain-specific abliteration is achievable” claim accordingly.
  3. [§3.3, Figure 1, Table 2] Refusal is detected by 31 string-matching patterns after stripping <think> blocks (§3.3). Pattern matchers systematically miss soft refusals, partial answers, and policy-compliant rephrasings, and can false-positive on educational disclaimers. For a paper whose primary dependent variable is per-domain refusal rate (Fig. 1, Tables 1–2), this is load-bearing measurement error. At minimum, a human or strong-LLM judge audit on a stratified sample (e.g., 50–100 responses per domain for Kimi K2 and 2–3 other tiers) should be reported, with inter-annotator agreement and sensitivity of the main deltas to the detection method.
minor comments (5)
  1. [Figure 1] Figure 1’s dual heatmaps are dense; the right panel’s sign convention (negative = refusal drop) should be stated in the caption, and the cybercrime column should be visually highlighted given the paper’s focus.
  2. [§5.1, Figure 5] Susceptibility tiers (High / Partial / Resistant) are defined by maximum refusal reduction at 30% but the exact cutoffs are not stated in the main text; add them near Figure 5.
  3. [§3] α > 1.0 is said to help MoE models (§3) but no systematic α-sweep results are shown for Kimi K2 or other MoEs; a short appendix table would strengthen the methodological claim.
  4. [§2] Related-work placement of RepIt and CAST is fair; clarify more explicitly that those are inference-time interventions while this work permanently edits weights, so robustness comparisons are not apples-to-apples.
  5. [Abstract, Appendix E] Typos / polish: “3abliteration” missing space (Abstract); “BOND+W ARM+W ARP” spacing inconsistency (Appendix E); arXiv id in header is 2607.02714 while some internal refs use 2026 dates—ensure consistency before camera-ready.

Circularity Check

0 steps flagged

Empirical intervention study with no derivation that reduces to its inputs by construction; susceptibility tiers and domain selectivity are measured outcomes, not fitted or self-defined predictions.

full rationale

The paper applies the standard Arditi-style mean-difference orthogonal projection (Section 3: r̂_ℓ = (m⁺_ℓ − m⁻_ℓ)/‖·‖, W′ = W − α r̂(r̂ᵀW)) using a cybersecurity-focused harmful/harmless extraction set, then measures refusal on two separate evaluation suites (scientific benchmark from HarmBench/AdvBench/PurpleLlama; custom cross-evaluation set) plus MMLU. Domain selectivity on Kimi K2 (Table 2), per-domain mean drops (Table 1, Fig. 2), susceptibility tiers (Fig. 5: labels by observed max refusal reduction at 30%), and safety-training/architecture correlations (Fig. 6, Appendix E) are all post-hoc summaries of those measurements. No parameter is fitted to a subset and then reported as a prediction of a closely related quantity; no uniqueness theorem or ansatz is imported from the authors’ own prior work; external geometric citations (Wollschläger et al., Pan et al.) are used only as motivation, not as load-bearing uniqueness constraints. Concerns about whether the extracted direction is a pure cybersecurity slice versus a mixture (distribution-specific compliance, misinfo co-drop, Appendix D cosines) are validity/interpretation issues, not circular reductions of claim to input. The work is self-contained empirical measurement against external benchmarks.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The central claim rests on the linear-representation hypothesis and the multi-dimensional refusal geometry established by prior work, plus a handful of experimental free choices (layer fraction, α, prompt-pair design). No new physical entities are postulated; the three susceptibility tiers are descriptive labels, not causal mechanisms. The ledger is therefore light: two domain assumptions, a few free experimental parameters, and one invented classification scheme.

free parameters (3)
  • abliteration strength α = 1.0 (main); >1.0 for MoE
    Set to 1.0 for exact orthogonal projection; >1.0 used for MoE compensation. Chosen by hand, not derived.
  • layer percentile range and fraction = 25–95 % depth, ≤30 % layers
    Uniform spread across 25–95 % of depth at 5–30 % of layers. Selected after comparing to norm-based selection; not theoretically fixed.
  • extraction-dataset composition = 52/73 CTF pairs
    52 harmful + 73 harmless CTF-style cybersecurity pairs for Kimi K2; generic harmful/harmless pairs produced global collapse. Hand-curated to isolate offensive intent.
axioms (3)
  • domain assumption Refusal (and other behavioral concepts) is linearly represented in residual-stream activations and can be removed by orthogonal projection of weight matrices.
    Taken from Arditi et al. 2024 and the linear-representation hypothesis literature (Park, Elhage, Zou); invoked throughout §3.
  • domain assumption The multi-dimensional refusal subspace admits domain-selective slices that can be targeted independently.
    Supported by Wollschläger et al. and Pan et al.; the paper's conjecture that cybersecurity can be carved out rests on this geometric claim.
  • ad hoc to paper Last-token hidden states under causal attention capture the full-sequence refusal signal better than mean pooling.
    Authors state they experimented with mean pooling and found it weaker; used as operational axiom for direction extraction (§3).
invented entities (1)
  • three abliteration-susceptibility tiers (High / Partial / Resistant) no independent evidence
    purpose: Taxonomy that organizes the 24-model results and links them to safety-training method and architecture.
    Descriptive labels defined by maximum refusal reduction at 30 % abliteration; no independent causal mechanism or external falsifiable prediction is supplied.

pith-pipeline@v1.1.0-grok45 · 19040 in / 3139 out tokens · 30632 ms · 2026-07-12T07:32:21.079820+00:00 · methodology

0 comments
read the original abstract

There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains and the level of potential harm of a query, which creates significant complications in the fields like cyber security, where a model should not be constrained by its safety circuits to accomplish the goals of legitimate, authorized operations. In this work, we share our findings from a large scale abliteration experiment on 24 open-source LLMs and show that domain-specific abliteration is achievable with standard methodology on the example of a 1T-parameter Kimi K2. Building on recent work showing that refusal in LLMs occupies a multi-dimensional subspace within layers, we find that it is also distributed widely across layers, especially in trillion-parameter MoE architectures, and so we aim to capture the part of it that represents harmful concepts in the cybersecurity domain exclusively. We also investigate the correlation between models' features and the effect of domain-specific abliteration, identifying that the type of safety training and architecture are the most reliable predictors. Finally, we classify the models into 3 abliteration susceptibility tiers and put forward a set of conjectures as to why a particular effect from this intervention might be observed in a given model.

Figures

Figures reproduced from arXiv: 2607.02714 by Artem Sorokin, Dario Pasquini, Vadym Hadetskyi.

Figure 1
Figure 1. Figure 1: Per-domain refusal rates at 30% abliteration (left) and change from baseline (right) across all [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Mean refusal reduction per domain at 30% abliteration. Left: all 24 models. Right: [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: MMLU vs. refusal rate trajectories (baseline [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Refusal change at different intensities of abliteration for a sample of models [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Maximum refusal reduction across 24 models, grouped by susceptibility class. Models [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Susceptibility predictors: MMLU baseline score vs. maximum cybercrime refusal drop. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Norm-based vs. uniform spread layer selection across 9 models. Uniform spread (red) [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Per-domain refusal direction similarity on the scientific benchmark, averaged across 5 [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Per-domain refusal direction similarity on the cross-evaluation dataset, averaged across 5 [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Maximum refusal reduction by safety training method. DPO-based methods are most [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Susceptibility by architecture type (left) and model family (right). Dense and MoE show [PITH_FULL_IMAGE:figures/full_fig_p014_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

20 extracted references · 17 linked inside Pith

  1. [1]

    An embarrassingly simple defense against LLM abliteration attacks.arXiv preprint arXiv:2505.19056,

    Abu Shairah, A., et al. An embarrassingly simple defense against LLM abliteration attacks.arXiv preprint arXiv:2505.19056,

  2. [2]

    PurpleLlama CyberSecEval: A secure coding benchmark for language models.arXiv preprint arXiv:2312.04724,

    Bhatt, S., Chennabasappa, S., et al. PurpleLlama CyberSecEval: A secure coding benchmark for language models.arXiv preprint arXiv:2312.04724,

  3. [3]

    Jailbreaking black box large language models in twenty queries.arXiv preprint arXiv:2310.08419,

    Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G.J., and Wong, E. Jailbreaking black box large language models in twenty queries.arXiv preprint arXiv:2310.08419,

  4. [4]

    A granular study of safety pretraining under model abliteration.arXiv preprint arXiv:2510.02768,

    Henderson, P., et al. A granular study of safety pretraining under model abliteration.arXiv preprint arXiv:2510.02768,

  5. [5]

    Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300,

    Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300,

  6. [6]

    Kimi K2: Open agentic intelligence.arXiv preprint arXiv:2507.20534,

    Kimi Team. Kimi K2: Open agentic intelligence.arXiv preprint arXiv:2507.20534,

  7. [7]

    A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity.arXiv preprint arXiv:2401.01967,

    Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J.K., and Mihalcea, R. A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity.arXiv preprint arXiv:2401.01967,

  8. [8]

    LoRA fine-tuning efficiently undoes safety training in Llama 2-Chat 70B.arXiv preprint arXiv:2310.20624,

    Lermen, S., Rogers-Smith, C., and Ladish, J. LoRA fine-tuning efficiently undoes safety training in Llama 2-Chat 70B.arXiv preprint arXiv:2310.20624,

  9. [9]

    AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.arXiv preprint arXiv:2310.04451,

    Liu, X., Xu, N., Chen, M., and Xiao, C. AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.arXiv preprint arXiv:2310.04451,

  10. [10]

    and Tegmark, M

    Marks, S. and Tegmark, M. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.arXiv preprint arXiv:2310.06824,

  11. [11]

    HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.arXiv preprint arXiv:2402.04249,

    Mazeika, M., Phan, L., Yin, X., Zou, A., et al. HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.arXiv preprint arXiv:2402.04249,

  12. [12]

    Steering Llama 2 via contrastive activation addition.arXiv preprint arXiv:2312.06681,

    Panickssery, N., Bowman, S.R., and Feng, S. Steering Llama 2 via contrastive activation addition.arXiv preprint arXiv:2312.06681,

  13. [13]

    The linear representation hypothesis and the geometry of large language models.arXiv preprint arXiv:2311.03658,

    Park, K., Choe, Y .J., and Veitch, V . The linear representation hypothesis and the geometry of large language models.arXiv preprint arXiv:2311.03658,

  14. [14]

    Linear representations of sentiment in large language models.arXiv preprint arXiv:2310.15154,

    Tigges, C., Hollinsworth, O.J., Geiger, A., and Nanda, N. Linear representations of sentiment in large language models.arXiv preprint arXiv:2310.15154,

  15. [15]

    Activation addition: Steering language models without optimization.arXiv preprint arXiv:2308.10248,

    Turner, A., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M. Activation addition: Steering language models without optimization.arXiv preprint arXiv:2308.10248,

  16. [16]

    SORRY-Bench: Systematically evaluating large language model safety refusal behaviors.arXiv preprint arXiv:2406.14598,

    Xie, T., et al. SORRY-Bench: Systematically evaluating large language model safety refusal behaviors.arXiv preprint arXiv:2406.14598,

  17. [17]

    Shadow alignment: The ease of subverting safely-aligned language models.arXiv preprint arXiv:2310.02949,

    Yang, X., et al. Shadow alignment: The ease of subverting safely-aligned language models.arXiv preprint arXiv:2310.02949,

  18. [18]

    Comparative analysis of LLM abliteration methods: A cross-architecture evaluation.arXiv preprint arXiv:2512.13655,

    Young, A. Comparative analysis of LLM abliteration methods: A cross-architecture evaluation.arXiv preprint arXiv:2512.13655,

  19. [19]

    Representation engineering: A top-down approach to AI transparency.arXiv preprint arXiv:2310.01405,

    Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., et al. Representation engineering: A top-down approach to AI transparency.arXiv preprint arXiv:2310.01405,

  20. [20]

    Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043,

    Zou, A., Wang, Z., Kolter, J.Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043,