Pith. sign in

REVIEW 3 major objections 6 minor 14 references

SAE feature ablations are not uniformly localized safety handles: their apparent efficiency over dense steering is largely a baseline-matching artifact and reverses under fair surface-matched comparison.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 13:22 UTC pith:3XODFGME

load-bearing objection Solid evaluation paper: SAE safety localization is baseline- and regime-dependent, and the Gemma-9B surface/basis reversal under two judges is the result that sticks. the 3 major comments →

arxiv 2607.10226 v1 pith:3XODFGME submitted 2026-07-11 cs.AI cs.CR

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

classification cs.AI cs.CR
keywords sparse autoencodersactivation steeringsafety interventionsrefusal directionsmatched evaluationcoherence gatingjailbreak metricsfeature ablation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks when sparse autoencoder features actually act as localized control handles for safety-relevant behavior such as refusal and jailbreak susceptibility. Apparent success can come from weak interventions, mismatched dense baselines, capability damage, or incoherent text that safety judges still mark unsafe. The authors introduce a matched coherence-gated evaluation protocol that compares sparse and dense interventions at equal target effect and equal perturbation, counts a jailbreak only when an output is both judge-unsafe and coherent, matches the intervention surface and subspace, and cross-checks with a second judge. On Gemma-2-9B, once dense steering is restricted to the same layer or to the SAE decoder span, the SAE efficiency advantage disappears and reverses: fair dense baselines produce more coherent harmful compliance at matched perturbation. High-strength SAE ablations mainly collapse coherence, and the same recipe fails on a small 2B model through capability collapse and single-judge inflation. The practical message is that sparse localization for safety must be treated as regime-dependent and evaluation-dependent, not assumed from sparsity alone.

Core claim

SAE feature ablation is not a uniformly localized safety control mechanism. Matching only total perturbation norm leaves the intervention surface unmatched and can make single-layer SAE look more efficient than all-layer dense steering. On Gemma-2-9B, same-layer and decoder-span-projected dense baselines reverse that advantage at every matched bin (true-jailbreak deficits up to about -0.29 under two judges), while high-k ablations mainly induce coherence collapse and small-model SAE jailbreaks are largely single-judge inflation with capability collapse. The useful medium-k regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank, and

What carries the argument

Matched Coherence-Gated (MCG) Evaluation: two complementary controls (matched target-effect and matched perturbation-norm), plus surface and basis matching (same-layer dense and dense projected onto the top-k SAE decoder span), a primary true-jailbreak metric that requires both judge-unsafe and coherent output, and a second behavior-completion judge. It turns localization into a measured behavior-per-perturbation comparison rather than a raw success claim.

Load-bearing premise

That equal per-token residual change, plus matching the layer and the SAE decoder subspace, is a fair enough proxy for localized control; if those still leave important differences unmeasured, the Gemma-2-9B reversal could partly reflect residual mismatch rather than true non-locality of the features.

What would settle it

Re-run the surface- and basis-matched perturbation-norm comparison on Llama-3.1-8B and Gemma-2-27B (not only all-layer dense): if SAE still wins after same-layer and decoder-span matching under two judges and retained capability, the baseline-dependence thesis weakens; if the Gemma-2-9B-style reversal appears there too, the claim strengthens.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Sparse-versus-dense safety claims that only match total perturbation can misread an all-layer-versus-single-layer mismatch as localization.
  • High-k SAE ablations should not be scored with unsafe-only judges; coherence gates and a second behavior-completion judge are required.
  • SAE steering results from small models should not be extrapolated to larger ones without capability and multi-judge checks.
  • The right intervention tool depends on direction: sparse ablation can be more surgical for removing refusal-related behavior, while dense steering better preserves capability when injecting refusal.
  • Feature diagnostics of rank-decay and head stability should accompany top-k selection rather than treating k as a free hyperparameter.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If surface and basis matching routinely flip SAE-versus-dense rankings, many published sparse-localization wins may need re-evaluation under the same protocol.
  • A practical safety stack may need hybrid controls: sparse heads for targeted removal, dense directions for broad refusal injection.
  • The rapid rank-decay of refusal alignment suggests diminishing returns and rising side-effect risk beyond a medium-k head, which could guide automatic k selection.
  • Judge disagreement at small scale is itself a diagnostic of off-distribution degeneration, not only a measurement nuisance.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper asks when SAE feature ablations act as localized control handles for safety-relevant behavior, and argues that apparent localization is highly sensitive to how dense baselines are matched. It introduces Matched Coherence-Gated (MCG) evaluation: complementary matched target-effect and matched perturbation-norm controls, a primary true-jailbreak metric requiring both judge-unsafe and coherent outputs, surface/basis-matched dense baselines (same-layer and decoder-span projected), and a second behavior-completion judge. On Gemma-2-9B with a Gemma Scope layer-20 SAE, the apparent SAE efficiency advantage over all-layer dense steering reverses under surface- and basis-matched dense baselines (true-jailbreak deficits up to about −0.29 under two judges), high-k ablations mainly induce coherence collapse, and a stable medium-k refusal-aligned feature head explains the useful regime. Against all-layer dense baselines the SAE advantage remains large on Llama-3.1-8B and Gemma-2-27B, while the same recipe fails on Gemma-2-2B via capability collapse and single-judge inflation. Refusal injection reverses the efficiency ranking in favor of dense steering.

Significance. If the results hold, the paper supplies a concrete evaluation standard for a claim that is currently often asserted rather than measured: that sparse SAE features are more localized safety handles than dense activation steering. The surface- and basis-matched reversal on Gemma-2-9B (Table 5), multi-seed paired bootstrap intervals, dual-judge cross-check, random-SAE and retain-set negative controls, feature-rank diagnostics, and the direction-dependent injection pairing are genuine strengths. The main contribution is methodological and operational rather than circuit-level: it shows that naive total-norm matching can manufacture an apparent localization advantage, and that SAE safety control should be treated as regime-, scale-, and direction-dependent. That is a useful corrective for both interpretability and safety-intervention evaluation.

major comments (3)
  1. [§5.2–5.3, Tables 5, 7–8] §5.2 Table 5 vs §5.3 Tables 7–8 and §8: The load-bearing result that surface/basis matching reverses the SAE advantage is shown only for Gemma-2-9B. Llama-3.1-8B and Gemma-2-27B retain a large SAE advantage only against all-layer dense steering—the baseline the paper itself argues is insufficient (§3, contribution 2). Because the central thesis is that baseline specification (surface/basis), not sparsity, decides whether SAE looks localized, the architecture/scale claims should either include same-layer and decoder-span projected dense baselines or be more tightly scoped throughout the abstract, §5.3, and conclusion as “advantage over all-layer dense,” without implying that the localization conclusion transfers. As written, the multi-model narrative partially reintroduces the unmatched-surface comparison the protocol was designed to eliminate.
  2. [§4.3, §5.2, §8] §4.3 and §5.2: Sharing a single refusal contrast to both rank SAE features and construct the dense direction is disclosed and intentional, but it means the comparison isolates intervention basis given a fixed target rather than testing whether SAE independently discovers a safety concept. That is fine for the stated claim, yet several passages (e.g., “refusal-aligned head,” “localized control handles”) can be read as stronger mechanistic localization. A short, explicit restatement near Table 5 and in the conclusion that no independent concept-discovery claim is made would prevent over-reading of the operational efficiency result.
  3. [§5.3 Table 8, §7] §5.3 Table 8 and §7: The 27B Pareto-dominance claim is restricted to safety/coherence/perturbation under 4-bit loading with capability at floor (GSM8K 0.12, MMLU 0.59). Quantization can interact differently with sparse feature ablation than with dense residual steering; without a full-precision check or a quantization-sensitivity note beyond the current limitation paragraph, the claim that the clean regime “grows” from 9B to 27B remains only partially supported. Either add a limited full-precision or higher-precision sanity run on a subset, or further qualify the 27B result as a quantized trend only.
minor comments (6)
  1. [§5.4] §5.4 human audit: The audit is correctly labeled targeted and single-annotator (n=101). For a primary metric definition, even a small second-annotator agreement subset (e.g., 30–40 items with reported κ) would make the gate more credible without requiring a full multi-annotator study.
  2. [Code and Data Availability] Code and Data Availability: The promised public artifact (protocol, same-layer/projected baselines, HarmBench rescorer, per-example metrics, pinned sae_lens 6.44.2) is central to reproducibility, especially given the manual Llama Scope JumpReLU/normalization steps in §4.1. Please ensure release coincides with revision or provide a stable anonymous repository link for review.
  3. [Table 2] Table 2 footnote on baseline MMLU 0.680 vs later 0.692 is helpful; consider moving that caveat into the table caption so readers do not compare cells across evaluation passes.
  4. [Figures 1, 3] Figure 1 and Figure 3: Low-coherence points and scale panels are informative; adding explicit matched-bin markers (or a small legend for coherence threshold) would make the Pareto and cross-scale plots easier to read in grayscale.
  5. [§4.1] §4.1 Llama Scope loading details (threshold 0.3555, norm √d, BOS exclusion) are important for reproduction; consider a short appendix box or checklist so they are not buried in prose.
  6. [Eq. (1), figure captions] Minor wording: “true jailbreak” is defined clearly in Eq. (1), but occasional use of “jailbreak” alone in figure captions can blur the gated vs unsafe-only distinction; prefer the gated term consistently in captions.

Circularity Check

0 steps flagged

No significant circularity: empirical matched-evaluation study whose claims are measured outcomes, not results forced by definition or self-citation.

full rationale

This paper is an empirical evaluation of SAE feature ablation versus dense refusal-direction steering under a matched coherence-gated protocol. Its load-bearing claims (sign-flip of SAE efficiency on Gemma-2-9B once surface/basis are matched; high-k coherence collapse; small-model single-judge inflation; direction-dependent injection result) are experimental measurements on held-out prompts with multi-judge and capability metrics, not derivations that reduce to their inputs by construction. True-jailbreak is defined independently as unsafe-and-coherent and then measured; it is not fitted from the intervention parameters. Sharing one refusal contrast to rank SAE features and build the dense direction is an explicit design choice to hold the target fixed and isolate the intervention basis—not a circular proof that sparse features control refusal. Ranking features by cosine similarity to the refusal direction and then reporting activation separation (Cohen’s d) on refused vs. complied prompts is a related diagnostic, not a self-definitional prediction of the main efficiency results; stability across splits, rank-decay geometry, single-feature ablation weakness, and random-SAE negative controls are additional empirical checks. Citations (Arditi, Bricken, Gemma Scope, Llama Scope, HarmBench, etc.) are external prior work, not load-bearing self-citation uniqueness theorems. The paper itself frames localization as an operational behavior-per-perturbation comparison rather than causal circuit isolation. No step reduces Eq. X to Eq. Y by construction, renames a fitted parameter as a prediction, or smuggles an ansatz via author-overlapping uniqueness claims. Score 0 is therefore the correct honest finding.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The paper is an empirical evaluation study. It inherits standard SAE and residual-stream intervention machinery, automated judges, and capability benchmarks from prior literature. Load-bearing modeling choices are the operational definition of localization (behavior-per-perturbation under matched surface/basis), the coherence gate, the shared refusal contrast for ranking, and the discrete k/beta grids. No new physical entities are postulated; free parameters are experimental design knobs rather than fitted universal constants.

free parameters (5)
  • SAE top-k grid (50–3200, focus on 400/800/1600/3200)
    Discrete ablation widths chosen by the authors; the 'useful regime' boundary is defined relative to these hand-selected ranks.
  • Dense scaling coefficients beta / gamma grids
    Intervention strengths swept and then binned for matching; bin centers (e.g. ~0.15/0.20/0.30 relative norm) are design choices that structure the comparison.
  • SAE layer choices (Gemma-9B L20, Gemma-2B L12, Llama L15)
    Primary layers selected from available SAE suites; layer scan partially mitigates but primary claims use these fixed layers.
  • Coherence-gate heuristic thresholds (length, lexical diversity, alphabetic ratio, repeated n-grams)
    Automatic gate parameters define which unsafe outputs count as true jailbreaks; human audit supports them but thresholds remain author-chosen.
  • Llama Scope JumpReLU threshold 0.3555 and activation normalization to sqrt(d)
    Loader corrections required for non-degenerate features; values taken from release details but critical to Llama results.
axioms (5)
  • domain assumption Per-token relative residual change is a valid proxy for intervention strength / locality when comparing methods.
    Stated in Sections 3 and 7 as operational, not causal isolation; central efficiency claims rest on this proxy.
  • domain assumption Unsafe-and-coherent (true jailbreak) better measures meaningful harmful compliance than unsafe-only judge labels.
    Protocol definition Eq. (1); supported by human audit but still a measurement axiom.
  • ad hoc to paper Sharing one refusal contrast to rank SAE features and build the dense direction fairly isolates intervention basis.
    Section 4.3 design choice; avoids confounding target discovery with basis comparison.
  • domain assumption Llama-Guard and HarmBench labels, gated by coherence, are adequate primary/secondary safety metrics for held-out prompts.
    Standard in the subfield; paper adds second judge and audit rather than proving judge correctness.
  • standard math Standard residual-stream SAE encode/decode and dense activation addition/ablation mathematics.
    Background from Bricken/Cunningham/Lieberum/He and activation steering literature.
invented entities (2)
  • Matched Coherence-Gated (MCG) Evaluation protocol no independent evidence
    purpose: Separate target strength, utility cost, and degeneration artifacts when comparing sparse vs dense safety interventions.
    Methodological construct defined in Section 3; not a physical entity, but the paper's main invented evaluation object.
  • true jailbreak metric (unsafe AND coherent) independent evidence
    purpose: Primary target metric that excludes incoherent unsafe-only artifacts.
    Defined in Eq. (1); validated by targeted human audit and second judge, not by external standard alone.

pith-pipeline@v1.1.0-grok45 · 22228 in / 3776 out tokens · 29150 ms · 2026-07-14T13:22:09.411410+00:00 · methodology

0 comments
read the original abstract

We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated safety judges mark as unsafe without representing meaningful harmful compliance. We introduce a matched coherence-gated evaluation protocol for runtime safety interventions: methods are compared at matched target-effect points, and the primary target metric counts harmful compliance only when an output is both judge-unsafe and coherent. Applying this protocol to three prompt splits on Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, we find that SAE feature ablation has a narrow useful regime. SAE top800 reaches a low-to-mid target effect with lower total perturbation and competitive utility, but SAE top1600 loses utility relative to a matched dense refusal-direction baseline, and SAE top3200 primarily induces coherence collapse. Human audit confirms that coherence gating removes unsafe-only artifacts, and feature diagnostics show that the useful regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank. These results argue that SAE-based safety interventions should be evaluated as regime-dependent control mechanisms rather than assumed to be uniformly localized.

Figures

Figures reproduced from arXiv: 2607.10226 by Daming Luo.

Figure 1
Figure 1. Figure 1: Target effect versus total perturbation on Gemma-2-9B. SAE reaches comparable low-to-mid target effects with lower total relative perturbation, but high-k SAE enters a coherence-collapse regime. 5.2 RQ2: Matched-TE Ties; the Matched-PN Advantage Reverses Under Sur￾face Matching Matched target-effect: the raw advantage ties. To control for intervention strength, we first compare dense and SAE methods at mat… view at source ↗
Figure 2
Figure 2. Figure 2: Utility at matched target-effect points on Gemma-2-9B. SAE top800 is competitive near the lower target-effect point, but SAE top1600 loses utility relative to dense β = 0.20 [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Matched perturbation-norm summary across model scale (Gemma-2 2B/9B/27B). (A) True￾jailbreak rate: SAE top1600 versus dense and the random-SAE control. (B) Coherence retention: SAE holds coherence while dense collapses as it is pushed, increasingly so at 27B. (C) SAE coherence advantage over dense β = 0.15: near zero at 2B and 9B, large and positive at 27B. The clean regime strengthens upward from 9B to 27… view at source ↗
Figure 4
Figure 4. Figure 4: Capability sensitivity on Gemma-2-9B. Random SAE top1600 strongly damages GSM8K, showing that large sparse-feature ablations can disrupt reasoning even when headline safety metrics look benign. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Targeted human audit of coherence and harmfulness labels (101 sampled Gemma-2-9B outputs, single annotator). The audit supports the coherence gate: high-strength SAE unsafe-incoherent examples are degenerate, while unsafe-coherent SAE examples are mostly meaningful harmful compliance. A second judge cross-checks the metric. A single safety judge could still systematically mis￾score harmful compliance, so w… view at source ↗
Figure 6
Figure 6. Figure 6: Feature-level diagnostics for the 9B layer-20 SAE. Top-ranked features are stable and refusal￾aligned, while lower-ranked additions have weak activation separation. This explains why medium-k inter￾vention is useful and high-k intervention becomes fragile. The top800 feature set is also stable across splits: 671 features appear in all three top800 sets, 131 appear in exactly two, and 125 appear in only one… view at source ↗
Figure 7
Figure 7. Figure 7: Direction-dependent matched-perturbation pairing on Gemma-2-9B in the safety-improving (refusal-injection) direction. (A) Harmful-compliance suppression: at matched perturbation SAE ampli￾fication suppresses harm more aggressively than dense addition. (B) Capability cost: dense addition holds GSM8K and MMLU flat across the range, while SAE amplification collapses reasoning. The safety gain of SAE injection… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

14 extracted references · 14 linked inside Pith

  1. [1]

    Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L

    arXiv:2406.11717. Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L. Turner, et al. Towards monosemanticity: Decomposing language models with dictio- nary learning,

  2. [2]

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey

    arXiv:2110.14168. Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoen- coders find highly interpretable features in language models,

  3. [3]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al

    arXiv:2309.08600. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. The llama 3 herd of models,

  4. [4]

    arXiv:2407.21783. Igor Fedorov, Kate Plawiak, Lemeng Wu, Tarek Elgamal, Naveen Suda, Eric Smith, Hongyuan Zhan, Jianfeng Chi, Yuriy Hulovatyy, Kimish Patel, Zechun Liu, Changsheng Zhao, Yangyang Shi, Tijmen Blankevoort, Mahesh Pasupuleti, Bilge Soran, Zacharie Delpierre Coudert, Rachad Alao, Raghuraman Krishnamoorthi, and Vikas Chandra. Llama guard 3-1b-i...

  5. [5]

    Gemma Team

    arXiv:2411.17713. Gemma Team. Gemma 2: Improving open language models at a practical size,

  6. [6]

    Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang, and Xipeng Qiu

    arXiv:2408.00118. Zhengfu He, Wentao Shu, Xuyang Ge, Lingjie Chen, Junxuan Wang, Yunhua Zhou, Frances Liu, Qipeng Guo, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang, and Xipeng Qiu. Llama scope: Ex- tracting millions of features from llama-3.1-8b with sparse autoencoders,

  7. [7]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt

    arXiv:2410.20526. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding,

  8. [8]

    Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri

    arXiv:2009.03300. Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models,

  9. [9]

    Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, Janos Kramar, Anca Dragan, Rohin Shah, and Neel Nanda

    arXiv:2406.18510. Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, Janos Kramar, Anca Dragan, Rohin Shah, and Neel Nanda. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2,

  10. [10]

    21 Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks

    arXiv:2408.05147. 21 Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: A standard- ized evaluation framework for automated red teaming and robust refusal,

  11. [11]

    Paul Rottger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy

    arXiv:2402.04249. Paul Rottger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models,

  12. [12]

    Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J

    arXiv:2308.01263. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering,

  13. [13]

    Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J

    arXiv:2308.10248. Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. Representation engineering: A top-...

  14. [14]

    arXiv:2310.01405. 22