Pith. sign in

REVIEW 13 cited by

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17148 v3 pith:GV7UZRHY submitted 2025-01-28 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords steeringfinetuningaxbenchpromptingautoencodersbenchmarkcompetitiveconcept
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability researchers have proposed a variety of representation-based techniques as well, including sparse autoencoders (SAEs), linear artificial tomography, supervised steering vectors, linear probes, and representation finetuning. At present, there is no benchmark for making direct comparisons between these proposals. Therefore, we introduce AxBench, a large-scale benchmark for steering and concept detection, and report experiments on Gemma-2-2B and 9B. For steering, we find that prompting outperforms all existing methods, followed by finetuning. For concept detection, representation-based methods such as difference-in-means, perform the best. On both evaluations, SAEs are not competitive. We introduce a novel weakly-supervised representational method (Rank-1 Representation Finetuning; ReFT-r1), which is competitive on both tasks while providing the interpretability advantages that prompting lacks. Along with AxBench, we train and publicly release SAE-scale feature dictionaries for ReFT-r1 and DiffMean.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance

    cs.CL 2026-08 conditional novelty 7.0 of 10

    ODRA synthesizes CBT sessions with a chain-of-thought CBT aligner and a dynamic patient resistance orchestrator, but its clinical robustness claims rest on evaluations using its own simulated patients.

  2. Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent

    cs.LG 2026-07 conditional novelty 7.0 of 10

    HARP shows that retrieval from an activation database plus linear-probe tools, driven by an LLM agent, matches or exceeds SAEs and activation oracles on four interpretability tasks with zero training.

  3. Temporal Preference Concepts and their Functions in a Large Language Model

    cs.LG 2026-05 unverdicted novelty 6.5 of 10

    Temporal preference in Qwen3-4B-Instruct-2507 localizes to layers 17–35 (especially L24 attention), has curved residual-stream geometry, is behaviorally unstable, and can be bidirectionally steered.

  4. The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail

    cs.LG 2025-12 conditional novelty 6.0 of 10

    Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.

  5. VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A single adversarially optimized image can reproduce activation-steering behavior in multiple VLMs and partially transfer to unseen models.

  6. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Incorrect L0 makes sparse autoencoders mix correlated features rather than disentangling them, and a decoder projection metric can identify the correct L0.

  7. Insights into a radiology-specialised multimodal large language model with sparse autoencoders

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Applying Matryoshka sparse autoencoders to a radiology-specialised multimodal LLM reveals a minority of interpretable clinical features, while steering them produces unreliable and often off-target report changes.

  8. Position: Use Sparse Autoencoders to Discover Unknowns

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Sparse autoencoders are best used to discover unknown concepts, not to act on known concepts.

  9. Resa: Transparent Reasoning Models via SAEs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SAE-Tuning, a sparse-autoencoder-guided SFT procedure, elicits RL-comparable reasoning in 1.5B models from CoT-free QA data at about $1 and 20 minutes of training.

  10. Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA

    cs.CL 2025-09 reject novelty 5.0 of 10

    Symbolic triggers like modifiers and named entities keep Gemma hallucination rates at 64-79% across model scales, suggesting larger models do not eliminate this failure mode.

  11. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  12. From Emergence to Control: Probing and Modulating Self-Reflection in Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    Self-reflection in LLMs can be steered up or down by a single activation-space vector, improving accuracy when amplified and cutting output length when suppressed.

  13. Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

    cs.AI 2026-05 reject novelty 4.0 of 10

    PCS improves steering direction accuracy by adaptively sampling the intervention coefficient from a cosine-similarity-conditioned Gaussian, but its evaluation is partly circular because the optimal coefficient is chos...

Pith tools