Pith. sign in

REVIEW 4 cited by

Representation Surgery: Theory and Practice of Affine Steering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09631 v7 pith:EVSXM3KO submitted 2024-02-15 cs.LG cs.CLcs.CY

Representation Surgery: Theory and Practice of Affine Steering

classification cs.LG cs.CLcs.CY
keywords behaviormodelsteeringundesirablelanguagerepresentationsaffineapproach
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Language models often exhibit undesirable behavior, e.g., generating toxic or gender-biased text. In the case of neural language models, an encoding of the undesirable behavior is often present in the model's representations. Thus, one natural (and common) approach to prevent the model from exhibiting undesirable behavior is to steer the model's representations in a manner that reduces the probability of it generating undesirable text. This paper investigates the formal and empirical properties of steering functions, i.e., transformation of the neural language model's representations that alter its behavior. First, we derive two optimal, in the least-squares sense, affine steering functions under different constraints. Our theory provides justification for existing approaches and offers a novel, improved steering approach. Second, we offer a series of experiments that demonstrate the empirical effectiveness of the methods in mitigating bias and reducing toxic generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Conditional Optimal Bridge for Riemannian Activation Steering

    cs.LG 2026-07 accept novelty 7.0

    Casting activation steering as a Schrödinger Bridge on the residual hypersphere derives the log-density-ratio objective and yields query-adaptive directions that beat fixed baselines without OOD collapse.

  2. ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

    cs.CL 2026-05 unverdicted novelty 6.0

    ASRU combines activation redirection and reward-optimized fine-tuning to unlearn cross-modal sensitive knowledge in MLLMs, reporting +24.6% better unlearning effectiveness and 5.8x higher generation quality on Qwen3-V...

  3. Conceptors for Semantic Steering

    cs.LG 2026-05 unverdicted novelty 6.0

    Conceptors as soft projection matrices from bipolar activations offer a multidimensional, compositional, and geometrically principled method for semantic steering in LLMs that outperforms single-vector baselines in mu...

  4. Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment

    cs.LG 2026-04 accept novelty 6.0

    Learning an input-conditioned mapping from embeddings to the best single steering layer (W2S) consistently beats fixed-layer CAA and L2S on 13 behaviors for two LLMs, in- and out-of-distribution.