REVIEW 5 cited by
Analyzing the Generalization and Reliability of Steering Vectors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Steering vectors (SVs) have been proposed as an effective approach to adjust language model behaviour at inference time by intervening on intermediate model activations. They have shown promise in terms of improving both capabilities and model alignment. However, the reliability and generalisation properties of this approach are unknown. In this work, we rigorously investigate these properties, and show that steering vectors have substantial limitations both in- and out-of-distribution. In-distribution, steerability is highly variable across different inputs. Depending on the concept, spurious biases can substantially contribute to how effective steering is for each input, presenting a challenge for the widespread use of steering vectors. Out-of-distribution, while steering vectors often generalise well, for several concepts they are brittle to reasonable changes in the prompt, resulting in them failing to generalise well. Overall, our findings show that while steering can work well in the right circumstances, there remain technical difficulties of applying steering vectors to guide models' behaviour at scale. Our code is available at https://github.com/dtch1997/steering-bench
Forward citations
Cited by 5 Pith papers
-
Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering
Additive steering directions survive chat-to-agent transfer representationally, but behavioral coupling is reset per model with no universal constant or sign.
-
Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong
Answer correctness splits CoT unfaithfulness detection into two regimes: behavioral signals work only on correct answers, and fail at chance on incorrect answers where most unfaithfulness lives.
-
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
A three-statistic Borda consensus over sparse-autoencoder features produces interpretable activation steering, but usable quality-preserving shifts are rare and highly localized.
-
Can Interpretation Predict Behavior on Unseen Data?
Presence of hierarchical attention heads on in-distribution data predicts hierarchical out-of-distribution generalization across 270 small transformers, independent of causal support.
-
CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
CircuitSteer builds cross-layer circuits from sparse autoencoder features using co-activation and decoder-direction alignment, then applies multi-layer steering vectors that it claims preserve fluency across all tested tasks.
Discussion (0). Continue with ORCID to comment.