Pith. sign in

Improving activation steering in language models with mean-centring

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 4 2024 1

roles

background 1

polarities

background 1

representative citing papers

When is Your LLM Steerable?

cs.CL · 2026-06-10 · unverdicted · novelty 6.0

Early hidden state features from the first few tokens allow a GBDT classifier to predict activation steering success, under-steering, or over-steering with 0.7 macro-F1 on unseen concepts.

The Cylindrical Representation Hypothesis for Language Model Steering

cs.CL · 2026-05-03 · unverdicted · novelty 6.0

The Cylindrical Representation Hypothesis (CRH) models LLM representations as a central axis for concept activation surrounded by a normal plane containing sensitive sectors that determine steering sensitivity and introduce intrinsic uncertainty.

citing papers explorer

Showing 5 of 5 citing papers.