Pith. sign in

hub

Improving steering vectors by targeting sparse autoencoder features.arXiv preprint arXiv:2411.02193

11 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

11 Pith papers citing it
1 external citations · external index

hub tools

years

2026 10 2025 1

verdicts

UNVERDICTED 11

representative citing papers

Sense Representations Are Inducible Interfaces

cs.CL · 2026-05-27 · unverdicted · novelty 7.0

ACROS induces explicit sense representations in frozen decoder LMs via gated residual addition, enabling competitive zero-shot WSD, lexical steering, and cross-lingual adaptation on SmolLM2-360M while preserving base quality.

Perplexity Can Miss SAE Feature Damage Under Quantization

cs.LG · 2026-06-02 · unverdicted · novelty 6.0

Quantization of LLMs can degrade many SAE features even when perplexity improves or stays similar, as shown by correlation measurements on frozen SAEs for Pythia-70M and Gemma-2-2B models across INT8 to INT4.

The Cylindrical Representation Hypothesis for Language Model Steering

cs.CL · 2026-05-03 · unverdicted · novelty 6.0

The Cylindrical Representation Hypothesis (CRH) models LLM representations as a central axis for concept activation surrounded by a normal plane containing sensitive sectors that determine steering sensitivity and introduce intrinsic uncertainty.

citing papers explorer

Showing 11 of 11 citing papers.