A single RL policy trained on twelve behaviors in GPT-2 small transfers to unseen behaviors and recovers their known circuits without retraining from scratch.
A circuit for Python docstrings in a 4-layer attention-only transformer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
A single RL policy trained on twelve behaviors in GPT-2 small transfers to unseen behaviors and recovers their known circuits without retraining from scratch.