A Transformer trained from scratch on symbolic variable-assignment programs develops a systematic dereferencing mechanism through three phases, building on early line-based heuristics rather than replacing them.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How Do Transformers Learn Variable Binding in Symbolic Programs?
A Transformer trained from scratch on symbolic variable-assignment programs develops a systematic dereferencing mechanism through three phases, building on early line-based heuristics rather than replacing them.