REVIEW 5 cited by
Discovering Variable Binding Circuitry with Desiderata
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-output circuits. Here, we introduce an approach which extends causal mediation experiments to automatically identify model components responsible for performing a specific subtask by solely specifying a set of \textit{desiderata}, or causal attributes of the model components executing that subtask. As a proof of concept, we apply our method to automatically discover shared \textit{variable binding circuitry} in LLaMA-13B, which retrieves variable values for multiple arithmetic tasks. Our method successfully localizes variable binding to only 9 attention heads (of the 1.6k) and one MLP in the final token's residual stream.
Forward citations
Cited by 5 Pith papers
-
One mechanism for many mental spaces: a shared router over a value slot in language models
A subspace trained to control one mental-space builder also controls others, indicating a shared router/slot mechanism across counterfactual, belief, fictional, and temporal spaces in LMs.
-
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
In-context entity retrieval in LMs is a mixture of positional, lexical, and reflexive mechanisms; the pure positional view fails in middle positions of long lists.
-
How Do Transformers Learn Variable Binding in Symbolic Programs?
A Transformer trained from scratch on symbolic variable-assignment programs develops a systematic dereferencing mechanism through three phases, building on early line-based heuristics rather than replacing them.
-
How Causal Abstraction Underpins Computational Explanation
Computational implementation is analyzed as abstraction-under-translation in causal models, with representation and generalization as further constraints.
-
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
A new survey organizes LLM interpretation methods by workflow stage and connects them to safety enhancement strategies and tools, covering around 70 works.
Discussion (0). Sign in to comment.