Pith. sign in

REVIEW 5 cited by

Discovering Variable Binding Circuitry with Desiderata

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.03637 v1 pith:HFM7HJHD submitted 2023-07-07 cs.AI

classification cs.AI
keywords variablebindingautomaticallycausalcircuitrycomponentsdesideratamethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that computation in language models may be human-understandable, with successful efforts to localize and intervene on both single-unit features and input-output circuits. Here, we introduce an approach which extends causal mediation experiments to automatically identify model components responsible for performing a specific subtask by solely specifying a set of \textit{desiderata}, or causal attributes of the model components executing that subtask. As a proof of concept, we apply our method to automatically discover shared \textit{variable binding circuitry} in LLaMA-13B, which retrieves variable values for multiple arithmetic tasks. Our method successfully localizes variable binding to only 9 attention heads (of the 1.6k) and one MLP in the final token's residual stream.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One mechanism for many mental spaces: a shared router over a value slot in language models

    cs.CL 2026-07 conditional novelty 7.5 of 10

    A subspace trained to control one mental-space builder also controls others, indicating a shared router/slot mechanism across counterfactual, belief, fictional, and temporal spaces in LMs.

  2. Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

    cs.CL 2025-10 conditional novelty 7.0 of 10

    In-context entity retrieval in LMs is a mixture of positional, lexical, and reflexive mechanisms; the pure positional view fails in middle positions of long lists.

  3. How Do Transformers Learn Variable Binding in Symbolic Programs?

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A Transformer trained from scratch on symbolic variable-assignment programs develops a systematic dereferencing mechanism through three phases, building on early line-based heuristics rather than replacing them.

  4. How Causal Abstraction Underpins Computational Explanation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Computational implementation is analyzed as abstraction-under-translation in causal models, with representation and generalization as further constraints.

  5. Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

    cs.SE 2025-06 accept novelty 5.0 of 10

    A new survey organizes LLM interpretation methods by workflow stage and connects them to safety enhancement strategies and tools, covering around 70 works.

Pith tools