Pith. sign in

REVIEW 2 major objections 2 minor 50 cited by

Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

T0 review · 2 major / 2 minor · reviewed 2026-05-17 · grok-4.3

Pith's one-line read Varying metrics and corruption methods in activation patching can produce conflicting pictures of which model components matter.

desk verdict Activation patching results shift with metric and corruption choices in the tested cases, but the best-practice recommendations rest on limited settings and need broader checks. read the letter →

arxiv 2309.16042 v2 pith:6W7OXTAI submitted 2023-09-27 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords activationpatchingmechanisticinterpretabilitylanguagemodelslocalizationcircuitdiscoveryevaluationmetricscorruptionmethodsbestpractices
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests how small changes in the way activation patching is performed affect results when researchers try to find important parts inside language models. Different choices for scoring the patch effect and for corrupting the original input often point to different components or circuits. A reader should care because activation patching is the main tool for causal localization, so inconsistent outcomes mean that published claims about model mechanisms may depend on unstated methodological decisions. The authors run controlled comparisons across localization and circuit-discovery tasks, give reasons why some metrics and corruption styles are more reliable than others, and close with concrete recommendations for future experiments.

What carries the argument

Activation patching (also called causal tracing), with its choices of evaluation metric and input-corruption method, used to measure the causal importance of model components.

What would settle it

A replication on the same models and tasks that finds the same localization and circuit results no matter which metric or corruption method is chosen would falsify the reported sensitivity.

Watch

Extended reading notes

Core claim

In multiple localization and circuit-discovery settings, changing the evaluation metric or the corruption procedure inside activation patching produces noticeably different rankings of important model components. The authors supply empirical evidence for these differences together with conceptual arguments favoring particular metrics and corruption strategies, and they distill the observations into a set of recommended practices for activation patching.

Load-bearing premise

The localization and circuit-discovery tasks and models the authors tested are representative of how activation patching is used more broadly.

Editorial extensions

If this is right

  • Standardizing on a small set of metrics and corruption methods would reduce conflicting localization claims across papers.
  • Some commonly used metrics give systematically different importance scores than others, so results are not interchangeable.
  • Reporting results under multiple metric choices would make circuit discoveries more robust.
  • The recommendations can be adopted immediately to make new localization studies more reproducible.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same hyperparameter sensitivity may appear in other intervention-based interpretability techniques.
  • Creating public benchmark suites that measure how much patching results change with metric choice would let the community test the recommendations on new models.
  • Authors may need to treat the choice of patching variant as an explicit experimental variable rather than a fixed detail.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript systematically studies the sensitivity of activation patching (also called causal tracing or interchange intervention) to choices of evaluation metrics and corruption methods when performing localization and circuit discovery in language models. Through experiments in several concrete settings, it reports that different methodological choices can produce inconsistent interpretability conclusions, supplies conceptual arguments for preferring particular metrics or corruption schemes, and distills these observations into concrete recommendations for future use of the technique.

Significance. If the reported sensitivities prove robust, the work addresses a genuine methodological gap: activation patching is widely used yet implemented with little standardization, so explicit guidance on metrics and corruption could improve reproducibility across mechanistic interpretability studies. The empirical comparisons themselves constitute a useful contribution even if the final recommendations require further qualification.

major comments (2)
  1. [§4] §4 (Experimental results on IOI and related tasks): the central empirical claim—that varying metrics and corruption methods yields disparate localization and circuit outcomes—is demonstrated only for the specific models, tasks, and corruption schemes examined. Because the derived best-practice recommendations are presented without additional validation on other architectures, scales, or interpretability targets, the extrapolation from these settings to general usage rests on an untested assumption of representativeness.
  2. [§5] §5 (Recommendations): the preference for certain metrics is supported by a combination of the reported empirical differences and conceptual arguments, yet the manuscript does not quantify the magnitude or statistical reliability of those differences (e.g., via confidence intervals or multiple-run statistics). This makes it difficult to judge whether the observed disparities are large enough to justify changing community practice.
minor comments (2)
  1. [Abstract / §1] The abstract and introduction would benefit from an explicit enumeration of the exact models, tasks, and corruption functions used in the main experiments so that readers can immediately assess the scope of the reported findings.
  2. [§3] Notation for the different patching metrics (e.g., the precise definitions of “direct effect,” “indirect effect,” or normalized variants) should be collected in a single table or subsection for easy reference when comparing results across sections.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the positive assessment of our work's significance and for the detailed, constructive major comments. We address each point below and have revised the manuscript to incorporate additional discussion and statistical quantification where feasible.

read point-by-point responses
  1. Referee: [§4] §4 (Experimental results on IOI and related tasks): the central empirical claim—that varying metrics and corruption methods yields disparate localization and circuit outcomes—is demonstrated only for the specific models, tasks, and corruption schemes examined. Because the derived best-practice recommendations are presented without additional validation on other architectures, scales, or interpretability targets, the extrapolation from these settings to general usage rests on an untested assumption of representativeness.

    Authors: We agree that the experiments are confined to specific, widely studied settings such as the IOI task and smaller-scale models like GPT-2. These were selected because they represent standard benchmarks in the mechanistic interpretability literature where activation patching is commonly applied. The inconsistencies we document even in these canonical cases already underscore the importance of methodological choices. In the revised manuscript we have added an explicit limitations and scope section that discusses the representativeness of our settings, notes that similar sensitivities are likely in other contexts, and recommends that future studies validate the proposed practices on additional architectures and tasks. We do not claim the recommendations are universally proven but present them as evidence-based guidance for prevalent use cases. revision: partial

  2. Referee: [§5] §5 (Recommendations): the preference for certain metrics is supported by a combination of the reported empirical differences and conceptual arguments, yet the manuscript does not quantify the magnitude or statistical reliability of those differences (e.g., via confidence intervals or multiple-run statistics). This makes it difficult to judge whether the observed disparities are large enough to justify changing community practice.

    Authors: We concur that quantifying the scale and reliability of the observed differences would strengthen the case for the recommendations. The revised version now includes results from multiple random seeds for the primary experiments, with error bars and approximate confidence intervals reported for key localization and circuit metrics. These additions allow readers to evaluate whether the disparities are sufficiently consistent and large to warrant changes in practice. The conceptual arguments remain as before but are now paired with this statistical context. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical comparisons and recommendations are self-contained

full rationale

The paper performs direct empirical experiments comparing activation patching metrics, corruption methods, and hyperparameters across specific localization and circuit discovery tasks in language models. It reports observed disparities in results and offers recommendations backed by those observations plus conceptual arguments, without any derivation chain, equations, fitted parameters renamed as predictions, or load-bearing self-citations that reduce the central claims to prior inputs by construction. The analysis stands on its own experimental data rather than self-referential logic.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is an empirical methodological study. It introduces no new free parameters, axioms, or invented entities; it relies on standard assumptions from the mechanistic interpretability literature about what activation patching measures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Best Practices of Activation Patching in Language Models: Metrics and Methods." pith.science (2026). https://pith.science/paper/6W7OXTAI

@misc{pith2026230916042,
  author       = {Pith},
  title        = {Pith review of: Towards Best Practices of Activation Patching in Language Models: Metrics and Methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6W7OXTAI}},
  note         = {Machine review of arXiv:2309.16042}
}
read the original abstract

Mechanistic interpretability seeks to understand the internal mechanisms of machine learning models, where localization -- identifying the important model components -- is a key step. Activation patching, also known as causal tracing or interchange intervention, is a standard technique for this task (Vig et al., 2020), but the literature contains many variants with little consensus on the choice of hyperparameters or methodology. In this work, we systematically examine the impact of methodological details in activation patching, including evaluation metrics and corruption methods. In several settings of localization and circuit discovery in language models, we find that varying these hyperparameters could lead to disparate interpretability results. Backed by empirical observations, we give conceptual arguments for why certain metrics or methods may be preferred. Finally, we provide recommendations for the best practices of activation patching going forwards.

Discussion (0). Sign in to comment.

Forward citations

Cited by 50 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Two-Process Theory of Machine Self-Report

    cs.CL 2026-07 conditional novelty 8.0 of 10

    The single 'Pinocchio Axis' of LLM self-report splits into two independent, training-dependent dimensions—persona installation (B) and attribution gating (A)—measurable with a reproducible 48-item inventory.

  2. WriteSAE: Sparse Autoencoders for Recurrent State

    cs.LG 2026-05 unverdicted novelty 8.0 of 10

    WriteSAE introduces sparse autoencoders with rank-1 matrix atoms for recurrent state updates, allowing replacement tests that outperform deletion on 92.4% of positions and a formula predicting logit changes with R²=0.98.

  3. WriteSAE: Sparse Autoencoders for Recurrent State

    cs.LG 2026-05 unverdicted novelty 8.0 of 10

    WriteSAE is the first sparse autoencoder that factors decoder atoms into the native d_k x d_v cache write shape of recurrent models and supplies a closed-form per-token logit shift for atom substitution.

  4. WriteSAE: Sparse Autoencoders for Recurrent State

    cs.LG 2026-05 unverdicted novelty 8.0 of 10

    WriteSAE decomposes recurrent model cache writes into substitutable atoms with a closed-form logit shift, achieving high substitution success and targeted behavioral installs on models like Qwen3.5 and Mamba-2.

  5. Do Audio-Visual Large Language Models Really See and Hear?

    cs.AI 2026-04 unverdicted novelty 8.0 of 10

    AVLLMs encode audio semantics in middle layers but suppress them in final text outputs when audio conflicts with vision, due to training that largely inherits from vision-language base models.

  6. The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning

    cs.LG 2026-07 conditional novelty 7.0 of 10

    In a chess latent-reasoning model, replacing or removing the silent thought vectors barely changes moves, so the RL improvement appears to be encoded in the weights, not in a consulted scratchpad.

  7. Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models

    cs.CL 2026-06 conditional novelty 7.0 of 10

    VLMs default to visual grounding but a sparse circuit of 2.5-4.8% attention heads in later layers mediates prior-knowledge overrides, identified causally via patching and ablation across three model families.

  8. The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

    cs.CL 2026-06 conditional novelty 7.0 of 10

    Shared chat-template tokens piggyback narrow finetuning behaviors onto out-of-domain queries; regularizing their KV states (TReFT) reduces emergent misalignment and other off-topic generalization.

  9. MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    MENTIS applies layerwise covariance torsion (T1), spectral torsion (T2), and ERA localization to paired IT/PA 7-8B models, finding selective larger shifts for normative concepts, negative correlation with entropy, and...

  10. Transformer Field Theory: A Response-Theoretic Approach to Mechanistic Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Transformer Field Theory frames the residual stream as a field, models patching as source insertion, and uses first-order sensitivities plus Green functions to predict and describe responses, with empirical tests on G...

  11. WriteSAE: Sparse Autoencoders for Recurrent State

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    WriteSAE factors sparse autoencoder decoder atoms to the native d_k x d_v cache write shape in recurrent models, provides a closed-form logit shift, and demonstrates high success in atom substitution and behavioral ed...

  12. How LLMs Are Persuaded: A Few Attention Heads, Rerouted

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    Persuasion in LLMs works by redirecting a small set of attention heads to copy the target option token instead of reasoning over evidence, via a rank-one routing feature that can be directly edited or removed.

  13. Data-driven Circuit Discovery for Interpretability of Language Models

    cs.AI 2026-05 unverdicted novelty 7.0 of 10

    Standard circuit discovery methods produce dataset-specific circuits rather than task-general ones, and a new clustering-based method discovers multiple more faithful circuits per dataset.

  14. Cell-Based Representation of Relational Binding in Language Models

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    Large language models encode relational bindings via a cell-based representation: a low-dimensional linear subspace in which each cell corresponds to an entity-relation index pair and attributes are retrieved from the...

  15. Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

    cs.CR 2026-04 unverdicted novelty 7.0 of 10

    HMNS is a new jailbreak method that uses causal head identification and nullspace-constrained injection to achieve higher attack success rates than prior techniques on aligned language models.

  16. CURE:Circuit-Aware Unlearning for LLM-based Recommendation

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    CURE disentangles LLM recommendation circuits into forget-specific, retain-specific, and task-shared modules with tailored update rules to achieve more effective unlearning than weighted baselines.

  17. V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models

    cs.CL 2025-09 conditional novelty 7.0 of 10

    V-SEAM combines concept-level visual semantic editing with attention head modulation to identify positive and negative contributors across object, attribute, and relationship levels, then uses this to improve VLM perf...

  18. From Attribution to Action: A Human-Centered Application of Activation Steering

    cs.AI 2026-04 conditional novelty 6.5 of 10

    Activation steering of SAE-attributed components lets practitioners move from correlational inspection to causal hypothesis testing on CLIP failures, with trust shifting to observed model responses (N=8 experts).

  19. HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A per-hop evidence-coverage verifier can safely stop or extend retrieval-augmented search agents, cutting redundant loops and flagging under-search without retraining the host agent.

  20. Emergent Latent-State Computation under Stochastic Volatility

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Volatility forecasters develop linearly decodable representations of the next hidden log-volatility state; in long cycles this appears immediately after the input projection and ℓ2 normalization.

  21. The Computational Basis of Confidence in Large Language Models

    cs.LG 2026-07 unverdicted novelty 6.0 of 10

    Answer-logit differences in multimodal LLMs satisfy statistical decision confidence signatures, behaving as monotonic readouts of a latent decision variable rather than heuristic preference scores.

  22. The Computational Basis of Confidence in Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Answer-logit differences in multimodal LMs behave as monotonic readouts of a latent decision variable in simple perceptual and memory tasks, but not in complex visual reasoning.

  23. Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models

    cs.AI 2026-06 conditional novelty 6.0 of 10

    Pre-trained LLMs on HMM next-token prediction appear to use finite-window Soft n-gram-like learned predictors rather than Bayes-optimal inference, as shown by a new activation-probing and causal-patching pipeline.

  24. Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Consistency training decodes verifier information from rationale representations but does not produce faithful natural-language explanations.

  25. Localizing Anchoring Pathways in Language Models

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Attribution methods localize anchoring signals in Qwen and Llama models; edge-level circuits transfer within a model but show sparse transfer from base to instruction-tuned variants.

  26. Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Standard tests for mechanistic roles in transformer attention heads are insufficient because heads that pass them fail to transfer computations across prompts under matched controls.

  27. When Attribution Patching Lies: Diagnosis and a Second-Order Correction

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Dominant error in attribution patching arises from downstream non-linearities; a single HVP correction removes the leading error term and matches Integrated Gradients accuracy at lower cost across 124M-9B models.

  28. The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    The Piggyback Hypothesis attributes emergent misalignment to chat-template tokens piggybacking finetuned behavior; Token-Regularized Finetuning (TReFT) mitigates it by regularizing prefix token representations.

  29. Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Base LLMs show multi-agent yield to peer pressure at rates equal to or higher than aligned models, localized by activation patching to mid-layers where attention dominates, with one dissenter cutting yield by 54-73 po...

  30. Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Pretrained base models exhibit higher yield to peer disagreement than RLHF instruct variants, with the effect localized to mid-layer attention and mitigated by structured dissent rather than prompt defenses.

  31. Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.

  32. Knowledge Vector of Logical Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    Distinct linear knowledge vectors for deductive, inductive, and abductive reasoning in LLMs can be refined via complementary subspace constraints to improve performance through mutual knowledge sharing.

  33. How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    LLMs implement a second-order confidence architecture where the PANL activation encodes both error likelihood and the ability to correct it, beyond verbal confidence or log-probabilities.

  34. Understanding the Mechanism of Altruism in Large Language Models

    econ.GN 2026-04 unverdicted novelty 6.0 of 10

    A small set of sparse autoencoder features in LLMs drives shifts between generous and selfish allocations in dictator games, with causal patching and steering confirming their role and generalization to other social games.

  35. Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation

    cs.SE 2026-04 unverdicted novelty 6.0 of 10

    Co-locating tests with implementation code yields substantially higher preservation and correctness in foundation-model-generated programs than separated test syntax.

  36. Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Transformers show limited adaptive depth use on relational reasoning, with clearer evidence after finetuning on the task.

  37. From Attribution to Action: A Human-Centered Application of Activation Steering

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Activation steering paired with attribution enables intervention-based debugging in vision models, as all 8 interviewed experts shifted to hypothesis testing, most trusted observed responses, and highlighted risks lik...

  38. How do LLMs Compute Verbal Confidence

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    Mechanistic experiments on Gemma 3 27B, Qwen 2.5 7B and Magistral Small 24B show verbal confidence is cached at post-answer positions from answer tokens and captures richer answer-quality information beyond token log-...

  39. Prototype Transformer: Towards Language Model Architectures Interpretable by Design

    cs.AI 2026-02 conditional novelty 6.0 of 10

    ProtoT is an autoregressive language model whose attention is replaced by learned prototype channels that are claimed to capture nameable concepts and allow targeted edits, at linear sequence cost but with slightly lo...

  40. Order Is Not Control

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Order is distinct from control, where control is defined as a local receiver-gated response law demonstrated across biological circuits and LLM response panels with reported prediction accuracies of 72-84%.

  41. Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Reasoning in large output spaces proceeds via shortlisting then fine-grained reasoning; this characterization enables a mechanistic distillation strategy that outperforms standard distillation.

  42. Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    Toxicity in language models is disproportionately encoded in early MLP layers and can be localized via activation differentials then suppressed at inference time without gradient descent.

  43. Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs

    cs.CL 2026-02 reject novelty 5.0 of 10

    A three-stage SAE-plus-SVD steering recipe claims to make Hindi or Spanish the default language of an LLM at inference time, but the visible manuscript reports only expected, not measured, outcomes.

  44. Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

    cs.CL 2025-10 conditional novelty 5.0 of 10

    The Knobe effect in fine-tuned LLMs is localized to mid-to-late transformer layers and can be removed by patching in pretrained activations at a single layer.

  45. How to use and interpret activation patching

    cs.LG 2024-04 accept novelty 5.0 of 10

    Activation patching provides evidence about neural network circuits when the choice of metric is aligned with the hypothesis and common interpretation errors are avoided.

  46. A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    A learned linear activation bridge achieves high alignment (cosine ~0.97) between Pythia-160M and Pythia-410M states but produces no improvement in downstream multi-hop answering when injected into the receiver.

  47. A Numerical PDEs Approach to Evolution Equations in Shape Analysis Based on Regularized Morphoelasticity

    math.NA 2026-04 unverdicted novelty 4.0 of 10

    Regularized morphoelasticity yields a high-order elliptic system for continuous shape evolution that is solved by mixed finite elements in FEniCSx within an LDDMM-style optimal-control growth model.

  48. NEAT: Concept driven Neuron Attribution in LLMs

    cs.CL 2025-08 reject novelty 4.0 of 10

    NEAT identifies concept neurons by feeding a single mean hidden-state vector through the model and ranking neurons by their effect on concept-word probabilities.

  49. Discovering Millions of Interpretable Features with Sparse Autoencoders

    cs.LG 2026-06 unverdicted novelty 3.0 of 10

    Trains and releases SAEs for Qwen3-1.7B/4B/8B models with layer-wise coverage and demonstrates causal steering of refusal via selected features.

  50. Towards a Foundation Model for the Martian Atmosphere

    astro-ph.EP 2026-05 unverdicted novelty 3.0 of 10

    The paper reviews data sources, physical models, downstream applications, and AI techniques to outline considerations for building a foundation model for the Martian atmosphere.

Reference graph

Works this paper leans on

108 extracted references · 108 canonical work pages · cited by 43 Pith papers

  1. [1]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Attention is all you need , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  2. [2]

    Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases , author=

  3. [3]

    Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors

    Yun, Zeyu and Chen, Yubei and Olshausen, Bruno and LeCun, Yann. Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors. Proceedings of Deep Learning Inside Out (DeeLIO): The 2nd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures

  4. [5]

    International Conference on Artificial Intelligence and Statistics (AISTATS) , year=

    Feature relevance quantification in explainable AI: A causal problem , author=. International Conference on Artificial Intelligence and Statistics (AISTATS) , year=

  5. [6]

    A circuit for

    Heimersheim, Stefan and Janiak, Jett , year =. A circuit for

  6. [7]

    Transformer

    Nanda, Neel and Bloom, Joseph , howpublished = ". Transformer

  7. [8]

    Wang, Ben and Komatsuzaki, Aran , title =

  8. [9]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Investigating gender bias in language models using causal mediation analysis , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

Show all 108 references
  1. [10]

    arXiv preprint arXiv:2307.03637 , year=

    Discovering Variable Binding Circuitry with Desiderata , author=. arXiv preprint arXiv:2307.03637 , year=

  2. [11]

    Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

    Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space , author=. Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

  3. [12]

    How does

    Hanna, Michael and Liu, Ollie and Variengien, Alexandre , booktitle=. How does

  4. [13]

    International Conference on Machine Learning (ICML) , year=

    How do transformers learn topic structure: Towards a mechanistic understanding , author=. International Conference on Machine Learning (ICML) , year=

  5. [14]

    Wen, Kaiyue and Li, Yuchen and Liu, Bingbin and Risteski, Andrej , booktitle=. (

  6. [15]

    International Conference on Learning Representations (ICLR) , year=

    Natural language descriptions of deep visual features , author=. International Conference on Learning Representations (ICLR) , year=

  7. [16]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Inference-Time Intervention: Eliciting Truthful Answers from a Language Model , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  8. [17]

    Hidden progress in deep learning:

    Barak, Boaz and Edelman, Benjamin and Goel, Surbhi and Kakade, Sham and Malach, Eran and Zhang, Cyril , booktitle=. Hidden progress in deep learning:

  9. [18]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance Explanations , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  10. [23]

    arXiv preprint arXiv:2303.08112 , year=

    Eliciting latent predictions from transformers with the tuned lens , author=. arXiv preprint arXiv:2303.08112 , year=

  11. [24]

    Interpreting Transformer's Attention Dynamic Memory and Visualizing the Semantic Information Flow of

    Katz, Shahar and Belinkov, Yonatan , journal=. Interpreting Transformer's Attention Dynamic Memory and Visualizing the Semantic Information Flow of

  12. [25]

    Advances in neural information processing systems (NeurIPS) , year=

    A benchmark for interpretability methods in deep neural networks , author=. Advances in neural information processing systems (NeurIPS) , year=

  13. [26]

    2022 , journal=

    In-context Learning and Induction Heads , author=. 2022 , journal=

  14. [28]

    Conference on Uncertainty and Artificial Intelligence (UAI) , year=

    Direct and indirect effects , author=. Conference on Uncertainty and Artificial Intelligence (UAI) , year=

  15. [29]

    International Conference on Learning Representations (ICLR) , year=

    Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task , author=. International Conference on Learning Representations (ICLR) , year=

  16. [30]

    International Conference on Learning Representations (ICLR) , year=

    Progress measures for grokking via mechanistic interpretability , author=. International Conference on Learning Representations (ICLR) , year=

  17. [31]

    International Conference on Machine Learning (ICML) , year=

    A toy model of universality: Reverse engineering how networks learn group operations , author=. International Conference on Machine Learning (ICML) , year=

  18. [32]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  19. [33]

    2021 , journal=

    A Mathematical Framework for Transformer Circuits , author=. 2021 , journal=

  20. [34]

    Cammarata, Nick and Carter, Shan and Goh, Gabriel and Olah, Chris and Petrov, Michael and Schubert, Ludwig and Voss, Chelsea and Egan, Ben and Lim, Swee Kiat , journal=. Thread:

  21. [35]

    International Conference on Machine Learning (ICML) , year=

    Inducing causal structure for interpretable neural networks , author=. International Conference on Machine Learning (ICML) , year=

  22. [36]

    Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP , year=

    Discovering the Compositional Structure of Vector Representations with Role Learning Networks , author=. Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP , year=

  23. [38]

    Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP , year=

    Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and Negation , author=. Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP , year=

  24. [39]

    Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models , author=. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP) , year=

  25. [40]

    Toward transparent

    Casper, Stephen and Rauker, Tilman and Ho, Anson and Hadfield-Menell, Dylan , booktitle=. Toward transparent

  26. [41]

    Interpretability in the Wild: a Circuit for Indirect Object Identification in

    Wang, Kevin Ro and Variengien, Alexandre and Conmy, Arthur and Shlegeris, Buck and Steinhardt, Jacob , booktitle=. Interpretability in the Wild: a Circuit for Indirect Object Identification in

  27. [42]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Language models are few-shot learners , author =. Advances in Neural Information Processing Systems (NeurIPS) , year=

  28. [43]

    Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

    Transformer Feed-Forward Layers Are Key-Value Memories , author=. Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

  29. [44]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Causal Abstractions of Neural Networks , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  30. [45]

    2023 , howpublished =

    Language models can explain neurons in language models , author=. 2023 , howpublished =

  31. [46]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Towards automated circuit discovery for mechanistic interpretability , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  32. [47]

    2022 , journal=

    Causal scrubbing, a method for rigorously testing interpretability hypotheses , author=. 2022 , journal=

  33. [48]

    Locating and editing factual associations in

    Meng, Kevin and Bau, David and Andonian, Alex and Belinkov, Yonatan , booktitle=. Locating and editing factual associations in

  34. [51]

    OpenAI blog , year=

    Language models are unsupervised multitask learners , author=. OpenAI blog , year=

  35. [52]

    Does localization inform editing?

    Hase, Peter and Bansal, Mohit and Kim, Been and Ghandeharioun, Asma , booktitle=. Does localization inform editing?

  36. [53]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Compositional explanations of neurons , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  37. [55]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Interpretability at scale: Identifying causal mechanisms in alpaca , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  38. [56]

    Annual Meeting of the Association for Computational Linguistics (ACL) , year=

    Analyzing Transformers in Embedding Space , author=. Annual Meeting of the Association for Computational Linguistics (ACL) , year=

  39. [58]

    2023 , booktitle =

    Hritik Bansal and Karthik Gopalakrishnan and Saket Dingliwal and Sravan Bodapati and Katrin Kirchhoff and Dan Roth , title =. 2023 , booktitle =

  40. [59]

    Lepori, Michael A and Pavlick, Ellie and Serre, Thomas , journal=. Neuro

  41. [60]

    Annual Meeting of the Association for Computational Linguistics (ACL) , year=

    Knowledge Neurons in Pretrained Transformers , author=. Annual Meeting of the Association for Computational Linguistics (ACL) , year=

  42. [63]

    Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale

    Hritik Bansal, Karthik Gopalakrishnan, Saket Dingliwal, Sravan Bodapati, Katrin Kirchhoff, and Dan Roth. Rethinking the role of scale for in-context learning: An interpretability-based case study at 66 billion scale. In Annual Meeting of the Association for Computational Lingu...

  43. [64]

    Hidden progress in deep learning: SGD learns parities near the computational limit

    Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. Hidden progress in deep learning: SGD learns parities near the computational limit. In Advances in Neural Information Processing Systems (NeurIPS), 2022

  44. [65]

    Language models can explain neurons in language models

    Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html , 2023

  45. [66]

    On privileged and convergent bases in neural network representations

    Davis Brown, Nikhil Vyas, and Yamini Bansal. On privileged and convergent bases in neural network representations. arXiv preprint arXiv:2307.12941, 2023

  46. [67]

    Thread: C ircuits

    Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim. Thread: C ircuits. Distill, 5 0 (3): 0 e24, 2020

  47. [68]

    Toward transparent AI : A survey on interpreting the inner structures of deep neural networks

    Stephen Casper, Tilman Rauker, Anson Ho, and Dylan Hadfield-Menell. Toward transparent AI : A survey on interpreting the inner structures of deep neural networks. In IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2022

  48. [69]

    Causal scrubbing, a method for rigorously testing interpretability hypotheses

    Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldwosky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas. Causal scrubbing, a method for rigorously testing interpretability hypotheses. AI Alignment Forum, 2022. https://www.alignme...

  49. [70]

    A toy model of universality: Reverse engineering how networks learn group operations

    Bilal Chughtai, Lawrence Chan, and Neel Nanda. A toy model of universality: Reverse engineering how networks learn group operations. In International Conference on Machine Learning (ICML), 2023

  50. [71]

    Towards automated circuit discovery for mechanistic interpretability

    Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri \`a Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  51. [72]

    Sparse autoencoders find highly interpretable features in language models

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023

  52. [73]

    Knowledge neurons in pretrained transformers

    Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. Knowledge neurons in pretrained transformers. In Annual Meeting of the Association for Computational Linguistics (ACL), 2022

  53. [74]

    Analyzing transformers in embedding space

    Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. Analyzing transformers in embedding space. In Annual Meeting of the Association for Computational Linguistics (ACL), 2023

  54. [75]

    A mathematical framework for transformer circuits

    Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dari...

  55. [76]

    Causal analysis of syntactic agreement mechanisms in neural language models

    Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart M Shieber, Tal Linzen, and Yonatan Belinkov. Causal analysis of syntactic agreement mechanisms in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and...

  56. [77]

    Neural natural language inference models partially embed theories of lexical entailment and negation

    Atticus Geiger, Kyle Richardson, and Christopher Potts. Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 2020

  57. [78]

    Causal abstractions of neural networks

    Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2021

  58. [79]

    Inducing causal structure for interpretable neural networks

    Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts. Inducing causal structure for interpretable neural networks. In International Conference on Machine Learning (ICML), 2022

  59. [80]

    Finding alignments between interpretable causal variables and distributed neural representations

    Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D Goodman. Finding alignments between interpretable causal variables and distributed neural representations. arXiv preprint arXiv:2303.02536, 2023

  60. [81]

    Transformer feed-forward layers are key-value memories

    Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021

  61. [82]

    Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space

    Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022

  62. [83]

    Dissecting recall of factual associations in auto-regressive language models

    Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767, 2023

  63. [84]

    Localizing model behavior with path patching

    Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. Localizing model behavior with path patching. arXiv preprint arXiv:2304.05969, 2023

  64. [85]

    Finding neurons in a haystack: Case studies with sparse probing

    Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. Finding neurons in a haystack: Case studies with sparse probing. arXiv preprint arXiv:2305.01610, 2023

  65. [86]

    How does GPT -2 compute greater-than?: I nterpreting mathematical abilities in a pre-trained language model

    Michael Hanna, Ollie Liu, and Alexandre Variengien. How does GPT -2 compute greater-than?: I nterpreting mathematical abilities in a pre-trained language model. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  66. [87]

    The out-of-distribution problem in explainability and search methods for feature importance explanations

    Peter Hase, Harry Xie, and Mohit Bansal. The out-of-distribution problem in explainability and search methods for feature importance explanations. In Advances in Neural Information Processing Systems (NeurIPS), 2021

  67. [88]

    Does localization inform editing? S urprising differences in causality-based localization vs

    Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. Does localization inform editing? S urprising differences in causality-based localization vs. knowledge editing in language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  68. [89]

    A circuit for P ython docstrings in a 4-layer attention-only transformer

    Stefan Heimersheim and Jett Janiak. A circuit for P ython docstrings in a 4-layer attention-only transformer. https://www.alignmentforum.org/posts/u6KXXmKFbXfWzoAXn/a-circuit-for-python-docstrings-in-a-4-layer-attention-only , 2023

  69. [90]

    Natural language descriptions of deep visual features

    Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas. Natural language descriptions of deep visual features. In International Conference on Learning Representations (ICLR), 2021

  70. [91]

    A benchmark for interpretability methods in deep neural networks

    Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. A benchmark for interpretability methods in deep neural networks. In Advances in neural information processing systems (NeurIPS), 2019

  71. [92]

    Feature relevance quantification in explainable ai: A causal problem

    Dominik Janzing, Lenon Minorics, and Patrick Bl \"o baum. Feature relevance quantification in explainable ai: A causal problem. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2020

  72. [93]

    Interpreting transformer's attention dynamic memory and visualizing the semantic information flow of GPT

    Shahar Katz and Yonatan Belinkov. Interpreting transformer's attention dynamic memory and visualizing the semantic information flow of GPT . arXiv preprint arXiv:2305.13417, 2023

  73. [94]

    Neuro S urgeon: A toolkit for subnetwork analysis

    Michael A Lepori, Ellie Pavlick, and Thomas Serre. Neuro S urgeon: A toolkit for subnetwork analysis. arXiv preprint arXiv:2309.00244, 2023

  74. [95]

    Emergent world representations: Exploring a sequence model trained on a synthetic task

    Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Vi \'e gas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. In International Conference on Learning Representations (ICLR), 2023 a

  75. [96]

    Inference-time intervention: Eliciting truthful answers from a language model

    Kenneth Li, Oam Patel, Fernanda Vi \'e gas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. In Advances in Neural Information Processing Systems (NeurIPS), 2023 b

  76. [97]

    How do transformers learn topic structure: Towards a mechanistic understanding

    Yuchen Li, Yuanzhi Li, and Andrej Risteski. How do transformers learn topic structure: Towards a mechanistic understanding. In International Conference on Machine Learning (ICML), 2023 c

  77. [98]

    Does circuit analysis interpretability scale? E vidence from multiple choice capabilities in C hinchilla

    Tom Lieberum, Matthew Rahtz, J \'a nos Kram \'a r, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik. Does circuit analysis interpretability scale? E vidence from multiple choice capabilities in C hinchilla. arXiv preprint arXiv:2307.09458, 2023

  78. [99]

    The hydra effect: Emergent self-repair in language model computations

    Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. The hydra effect: Emergent self-repair in language model computations. arXiv preprint arXiv:2307.15771, 2023

  79. [100]

    Locating and editing factual associations in GPT

    Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT . In Advances in Neural Information Processing Systems (NeurIPS), 2022

  80. [101]

    Language models implement simple word2vec-style vector arithmetic

    Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. Language models implement simple word2vec-style vector arithmetic. arXiv preprint arXiv:2305.16130, 2023

  81. [102]

    Compositional explanations of neurons

    Jesse Mu and Jacob Andreas. Compositional explanations of neurons. In Advances in Neural Information Processing Systems (NeurIPS), 2020

  82. [103]

    Transformer L ens

    Neel Nanda and Joseph Bloom. Transformer L ens. https://github.com/neelnanda-io/TransformerLens , 2022

  83. [104]

    Progress measures for grokking via mechanistic interpretability

    Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. In International Conference on Learning Representations (ICLR), 2023 a

  84. [105]

    Emergent linear representations in world models of self-supervised sequence models

    Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023 b

  85. [106]

    Mechanistic interpretability, variables, and the importance of interpretable bases

    Chris Olah. Mechanistic interpretability, variables, and the importance of interpretable bases. https://transformer-circuits.pub/2022/mech-interp-essay/index.html , 2022

  86. [107]

    In-context learning and induction heads

    Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kam...

  87. [108]

    Direct and indirect effects

    Judea Pearl. Direct and indirect effects. In Conference on Uncertainty and Artificial Intelligence (UAI), 2001

  88. [109]

    Language models are unsupervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 2019

  89. [110]

    Polysemanticity and capacity in neural networks

    Adam Scherlis, Kshitij Sachan, Adam S Jermyn, Joe Benton, and Buck Shlegeris. Polysemanticity and capacity in neural networks. arXiv preprint arXiv:2210.01892, 2022

  90. [111]

    Discovering the compositional structure of vector representations with role learning networks

    Paul Soulos, R Thomas McCoy, Tal Linzen, and Paul Smolensky. Discovering the compositional structure of vector representations with role learning networks. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 2020

  91. [112]

    Understanding arithmetic reasoning in language models using causal mediation analysis

    Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan. Understanding arithmetic reasoning in language models using causal mediation analysis. arXiv preprint arXiv:2305.15054, 2023

  92. [113]

    Explaining grokking through circuit efficiency

    Vikrant Varma, Rohin Shah, Zachary Kenton, J \'a nos Kram \'a r, and Ramana Kumar. Explaining grokking through circuit efficiency. arXiv preprint arXiv:2309.02390, 2023

  93. [114]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), 2017

  94. [115]

    Investigating gender bias in language models using causal mediation analysis

    Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems (NeurIPS), 2020

  95. [116]

    GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model

    Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model . https://github.com/kingoflolz/mesh-transformer-jax, May 2021

  96. [117]

    Interpretability in the wild: a circuit for indirect object identification in GPT -2 small

    Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in GPT -2 small. In International Conference on Learning Representations (ICLR), 2023

  97. [118]

    ( U n)interpretability of transformers: a case study with D yck grammars

    Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski. ( U n)interpretability of transformers: a case study with D yck grammars. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  98. [119]

    Interpretability at scale: Identifying causal mechanisms in alpaca

    Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D Goodman. Interpretability at scale: Identifying causal mechanisms in alpaca. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  99. [120]

    Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors

    Zeyu Yun, Yubei Chen, Bruno Olshausen, and Yann LeCun. Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors. In Proceedings of Deep Learning Inside Out (DeeLIO): The 2nd Workshop on Knowledge Extraction an...

  100. [121]

    The clock and the pizza: Two stories in mechanistic explanation of neural networks

    Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas. The clock and the pizza: Two stories in mechanistic explanation of neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2023

Pith tools

Reviewed May 17, 2026 · model on record in the stance chip above.