REVIEW 3 major objections 6 minor 70 references
Can Graph Learning Learn Circuits?
T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims circuit localization can be framed as a supervised amortized graph learning problem; the best configuration reaches median edge AUROC 0.902 on 16 held-out InterpBench cases.
desk verdict Careful and honestly-scoped pilot for amortized circuit localization; the headline AUROC is plausible but the leakage-control checks don't cover task behavior, so treat 0.902 as a promising estimate rather than a proven transfer result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the component-level computation graph of a transformer, a directed acyclic graph whose nodes are attention heads and MLP blocks and whose edges are residual-stream-mediated pathways. GCL transforms it into a directed line graph (one node per candidate edge) or an incidence graph (component nodes plus edge nodes) so that circuit membership becomes node classification on the transformed graph. Each edge node carries six context-specific features—clean and corrupted source states, target gradient, clean and corrupted edge messages, and the gradient mapped back through the edge—aligned across heterogeneous dimensions by a 1D feature-attention layer borrowed from tabular foundation models, processed by either DirGNN or DAGformer, pooled over token positions and prompt pairs, and read out to edge scores. The motivating identity is the shared-target interaction formula, which shows that two edges entering the same read input interact through the Hessian of the task metric, giving the graph learner a principled reason to share information between computationally related pathways.
What would settle it
Train the same GCL configuration on the 50-pair pool and evaluate it on real transformers using a faithfulness metric: does the predicted circuit reproduce the target behavior when kept edges are clean and masked edges are corrupted? If faithfulness is no better than the no-message-passing control, or if replacing the three related-pair checks with a stricter program-template match drops the held-out median edge AUROC from 0.902 toward 0.825, the paper's central transfer claim is refuted.
Extended reading notes
Core claim
The paper's central claim is that circuit localization can be framed as supervised amortized graph learning: build a component-level directed graph of a transformer whose edges are residual-stream pathways, train a GNN on labeled model–task pairs to predict circuit membership of each edge, and apply the same predictor to unseen pairs. On the 16 original held-out InterpBench cases, the best of 14 configurations (a directed line graph processed by a DirGNN GraphConv backend) scores a median edge AUROC of 0.902, close to the published EAP-IG median of 0.910 and below ACDC's 0.959; with all message-passing edges removed the median falls to 0.825, and an adapted PGExplainer reaches 0.858. The paper reports the top score as a post-selection estimate, since the configuration was chosen after evaluating all 14 on the held-out cases. It also proves a local interaction identity: for two edges feeding the same component input, the leading interaction equals $\Delta m_{e_1,p}^\top Q_{t,p} \Delta m_{e_2,p}$, a Hessian-weighted bilinear form of the two message changes, so edge relevance cannot in general be decomposed into independent scores.
Load-bearing premise
The claim assumes that circuits learned on small semi-synthetic transformers—whose ground-truth circuits are known by construction and which have no normalization layers—transfer to real transformers, and that the three related-pair checks used in the data split catch enough leakage that the held-out scores are not inflated by training/test similarity.
Editorial extensions
If this is right
- Circuit localization can be amortized: after one training phase, a learned predictor localizes circuits for new model–task pairs without per-case optimization, reducing the sequential per-case cost.
- Message passing over the observed computation graph carries signal: stripping all message-passing edges lowers median held-out AUROC from 0.902 to 0.825 while keeping nodes and features.
- GNN explainability methods transfer to mechanistic interpretability: the PGExplainer adaptation reaches 0.858 median edge AUROC without using ground-truth circuits at fit time.
- The framework produces edge-level scores that can be evaluated with the same head-promotion rule as existing baselines, making cross-case methods directly comparable to per-case circuit discovery.
Reading between the lines
- If GCL transfers to real transformers, circuit localization could become a reusable infrastructure task, but this depends on whether the regularity of Tracr-derived synthetic circuits survives contact with LayerNorm and learned representations.
- The interaction identity suggests scalar edge scores are incomplete descriptors, yet GCL's readout still emits one score per edge; a testable extension is to make the predictor explicitly pair-aware, for example by predicting interaction terms or using the Hessian form as an auxiliary loss.
- The paper's leakage-control split is explicitly acknowledged as incomplete; a stronger test would match latent program templates or circuit motifs rather than the three syntactic checks, and would likely lower the reported transfer numbers.
- One could test the amortization hypothesis directly on real models by using faithfulness instead of AUROC: if predicted circuits reproduce behavior under causal masking as well as per-case baselines do, the synthetic-to-real gap is smaller than feared.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Graph Circuit Learning (GCL), a supervised, amortized framework that trains a graph neural network across multiple model-task pairs to predict circuit masks on unseen pairs, and also adapts PGExplainer to the per-model-task circuit localization setting. To support cross-case training, the authors augment InterpBench with 30 TracrBench-derived, SIIT-trained transformer-task pairs, reserve the 16 original InterpBench evaluation pairs as a held-out test set, and use grouped 5-fold cross-validation on the remaining 50 pairs. The best GCL configuration achieves a median edge AUROC of 0.902 on the held-out cases, close to the published EAP-IG median of 0.910 and below ACDC's 0.959; removing all message-passing edges reduces the median to 0.825, and the PGExplainer adaptation reaches 0.858. The paper also derives a local second-order interaction formula for two edge interventions at a shared component input (Proposition E.1) as motivation for graph-structured modeling of pathway dependencies.
Significance. If the transfer result survives closer scrutiny, this is a valuable step for mechanistic interpretability: it casts circuit localization as an amortized graph learning problem, contributes a new benchmark of 30 SIIT-trained TracrBench cases with a grouped split, provides an open implementation, and gives a clean interaction identity that justifies modeling edge interactions. The authors are commendably explicit about several limitations: the best configuration is flagged as a post-selection estimate (Section F.2), the benchmark is restricted to synthetic models without normalization layers, and the relatedness checks are acknowledged to be incomplete (Section D.2). The central scientific contribution is the framing and the controlled evidence that message passing over computation graphs can help predict circuits, rather than a claim of state-of-the-art performance on real transformers.
major comments (3)
- [D.2] The split's relatedness checks compare RASP program templates, exact ground-truth circuits, and directed circuit isomorphism, but none of the three checks compares task identity or input-output behavior. Since Table 1 shows that 26 original InterpBench pairs remain in the cross-validation pool while all 16 held-out cases come from InterpBench, a pair implementing the same task through a different RASP program or a non-isomorphic circuit can cross the split, allowing the GNN to memorize task-to-circuit mappings rather than learn a general localizer. The manuscript's own admission that these checks 'do not give a complete theoretical account' (Section D.2) makes this a known gap. To make the transfer claim load-bearing, I ask for a task-level relatedness check (e.g., normalized task templates, prompt-pair similarity, or behavioral equivalence on a shared probe set) and/or a sensitivity analysis that removes from the training pool any pair sharing a task family with a test case. Without this, both the observed 0.902 and the no-message-passing control 0.825 are upper bounds whose inflation is unquantified.
- [F.2] The abstract's headline median of 0.902 is the best of 14 configurations after evaluating all of them on the 16 held-out cases, as Section F.2 states: the result 'should be read as a post-selection estimate.' Cross-validation ranks this configuration second, and the configuration that would have been selected by cross-validation (DirGNN on the incidence graph) has a held-out median of 0.876. Presenting the post-selected number as the main result overstates the method's expected performance on unseen pairs. I ask the authors to report the cross-validation-selected configuration as the primary estimate, to present 0.902 as an exploratory upper bound, and to use a nested selection procedure or a separate validation split if they want to claim competitive parity with EAP-IG and ACDC.
- [F.2, Table 4] The claim that message passing over the observed computation graph contributes to performance rests on the comparison between the DirGraphConv line-graph configuration (median 0.902) and its no-message-passing control (median 0.825). The paper reports seed-level standard deviations of the medians but no paired comparison across the 16 cases; with only 16 cases and a 0.077 median gap, it is not clear how many cases actually improve and by how much. I ask for per-case paired differences (e.g., median of within-case differences, and a sign test or Wilcoxon signed-rank test). For the incidence-graph rows, the paper should also state explicitly what the control retains, because three of the six feature roles are stored on component nodes and become inaccessible to the edge readout when message-passing edges are removed.
minor comments (6)
- [Abstract] Consider qualifying the headline number in the abstract as a post-selection estimate so that readers do not mistake 0.902 for an unbiased evaluation of a pre-registered configuration.
- [D.1] The paper reports that 48 candidate pairs were constructed and 30 passed both behavior and compilation checks; please state how the 18 failures are distributed between compilation failures and reconstruction-check failures, since this affects the representativeness of the added benchmark cases.
- [E.4] Proposition E.1 is derived for infinitesimally small message changes, whereas the benchmark uses binary clean-versus-corrupted patching with potentially large deltas; the main text should explicitly say that the proposition is a motivation for graph structure rather than a guarantee for the actual benchmark interventions.
- [F.1, Table 4] The 18 hyperparameter settings per configuration are not enumerated; for reproducibility, include the full grid (learning rate, hidden dimension, number of layers, dropout, etc.) in an appendix.
- [D.2] The sentence 'so 84 of its 86 pairs enter the pool' is confusing because only 50 pairs remain after exclusion; rephrase to '84 are candidates; after removing groups containing test pairs, 50 remain.'
- [Figure 2] The caption states that baseline box plots are dashed because seed-level variability is unavailable; add a legend or note explaining that the published baseline values are point estimates from InterpBench and may not be directly comparable in their uncertainty.
Circularity Check
No significant circularity: the central transfer claim is evaluated against external ground-truth circuits, and the motivation theorem is a genuine Taylor-expansion identity.
full rationale
The paper's central claim is not circular: the 16 held-out InterpBench cases have ground-truth circuits fixed by external SIIT/Tracr constructions, and GCL's features (activations, gradients, and virtual-weight messages) are computed from the model and task, not from the label masks; no held-out mask enters training. The augmented 30 TracrBench pairs are added before the split and are excluded from the test group by the relatedness checks, so the test set is not used to fit the predictor. The interaction result in Section E is a Taylor-expansion identity used as motivation for graph structure, not as an assumption feeding the AUROC numbers. The paper's own disclosures, that the relatedness checks 'do not give a complete theoretical account' (Section D.2) and that the headline 0.902 is a 'post-selection estimate' (Section F.2), are limitations on external validity and statistical interpretation, not reductions of the claimed result to its inputs. There are no load-bearing self-citations: the benchmark, Tracr compilation, and baseline sources are external. No self-definitional, fitted-input-as-prediction, uniqueness-import, or ansatz-smuggling pattern is present.
Assumptions & free parameters
free parameters (4)
- Feature alignment dimension d_align =
not reported in preprint
- DirGNN aggregation weight alpha =
0.5 (stated)
- Per-case BCE positive-class weight =
max(1, n_other / max(n_circuit, 1))
- Hyperparameter settings (18 settings, 14 configurations) =
selected via grouped 5-fold cross-validation
assumptions (5)
- domain assumption The component-level computation graph with virtual weight W_e = W_I^t W_O^s exactly represents computational pathways in the studied transformers.
- domain assumption Edge-wise additive masking (replacing clean messages with corrupted ones) is a valid intervention model for circuit localization.
- domain assumption The three related-pair similarity checks (program template, exact circuit, circuit isomorphism) sufficiently prevent leakage between training and held-out pairs.
- standard math Proposition E.1 requires L_t,p to be twice continuously differentiable with locally Lipschitz Hessian over the intervention surface.
- domain assumption Ground-truth circuits from Tracr compilation and SIIT training are correct and unique enough to serve as supervision.
Cite this review
Pith. "Pith review of Can Graph Learning Learn Circuits?." pith.science (2026). https://pith.science/paper/QL2VMVNZ
@misc{pith2026260808536,
author = {Pith},
title = {Pith review of: Can Graph Learning Learn Circuits?},
year = {2026},
howpublished = {\url{https://pith.science/paper/QL2VMVNZ}},
note = {Machine review of arXiv:2608.08536}
}
abstract
Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient to reproduce a particular behavior. Most established methods localize circuits independently for each model--task pair. We instead frame circuit localization as a graph machine learning problem in which the edges of a computation graph represent computational pathways, and graph neural networks (GNNs) model interactions among these pathways. We introduce Graph Circuit Learning (GCL), a supervised, amortized framework that trains a GNN across multiple model--task pairs and applies it to unseen cases. To provide sufficient data, we augment the InterpBench benchmark with additional cases derived from the TracrBench programs. Of the 14 evaluated GCL configurations, the highest scored a median edge AUROC of $0.902$ (interquartile interval $[0.861, 0.942]$) on the 16 original held-out InterpBench cases. This is close to the published InterpBench median of $0.910$ for EAP-IG while remaining below ACDC's $0.959$. Removing all message-passing edges reduces the median to $0.825$. We also adapt PGExplainer, a GNN explainability method, to circuit localization, obtaining a median edge AUROC of $0.858$ on the same cases. These preliminary results suggest that graph machine learning offers a natural and potentially powerful perspective on circuit localization, and we hope this perspective encourages closer exchange between the two communities.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Attribution Patching Outperforms Automated Circuit Discovery
Aaquib Syed, Can Rager, and Arthur Conmy. Attribution Patching Outperforms Automated Circuit Discovery. In Yonatan Belinkov, Najoung Kim, Jaap Jumelet, Hosein Mohebbi, Aaron Mueller, and Hanjie Chen, editors,Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 407–416, Miami, Florida, US, November
-
[2]
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov. Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms. InFirst Conference on Language Modeling, 2024. URLhttps://openreview.net/forum?id=TZ0CCGDcuT. 22, 23, 25
work page 2024
-
[3]
MIB: A Mechanistic Interpretability Benchmark
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fried Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, and Yonatan Belinkov. MIB: A...
work page 2025
-
[4]
Finding Trans- former Circuits With Edge Pruning
Adithya Bhaskar, Alexander Wettig, Dan Friedman, and Danqi Chen. Finding Trans- former Circuits With Edge Pruning. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Pro- cessing Systems, volume 37, pages 18506–18534. Curran Associates, Inc., 2024. doi: 10.52202/079017-0587. URL htt...
-
[5]
Low-Complexity Probing via Finding Sub- networks
Steven Cao, Victor Sanh, and Alexander Rush. Low-Complexity Probing via Finding Sub- networks. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computat...
work page 2021
-
[6]
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
Philipp Mondorf, Mingyang Wang, Sebastian Gerstner, Ahmad Dawar Hakimi, Yihong Liu, Leonor Veloso, Shijia Zhou, Hinrich Schuetze, and Barbara Plank. BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods. In Yonatan Be- linkov, Aaron Mueller, Najoung Kim, Hosein Mohebbi, Hanjie Chen, Dana Arad, and Gabriele Sarti,...
work page 2025
-
[7]
Towards Automated Circuit Discovery for Mechanistic Interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. Towards Automated Circuit Discovery for Mechanistic Interpretability. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in 5 Can Graph Learning Learn Circuits? Neural Information Processing Systems, volume 36, pages 1631...
work page 2023
-
[8]
Graph Metanetworks for Processing Diverse Neural Architectures
Derek Lim, Haggai Maron, Marc T Law, Jonathan Lorraine, and James Lu- cas. Graph Metanetworks for Processing Diverse Neural Architectures. In B. Kim, Y . Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y . Sun, editors, International Conference on Learning Representations, volume 2024, pages 11430– 11458, 2024. URL https://proceedings.iclr.cc/paper_files/...
work page 2024
Show all 70 references
-
[9]
A Survey of Weight Space Learning: Understanding, Representation, and Generation, 2026
Xiaolong Han, Zehong Wang, Bo Zhao, Binchi Zhang, Jundong Li, Damian Borth, Rose Yu, Haggai Maron, Yanfang Ye, Lu Yin, and Ferrante Neri. A Survey of Weight Space Learning: Understanding, Representation, and Generation, 2026. URL https://arxiv.org/abs/2603. 10090. 13, 25
2026
-
[10]
Zico Kolter, and Chelsea Finn
Allan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace, Yiding Jiang, Samuel Sokota, J. Zico Kolter, and Chelsea Finn. Permutation Equivariant Neural Functionals. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neu- ral Information Pr...
2023
-
[11]
Parameterized Explainer for Graph Neural Network
Dongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu, Bo Zong, Haifeng Chen, and Xiang Zhang. Parameterized Explainer for Graph Neural Network. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors,Advances in Neural Infor- mation Processing Systems, volume 3...
-
[12]
InterpBench: Semi- Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
Rohan Gupta, Iván Arcuschin, Thomas Kwa, and Adrià Garriga-Alonso. InterpBench: Semi- Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural...
2024 doi
-
[13]
Frank Harary and Robert Z. Norman. Some properties of line digraphs.Rendiconti del Circolo Matematico di Palermo, 9(2):161–168, May 1960. ISSN 1973-4409. doi: 10.1007/bf02854581. URLhttp://dx.doi.org/10.1007/BF02854581. 3, 14
1960 doi
-
[14]
Springer International Publishing, 2013
Alain Bretto.Hypergraph Theory: An Introduction. Springer International Publishing, 2013. ISBN 9783319000800. doi: 10.1007/978-3-319-00080-0. URL http://dx.doi.org/10. 1007/978-3-319-00080-0. 3, 14
2013 doi
-
[15]
Improving Graph Neural Networks on Multi-node Tasks with the Labeling Trick.Journal of Machine Learning Research, 26(23):1–44, 2025
Xiyuan Wang, Pan Li, and Muhan Zhang. Improving Graph Neural Networks on Multi-node Tasks with the Labeling Trick.Journal of Machine Learning Research, 26(23):1–44, 2025. URLhttp://jmlr.org/papers/v26/23-0560.html. 3, 14
2025
-
[16]
Accurate predictions on small data with a tabular foundation model.Nature, 637(8045):319–326, January 2025
Noah Hollmann, Samuel Müller, Lennart Purucker, Arjun Krishnakumar, Max Körfer, Shi Bin Hoo, Robin Tibor Schirrmeister, and Frank Hutter. Accurate predictions on small data with a tabular foundation model.Nature, 637(8045):319–326, January 2025. ISSN 1476-4687. doi: 10. 1038/s...
2025 doi
-
[17]
Bronstein
Emanuele Rossi, Bertrand Charpentier, Francesco Di Giovanni, Fabrizio Frasca, Stephan Günnemann, and Michael M. Bronstein. Edge Directionality Improves Learning on Heterophilic Graphs. In Soledad Villar and Benjamin Chamberlain, editors,Proceedings of the Second Learning on Gr...
2023
-
[18]
Transformers over Directed Acyclic Graphs
Yuankai Luo, Veronika Thost, and Lei Shi. Transformers over Directed Acyclic Graphs. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Processing Systems, volume 36, pages 47764–47782. Curran Associates, 6 Can Graph ...
2023
-
[19]
Deep Sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander Smola. Deep Sets. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors,Ad- vances in Neural Information Processing Systems...
-
[20]
A Mathematical Framework for Transformer Circuits.Transformer Circuits Thread,
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dari...
-
[21]
Graph Foundation Models: A Comprehensive Survey, 2025
Zehong Wang, Zheyuan Liu, Tianyi Ma, Jiazheng Li, Zheyuan Zhang, Xingbo Fu, Yiyang Li, Zhengqing Yuan, Wei Song, Yijun Ma, Qingkai Zeng, Xiusi Chen, Jianan Zhao, Jundong Li, Meng Jiang, Pietro Lio, Nitesh Chawla, Chuxu Zhang, and Yanfang Ye. Graph Foundation Models: A Comprehe...
2025 arXiv
-
[22]
Thinking Like Transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav. Thinking Like Transformers. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 11080–11090. PMLR, 18–24 Jul 2021. ...
2021
-
[23]
Tracr: Compiled Transformers as a Laboratory for Interpretability
David Lindner, Janos Kramar, Sebastian Farquhar, Matthew Rahtz, Tom McGrath, and Vladimir Mikulik. Tracr: Compiled Transformers as a Laboratory for Interpretability. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neu- ral Informatio...
2023
-
[24]
TracrBench: Generating Interpretability Testbeds with Large Language Models
Hannes Thurnherr and Jérémy Scheurer. TracrBench: Generating Interpretability Testbeds with Large Language Models. InICML 2024 Workshop on Mechanistic Interpretability, 2024. URL https://openreview.net/forum?id=vNubZ5zK8h. 4, 16, 26
2024
-
[25]
2, 10, 11, 12
https://transformer-circuits.pub/2021/framework/index.html. 2, 10, 11, 12
2021
-
[26]
Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models. In A. Globerson, L. Mackey, D. Bel- grave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Processing S...
2024 doi
-
[27]
Is This the Sub- space You Are Looking for? An Interpretability Illusion for Subspace Activation Patch- ing
Aleksandar Makelov, Georg Lange, Atticus Geiger, and Neel Nanda. Is This the Sub- space You Are Looking for? An Interpretability Illusion for Subspace Activation Patch- ing. In B. Kim, Y . Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y . Sun, edi- tors,International Confere...
2024
-
[28]
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
Areeb Ahmad, Abhinav Joshi, and Ashutosh Modi. Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors,Advances in Neural In- formation Processing Systems, vo...
2025
-
[29]
PyTorch Geometric Signed Directed: A Software Package on Graph Neural Networks for Signed and Directed Graphs
Yixuan He, Xitong Zhang, Junjie Huang, Benedek Rozemberczki, Mihai Cucuringu, and Gesine Reinert. PyTorch Geometric Signed Directed: A Software Package on Graph Neural Networks for Signed and Directed Graphs. In Soledad Villar and Benjamin Chamberlain, editors,Proceedings of t...
2023
-
[30]
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is All you Need. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, ed- itors,Advances in Neural...
2017
-
[31]
Nonlinear Laplacians Improve Signed-Directed Graph Learning
Ali Parviz and Yuichi Yoshida. Nonlinear Laplacians Improve Signed-Directed Graph Learning. InNew Perspectives in Graph Machine Learning, 2025. URL https://openreview.net/ forum?id=3CcP3VI6LW
2025
-
[32]
Billion-Scale Graph Foundation Models, 2026
Maya Bechler-Speicher, Yoel Gottlieb, Andrey Isakov, David Abensur, Ami Tavory, Daniel Haimovich, Ido Guy, and Udi Weinsberg. Billion-Scale Graph Foundation Models, 2026. URL https://arxiv.org/abs/2602.04768. 12
2026 arXiv
-
[33]
Generalized Linear Mode Connectivity for Transformers
Alexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto, Sidak Pal Singh, and Valentina Boeva. Generalized Linear Mode Connectivity for Transformers. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors,Advances in Neura...
2025
-
[34]
Fast Graph Representation Learning with PyTorch Geomet- ric, 2019
Matthias Fey and Jan Eric Lenssen. Fast Graph Representation Learning with PyTorch Geomet- ric, 2019. URLhttps://arxiv.org/abs/1903.02428. 15
2019 arXiv
-
[35]
MSGNN: A Spectral Graph Neural Network Based on a Novel Magnetic Signed Laplacian
Yixuan He, Michael Perlmutter, Gesine Reinert, and Mihai Cucuringu. MSGNN: A Spectral Graph Neural Network Based on a Novel Magnetic Signed Laplacian. In Bastian Rieck and Razvan Pascanu, editors,Proceedings of the First Learning on Graphs Conference, volume 198 ofProceedings ...
2022
-
[36]
Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe
Christopher Morris, Martin Ritzert, Matthias Fey, William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and Leman Go Neural: Higher-Order Graph Neural Networks. InProceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty...
2019
-
[37]
Scalable Message Passing Neural Networks: No Need for Attention in Large Graph Representation Learning, 2026
Haitz Sáez de Ocáriz Borde, Artem Lukoianov, Anastasis Kratsios, Michael Bronstein, and Xiaowen Dong. Scalable Message Passing Neural Networks: No Need for Attention in Large Graph Representation Learning, 2026. URLhttps://arxiv.org/abs/2411.00835. 15
2026
-
[38]
Rethinking Attention with Performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. Rethinking Attention with Performers. InIn...
2021
-
[39]
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In Hal Daumé III and Aarti Singh, editors,Proceedings of the 37th International Conference on Machine Learning, volume 119 ...
2020
-
[40]
PyG 2.0: Scalable Learning on Real World Graphs, 2025
Matthias Fey, Jinu Sunil, Akihiro Nitta, Rishi Puri, Manan Shah, Blaž Stojanovi ˇc, Ramona Bendias, Alexandria Barghi, Vid Kocijan, Zecheng Zhang, Xinwei He, Jan Eric Lenssen, and Jure Leskovec. PyG 2.0: Scalable Learning on Real World Graphs, 2025. URL https: //arxiv.org/abs/...
2025 arXiv
-
[41]
Rethinking Circuit Complete- ness in Language Models: AND, OR, and ADDER Gates
Hang Chen, Jiaying Zhu, Xinyu Yang, and Wenya Wang. Rethinking Circuit Complete- ness in Language Models: AND, OR, and ADDER Gates. In D. Belgrave, C. Zhang, 8 Can Graph Learning Learn Circuits? H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors,Advances in Neu-...
2025
-
[42]
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. In The Eleventh International Conference on Learning Representations, 2023. URL https: //openrevi...
2023
-
[43]
The Hydra Effect: Emergent Self-repair in Language Model Computations, 2023
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. The Hydra Effect: Emergent Self-repair in Language Model Computations, 2023. URL https: //arxiv.org/abs/2307.15771
2023 arXiv
-
[44]
Explorations of Self-Repair in Language Models
Cody Rushing and Neel Nanda. Explorations of Self-Repair in Language Models. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,Proceedings of the 41st International Conference on Machine Learn-...
2024
-
[45]
Janizek, Pascal Sturmfels, and Su-In Lee
Joseph D. Janizek, Pascal Sturmfels, and Su-In Lee. Explaining Explanations: Axiomatic Feature Interactions for Deep Networks.Journal of Machine Learning Research, 22(104):1–54,
-
[46]
URLhttp://jmlr.org/papers/v22/20-1223.html. 22
-
[47]
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability, 2026
Barsat Khadka. MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability, 2026. URLhttps://arxiv.org/abs/2605.26343. 25
2026 arXiv
-
[48]
Cohen, Anna Korhonen, and Yonatan Belinkov
Shun Shao, Binxu Wang, Shay B. Cohen, Anna Korhonen, and Yonatan Belinkov. Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer, 2026. URL https://arxiv.org/ abs/2604.24302. 25
2026 arXiv
-
[49]
Derpanis
Harrish Thasarathan, Matthew Kowal, Thomas Fel, and Konstantinos G. Derpanis. Cross- Model Circuit Discovery. InMechanistic Interpretability Workshop at ICML 2026, 2026. URL https://openreview.net/forum?id=SliNacpKqc. 25
2026
-
[50]
Scalable Circuit Learning for Interpreting Large Language Models, 2026
Naiyu Yin, Dennis Wei, Tian Gao, Amit Dhurandhar, Karthikeyan Natesan Ramamurthy, and Yue Yu. Scalable Circuit Learning for Interpreting Large Language Models, 2026. URL https://arxiv.org/abs/2606.16939. 25
2026
-
[51]
Copy Suppression: Comprehensively Understanding an Attention Head, 2023
Callum McDougall, Arthur Conmy, Cody Rushing, Thomas McGrath, and Neel Nanda. Copy Suppression: Comprehensively Understanding an Attention Head, 2023. URLhttps://arxiv. org/abs/2310.04625. 22
2023 arXiv
-
[52]
Blelloch
Guy E. Blelloch. Programming parallel algorithms.Communications of the ACM, 39(3):85–97, March 1996. ISSN 1557-7317. doi: 10.1145/227234.227246. URL http://dx.doi.org/10. 1145/227234.227246. 23
1996
-
[53]
Interpreting Graph Neural Net- works for NLP With Differentiable Edge Masking
Michael Sejr Schlichtkrull, Nicola De Cao, and Ivan Titov. Interpreting Graph Neural Net- works for NLP With Differentiable Edge Masking. InInternational Conference on Learning Representations, 2021. URLhttps://openreview.net/forum?id=WznmQa42ZAx. 25
2021
-
[54]
Generative Causal Explanations for Graph Neural Networks
Wanyu Lin, Hao Lan, and Baochun Li. Generative Causal Explanations for Graph Neural Networks. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 6666–6679. P...
2021
-
[55]
GFT: Graph Foundation Model with Transferable Tree V ocabulary
Zehong Wang, Zheyuan Zhang, Nitesh V Chawla, Chuxu Zhang, and Yanfang Ye. GFT: Graph Foundation Model with Transferable Tree V ocabulary. In A. Globerson, L. Mackey, D. Bel- grave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information 9 Can Graph ...
2024 doi
-
[56]
Equivariance Every- where All At Once: A Recipe for Graph Foundation Models
Ben Finkelshtein, Ismail Ilkan Ceylan, Michael Bronstein, and Ron Levie. Equivariance Every- where All At Once: A Recipe for Graph Foundation Models. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors,Advances in Neu- ral Information Pr...
2025
-
[57]
GNNExplainer: Generating Explanations for Graph Neural Networks
Zhitao Ying, Dylan Bourgeois, Jiaxuan You, Marinka Zitnik, and Jure Leskovec. GNNExplainer: Generating Explanations for Graph Neural Networks. In H. Wal- lach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, edi- tors,Advances in Neural Information Proce...
2019
-
[58]
On Explainability of Graph Neural Networks via Subgraph Explorations
Hao Yuan, Haiyang Yu, Jie Wang, Kang Li, and Shuiwang Ji. On Explainability of Graph Neural Networks via Subgraph Explorations. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Lea...
2021
-
[59]
Discovering Interpretable Algorithms by Decompiling Transformers to RASP, 2026
Xinting Huang, Aleksandra Bakalova, Satwik Bhattamishra, William Merrill, and Michael Hahn. Discovering Interpretable Algorithms by Decompiling Transformers to RASP, 2026. URLhttps://arxiv.org/abs/2602.08857. 26
2026 arXiv
-
[60]
Compact Proofs of Model Performance via Mechanistic Interpretability
Jason Gross, Rajashree Agrawal, Thomas Kwa, Euan Ong, Chun Hei Yip, Alex Gib- son, Soufiane Noubir, and Lawrence Chan. Compact Proofs of Model Performance via Mechanistic Interpretability. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Pa- quet, J. Tomczak, and C. Zhang, ...
2024
-
[61]
reads” information from the residual stream through a projection W ∗ I , and (ii) “writes
Itamar Hadad, Guy Katz, and Shahaf Bassan. Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees. InThe Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=Timsb74vIY. 26 A A Transformer...
2026
-
[63]
Learning Transformer Programs
Dan Friedman, Alexander Wettig, and Danqi Chen. Learning Transformer Programs. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Processing Systems, volume 36, pages 49044–49067. Curran Associates, Inc., 2023. URL ht...
2023
-
[64]
Neural Decompiling of Tracr Transformers
Hannes Thurnherr and Kaspar Riesen. Neural Decompiling of Tracr Transformers. In Ching Yee Suen, Adam Krzyzak, Mirco Ravanelli, Edmondo Trentin, Cem Subakan, and Nicola Nobile, editors,Artificial Neural Networks in Pattern Recognition, pages 25–36, Cham, 2024. Springer Nature ...
2024
-
[68]
RASP program templates.After standardizing function, helper, argument, and local names; removing docstrings, decorators, annotations, and type comments; and replacing string and numeric literal values with placeholders, two pairs match when their program templates are identical
-
[69]
17 Can Graph Learning Learn Circuits?
Exact ground-truth circuits.Two pairs match when their ground-truth circuits contain the same named model components, including components with no circuit connection, and the same directed component edges labeled as part of the circuit. 17 Can Graph Learning Learn Circuits?
-
[70]
median [Q1–Q3]
Circuit structure.When two circuits are not exact matches, we use directed graph isomorphism. Two pairs match when their circuit graphs can be put into one-to-one correspondence while preserving edge directions and broad component categories, but ignoring individual component ...
-
[2017]
URL https://proceedings.neurips.cc/paper_files/paper/2017/file/ f22e4747da1aa27e363d86d40ff442fe-Paper.pdf. 3, 16
2017
-
[2020]
2, 3, 16, 17, 24, 25
URL https://proceedings.neurips.cc/paper_files/paper/2020/file/ e37b08dd3015330dcbb5d6663667b8b8-Paper.pdf. 2, 3, 16, 17, 24, 25
2020
-
[2021]
doi: 10.18653/v1/2021.naacl-main.74
Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.74. URL https://aclanthology.org/2021.naacl-main.74/. 24, 25
2021 doi
-
[2024]
doi: 10.18653/v1/2024.blackboxnlp-1.25
Association for Computational Linguistics. doi: 10.18653/v1/2024.blackboxnlp-1.25. URLhttps://aclanthology.org/2024.blackboxnlp-1.25/. 1, 22, 23, 25
2024 doi
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.