REVIEW 3 major objections 48 references
A learned dynamics model can decode a conserved quantity at near-perfect accuracy and still not use it; whether the quantity is causally deployed is set by the training objective and the invariant’s algebra relative to the output, not by th
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 00:20 UTC pith:DN53EBVU
load-bearing objection Clean presence≠use result for continuous invariants; the slogan over-weights a frozen two-readout protocol, but the core finding and instrument still hold. the 3 major comments →
The Objective Decides: When a Learned Dynamics Model Uses a Conserved Quantity
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Linear decodability of a conserved quantity does not imply causal use. Single-step activation interchange shows a direction decoded at R² ≈ 1 can be inert on next-state prediction (transfer-corr τ ≈ 0) while becoming load-bearing (τ → +1) under an invariant-relevant objective, with representation and probe held fixed. Deployment is therefore a property of the objective and of the invariant–output algebra, not of presence. The deployment gap forecasts OOD accuracy where every model decodes the target at R² = 1.00.
What carries the argument
The single-step deployment probe: fit a linear invariant direction, overwrite that component with a donor state’s value, complete one forward step, and measure the scale-free transfer-corr τ between the change in the donor’s invariant and the change in a scalar readout of the prediction. τ ≈ 0 means present but inert; τ → +1 means load-bearing. The deployment gap is τ under an invariant-relevant readout minus τ under the model’s own next-state output.
Load-bearing premise
That flipping the post-interchange readout—or training a head on the same frozen features to output the invariant—fairly measures objective-dependent causal deployment, rather than mainly showing that the probe direction is accessible once the experimenter chooses an invariant-aligned target.
What would settle it
On a next-state-trained model with a nonlinear, output-separated invariant (e.g. pendulum energy), multi-direction or SAE-aligned interchange that yields τ near +1 on the model’s own next-state output would refute inertness; end-to-end training on an invariant objective that leaves τ near 0 for that same invariant on its own objective would refute the claim that the objective decides deployment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that linear decodability of a conserved quantity from a learned dynamics model does not imply causal use. Across mechanical, circuit, and PDE systems (and a frozen 158M Poseidon-B checkpoint on real Navier–Stokes), energy and related invariants are recoverable at R²≈1 yet causally inert under single-step activation interchange on next-state prediction (transfer-corr τ≈0), while the same direction becomes load-bearing under an invariant-relevant readout or objective (τ→+1). Deployment is further tied to an algebraic relation between the invariant and the output coordinates, demonstrated by flipping linear momentum from confounded to load-bearing via an asinh output map alone. The deployment gap γ=τ_inv−τ_next correlates with OOD accuracy at r=+0.97 across twelve models that all decode at R²=1.00. The authors propose a cheap single-step interchange instrument (Algorithm 1) and argue that interpretability should report causal deployment, not presence, when the question is whether a model uses a quantity.
Significance. If the results hold, the paper cleanly separates two claims that the scientific-ML and probing literatures routinely conflate—presence versus causal use of physical invariants—and supplies a practical, scale-free instrument with matched nulls. Strengths include broad substrate/architecture coverage, bootstrap and permutation statistics (App. D), explicit redundancy lower-bound analysis, an unsupervised SAE localization check, a depth-resolved foundation-model replication on real trajectories, and a downstream OOD forecast where decodability is saturated. The algebraic flip (same invariant, only output algebra changed) is a particularly clean control. These are falsifiable, reproducible claims of direct interest to interpretability of dynamics surrogates and PDE foundation models.
major comments (3)
- Abstract, §1 Move 2, and §5 pendulum anchor: the claim that “the same direction in the same representation becomes causally load-bearing the moment the training objective rewards the invariant” is not fully supported by the primary two-readout protocol. In Algorithm 1 and the pendulum anchor, f_θ is frozen after next-state training; only the post-interchange scalar (or a small head g_inv fit on frozen features to output φ) changes. That design shows linear accessibility of û_ℓ to a φ-head, not that retraining under an invariant-rewarding objective rewires deployment of that direction. End-to-end objective comparisons do exist (wave PDE presence R² 0.13 vs 0.99; architecture-graded load-bearing; asinh(P) flip; Poseidon depth), but the headline framing treats the frozen protocol as primary. Please either (i) retrain end-to-end invariant-objective models for the main matrix and report τ_nex
- §5 “The payoff” / Figure 4: the OOD result (r=+0.97, 7.2× accuracy gap) is load-bearing for the claim that the deployment gap “has teeth,” yet the construction of the twelve shortcut vs robust models is underspecified in the main text (only “a task with a spurious shortcut”). Free design choices in how the shortcut is injected and how “robust” models are induced could inflate the correlation. Please specify the task, the shortcut mechanism, the training differences that produce the two clusters, and whether the deployment gap was computed before or after OOD evaluation (pre-registration / held-out protocol).
- §3 algebraic predicate and §5 taxonomy: the predicate “output-entangled iff approximately affine in the readout coordinates on-distribution” predicts next-state τ by construction for linear/extensive invariants (Table 2 shaded row; spring/wave momentum τ≈−0.7 to −1.02). The asinh(P) flip is the decisive non-tautological test and should be elevated as the primary evidence for the taxonomy; the confounded linear cases should be labeled more clearly as algebraic leakage rather than as independent support for the predicate. A short formal statement of the on-distribution linear part (e.g., ∂φ/∂y ≈ 0 vs ≠ 0) would make the claim falsifiable rather than post-hoc.
Circularity Check
Frozen-features + φ-head τ_inv is near a positive control for linear accessibility, so part of the “objective decides / same representation becomes load-bearing” slogan is by construction; next-state inertness, algebraic flip, and OOD link remain independent.
specific steps
-
self definitional
[Algorithm 1; §3 Eqs. 4–5; §5 pendulum anchor / Fig. 2b]
"g_inv (head trained to output φ) ... Re-reading out the same frozen features against an invariant-relevant target ... flips the verdict to load-bearing: ... τ=+0.97. ... only the quantity we regress the post-step change onto differs. ... What moves it is the objective: ... Thesame direction in the same representation becomes causally load-bearing (τ→+1) the moment the training objective rewards the invariant, so deployment is a property of the objective"
Under the frozen two-readout protocol, f_θ and h_ℓ are fixed after next-state training; g_inv is fit to predict φ from features that already admit a linear decode of φ at R²≈1. Interchanging along that decode direction and then reading τ=corr(Δg_inv, Δφ) therefore largely checks that the φ-head is sensitive to the φ direction—i.e. accessibility of a fitted linear code—not that the dynamics model was retrained under an invariant-rewarding objective. Calling this “the training objective rewards the invariant” and using it as primary evidence that “the objective decides” deployment of the same representation equates a near-positive control with objective-dependent causal use. (End-to-end invariant-trained models and the asinh flip are separate and non-circular.)
-
fitted input called prediction
[Algorithm 1 steps 1–7; Appendix A readout targets; Table 2 inv-obj τ column]
"w_ℓ ← z-scored CV ridge of φ(s) on h_ℓ(s); û_ℓ ← w_ℓ/∥w_ℓ∥ ... g_inv is a small head trained to output φ from the frozen features (held-out fit R² ≥0.99). ... τ_g ← corr(Δg, Δφ) ... invariant obj. τ (load-bearing) ... +0.97 / +0.31 / +1.00 / +0.97 ..."
The decode fit finds the linear direction of φ; g_inv is then supervised to output the same φ from the same frozen activations. Reporting high inv-obj τ after patching û_ℓ is statistically forced to the extent g_inv relies on that (or redundant parallel) linear code—analogous to fitting a parameter then “predicting” a quantity that is the fit. The paper’s load-bearing column for frozen next-state backbones therefore does not independently establish objective-driven deployment; it largely reconfirms that a φ-supervised head can use a φ-decodable direction.
full rationale
The paper’s central negative result—perfect linear decodability of energy/invariants with next-state transfer-corr τ≈0 under single-step interchange—is an empirical measurement on frozen next-state models and is not forced by the training loss or by the definition of R². The algebraic taxonomy and asinh(P) output flip, Poseidon depth structure, end-to-end presence differences (e.g. wave PDE decode R² 0.13 vs 0.99), and the OOD correlation are likewise experimental, not definitional. Mild circularity appears only in one arm of the primary two-readout instrument: when g_inv is a head trained to output φ from frozen features that already linearly encode φ at R²≈1, high τ_inv after patching the decode direction is largely expected if that head uses the fitted direction. Framing that arm as “the same direction becomes causally load-bearing the moment the training objective rewards the invariant” and as proof that “deployment is a property of the objective” overstates a near-positive-control for accessibility as objective-driven redeployment of f_θ (whose weights stay next-state-trained). That does not collapse the whole paper: stronger end-to-end and algebra-flip evidence still supports objective/algebra dependence. No self-citation uniqueness chain, ansatz smuggling, or renaming of a known closed-form result was found. Score 3 reflects partial, localized by-construction load on the frozen τ_inv rhetoric, not a fully circular derivation.
Axiom & Free-Parameter Ledger
free parameters (4)
- Ridge / CV probe hyperparameters and layer choice
- Interchange pair sampling and readout head training for g_inv
- Model widths, seeds, and capacity-sweep grid {4,8,16,32,64}
- OOD shortcut-task construction (twelve models)
axioms (4)
- domain assumption Single-step activation interchange along a unit linear direction is a valid lower-bound test of whether the forward computation uses a continuous invariant.
- domain assumption Ground-truth conserved quantities of the simulated systems (energy, angular momentum, etc.) are correctly specified and approximately conserved under the integrators used.
- ad hoc to paper An invariant is output-entangled iff approximately affine in the readout coordinates on-distribution, else output-separated; this algebraic relation predicts next-state interchange verdicts.
- standard math Standard linear probing, OLS slope, and Pearson correlation are appropriate summary statistics for presence and causal transfer.
invented entities (2)
-
Transfer-corr τ and deployment gap γ = τ_inv − τ_next
no independent evidence
-
Algebraic taxonomy of output-entangled vs output-separated invariants
independent evidence
read the original abstract
A linear probe that recovers a conserved quantity from a learned dynamics model's activations is routinely read as evidence that the model uses that quantity. We show this inference is unsound. Across mechanical, circuit, and partial-differential-equation (PDE) systems, and on a 158M-parameter pretrained PDE foundation model, energy and other invariants are linearly decodable at $R^2 \approx 1$ yet causally inert on next-state prediction: overwriting the decoded direction with a donor state's value (single-step activation interchange) leaves the forward pass essentially unchanged (transfer-corr $\tau \approx 0$). The same direction in the same representation becomes causally load-bearing ($\tau \to +1$) the moment the training objective rewards the invariant, so deployment is a property of the objective, not of the representation or the probe. We further show that when an invariant is deployed is governed by a precise algebraic predicate (its relation to the prediction output), by flipping a single invariant from inert to load-bearing by changing only the output's algebra. Finally, the gap has teeth: across models that all decode the target at $R^2 = 1.00$, the deployment gap forecasts out-of-distribution (OOD) accuracy ($r = +0.97$) where decodability is blind. We argue that causal deployment, not decodability, is what interpretability should measure when the question is whether a model uses a piece of knowledge, and we give a cheap instrument for measuring it.
Figures
Reference graph
Works this paper leans on
-
[1]
Understanding intermediate layers using linear classifier probes.ICLR Workshop, 2017
Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes.ICLR Workshop, 2017
2017
-
[2]
Noether networks: Meta-learning useful conserved quantities
Ferran Alet, Dylan Doblar, Allan Zhou, Josh Tenenbaum, Kenji Kawaguchi, and Chelsea Finn. Noether networks: Meta-learning useful conserved quantities. InAdvances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[3]
Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, and Koray Kavukcuoglu
Peter W. Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, and Koray Kavukcuoglu. Interaction networks for learning about objects, relations and physics. In Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[4]
Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022
Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022
2022
-
[5]
Worrall, and Max Welling
Johannes Brandstetter, Daniel E. Worrall, and Max Welling. Message passing neural PDE solvers. InInternational Conference on Learning Representations (ICLR), 2022
2022
-
[6]
Towards monosemanticity: Decompos- ing language models with dictionary learning.Transformer Circuits Thread, 2023
Trenton Bricken, Adly Templeton, Joshua Batson, et al. Towards monosemanticity: Decompos- ing language models with dictionary learning.Transformer Circuits Thread, 2023
2023
-
[7]
Brunton, Joshua L
Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Discovering governing equations from data by sparse identification of nonlinear dynamical systems.Proceedings of the National Academy of Sciences, 113(15):3932–3937, 2016
2016
-
[8]
Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri`a Garriga-Alonso
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri`a Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[9]
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Lo¨ıc Barrault, and Marco Baroni. What you can cram into a single vector: Probing sentence embeddings for linguistic properties. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), 2018
2018
-
[10]
Lagrangian neural networks
Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho. Lagrangian neural networks. InICLR Workshop on Integration of Deep Neural Models and Differential Equations, 2020
2020
-
[11]
Discovering symbolic models from deep learning with inductive biases
Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho. Discovering symbolic models from deep learning with inductive biases. InAdvances in Neural Information Processing Systems (NeurIPS), 2020. 10 Preprint
2020
-
[12]
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. InInternational Conference on Learning Representations (ICLR), 2024
2024
-
[13]
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, et al. Toy models of superposition. In Transformer Circuits Thread, 2022
2022
-
[14]
Scaling and evaluating sparse autoencoders.arXiv preprint arXiv:2406.04093, 2024
Leo Gao, Tom Dupr ´e la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders.arXiv preprint arXiv:2406.04093, 2024
Pith/arXiv arXiv 2024
-
[15]
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[16]
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman. Finding alignments between interpretable causal variables and distributed neural representations. InConference on Causal Learning and Reasoning (CLeaR), 2024
2024
-
[17]
Wichmann
Robert Geirhos, J¨orn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
2020
-
[18]
Hamiltonian neural networks
Samuel Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2019
2019
-
[19]
Language models represent space and time
Wes Gurnee and Max Tegmark. Language models represent space and time. InInternational Conference on Learning Representations (ICLR), 2024
2024
-
[20]
Seungwoong Ha and Hawoong Jeong. Discovering conservation laws from trajectories via machine learning.arXiv preprint arXiv:2102.04008, 2021
Pith/arXiv arXiv 2021
-
[21]
Poseidon: Efficient foundation models for PDEs
Maximilian Herde, Bogdan Raoni ´c, Tobias Rohner, Roger K¨appeli, Roberto Molinaro, Em- manuel de B´ezenac, and Siddhartha Mishra. Poseidon: Efficient foundation models for PDEs. InAdvances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[22]
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In Proceedings of EMNLP-IJCNLP, 2019
2019
-
[23]
John Hewitt and Christopher D. Manning. A structural probe for finding syntax in word representations. InProceedings of NAACL-HLT, 2019
2019
-
[24]
Discovering physical concepts with neural networks.Physical Review Letters, 124(1):010508, 2020
Raban Iten, Tony Metger, Henrik Wilming, L´ıdia del Rio, and Renato Renner. Discovering physical concepts with neural networks.Physical Review Letters, 124(1):010508, 2020
2020
-
[25]
Emergent representations of program semantics in language models trained on programs
Charles Jin and Martin Rinard. Emergent representations of program semantics in language models trained on programs. InInternational Conference on Machine Learning (ICML), 2024
2024
-
[26]
Rediscovering orbital mechanics with machine learning.Machine Learning: Science and Technology, 4(4): 045002, 2023
Pablo Lemos, Niall Jeffrey, Miles Cranmer, Shirley Ho, and Peter Battaglia. Rediscovering orbital mechanics with machine learning.Machine Learning: Science and Technology, 4(4): 045002, 2023
2023
-
[27]
Hopkins, David Bau, Fernanda Vi´egas, Hanspeter Pfister, and Martin Wattenberg
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Vi´egas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[28]
Fourier neural operator for parametric partial differen- tial equations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. Fourier neural operator for parametric partial differen- tial equations. InInternational Conference on Learning Representations (ICLR), 2021
2021
-
[29]
Machine learning conservation laws from trajectories.Physical Review Letters, 126(18):180604, 2021
Ziming Liu and Max Tegmark. Machine learning conservation laws from trajectories.Physical Review Letters, 126(18):180604, 2021
2021
-
[30]
Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators
Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, 2021. 11 Preprint
2021
-
[31]
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. InAdvances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[32]
Corrado, and Jeff Dean
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. InAdvances in Neural Information Processing Systems (NeurIPS), 2013
2013
-
[33]
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[34]
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. InBlackboxNLP Workshop at EMNLP, 2023
2023
-
[35]
Zoom in: An introduction to circuits.Distill, 2020
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits.Distill, 2020
2020
-
[36]
Kiho Park, Yo Joong Choe, and Victor Veitch. The linear representation hypothesis and the geometry of large language models.arXiv preprint arXiv:2311.03658, 2023
Pith/arXiv arXiv 2023
-
[37]
Karniadakis
Maziar Raissi, Paris Perdikaris, and George E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.Journal of Computational Physics, 378:686–707, 2019
2019
-
[38]
Toward transparency in AI: Survey on interpreting the inner structures of deep neural networks.IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2023
Tilman R¨auker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. Toward transparency in AI: Survey on interpreting the inner structures of deep neural networks.IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2023
2023
-
[39]
Battaglia
Alvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying, Jure Leskovec, and Peter W. Battaglia. Learning to simulate complex physics with graph networks. InInternational Conference on Machine Learning (ICML), 2020
2020
-
[40]
Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, et al. Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet. Transformer Circuits Thread, 2024
2024
-
[41]
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. BERT rediscovers the classical NLP pipeline. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), 2019
2019
-
[42]
Chess as a testbed for language model state tracking
Shubham Toshniwal, Sam Wiseman, Karen Livescu, and Kevin Gimpel. Chess as a testbed for language model state tracking. InProceedings of the AAAI Conference on Artificial Intelligence, 2022
2022
-
[43]
AI Feynman: A physics-inspired method for symbolic regression.Science Advances, 6(16):eaay2631, 2020
Silviu-Marian Udrescu and Max Tegmark. AI Feynman: A physics-inspired method for symbolic regression.Science Advances, 6(16):eaay2631, 2020
2020
-
[44]
Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan
Keyon Vafa, Justin Y . Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the world model implicit in a generative model. InAdvances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[45]
Chang, Ashesh Rambachan, and Sendhil Mullainathan
Keyon Vafa, Peter G. Chang, Ashesh Rambachan, and Sendhil Mullainathan. What has a foundation model found? Using inductive bias to probe for world models. InInternational Conference on Machine Learning (ICML), 2025
2025
-
[46]
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Investigating gender bias in language models using causal mediation analysis. InAdvances in Neural Information Processing Systems (NeurIPS), 2020
2020
-
[47]
Interpretability in the wild: A circuit for indirect object identification in GPT-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: A circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR), 2023
2023
-
[48]
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman. Interpretability at scale: Identifying causal mechanisms in Alpaca. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. 12 Preprint A INSTRUMENT DETAILS Decode fit.Unless noted, wℓ is fit by ridge regression on z-scored activations with 5-fold cross- valid...
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.