Pith. sign in

REVIEW 3 major objections 48 references

A learned dynamics model can decode a conserved quantity at near-perfect accuracy and still not use it; whether the quantity is causally deployed is set by the training objective and the invariant’s algebra relative to the output, not by th

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 00:20 UTC pith:DN53EBVU

load-bearing objection Clean presence≠use result for continuous invariants; the slogan over-weights a frozen two-readout protocol, but the core finding and instrument still hold. the 3 major comments →

arxiv 2607.03728 v1 pith:DN53EBVU submitted 2026-07-04 cs.CE

The Objective Decides: When a Learned Dynamics Model Uses a Conserved Quantity

classification cs.CE
keywords interpretabilityprobingconserved quantitiesactivation interchangecausal deploymentlearned dynamics modelsPDE foundation modelsout-of-distribution generalization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When a linear probe recovers energy or another conserved quantity from a dynamics model’s activations, the field often treats that as evidence the model uses the quantity. This paper shows that inference is unsound. Across mechanical systems, circuits, PDEs, and a 158M-parameter pretrained PDE foundation model, invariants are linearly decodable at R² near 1 yet causally inert on next-state prediction: overwriting the decoded direction with a donor state’s value leaves the one-step forecast essentially unchanged. The same direction in the same representation becomes load-bearing the moment the training objective rewards the invariant, and a precise algebraic rule—whether the invariant is affine in the output coordinates—predicts when next-state prediction deploys it. The gap has downstream force: across models that all decode the target at R² = 1.00, the deployment gap forecasts out-of-distribution accuracy at r = +0.97, so decodability would certify failures that the causal test separates. The argument is that causal deployment, not decodability, is what interpretability should measure when the question is whether a model uses a piece of knowledge.

Core claim

Linear decodability of a conserved quantity does not imply causal use. Single-step activation interchange shows a direction decoded at R² ≈ 1 can be inert on next-state prediction (transfer-corr τ ≈ 0) while becoming load-bearing (τ → +1) under an invariant-relevant objective, with representation and probe held fixed. Deployment is therefore a property of the objective and of the invariant–output algebra, not of presence. The deployment gap forecasts OOD accuracy where every model decodes the target at R² = 1.00.

What carries the argument

The single-step deployment probe: fit a linear invariant direction, overwrite that component with a donor state’s value, complete one forward step, and measure the scale-free transfer-corr τ between the change in the donor’s invariant and the change in a scalar readout of the prediction. τ ≈ 0 means present but inert; τ → +1 means load-bearing. The deployment gap is τ under an invariant-relevant readout minus τ under the model’s own next-state output.

Load-bearing premise

That flipping the post-interchange readout—or training a head on the same frozen features to output the invariant—fairly measures objective-dependent causal deployment, rather than mainly showing that the probe direction is accessible once the experimenter chooses an invariant-aligned target.

What would settle it

On a next-state-trained model with a nonlinear, output-separated invariant (e.g. pendulum energy), multi-direction or SAE-aligned interchange that yields τ near +1 on the model’s own next-state output would refute inertness; end-to-end training on an invariant objective that leaves τ near 0 for that same invariant on its own objective would refute the claim that the objective decides deployment.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper argues that linear decodability of a conserved quantity from a learned dynamics model does not imply causal use. Across mechanical, circuit, and PDE systems (and a frozen 158M Poseidon-B checkpoint on real Navier–Stokes), energy and related invariants are recoverable at R²≈1 yet causally inert under single-step activation interchange on next-state prediction (transfer-corr τ≈0), while the same direction becomes load-bearing under an invariant-relevant readout or objective (τ→+1). Deployment is further tied to an algebraic relation between the invariant and the output coordinates, demonstrated by flipping linear momentum from confounded to load-bearing via an asinh output map alone. The deployment gap γ=τ_inv−τ_next correlates with OOD accuracy at r=+0.97 across twelve models that all decode at R²=1.00. The authors propose a cheap single-step interchange instrument (Algorithm 1) and argue that interpretability should report causal deployment, not presence, when the question is whether a model uses a quantity.

Significance. If the results hold, the paper cleanly separates two claims that the scientific-ML and probing literatures routinely conflate—presence versus causal use of physical invariants—and supplies a practical, scale-free instrument with matched nulls. Strengths include broad substrate/architecture coverage, bootstrap and permutation statistics (App. D), explicit redundancy lower-bound analysis, an unsupervised SAE localization check, a depth-resolved foundation-model replication on real trajectories, and a downstream OOD forecast where decodability is saturated. The algebraic flip (same invariant, only output algebra changed) is a particularly clean control. These are falsifiable, reproducible claims of direct interest to interpretability of dynamics surrogates and PDE foundation models.

major comments (3)
  1. Abstract, §1 Move 2, and §5 pendulum anchor: the claim that “the same direction in the same representation becomes causally load-bearing the moment the training objective rewards the invariant” is not fully supported by the primary two-readout protocol. In Algorithm 1 and the pendulum anchor, f_θ is frozen after next-state training; only the post-interchange scalar (or a small head g_inv fit on frozen features to output φ) changes. That design shows linear accessibility of û_ℓ to a φ-head, not that retraining under an invariant-rewarding objective rewires deployment of that direction. End-to-end objective comparisons do exist (wave PDE presence R² 0.13 vs 0.99; architecture-graded load-bearing; asinh(P) flip; Poseidon depth), but the headline framing treats the frozen protocol as primary. Please either (i) retrain end-to-end invariant-objective models for the main matrix and report τ_nex
  2. §5 “The payoff” / Figure 4: the OOD result (r=+0.97, 7.2× accuracy gap) is load-bearing for the claim that the deployment gap “has teeth,” yet the construction of the twelve shortcut vs robust models is underspecified in the main text (only “a task with a spurious shortcut”). Free design choices in how the shortcut is injected and how “robust” models are induced could inflate the correlation. Please specify the task, the shortcut mechanism, the training differences that produce the two clusters, and whether the deployment gap was computed before or after OOD evaluation (pre-registration / held-out protocol).
  3. §3 algebraic predicate and §5 taxonomy: the predicate “output-entangled iff approximately affine in the readout coordinates on-distribution” predicts next-state τ by construction for linear/extensive invariants (Table 2 shaded row; spring/wave momentum τ≈−0.7 to −1.02). The asinh(P) flip is the decisive non-tautological test and should be elevated as the primary evidence for the taxonomy; the confounded linear cases should be labeled more clearly as algebraic leakage rather than as independent support for the predicate. A short formal statement of the on-distribution linear part (e.g., ∂φ/∂y ≈ 0 vs ≠ 0) would make the claim falsifiable rather than post-hoc.

Circularity Check

2 steps flagged

Frozen-features + φ-head τ_inv is near a positive control for linear accessibility, so part of the “objective decides / same representation becomes load-bearing” slogan is by construction; next-state inertness, algebraic flip, and OOD link remain independent.

specific steps
  1. self definitional [Algorithm 1; §3 Eqs. 4–5; §5 pendulum anchor / Fig. 2b]
    "g_inv (head trained to output φ) ... Re-reading out the same frozen features against an invariant-relevant target ... flips the verdict to load-bearing: ... τ=+0.97. ... only the quantity we regress the post-step change onto differs. ... What moves it is the objective: ... Thesame direction in the same representation becomes causally load-bearing (τ→+1) the moment the training objective rewards the invariant, so deployment is a property of the objective"

    Under the frozen two-readout protocol, f_θ and h_ℓ are fixed after next-state training; g_inv is fit to predict φ from features that already admit a linear decode of φ at R²≈1. Interchanging along that decode direction and then reading τ=corr(Δg_inv, Δφ) therefore largely checks that the φ-head is sensitive to the φ direction—i.e. accessibility of a fitted linear code—not that the dynamics model was retrained under an invariant-rewarding objective. Calling this “the training objective rewards the invariant” and using it as primary evidence that “the objective decides” deployment of the same representation equates a near-positive control with objective-dependent causal use. (End-to-end invariant-trained models and the asinh flip are separate and non-circular.)

  2. fitted input called prediction [Algorithm 1 steps 1–7; Appendix A readout targets; Table 2 inv-obj τ column]
    "w_ℓ ← z-scored CV ridge of φ(s) on h_ℓ(s); û_ℓ ← w_ℓ/∥w_ℓ∥ ... g_inv is a small head trained to output φ from the frozen features (held-out fit R² ≥0.99). ... τ_g ← corr(Δg, Δφ) ... invariant obj. τ (load-bearing) ... +0.97 / +0.31 / +1.00 / +0.97 ..."

    The decode fit finds the linear direction of φ; g_inv is then supervised to output the same φ from the same frozen activations. Reporting high inv-obj τ after patching û_ℓ is statistically forced to the extent g_inv relies on that (or redundant parallel) linear code—analogous to fitting a parameter then “predicting” a quantity that is the fit. The paper’s load-bearing column for frozen next-state backbones therefore does not independently establish objective-driven deployment; it largely reconfirms that a φ-supervised head can use a φ-decodable direction.

full rationale

The paper’s central negative result—perfect linear decodability of energy/invariants with next-state transfer-corr τ≈0 under single-step interchange—is an empirical measurement on frozen next-state models and is not forced by the training loss or by the definition of R². The algebraic taxonomy and asinh(P) output flip, Poseidon depth structure, end-to-end presence differences (e.g. wave PDE decode R² 0.13 vs 0.99), and the OOD correlation are likewise experimental, not definitional. Mild circularity appears only in one arm of the primary two-readout instrument: when g_inv is a head trained to output φ from frozen features that already linearly encode φ at R²≈1, high τ_inv after patching the decode direction is largely expected if that head uses the fitted direction. Framing that arm as “the same direction becomes causally load-bearing the moment the training objective rewards the invariant” and as proof that “deployment is a property of the objective” overstates a near-positive-control for accessibility as objective-driven redeployment of f_θ (whose weights stay next-state-trained). That does not collapse the whole paper: stronger end-to-end and algebra-flip evidence still supports objective/algebra dependence. No self-citation uniqueness chain, ansatz smuggling, or renaming of a known closed-form result was found. Score 3 reflects partial, localized by-construction load on the frozen τ_inv rhetoric, not a fully circular derivation.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The paper is empirical interpretability, not a derivation from new physics. Load-bearing content is experimental protocol plus standard dynamical systems and causal-intervention machinery. Free parameters are ordinary ML/probe choices; axioms are domain and method assumptions; invented entities are named metrics (τ, γ), not new physical objects. No graviton-like mediators.

free parameters (4)
  • Ridge / CV probe hyperparameters and layer choice
    Decode direction û_ℓ from z-scored CV ridge; layer chosen as decode-maximizing. Affects which direction is patched; standard but free.
  • Interchange pair sampling and readout head training for g_inv
    Number of (target, donor) pairs, and training of the invariant head on frozen features, set the measured τ_inv and gap.
  • Model widths, seeds, and capacity-sweep grid {4,8,16,32,64}
    Architecture and capacity choices change redundancy and single-direction under-reading of deployment.
  • OOD shortcut-task construction (twelve models)
    The r=+0.97 result depends on how shortcut vs robust models and the spurious feature are defined; not a fixed public benchmark fully specified in the main text.
axioms (4)
  • domain assumption Single-step activation interchange along a unit linear direction is a valid lower-bound test of whether the forward computation uses a continuous invariant.
    §3–4; standard causal abstraction / activation patching applied to continuous φ; redundancy makes it a lower bound (§5).
  • domain assumption Ground-truth conserved quantities of the simulated systems (energy, angular momentum, etc.) are correctly specified and approximately conserved under the integrators used.
    Table 3 / App. B; RK4 with reported energy drift; answer key for decode and Δφ.
  • ad hoc to paper An invariant is output-entangled iff approximately affine in the readout coordinates on-distribution, else output-separated; this algebraic relation predicts next-state interchange verdicts.
    §3 algebraic predicate and §5 taxonomy; empirically supported by the asinh(P) flip but is a paper-introduced organizing rule.
  • standard math Standard linear probing, OLS slope, and Pearson correlation are appropriate summary statistics for presence and causal transfer.
    Eqs. 1–4; conventional statistics.
invented entities (2)
  • Transfer-corr τ and deployment gap γ = τ_inv − τ_next no independent evidence
    purpose: Scale-free metric of causal deployment and the presence–use gap used as the paper’s main instrument and OOD predictor.
    Defined in Eqs. 4–5; metrics rather than physical entities; independent_evidence false as named constructs, though operationally measurable.
  • Algebraic taxonomy of output-entangled vs output-separated invariants independent evidence
    purpose: Predict when next-state interchange is confounded vs inert, independent of invariant identity.
    §3 and §5; organizing predicate introduced and tested via asinh flip.

pith-pipeline@v1.1.0-grok45 · 23290 in / 3943 out tokens · 46295 ms · 2026-07-12T00:20:24.756296+00:00 · methodology

0 comments
read the original abstract

A linear probe that recovers a conserved quantity from a learned dynamics model's activations is routinely read as evidence that the model uses that quantity. We show this inference is unsound. Across mechanical, circuit, and partial-differential-equation (PDE) systems, and on a 158M-parameter pretrained PDE foundation model, energy and other invariants are linearly decodable at $R^2 \approx 1$ yet causally inert on next-state prediction: overwriting the decoded direction with a donor state's value (single-step activation interchange) leaves the forward pass essentially unchanged (transfer-corr $\tau \approx 0$). The same direction in the same representation becomes causally load-bearing ($\tau \to +1$) the moment the training objective rewards the invariant, so deployment is a property of the objective, not of the representation or the probe. We further show that when an invariant is deployed is governed by a precise algebraic predicate (its relation to the prediction output), by flipping a single invariant from inert to load-bearing by changing only the output's algebra. Finally, the gap has teeth: across models that all decode the target at $R^2 = 1.00$, the deployment gap forecasts out-of-distribution (OOD) accuracy ($r = +0.97$) where decodability is blind. We argue that causal deployment, not decodability, is what interpretability should measure when the question is whether a model uses a piece of knowledge, and we give a cheap instrument for measuring it.

Figures

Figures reproduced from arXiv: 2607.03728 by Chih-Ting Liao, Xin Cao.

Figure 1
Figure 1. Figure 1: The instrument (formalized in Algorithm 1). Decode fits a linear invariant direction uˆℓ; intervene overwrites its component with a donor’s value and runs one forward step; verdict regresses the readout change on the donor’s invariant change, flat, τ ≈ 0 means present but inert; positive, τ →+1 means causally load-bearing. Regress the induced change in g on the change in the donor’s invariant, ∆g = g [PIT… view at source ↗
Figure 2
Figure 2. Figure 2: The decode–deploy dissociation, at a glance. (a) Across six systems, every invariant is decodable yet inert on next-state (left column, τ ≈0) and load-bearing under an invariant-relevant objective (right column, τ →+1). (b) The pendulum anchor (§5): decode R2 = 1 under both readouts; next-state inert, invariant-target load-bearing. (c) The algebraic taxonomy (§5): the next-state verdict tracks the invarian… view at source ↗
Figure 3
Figure 3. Figure 3: Decodability does not predict deployment. Every system decodes its invariant at R2 ≥0.8 (all points at high x), yet the same high decodability is compatible with both a next-state-inert reading (τ ≈0, lower band) and a load-bearing reading (τ spread 0.3–1.0, upper band). The vertical link for each system connects its two objectives: presence is fixed and high; deployment is what varies. This is the decode–… view at source ↗
Figure 4
Figure 4. Figure 4: Deployment predicts generalization; decodability does not. (a) Twelve models all decode the target invariant at R2 = 1.00 (flat, left), so decode is blind to a 7.2× spread in OOD accuracy. (b) The deployment gap forecasts OOD accuracy with r=+0.97: shortcut models (gap ≈−0.13) collapse OOD (0.14), while robust models (gap ≈+0.66) generalize (1.00). The foundation model also exposes a depth axis that the co… view at source ↗
Figure 5
Figure 5. Figure 5: The deployed invariant is redundantly encoded (so our numbers are lower bounds). (a) Deployment strength is non-monotonic in width, peaking at intermediate capacity, evidence of distributed, redundant coding rather than a single privileged unit. (b) Inertness on next-state is architecture-agnostic; deployment magnitude is graded. (c) A sparse-autoencoder latent aligns with the invariant at |corr|= 0.84 and… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

48 extracted references · 3 linked inside Pith

  1. [1]

    Understanding intermediate layers using linear classifier probes.ICLR Workshop, 2017

    Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes.ICLR Workshop, 2017

  2. [2]

    Noether networks: Meta-learning useful conserved quantities

    Ferran Alet, Dylan Doblar, Allan Zhou, Josh Tenenbaum, Kenji Kawaguchi, and Chelsea Finn. Noether networks: Meta-learning useful conserved quantities. InAdvances in Neural Information Processing Systems (NeurIPS), 2021

  3. [3]

    Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, and Koray Kavukcuoglu

    Peter W. Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, and Koray Kavukcuoglu. Interaction networks for learning about objects, relations and physics. In Advances in Neural Information Processing Systems (NeurIPS), 2016

  4. [4]

    Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022

    Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1):207–219, 2022

  5. [5]

    Worrall, and Max Welling

    Johannes Brandstetter, Daniel E. Worrall, and Max Welling. Message passing neural PDE solvers. InInternational Conference on Learning Representations (ICLR), 2022

  6. [6]

    Towards monosemanticity: Decompos- ing language models with dictionary learning.Transformer Circuits Thread, 2023

    Trenton Bricken, Adly Templeton, Joshua Batson, et al. Towards monosemanticity: Decompos- ing language models with dictionary learning.Transformer Circuits Thread, 2023

  7. [7]

    Brunton, Joshua L

    Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Discovering governing equations from data by sparse identification of nonlinear dynamical systems.Proceedings of the National Academy of Sciences, 113(15):3932–3937, 2016

  8. [8]

    Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri`a Garriga-Alonso

    Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri`a Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems (NeurIPS), 2023

  9. [9]

    What you can cram into a single vector: Probing sentence embeddings for linguistic properties

    Alexis Conneau, German Kruszewski, Guillaume Lample, Lo¨ıc Barrault, and Marco Baroni. What you can cram into a single vector: Probing sentence embeddings for linguistic properties. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), 2018

  10. [10]

    Lagrangian neural networks

    Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho. Lagrangian neural networks. InICLR Workshop on Integration of Deep Neural Models and Differential Equations, 2020

  11. [11]

    Discovering symbolic models from deep learning with inductive biases

    Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho. Discovering symbolic models from deep learning with inductive biases. InAdvances in Neural Information Processing Systems (NeurIPS), 2020. 10 Preprint

  12. [12]

    Sparse autoencoders find highly interpretable features in language models

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. InInternational Conference on Learning Representations (ICLR), 2024

  13. [13]

    Toy models of superposition

    Nelson Elhage, Tristan Hume, Catherine Olsson, et al. Toy models of superposition. In Transformer Circuits Thread, 2022

  14. [14]

    Scaling and evaluating sparse autoencoders.arXiv preprint arXiv:2406.04093, 2024

    Leo Gao, Tom Dupr ´e la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders.arXiv preprint arXiv:2406.04093, 2024

  15. [15]

    Causal abstractions of neural networks

    Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), 2021

  16. [16]

    Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman. Finding alignments between interpretable causal variables and distributed neural representations. InConference on Causal Learning and Reasoning (CLeaR), 2024

  17. [17]

    Wichmann

    Robert Geirhos, J¨orn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

  18. [18]

    Hamiltonian neural networks

    Samuel Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2019

  19. [19]

    Language models represent space and time

    Wes Gurnee and Max Tegmark. Language models represent space and time. InInternational Conference on Learning Representations (ICLR), 2024

  20. [20]

    Discovering conservation laws from trajectories via machine learning.arXiv preprint arXiv:2102.04008, 2021

    Seungwoong Ha and Hawoong Jeong. Discovering conservation laws from trajectories via machine learning.arXiv preprint arXiv:2102.04008, 2021

  21. [21]

    Poseidon: Efficient foundation models for PDEs

    Maximilian Herde, Bogdan Raoni ´c, Tobias Rohner, Roger K¨appeli, Roberto Molinaro, Em- manuel de B´ezenac, and Siddhartha Mishra. Poseidon: Efficient foundation models for PDEs. InAdvances in Neural Information Processing Systems (NeurIPS), 2024

  22. [22]

    Designing and interpreting probes with control tasks

    John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In Proceedings of EMNLP-IJCNLP, 2019

  23. [23]

    John Hewitt and Christopher D. Manning. A structural probe for finding syntax in word representations. InProceedings of NAACL-HLT, 2019

  24. [24]

    Discovering physical concepts with neural networks.Physical Review Letters, 124(1):010508, 2020

    Raban Iten, Tony Metger, Henrik Wilming, L´ıdia del Rio, and Renato Renner. Discovering physical concepts with neural networks.Physical Review Letters, 124(1):010508, 2020

  25. [25]

    Emergent representations of program semantics in language models trained on programs

    Charles Jin and Martin Rinard. Emergent representations of program semantics in language models trained on programs. InInternational Conference on Machine Learning (ICML), 2024

  26. [26]

    Rediscovering orbital mechanics with machine learning.Machine Learning: Science and Technology, 4(4): 045002, 2023

    Pablo Lemos, Niall Jeffrey, Miles Cranmer, Shirley Ho, and Peter Battaglia. Rediscovering orbital mechanics with machine learning.Machine Learning: Science and Technology, 4(4): 045002, 2023

  27. [27]

    Hopkins, David Bau, Fernanda Vi´egas, Hanspeter Pfister, and Martin Wattenberg

    Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Vi´egas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. InInternational Conference on Learning Representations (ICLR), 2023

  28. [28]

    Fourier neural operator for parametric partial differen- tial equations

    Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. Fourier neural operator for parametric partial differen- tial equations. InInternational Conference on Learning Representations (ICLR), 2021

  29. [29]

    Machine learning conservation laws from trajectories.Physical Review Letters, 126(18):180604, 2021

    Ziming Liu and Max Tegmark. Machine learning conservation laws from trajectories.Physical Review Letters, 126(18):180604, 2021

  30. [30]

    Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators

    Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, 2021. 11 Preprint

  31. [31]

    Locating and editing factual associations in GPT

    Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. InAdvances in Neural Information Processing Systems (NeurIPS), 2022

  32. [32]

    Corrado, and Jeff Dean

    Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. InAdvances in Neural Information Processing Systems (NeurIPS), 2013

  33. [33]

    Progress measures for grokking via mechanistic interpretability

    Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. InInternational Conference on Learning Representations (ICLR), 2023

  34. [34]

    Emergent linear representations in world models of self-supervised sequence models

    Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. InBlackboxNLP Workshop at EMNLP, 2023

  35. [35]

    Zoom in: An introduction to circuits.Distill, 2020

    Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits.Distill, 2020

  36. [36]

    The linear representation hypothesis and the geometry of large language models.arXiv preprint arXiv:2311.03658, 2023

    Kiho Park, Yo Joong Choe, and Victor Veitch. The linear representation hypothesis and the geometry of large language models.arXiv preprint arXiv:2311.03658, 2023

  37. [37]

    Karniadakis

    Maziar Raissi, Paris Perdikaris, and George E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.Journal of Computational Physics, 378:686–707, 2019

  38. [38]

    Toward transparency in AI: Survey on interpreting the inner structures of deep neural networks.IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2023

    Tilman R¨auker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. Toward transparency in AI: Survey on interpreting the inner structures of deep neural networks.IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2023

  39. [39]

    Battaglia

    Alvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying, Jure Leskovec, and Peter W. Battaglia. Learning to simulate complex physics with graph networks. InInternational Conference on Machine Learning (ICML), 2020

  40. [40]

    Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet

    Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, et al. Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet. Transformer Circuits Thread, 2024

  41. [41]

    BERT rediscovers the classical NLP pipeline

    Ian Tenney, Dipanjan Das, and Ellie Pavlick. BERT rediscovers the classical NLP pipeline. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), 2019

  42. [42]

    Chess as a testbed for language model state tracking

    Shubham Toshniwal, Sam Wiseman, Karen Livescu, and Kevin Gimpel. Chess as a testbed for language model state tracking. InProceedings of the AAAI Conference on Artificial Intelligence, 2022

  43. [43]

    AI Feynman: A physics-inspired method for symbolic regression.Science Advances, 6(16):eaay2631, 2020

    Silviu-Marian Udrescu and Max Tegmark. AI Feynman: A physics-inspired method for symbolic regression.Science Advances, 6(16):eaay2631, 2020

  44. [44]

    Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan

    Keyon Vafa, Justin Y . Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the world model implicit in a generative model. InAdvances in Neural Information Processing Systems (NeurIPS), 2024

  45. [45]

    Chang, Ashesh Rambachan, and Sendhil Mullainathan

    Keyon Vafa, Peter G. Chang, Ashesh Rambachan, and Sendhil Mullainathan. What has a foundation model found? Using inductive bias to probe for world models. InInternational Conference on Machine Learning (ICML), 2025

  46. [46]

    Investigating gender bias in language models using causal mediation analysis

    Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Investigating gender bias in language models using causal mediation analysis. InAdvances in Neural Information Processing Systems (NeurIPS), 2020

  47. [47]

    Interpretability in the wild: A circuit for indirect object identification in GPT-2 small

    Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: A circuit for indirect object identification in GPT-2 small. In International Conference on Learning Representations (ICLR), 2023

  48. [48]

    Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman. Interpretability at scale: Identifying causal mechanisms in Alpaca. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. 12 Preprint A INSTRUMENT DETAILS Decode fit.Unless noted, wℓ is fit by ridge regression on z-scored activations with 5-fold cross- valid...