Pith. sign in

REVIEW 3 major objections 5 minor 50 references

The paper claims materials-science mechanisms in a language model are best detected as controlled changes between internal states, and that these state changes order direct, neutral, and inverse constitutive laws nearly perfectly.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 10:53 UTC pith:ZFIGRGHA

load-bearing objection A careful, honest interpretability study whose headline 60-law result is likely confounded by the answer-word axis; the graph identifiability audit is the real contribution. the 3 major comments →

arxiv 2607.20058 v1 pith:ZFIGRGHA submitted 2026-07-22 cs.AI cond-mat.mes-hallcond-mat.mtrl-scics.CL

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

classification cs.AI cond-mat.mes-hallcond-mat.mtrl-scics.CL
keywords mechanistic interpretabilitylarge language modelsmaterials scienceconstitutive relationsJacobian lenscausal interventioncounterfactual benchmarkrepresentation learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that in an open-weight 42-layer language model, materials-science mechanisms exist in three separable forms, and that the most informative evidence is relational rather than absolute. It argues that when a prompt and its near-verbatim reverse differ only in the direction of numerical change, the model's hidden states move in a way that tracks whether the governing law is direct, inverse, or neutral. In a 60-law benchmark this matched state change orders the three law classes with direct-versus-inverse AUC 1.000, while lexical controls are near chance. The paper also argues that absolute-state geometry and similarity graphs, though statistically organized, are not identifiable as physics because a numerical-direction label is exactly aliased with the physical outcome label. It concludes that causal interventions can steer answer decisions in constrained contexts but do not transfer across answer vocabularies or physical regimes.

Core claim

The paper's central claim is that the physical abstraction of monotonic constitutive orientation—whether a stated increase in an input quantity raises or lowers the response under a supplied law—is carried by controlled transformations between hidden states, not by the absolute geometry of those states. Using a frozen direction fitted on 16 development laws and a neutral class calibrated on 10 independent laws, the author compares otherwise identical prompts that differ only in the sign of the numerical change. Across 60 laws, the resulting state contrasts order 20 inverse, 20 neutral, and 20 direct relations with direct-versus-inverse AUC 1.000, validation-neutral-versus-inverse AUC 0.935,

What carries the argument

The main positive result is built on a matched-reversal state contrast. At layer 34, a unit direction d = (mu+ - mu-)/||mu+ - mu-|| is fitted on development-law centroids; for each law and cell, the projected score of the numerical-decrease prompt is subtracted from the numerical-increase prompt, holding equation surface, material wording, endpoint values, and answer order fixed, with ten calibration-neutral laws defining the empirical zero and scale. This converts the relation 'response sign = law sign × change sign' into a measurable hidden-state displacement. The negative result's machinery is the algebraic identity y = sx: within one mechanism family, same/different physical outcome is e

Load-bearing premise

The main result rests on the assumption that a single internal direction fitted on 16 development laws, plus an empirical zero calibrated on 10 neutral laws, carries over to 60 new laws without being refit; if those development laws are not representative of the 13 domains, the measured ordering could be renormalized or even inverted.

What would settle it

Rerun the 60-law matched-reversal benchmark with a leave-one-family-out direction: fit the layer-34 centroid on 15 of the development-law families and test on the held-out family plus the 60-law cohort. If direct-versus-inverse AUC falls far below 0.935 or the Spearman rho drops substantially, the transferability claim fails. A second check: swap the ten calibration-neutral laws with the ten validation-neutral laws; if the neutral scale shifts enough that the 39/40 classification degrades, the claim of a stable empirical zero fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the 60-law result holds, controlled state differences are a stronger evidence type than absolute-state similarity for whether a model tracks a physical relation.
  • Interpretability studies should run an aliasing audit before claiming a similarity graph encodes physics; in this design, 'same physical outcome' was exactly 'same numerical direction' within a family.
  • Causal steering is real but format-limited: the same grain-size direction reverses answers correctly for refinement and coarsening in one answer vocabulary, yet fails when answer words or physical regime change.
  • A practical test for physics representation in LLMs can be built from matched counterfactual prompts with a neutral class as an empirical zero, without requiring the model to output correct inverse-law text, which lagged behind the hidden-state signal.
  • The result supports exploring representation-aware training rewards that penalize lexical shortcuts and reward counterfactual consistency, while keeping the reward lens separate from the evaluation lens.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • I would expect the matched-reversal contrast to also order laws in other scientific domains, since the y = sx structure is generic, but the specific layer and direction would likely need refitting per domain.
  • The single sign error—the classical nucleation-barrier relation versus undercooling—suggests a testable boundary: monotonic-sign internalization may fail for laws with genuinely non-monotonic behavior, and a future benchmark could deliberately include such laws.
  • The answer-scaffold audit implies that readability of a scientific word after answer options are shown can be misleading; reward or evaluation readouts should be taken before choices appear.
  • The contrast between near-perfect internal ordering and only 53.8% correct inverse-law output hints that failure to verbalize a law does not mean the model lacks it—an inference the paper supports but does not fully explore.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper investigates whether materials-science mechanism information in an open-weight language model (google/gemma-4-E4B-it) is (i) readable in hidden states, (ii) carried by transformations between states, and (iii) causally usable for engineering answers. It combines direct and Jacobian vocabulary readouts, target-free decoding, state-space geometry, graph analyses, a 60-law neutral-anchored matched-reversal benchmark, and causal interventions. The main positive claims are that controlled state transformations order direct, neutral, and inverse constitutive relations nearly perfectly (AUC 1.000 direct vs. inverse, Spearman ρ = 0.910, 39/40 directional laws), and that one frozen grain-size direction steers answers in a relation-appropriate way in a prospectively confirmed cohort, while transfer to other answer vocabularies and regimes fails. The paper also reports a series of negative or limiting results: the broad 24-triplet physical-equivalence endpoint failed, the 80–96% late-window primary failed, the graph's physical polarity is not identifiable, and transfer cohorts fail. The authors are unusually explicit about preregistration, discovery-versus-replication chronology, and limitations.

Significance. If the neutral-anchored claim is sustained, the paper is a valuable methodological contribution: it demonstrates a framework for separating readable concepts, relational transformations, and causal use in LLM representations, and it provides a clear identifiability proof (Eq. 1) showing that a graph endpoint can collapse exactly to a prompt variable. The paper is also a model of scientific candor: endpoints are frozen, discovery cohorts are separated from replication cohorts, exact nulls and exhaustive partitions are used, and several high-profile hypotheses are reported as failures. The causal steering results are carefully bounded and the transfer failures are reported rather than hidden. The main unresolved risk is whether the headline 60-law result measures constitutive orientation or, more shallowly, the model's movement toward the expected answer word. That issue is central enough that the claim, as currently stated, is not yet fully established.

major comments (3)
  1. [§5.12, Eq. (12)–(13), Fig. 10C] The centroid direction d is fitted on positive-versus-negative physical-outcome labels in a scaffold where the physical outcome sign directly determines the model's allowed answer word (higher/lower). A direction fitted on those labels can therefore be an answer-token direction, and the matched contrast Δr may simply measure whether the hidden state moves toward the correct output word. The output-logit control reaching AUC 1.000 in Fig. 10C is consistent with this alternative reading, and §2.11.3 shows that a different direction can be answer-vocabulary specific. Word/char TF-IDF controls do not address this because they model prompt surface, not the output token. To support the constitutive-orientation interpretation, please add a control that breaks the alignment: e.g., fit d on a development set containing a balanced mix of direct and inverse laws (so positive/negative labels are not
  2. [§2.8, Fig. 10C and §3.2] The paper reports that exact inverse-law answer accuracy is only 53.8% while the output-logit control reaches AUC 1.000. This discrepancy needs explicit reconciliation. If the layer-34 hidden direction is essentially the same axis as the final output logits, then the claim that the hidden state is 'more stable than its conversion into the requested discrete answer' is weakened rather than supported: the hidden-state result may be an earlier manifestation of answer planning, not an independent representation of constitutive orientation. Please either quantify the correlation between r(h) and the output-logit contrast per law, or show that the ordering survives after removing the component of d that is aligned with the decoder's higher/lower direction.
  3. [§2.8, §5.12] The 60-law benchmark uses an explicit two-stage scaffold that instructs the model to compose the sign of the supplied equation with the sign of the numerical change. This makes the result task-elicited, as the authors acknowledge. The concern is not that task-elicitation is worthless, but that the headline phrasing 'physical abstraction of monotonic constitutive orientation' overstates what the design can show without the answer-token control above. Please either soften the abstract and conclusion claims accordingly, or provide the requested controls. The current limitation statement in §3.2 ('does not show spontaneous use...') is helpful but does not address the answer-token confound specifically.
minor comments (5)
  1. [§2.6] Typo: 'We find it ito be strong' should be 'We find it to be strong'.
  2. [§2.8] Typo: 'V ocabulary readouts' should be 'Vocabulary readouts'.
  3. [§4] Typo: 'addiitonal insights' should be 'additional insights'.
  4. [Abstract / §5.20] The phrase 'blinded identification of 9 of 10 mechanism families' should clarify that the interpreter is one automated model (gpt-5.5) with five order-randomized passes, not a panel of independent experts. The limitation is stated later, but the abstract could be more precise.
  5. [§5.12] The description of the output-head control and the exact-answer accuracy would benefit from a clearer statement of where the 'clean final higher-minus-lower logit difference' is measured (at the final token position? before generation?) so that the reader can see why the 53.8% inverse-law exact accuracy can coexist with AUC 1.000 on the logit contrast.

Circularity Check

0 steps flagged

No significant circularity: the core 60-law result is an out-of-sample supervised readout with a separate neutral calibration, and the paper explicitly audits the y=sx alias rather than hiding it.

full rationale

The paper's central relational claim is not circular. The layer-34 direction d=(mu+ - mu-)/||...|| is fitted only on 16 development laws and then frozen; no 60-law state or label is used to refit it. The matched contrast subtracts two prompts that are identical except for the direction of the numerical change, while equation surface, material wording, endpoint values, and answer order are held fixed. Ten neutral laws define the empirical zero and scale, and ten different neutral laws are used for validation, so the neutral class does not enter the fit. The paper also explicitly proves the exact alias y=sx for the earlier graph benchmark (Eq. 1) and reports the graph as non-identifiable, rather than presenting it as evidence of physics. The potential concern that the fitted direction could track the expected higher/lower output token rather than an abstract constitutive orientation is a real interpretive confound and is transparently acknowledged via the output-logit control reaching 1.000; but it is not a case where the prediction is equivalent to the fit by construction. The experiment could have failed (in fact the earlier exact-behavior and rearranged-formula endpoints did fail), and the direction transfers to a new 60-law cohort without refitting. Self-citations are contextual and not load-bearing for the central derivation. No circular step meeting the required quote-and-reduction standard is present.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central claim depends on several fitted parameters (layer-34 direction, midpoint, steering layer), but most were frozen on development/disjoint data before the final cohorts, and the paper honestly discloses this. The main unstated assumption is the validity of the linear readout and the transferability of development-fitted directions.

free parameters (4)
  • layer-34 centroid direction d = fitted on 16 development laws
    The 60-law benchmark projects states onto d = (mu+ - mu-)/||...|| fitted on a development set. This is a fitted parameter, but it is fixed before the 60-law cohort, so the central claim is not a direct in-sample fit.
  • layer choice (layer 34) and direction midpoint m = chosen on development/disjoint cohorts
    The layer and midpoint were selected on earlier cohorts (16 laws and a disjoint 12-law confirmation), not on the 60-law cohort. Still, the number of degrees of freedom in protocol selection is not fully accounted for in the reported p-values.
  • grain steering layer 16 and direction vector = selected in a preliminary study on disjoint calibration conditions
    The layer and direction were selected on earlier data and then frozen, but the exact selection rule (best mean signed endpoint minus one SD) is a hyperparameter. The prospective confirmation is a genuine out-of-sample test, but the selection rule itself is a free choice.
  • perturbation doses +/-4%, +/-2% = chosen by protocol
    Dose is a hyperparameter; the paper shows robustness at half-dose, so the central steering result does not depend critically on the exact dose.
axioms (5)
  • domain assumption The fixed decoder (unembedding) is a valid linear readout instrument for intermediate states.
    The Jacobian lens and direct lens both use Gemma's final decoder to map hidden states to vocabulary scores. If intermediate states are not in the decoder's coordinate system, the readout may be meaningless. The paper relies on the lens literature for this.
  • domain assumption The Jacobian estimator (Eq. 4) correctly approximates the average downstream transport.
    The Jacobian lens is defined as an expectation over positions and records, computed with reverse-mode autodiff. This is a linear approximation that may not capture nonlinear downstream dynamics, but the paper uses it only to rank vocabulary.
  • domain assumption The 16 development laws are representative of the 60-law benchmark's 13 domains.
    The layer-34 direction is fitted on development laws; if the 60-law domains differ in distribution, the projection would not transfer. The paper does not test this directly.
  • domain assumption Matched prompt pairs with only the numerical direction reversed isolate the physical relation change.
    The matched-reversal design holds equation, material wording, endpoints, and answer order fixed. It assumes that the model does not respond to small positional or formatting changes caused by reordering numerical endpoints. The paper's TF-IDF controls support this for the central relation, but not exhaustively.
  • standard math The graph-identifiability claim y=sx is an exact reading of the prompt manifest.
    The paper proves the algebraic identity y = s*x for the benchmark, which means within-family physical outcome is aliased with numerical direction. This is a deterministic property of the prompt construction, not an empirical assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 4197 in / 4846 out tokens · 71356 ms · 2026-08-01T10:53:08.603247+00:00 · methodology

0 comments
read the original abstract

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

Figures

Figures reproduced from arXiv: 2607.20058 by Markus J. Buehler.

Figure 1
Figure 1. Figure 1: Overview of the study reported in this article. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Can either readout recover an engineering term that was deliberately omitted from the description? [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Target-free family vocabularies. Each row is one physical family and each cell is a word selected without the declared concept list. Support is the number of independent phrasings, out of five, in which the word appeared under the three-fit consensus rule; darker blue-green cells indicate stronger background-corrected support. Asterisks mark exact declared-word overlap added only after ranking. Read across… view at source ↗
Figure 4
Figure 4. Figure 4: Do the freely discovered words identify the materials mechanism without a supplied answer list? [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: What does the stricter target-free filter remove, and what materials vocabulary remains? [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Do differently worded descriptions of the same mechanism remain close in the model’s full state space? [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Does physical equivalence overcome deliberately stronger wording similarity? [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Where does option-free comparative graph structure appear, and where does it fail? [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: What can the option-free graph identify? [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: A physical abstraction appears in matched state transformations. [PITH_FULL_IMAGE:figures/full_fig_p018_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Can an automated interpreter recognize a materials mechanism from discovered words alone? [PITH_FULL_IMAGE:figures/full_fig_p019_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Two magnified examples of the controlled-rank measurement. [PITH_FULL_IMAGE:figures/full_fig_p021_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Broad steering is mechanism dependent. Each panel averages ten new physical conditions and plots how an intermediate-state intervention changes the log odds of the first scientific answer relative to the second: grooves versus clean for corrosion (A), hard versus soft for transformation (B), and higher versus lower for grain size (C). Zero is the unperturbed model; the horizontal coordinate is the added d… view at source ↗
Figure 14
Figure 14. Figure 14: Can one frozen grain-size mechanism direction change a scientific answer in the relation-appropriate direction? [PITH_FULL_IMAGE:figures/full_fig_p023_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Can a hidden state from the opposite grain-size relation causally transfer its answer? [PITH_FULL_IMAGE:figures/full_fig_p024_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: What does a natural-question hidden-state patch transfer across mechanisms? [PITH_FULL_IMAGE:figures/full_fig_p026_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 1 canonical work pages

  1. [1]

    O., Lehman, J

    Stanley, K. O., Lehman, J. & Soros, L. Open-endedness: The last grand challenge you’ve never heard of.O’Reilly Online(2017)

  2. [2]

    & Stanley, K

    Wang, R., Lehman, J., Clune, J. & Stanley, K. O. Paired open-ended trailblazer (POET): Endlessly gen- erating increasingly complex and diverse learning environments and their solutions.arXiv preprint(2019). ArXiv:1901.01753

  3. [3]

    M., Schwaller, P., Ortega-Guerrero, A

    Jablonka, K. M., Schwaller, P., Ortega-Guerrero, A. & Smit, B. Leveraging large language models for predictive chemistry.Nature Machine Intelligence6, 161–169 (2024)

  4. [4]

    & Buehler, M

    Ghafarollahi, A. & Buehler, M. J. SciAgents: Automating Scientific Discovery Through Bioinspired Multi- Agent Intelligent Graph Reasoning.Advanced Materials37, 2413523 (2025). URL https://advanced. onlinelibrary.wiley.com/doi/10.1002/adma.202413523

  5. [5]

    & Buehler, M

    Ghafarollahi, A. & Buehler, M. J. Sparks: Multi-Agent Artificial Intelligence Model Discovers Protein Design Principles (2025). URLhttp://arxiv.org/abs/2504.19017. ArXiv:2504.19017 [cs]

  6. [6]

    ArXiv:2408.06292

    Lu, C.et al.The AI scientist: Towards fully automated open-ended scientific discovery.arXiv preprint(2024). ArXiv:2408.06292

  7. [7]

    ArXiv:2504.08066

    Lu, C.et al.The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search.arXiv preprint(2025). ArXiv:2504.08066. 41 Reading and Steering Materials Science-Mechanism Representations M.J. Buehler

  8. [8]

    InAdvances in Neural Information Processing Systems (NeurIPS 2025)(2025)

    Agarwal, D.et al.AutoDiscovery: Open-ended scientific discovery via bayesian surprise. InAdvances in Neural Information Processing Systems (NeurIPS 2025)(2025). URL https://arxiv.org/abs/2507.00310. 2507.00310

  9. [9]

    Formal theory of creativity, fun, and intrinsic motivation.IEEE Transactions on Autonomous Mental Development2, 230–247 (2010)

    Schmidhuber, J. Formal theory of creativity, fun, and intrinsic motivation.IEEE Transactions on Autonomous Mental Development2, 230–247 (2010)

  10. [10]

    Griffith, A. A. The phenomena of rupture and flow in solids.Philosophical Transactions of the Royal Society A 221, 163–198 (1921)

  11. [11]

    Hall, E. O. The deformation and ageing of mild steel: III discussion of results.Proceedings of the Physical Society. Section B64, 747–753 (1951)

  12. [12]

    Petch, N. J. The cleavage strength of polycrystals.Journal of the Iron and Steel Institute174, 25–28 (1953)

  13. [13]

    Diffusional viscosity of a polycrystalline solid.Journal of Applied Physics21, 437–445 (1950)

    Herring, C. Diffusional viscosity of a polycrystalline solid.Journal of Applied Physics21, 437–445 (1950)

  14. [14]

    Callister, W. D. & Rethwisch, D. G.Materials Science and Engineering: An Introduction(Wiley, 2018), 10 edn

  15. [15]

    URL https://www.nature.com/articles/s41563-022-01384-1

    Nepal, D.et al.Hierarchically structured bioinspired nanocomposites.Nature Materials22, 18–35 (2023). URL https://www.nature.com/articles/s41563-022-01384-1

  16. [16]

    Wegst, U. G. K., Bai, H., Saiz, E., Tomsia, A. P. & Ritchie, R. O. Bioinspired structural materials.Nature Materials14, 23–36 (2015). URLhttps://www.nature.com/articles/nmat4089

  17. [17]

    Jain, A.et al.Commentary: The materials project: A materials genome approach to accelerating materials innovation.APL Materials1, 011002 (2013)

  18. [18]

    Merchant, A.et al.Scaling deep learning for materials discovery.Nature624, 80–85 (2023)

  19. [19]

    Buehler, M. J. Accelerating scientific discovery with generative knowledge extraction, graph-based representation, and multimodal intelligent graph reasoning.Machine Learning: Science and Technology5, 035083 (2024). URL https://doi.org/10.1088/2632-2153/ad7228

  20. [20]

    & Buehler, M

    Pal, S., Sourav, S., Ghosal, T. & Buehler, M. J. Graph-native reinforcement learning enables traceable scientific hypothesis generation through conceptual recombination (2026). URL https://arxiv.org/abs/2607.00924. 2607.00924

  21. [21]

    Buehler, M. J. PRefLexOR: Preference-Based Recursive Language Modeling for Exploratory Optimization of Reasoning and Agentic Thinking.npj Artificial Intelligence1(2025). URL https://doi.org/10.1038/ s44387-025-00003-z

  22. [22]

    Buehler, M. J. Generative retrieval-augmented ontologic graph and multiagent strategies for interpretive large language model-based materials design.ACS Engineering Au4, 133–152 (2024). URLhttps://doi.org/10. 1021/acsengchemau.3c00053.https://doi.org/10.1021/acsengchemau.3c00053

  23. [23]

    Nature571, 95–98 (2019)

    Tshitoyan, V .et al.Unsupervised word embeddings capture latent knowledge from materials science literature. Nature571, 95–98 (2019)

  24. [24]

    Gupta, T., Zaki, M., Krishnan, N. M. A. & Mausam. MatSciBERT: A materials domain language model for text mining and information extraction.npj Computational Materials8, 102 (2022)

  25. [25]

    Luu, R. K. & Buehler, M. J. BioinspiredLLM: Conversational Large Language Model for the Mechanics of Biological and Bio-Inspired Materials.Advanced Science11, 2306724 (2024). URL https://advanced. onlinelibrary.wiley.com/doi/10.1002/advs.202306724

  26. [26]

    Hage, T. P. & Buehler, M. J. BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Com- pact LLMs for Structured Beam Mechanics Reasoning (2026). URL http://arxiv.org/abs/2603.04124. ArXiv:2603.04124 [cs.AI]

  27. [27]

    & Buehler, M

    Ghafarollahi, A. & Buehler, M. J. ProtAgents: Protein discovery via large language model multi-agent col- laborations combining physics and machine learning (2024). URL http://arxiv.org/abs/2402.04268. ArXiv:2402.04268 [cond-mat]

  28. [28]

    Buehler, M. J. MechGPT, a Language-Based Strategy for Mechanics and Materials Modeling That Connects Knowledge Across Scales, Disciplines, and Modalities.Applied Mechanics Reviews76(2024). URL https: //doi.org/10.1115/1.4063843

  29. [29]

    Buehler, M. J. MeLM, a generative pretrained language modeling framework that solves forward and inverse mechanics problems.Journal of the Mechanics and Physics of Solids181, 105454 (2023). URL https: //www.sciencedirect.com/science/article/pii/S0022509623002582

  30. [30]

    Y .et al.Autonomous agents coordinating distributed discovery through emergent artifact exchange (2026)

    Wang, F. Y .et al.Autonomous agents coordinating distributed discovery through emergent artifact exchange (2026). URLhttps://arxiv.org/abs/2603.14312.2603.14312. 42 Reading and Steering Materials Science-Mechanism Representations M.J. Buehler

  31. [31]

    N., Arnold, C., Rand, B

    Rubungo, A. N., Arnold, C., Rand, B. P. & Dieng, A. B. LLM-Prop: Predicting physical and electronic properties of crystalline solids from their text descriptions.arXiv preprint arXiv:2310.14029(2023). URL https://arxiv.org/abs/2310.14029

  32. [32]

    & Nagato, K

    Yoshitake, M., Suzuki, Y ., Igarashi, R., Ushiku, Y . & Nagato, K. MaterialBENCH: Evaluating college-level materials science problem-solving abilities of large language models.arXiv preprint arXiv:2409.03161(2024). URLhttps://arxiv.org/abs/2409.03161

  33. [33]

    URL https://proceedings.neurips.cc/paper_files/paper/2017/hash/ 3f5ee243547dee91fbd053c1c4a845aa-Abstract.html

    Vaswani, A.et al.Attention is all you need.Advances in Neural Information Process- ing Systems30(2017). URL https://proceedings.neurips.cc/paper_files/paper/2017/hash/ 3f5ee243547dee91fbd053c1c4a845aa-Abstract.html

  34. [34]

    URLhttps://transformer-circuits.pub/2021/framework/index.html

    Elhage, N., Nanda, N., Olsson, C.et al.A mathematical framework for transformer circuits.Transformer Circuits Thread(2021). URLhttps://transformer-circuits.pub/2021/framework/index.html

  35. [35]

    & Levy, O

    Geva, M., Schuster, R., Berant, J. & Levy, O. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 5484–5495 (2021)

  36. [36]

    URL https://transformer-circuits.pub/2023/ monosemantic-features/index.html

    Bricken, T., Templeton, A., Batson, J.et al.Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread(2023). URL https://transformer-circuits.pub/2023/ monosemantic-features/index.html

  37. [37]

    URL https://transformer-circuits.pub/2024/ scaling-monosemanticity/index.html

    Templeton, A., Conerly, T., Marcus, J.et al.Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet.Transformer Circuits Thread(2024). URL https://transformer-circuits.pub/2024/ scaling-monosemanticity/index.html

  38. [38]

    URLhttps://transformer-circuits.pub/2025/attribution-graphs/biology.html

    Lindsey, J., Gurnee, W., Ameisen, E.et al.On the biology of a large language model.Transformer Circuits Thread (2025). URLhttps://transformer-circuits.pub/2025/attribution-graphs/biology.html

  39. [39]

    & Steinhardt, J

    Burns, C., Ye, H., Klein, D. & Steinhardt, J. Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827(2022). URLhttps://arxiv.org/abs/2212.03827

  40. [40]

    URLhttps://arxiv.org/abs/2310.14491

    Hou, Y .et al.Towards a mechanistic interpretation of multi-step reasoning capabilities of language models.arXiv preprint arXiv:2310.14491(2023). URLhttps://arxiv.org/abs/2310.14491

  41. [41]

    & Liang, P

    Hewitt, J. & Liang, P. Designing and interpreting probes with control tasks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 2733–2743 (2019)

  42. [42]

    URLhttps://arxiv.org/abs/2303.08112

    Belrose, N.et al.Eliciting latent predictions from transformers with the tuned lens.arXiv preprint arXiv:2303.08112(2023). URLhttps://arxiv.org/abs/2303.08112

  43. [43]

    URL https://arxiv.org/abs/2607.15495.2607.15495

    Gurnee, W.et al.Verbalizable representations form a global workspace in language models (2026). URL https://arxiv.org/abs/2607.15495.2607.15495

  44. [44]

    & Melville, J

    McInnes, L., Healy, J. & Melville, J. UMAP: Uniform manifold approximation and projection for dimension reduction.arXiv preprint arXiv:1802.03426(2018). URLhttps://arxiv.org/abs/1802.03426

  45. [45]

    F.et al.Deep reinforcement learning from human preferences

    Christiano, P. F.et al.Deep reinforcement learning from human preferences. InAdvances in Neural Infor- mation Processing Systems, vol. 30 (2017). URLhttps://proceedings.neurips.cc/paper/2017/hash/ d5e2c0adad503c91f91df240d0cd4e49-Abstract.html

  46. [46]

    In Advances in Neural Information Processing Systems, vol

    Ouyang, L., Wu, J., Jiang, X.et al.Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, vol. 35 (2022). URL https://arxiv.org/abs/2203. 02155

  47. [47]

    URL https://arxiv

    Lightman, H.et al.Let’s verify step by step.arXiv preprint arXiv:2305.20050(2023). URL https://arxiv. org/abs/2305.20050

  48. [48]

    URLhttps://arxiv.org/abs/2310.01405

    Zou, A., Phan, L., Chen, S.et al.Representation engineering: A top-down approach to AI transparency.arXiv preprint arXiv:2310.01405(2023). URLhttps://arxiv.org/abs/2310.01405

  49. [49]

    M.et al.Steering language models with activation engineering.arXiv preprint arXiv:2308.10248 (2023)

    Turner, A. M.et al.Steering language models with activation engineering.arXiv preprint arXiv:2308.10248 (2023). URLhttps://arxiv.org/abs/2308.10248

  50. [50]

    Gemma 4 e4b it model card

    Google. Gemma 4 e4b it model card. Hugging Face model repository (2026). URL https://huggingface. co/google/gemma-4-E4B-it. 43