Pith. sign in

REVIEW 4 major objections 6 minor 62 references

TRACE claims that one fixed-environment architecture—with no learned state passed between atoms—can reproduce crystalline phase behavior, liquid water structure, and chemical reactivity, including a CsPbI3 phase crossing near 580 K and a me

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 01:50 UTC pith:PIXUOTMW

load-bearing objection Genuinely new local MLIP architecture—fixed-environment cross-attention with no sender-state passing—worth refereeing, but the CsPbI3 phase-crossing claim needs a cell-size fix and a cutoff test. the 4 major comments →

arxiv 2607.25652 v1 pith:PIXUOTMW submitted 2026-07-28 cond-mat.mtrl-sci

Transformer Atomic Cluster Expansion: TRACE

classification cond-mat.mtrl-sci
keywords machine-learning interatomic potentialsatomic cluster expansionequivariant attentionCsPbI3 polymorphsliquid water structuremethyl migrationmolecular dynamicsfree-energy sampling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

TRACE is a machine-learning interatomic potential built on a deliberately restrictive design: each atom's energy is computed from its local environment only, with the center state built from atomic-cluster-expansion (ACE) density correlations and then used as a query to attend to fixed, equivariant edge features. Crucially, no learned state is passed between atoms—no message passing. The paper's central claim is that this fixed-environment construction is enough to describe dramatically different physics: the relative energies and finite-temperature phase stability of CsPbI3 polymorphs (a δ→α Gibbs crossing near 580 K, close to the experimental ~600 K), the liquid structure of water (first O–O peak at 2.85 Å vs 2.80 Å experimental), and the activation free energy of a methyl-migration reaction (27.92±0.03 kcal/mol vs 29.2±1.1 kcal/mol experimental). All three systems are handled by the same unified architecture, trained separately. If true, this means that learned propagation between atoms is not necessary for a broadly useful interatomic potential.

Core claim

On its own terms, the paper establishes that a single architecture—one local cross-attention block in which an ACE-correlated center state queries non-updated, geometry-fixed tensorial edge features—can reproduce experimentally meaningful observables across three regimes. For CsPbI3, independently relaxed relative energies of four polymorphs track r2SCAN+rVV10 DFT to within about 1.4 kJ/mol, and free-energy integration from an Einstein-crystal reference gives a δ/α Gibbs crossing at ~580 K with a 95% bootstrap interval of 553–599 K. The same frozen potential, when biased with a collective structure-factor coordinate, drives a 640-atom δ-to-perovskite transformation in which corner-sharing Pb

What carries the argument

The central object is the fixed-environment cross-attention block: a center state h_i, built by summing ACE density correlations over the cutoff-wide neighbor list and learning Clebsch–Gordan products, supplies the query, while the keys and equivariant values come from fixed edge tensors that depend only on input species and geometry. A cutoff-preserving softmax with a null channel gates the attention weights, and a scalar feed-forward updates only scalar channels using norms of higher-l tensors; forces and stress are exact derivatives of one invariant energy. This mechanism lets the model refine the nonlinear response within one environment—without ever reading a neighbor's updated hidden s

Load-bearing premise

The design rests on the assumption that a 6 Å cutoff with no Coulomb, Ewald, charge-equilibration, or dispersion term adequately captures the energetics of ionic CsPbI3; if truncation shifts relative phase free energies by more than about 2 kJ/mol per formula unit, the ~580 K crossing and the δ→α transformation results would be undermined.

What would settle it

Compute the relative Gibbs free energy of δ- and α-CsPbI3 with a long-range-corrected version of the same potential—adding an Ewald/Coulomb term or extending the cutoff well beyond 6 Å—and check whether the crossing temperature moves outside the 553–599 K bootstrap interval and whether the 0.7–4.7 meV/atom forward-reverse hysteresis grows beyond the per-formula-unit error budget.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The same TRACE potential can reproduce polymorph ordering, finite-temperature phase stability, and a collective structural transformation in CsPbI3 without any active learning or retraining.
  • A potential with no message passing can still drive rare events: biased sampling with a fixed checkpoint crossed the δ-to-perovskite network rewiring, meaning learned state propagation is not required for collective phase changes.
  • Reduced data sets (417 water configurations, 274 reactive structures) are enough for the architecture to match experimental structure and kinetics to within a few percent.
  • Because each atomic energy's spatial support is exactly the cutoff, domain decomposition requires only one ghost halo, and scaling is linear in the number of atoms for bounded neighbor counts.
  • Conservative forces and stress from a single invariant energy make the model directly usable for variable-cell molecular dynamics without separate force/stress heads.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's strongest implicit claim is that long-range electrostatics can be absorbed by a short-range learned potential for ionic systems; a direct test is to add an explicit Coulomb term or enlarge the cutoff and see whether the CsPbI3 phase crossing shifts by more than the reported ~2 kJ/mol error budget.
  • If the fixed-environment result generalizes, then much of the complexity in equivariant message-passing architectures may be unnecessary for local chemical accuracy; one could test this by comparing learning curves on a common dataset against models that do exchange hidden states.
  • The water model's short trajectory and classical treatment leave density and pressure unconverged, so the structural agreement is a necessary but not sufficient check; a pressure-consistent equation of state would be a harder target for the no-message-passing design.
  • The architecture's explicit separation of fixed edge memory and updated center state suggests it could be extended to longer-range physics by adding a small number of nonlocal tokens without falling back to full message passing—for example, global charge or dielectric channels.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript introduces TRACE, a local O(3)-equivariant machine-learning interatomic potential in which an ACE-correlated center state acts as the query in multihead cross-attention over fixed, geometry-and-species-dependent edge tensors; no learned state is propagated between atoms. The architecture is tested on three systems: CsPbI3 polymorph energetics and the δ–α Gibbs free-energy crossing, liquid water structure from a reduced CCSD(T)-derived data set, and methyl-migration activation free energy. Headline results are: relaxed δ/γ/β/α relative energies of 0/12.55/16.17/26.05 kJ/mol per formula unit versus r2SCAN+rVV10 values of 0/12.1/17.1/27.4; a δ–α Gibbs crossing at ~580 K with a nominal 95% bootstrap interval of 553–599 K, compared with experimental δ/cubic coexistence near 563–602 K; a first O–O peak at 2.85 Å versus 2.80 Å in experiment; and ΔG‡ = 27.92 ± 0.03 kcal/mol from seven umbrella-sampling replicas, versus an experimental value of 29.2 ± 1.1 kcal/mol.

Significance. The fixed-environment factorization is a genuinely distinctive architectural proposal: it separates attention depth from spatial propagation, so that additional attention blocks refine the local environment without enlarging the cutoff halo. The paper ships implementation, training configurations, and tests, and the benchmarks are largely well executed: blocked data splits, independent relaxations for all polymorphs, seven umbrella-sampling replicas with reported window overlap, and explicit checks of geometric stability. If the CsPbI3 finite-temperature phase-stability result is confirmed, this would be a strong demonstration that an attention-based local description without message passing can capture phase behavior. However, the phase-diagram claim currently rests on an explicitly preliminary classical estimate with unresolved cell-size and truncation issues, so the central claim is not yet established at the level of the abstract.

major comments (4)
  1. [§III.A and Fig. 3 caption] There is a direct inconsistency in the reported cell size for the Frenkel–Ladd calculation. The main text states that 'absolute free energies of 480-atom edge-sharing δ and cubic α cells' were computed, while the Fig. 3 caption says 'The calculation uses 240-atom cells', and a later sentence refers to 'the error from the single 480-atom cell.' This is not a cosmetic issue: finite-size errors in the Frenkel–Ladd path and Gibbs–Helmholtz integration can shift the δ–α free-energy difference by an amount comparable to the ~2 kJ/mol margin that separates the calculated crossing from the experimental window. The published number is therefore not reproducible as written. Please state the exact cell size, rerun or re-analyze the free-energy calculation for at least two cell sizes, and quantify the finite-size uncertainty in the crossing temperature.
  2. [Table I, §II.C, and §IV] The CsPbI3 potential uses r_c = 6.0 Å with no Coulomb, Ewald, charge-equilibration, dispersion-tail, or reciprocal-space term, and Section IV explicitly concedes that for ionic CsPbI3 this 'must be tested against cutoff and cell size'. This assumption is load-bearing for the phase-diagram claim: the δ–α energy difference is only ~26 kJ/mol per formula unit, and the finite-temperature free-energy difference changes by about 6.8 kJ/mol per formula unit between 400 and 650 K. A truncation error of even 2 kJ/mol per formula unit would shift the crossing by tens of kelvin, well outside the nominal 553–599 K bootstrap interval. Please provide a cutoff-convergence study (e.g., r_c = 5, 6, 7, 8 Å for the δ and α free energies) or add a controlled long-range correction and report the resulting crossing temperature.
  3. [§III.A text following Eq. (54)] The crossing is explicitly labeled a 'preliminary classical estimate', yet the paper reports forward–reverse hysteresis of 0.7–4.7 meV/atom. Converting per atom to per CsPbI3 formula unit (5 atoms), this is ~0.34–2.3 kJ/mol per formula unit, which is the same order of magnitude as the ΔG margin used to locate the crossing. The bootstrap interval of 553–599 K captures statistical trajectory error but not this hysteresis or the systematic finite-size and truncation errors discussed above. Please quantify the hysteresis contribution to the crossing uncertainty, for example by reporting the spread of the symmetric-work estimates over the five pairs of forward/reverse paths and, if feasible, reducing the error with slower switching or additional replicas.
  4. [Abstract vs. §III.A] The abstract states that TRACE 'captures ... phase diagrams' and gives a crossing 'near the experimental observations of ≃600K', while the body text calls the number a 'preliminary classical estimate' with unquantified systematic errors. Given the issues above, the abstract overstates the weight of evidence for the phase-diagram claim. Please qualify the abstract or move the phase-diagram language to a claim that is commensurate with the currently demonstrated accuracy.
minor comments (6)
  1. [Main text and Fig. 3] The 480-atom versus 240-atom inconsistency also appears between the text and the figure caption; this needs a single unambiguous statement in the revised manuscript.
  2. [§I, Fig. 1 caption] The caption uses 'e' for edge tensors and the text uses 'a_ij'; please unify notation so that the fixed edge memory is consistently denoted.
  3. [§III.B and Fig. 5] The O–H and H–H comparisons are read from a published figure rather than tabulated data, and only the O–O comparison is quantified. Please provide numerical pointwise RMSE values for all three partials or state explicitly that the O–H/H–H comparisons are qualitative.
  4. [§II.J and Table I] The Muon optimizer reference [27] is a methods note without a DOI or journal anchor; if a permanent version exists, please cite it.
  5. [§IV] The 'Scope and Limitations' section is a useful addition. Consider moving the two most relevant limitations—the no-electrostatics/cutoff issue and the preliminary nature of the crossing—into the main text near the CsPbI3 results, rather than only in a separate section, so that casual readers do not miss them.
  6. [Eq. (20)] The notation of the scaled logit (the second occurrence of 'es') is difficult to read in the typeset text; please use a distinct symbol such as \tilde{s}_{ijp} and define it cleanly.

Circularity Check

0 steps flagged

No significant circularity: the headline predictions are forward simulations against external references; self-citations are data/tool sources, not load-bearing derivations.

full rationale

The paper's derivation chain is: define a local energy E = sum_i E_i + E_ref (Eqs. 2 and 31), fit E_i to external energies/forces/stresses via the losses (37)-(40), then evaluate physical observables by relaxation, thermodynamic integration, MD, and umbrella sampling. None of the three headline targets — the CsPbI3 delta-alpha Gibbs crossing (~580 K), the water O-O peak (2.85 Å), or the methyl-migration barrier (27.92 kcal/mol) — appears in any loss function, and none is defined in terms of the fitted parameters by construction. The relaxed polymorph energies compare independently relaxed TRACE and DFT structures, not single points at training geometries. Self-citations ([31] for r2SCAN+rVV10 training data and the S_p reaction coordinate; [5] for a smoothness statement) are not used to justify the central architecture claim: the training labels are external electronic-structure calculations, and the reaction coordinate is explicitly parameterized in Eqs. (55)-(57). No uniqueness theorem or ansatz is imported from prior work. The Section IV limitations — finite cutoff, no electrostatics, and the 'preliminary classical estimate' label for ~580 K — are physics/reproducibility risks (including the 480- vs 240-atom cell discrepancy), but they do not make any prediction equivalent to its input. The predictions could be wrong without disturbing the architecture definition; hence no circular step is present.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The paper introduces no new physical entities; it fits a parametric energy to existing DFT/CCSD(T)/PBE0 labels. The ledger therefore lists hand-chosen modeling parameters (cutoffs, loss weights, order-parameter coefficients, sampling design) and domain assumptions (label fidelity, classical nuclei, locality without electrostatics, natural-parity truncation). The most consequential un-validated assumption is the electrostatics-free finite cutoff for ionic CsPbI3, which the paper itself flags (Section IV).

free parameters (6)
  • Architecture hyperparameters (r_c = 6.0 Å, ℓ_max = 2, node irreps 64×0e+32×1o+16×2e, 12 Bessel functions, correlation de = see Table I
    Hand-chosen capacity settings; no ablation or cutoff-convergence study is reported (Table I; Section IV).
  • Loss weights and stress ramp (w_E, w_F, w_σ), local-linearization weight = (1, 10, 10³→10⁵), 20-epoch ramp; w_Sob = 10⁻³
    Training hyperparameters; the reported checkpoint is picked by lowest validation loss (epoch 76/77), so stated errors depend on these choices (Section II.I, Table I).
  • S_p order-parameter coefficients (q_α, s*_α) = q = (21.99, 23.78, 14.36) nm⁻¹; s* = (1.5, 1.3, 1.5)
    Hand-tuned to the perovskite length scale and response (Eq. 57); affects the biased-transition narrative in Fig. 4, not the free-energy result.
  • Pb–I connectivity cutoff for corner/edge/higher sharing counts = 4.2 Å
    Chosen threshold for classifying PbI6 connectivity in Fig. 4(b); not derived from data.
  • Umbrella bias κ and 39 window centers = κ = 20 eV Å⁻²; 9 midpoint windows added to 30 published centers
    Sampling design choices; reported window overlap 0.399–0.413 supports convergence (Section III.C).
  • Frenkel–Ladd spring constants κ_{Z,s} = from per-phase MSD
    Standard protocol rather than an ad hoc physical parameter, but the reference Einstein crystal depends on TRACE's own mean-square displacements (Eq. 52).
axioms (6)
  • domain assumption Reference electronic-structure labels are accurate enough (r2SCAN+rVV10 for CsPbI3; CCSD(T)/MB-pol for water; PBE0-D3BJ/def2-SVP for methyl migration)
    The MLIP is supervised on these labels; systematic DFT error (e.g., in the δ/α free-energy balance) is inherited by the 'agreement with experiment' claims (Sections III.A–III.C).
  • domain assumption Classical treatment of nuclei is adequate for the reported thermodynamic comparisons
    Water NPT and CsPbI3 free-energy calculations use classical dynamics; Section III.B concedes nuclear quantum effects would soften the water structure.
  • domain assumption A 6.0 Å local cutoff with no explicit electrostatics describes CsPbI3 energetics
    Section II.C/Table I; the paper flags in Section IV that this is untested for ionic CsPbI3 against cutoff/cell size, long-wavelength distortions, and charged defects. Load-bearing for the δ→α claims.
  • domain assumption Natural-parity O(3) irreps with ℓ_max = 2 and a finite learned channel space suffice
    Section II.A calls this 'a compact truncation of the O(3) tensor space, not a complete set'; Section IV says the effect should be tested by varying parity content.
  • domain assumption The fixed Bessel radial basis (Eq. 9, ω_n = nπ) is sufficiently expressive
    Section II.C: 'All calculations reported here use fixed-frequency Bessel functions'; trainable frequencies and a Gaussian basis are implemented but not evaluated.
  • ad hoc to paper The S_p order parameter (Eqs. 55–57) separates δ and perovskite basins under bias
    Reaction-coordinate parameters are tuned for this system (Eq. 57); the Fig. 4 transformation narrative depends on this coordinate driving the intended collective change without being trained on intermediate structures.

pith-pipeline@v1.3.0-alltime-deepseek · 22785 in / 23011 out tokens · 225593 ms · 2026-08-01T01:50:09.839791+00:00 · methodology

0 comments
read the original abstract

Designing machine-learning interatomic potentials involves achieving the precise representation of complex many-body interactions alongside the efficiency required for scalable molecular dynamics. We introduce Transformer Atomic Cluster Expansion (TRACE), an energy-conserving architecture that combines atomic cluster expansion density correlations with local multihead cross-attention. The correlations form an O(3)-equivariant state for each center, which queries tensorial neighbor features that remain fixed functions of species and geometry. No learned state is passed between atoms. On a laptop MacBook-M1, we train and test TRACE for polymorphic cesium lead iodide, liquid water, and intramolecular methyl migration against experiments. For cesium lead iodide, TRACE reproduces the r$^2$SCAN+rVV10 ordering of four polymorphs and gives a classical edge-sharing hexagonal non-perovskite($\delta$) to corner-sharing cubic perovskite($\alpha$) Gibbs-free-energy crossing $\simeq$580K near the experimental observations of $\simeq$600K. By employing enhanced sampling to cross high energy barriers, the same TRACE potential successfully captures the $\delta$-to-$\alpha$ perovskite transformation without any reinforcement learning. A water potential trained on a reduced set of CCSD(T) configurations places the first oxygen--oxygen maximum at 2.85~\AA, compared to the experimental value of 2.80~\AA{}. For the gas-phase methyl migration in 2,2-dimethylisoindene, umbrella sampling yields an activation free energy of $27.92\pm0.03$~kcal~mol$^{-1}$, in close agreement with the experimental measurement of $29.2\pm1.1$~kcal~mol$^{-1}$. Across these diverse benchmarks, a single unified architecture successfully captures multi-species crystallization, liquid structures, phase diagrams, and chemical reactivity.

Figures

Figures reproduced from arXiv: 2607.25652 by Paramvir Ahlawat.

Figure 1
Figure 1. Figure 1: FIG. 1. TRACE architecture: (a) For receiver [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Relative energies of independently relaxed CsPbI [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Finite-temperature stability of edge-sharing [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Collective phase transformation during the enhanced-sampling transition in CsPbI [PITH_FULL_IMAGE:figures/full_fig_p021_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Partial radial distribution functions from a 144-molecule classical TRACE trajectory at [PITH_FULL_IMAGE:figures/full_fig_p024_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Description of methyl migration from 2,2-dimethylisoindene to 1,2-dimethylindene. (a) [PITH_FULL_IMAGE:figures/full_fig_p026_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

62 extracted references · 12 canonical work pages

  1. [1]

    Generalized Neural-Network Representation of High-Dimensional Potential-Energy Surfaces,

    J. Behler and M. Parrinello, “Generalized Neural-Network Representation of High-Dimensional Potential-Energy Surfaces,”Physical Review Letters98, 146401 (2007). doi:10.1103/PhysRevLett.98.146401

  2. [2]

    Gaussian Approximation Potentials: The Accuracy of Quantum Mechanics, without the Electrons,

    A. P. Bart´ ok, M. C. Payne, R. Kondor, and G. Cs´ anyi, “Gaussian Approximation Potentials: The Accuracy of Quantum Mechanics, without the Electrons,”Physical Review Letters104, 136403 (2010). doi:10.1103/PhysRevLett.104.136403

  3. [3]

    Spectral neighbor analysis method for automated generation of quantum-accurate interatomic potentials,

    A. P. Thompson, L. P. Swiler, C. R. Trott, S. M. Foiles, and G. J. Tucker, “Spectral neighbor analysis method for automated generation of quantum-accurate interatomic potentials,” Journal of Computational Physics285, 316–330 (2015). doi:10.1016/j.jcp.2014.12.018. 29

  4. [4]

    Moment Tensor Potentials: A Class of Systematically Improvable Interatomic Potentials,

    A. V. Shapeev, “Moment Tensor Potentials: A Class of Systematically Improvable Interatomic Potentials,”Multiscale Modeling & Simulation14, 1153–1173 (2016). doi:10.1137/15M1054183

  5. [5]

    Thermal Conductivity Predictions with Foundation Atomistic Models,

    B. P´ ota, P. Ahlawat, G. Cs´ anyi, and M. Simoncelli, “Thermal Conductivity Predictions with Foundation Atomistic Models,” arXiv:2408.00755 (2024). arXiv:2408.00755

  6. [6]

    Atomic cluster expansion for accurate and transferable interatomic potentials,

    R. Drautz, “Atomic cluster expansion for accurate and transferable interatomic potentials,” Physical Review B99, 014104 (2019). doi:10.1103/PhysRevB.99.014104

  7. [7]

    e3nn: Euclidean Neural Networks,

    M. Geiger and T. Smidt, “e3nn: Euclidean Neural Networks,” arXiv:2207.09453 (2022). arXiv:2207.09453

  8. [8]

    E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials,

    S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, “E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials,”Nature Communications13, 2453 (2022). doi:10.1038/s41467-022-29939-5

  9. [9]

    A universal graph deep learning interatomic potential for the periodic table,

    C. Chen and S. P. Ong, “A universal graph deep learning interatomic potential for the periodic table,”Nature Computational Science2, 718–728 (2022). doi:10.1038/s43588-022-00349-3

  10. [10]

    A framework to evaluate machine learning crystal stability predictions,

    J. Riebesell, R. E. A. Goodall, P. Benner, Y. Chiang, B. Deng, G. Ceder, M. Asta, A. A. Lee, A. Jain, and K. A. Persson, “A framework to evaluate machine learning crystal stability predictions,”Nature Machine Intelligence7, 836–847 (2025). doi:10.1038/s42256-025-01055-1

  11. [11]

    MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields,

    I. Batatia, D. P. Kovacs, G. Simm, C. Ortner, and G. Csanyi, “MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields,” in Advances in Neural Information Processing Systems35(2022). NeurIPS proceedings

  12. [12]

    Graph Atomic Cluster Expansion for Semilocal Interactions beyond Equivariant Message Passing,

    A. Bochkarev, Y. Lysogorskiy, and R. Drautz, “Graph Atomic Cluster Expansion for Semilocal Interactions beyond Equivariant Message Passing,”Physical Review X14, 021036 (2024). doi:10.1103/PhysRevX.14.021036

  13. [13]

    Graph atomic cluster expansion for foundational machine learning interatomic potentials,

    Y. Lysogorskiy, A. Bochkarev, and R. Drautz, “Graph atomic cluster expansion for foundational machine learning interatomic potentials,” arXiv:2508.17936 (2025). arXiv:2508.17936

  14. [14]

    Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction,

    X. Fu, B. M. Wood, L. Barroso-Luque, D. S. Levine, M. Gao, M. Dzamba, and C. L. Zitnick, “Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction,” 30 arXiv:2502.12147 (2025). arXiv:2502.12147

  15. [15]

    Neural Machine Translation by Jointly Learning to Align and Translate,

    D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” inInternational Conference on Learning Representations(2015). arXiv:1409.0473

  16. [16]

    Attention Is All You Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” inAdvances in Neural Information Processing Systems30(2017). NeurIPS proceedings

  17. [17]

    Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks,

    J. Lee, Y. Lee, J. Kim, A. R. Kosiorek, S. Choi, and Y. W. Teh, “Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks,” inProceedings of the 36th International Conference on Machine Learning, Proceedings of Machine Learning Research97, 3744–3753 (2019). PMLR proceedings

  18. [18]

    Pretraining of attention-based deep learning potential model for molecular simulation,

    D. Zhanget al., “Pretraining of attention-based deep learning potential model for molecular simulation,”npj Computational Materials10, 94 (2024). doi:10.1038/s41524-024-01278-7

  19. [19]

    SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks,

    F. B. Fuchs, D. E. Worrall, V. Fischer, and M. Welling, “SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks,” inAdvances in Neural Information Processing Systems33, 1970–1981 (2020). arXiv:2006.10503

  20. [20]

    TorchMD-NET: Equivariant Transformers for Neural Network Based Molecular Potentials,

    P. Th¨ olke and G. De Fabritiis, “TorchMD-NET: Equivariant Transformers for Neural Network Based Molecular Potentials,” inInternational Conference on Learning Representations(2022). arXiv:2202.02541

  21. [21]

    EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers,

    Y.-L. Liao, A. J. Hoffman, S. C. Shen, A. Duval, S. W. Norwood, and T. E. Smidt, “EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers,” arXiv:2604.09130 (2026). arXiv:2604.09130

  22. [22]

    SO3krates: Equivariant Attention for Interactions on Arbitrary Length-Scales in Molecular Systems,

    J. T. Frank, O. T. Unke, and K.-R. M¨ uller, “SO3krates: Equivariant Attention for Interactions on Arbitrary Length-Scales in Molecular Systems,” inAdvances in Neural Information Processing Systems35, 29400–29413 (2022). NeurIPS proceedings

  23. [23]

    Smooth, exact rotational symmetrization for deep learning on point clouds,

    S. N. Pozdnyakov and M. Ceriotti, “Smooth, exact rotational symmetrization for deep learning on point clouds,” inAdvances in Neural Information Processing Systems36(2023). NeurIPS proceedings

  24. [24]

    The Importance of Being Scalable: Improving the Speed and Accuracy of Neural Network Interatomic Potentials Across Chemical Domains,

    E. Qu and A. S. Krishnapriyan, “The Importance of Being Scalable: Improving the Speed and Accuracy of Neural Network Interatomic Potentials Across Chemical Domains,” in Advances in Neural Information Processing Systems37(2024). doi:10.52202/079017-4412. 31

  25. [25]

    Orb-v3: atomistic simulation at scale,

    B. Rhodes, S. Vandenhaute, V. ˇSimkus, J. Gin, J. Godwin, T. Duignan, and M. Neumann, “Orb-v3: atomistic simulation at scale,”arXiv preprintarXiv:2504.06231 (2025). doi:10.48550/arXiv.2504.06231

  26. [26]

    Learning local equivariant representations for large-scale atomistic dynamics,

    A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, “Learning local equivariant representations for large-scale atomistic dynamics,” Nature Communications14, 579 (2023). doi:10.1038/s41467-023-36329-y

  27. [27]

    Muon: An optimizer for hidden layers in neural networks

    K. Jordan, Y. Jin, V. Boza, You Jiacheng, F. Cesista, L. Newhouse, and J. Bernstein, “Muon: An optimizer for hidden layers in neural networks” (2024), online methods note

  28. [28]

    Training a Foundation Model for Materials on a Budget,

    T. Koker, M. Kotak, and T. Smidt, “Training a Foundation Model for Materials on a Budget,” arXiv:2508.16067 (2025). arXiv:2508.16067; associated Nequix implementation: github.com/atomicarchitects/nequix

  29. [29]

    Optimum band gap combinations to make best use of new photovoltaic materials,

    S. P. Bremner, C. Yi, I. Almansouri, A. Ho-Baillie, and M. A. Green, “Optimum band gap combinations to make best use of new photovoltaic materials,”Solar Energy135, 750–757 (2016). doi:10.1016/j.solener.2016.06.042

  30. [30]

    Efficiency Limit of Perovskite/Si Tandem Solar Cells,

    M. H. Futscher and B. Ehrler, “Efficiency Limit of Perovskite/Si Tandem Solar Cells,”ACS Energy Letters1, 863–868 (2016). doi:10.1021/acsenergylett.6b00405

  31. [31]

    Size dependent solid-solid crystallization of halide perovskites,

    P. Ahlawat, “Size dependent solid-solid crystallization of halide perovskites,” arXiv:2404.05644 (2024). arXiv:2404.05644

  32. [32]

    Molecule-induced ripening control in perovskite quantum dots for efficient and stable light-emitting diodes,

    J. Chen, S. Chen, X. Liu, D. Zhu, B. Cai, X. Luo, W. Feng, Y. Cheng, Y. Xiong, J. Du, Z. Li, D. Zhang, L. Duan, and D. Ma, “Molecule-induced ripening control in perovskite quantum dots for efficient and stable light-emitting diodes,”Science Advances11, eads7159 (2025). doi:10.1126/sciadv.ads7159

  33. [33]

    Nanoscale heterophase regulation enables sunlight-like full-spectrum white electroluminescence,

    J. Chen, K. Ji, L. Dai, H. Xiang, Z. Yu, A. N. Iqbal, J. Wang, X. Ma, R. Guo, M. Anaya, X. Song, Y. Lu, Y.-H. Chiang, W. Li, Y. Shen, X. Luo, A. Mirabelli, Y. Cheng, X. Chen, D. Ma, Z. Fan, Y. Yang, L. Duan, S. D. Stranks, and H. Zeng, “Nanoscale heterophase regulation enables sunlight-like full-spectrum white electroluminescence,”Nature Communications16,...

  34. [34]

    Intragrain 3D perovskite heterostructure for high-performance pure-red perovskite LEDs,

    Y.-H. Song, B. Li, Z.-J. Wang, X.-L. Tai, G.-J. Ding, Z.-D. Li, H. Xu, J.-M. Hao, K.-H. Song, L.-Z. Feng, Y.-L. Hu, Y.-C. Yin, B.-S. Zhu, G. Zhang, H. Ju, G. Zheng, W. Hu, Y. Lin, F. Fan, and H.-B. Yao, “Intragrain 3D perovskite heterostructure for high-performance pure-red perovskite LEDs,”Nature641, 352–357 (2025). doi:10.1038/s41586-025-08867-6. 32

  35. [35]

    Intermediate phase evolution for stable and oriented evaporated wide-bandgap perovskite solar cells,

    Z. Dong, J. Hu, X. Guo, Z. Shi, H. Chen, Y. Wang, R. Luo, J. A. Steele, Z. Degnan, E. Solano, Q. Zhou, N. Kalasariya, N. Li, T. Wang, J. Chen, L. K. Lee, Y. Wang, J. Li, M. Stolterfoht, M. Sui, Y. Lu, and Y. Hou, “Intermediate phase evolution for stable and oriented evaporated wide-bandgap perovskite solar cells,”Nature Materials25, 635–642 (2026). doi:10...

  36. [36]

    New Monte Carlo method to compute the free energy of arbitrary solids. Application to the fcc and hcp phases of hard spheres,

    D. Frenkel and A. J. C. Ladd, “New Monte Carlo method to compute the free energy of arbitrary solids. Application to the fcc and hcp phases of hard spheres,”Journal of Chemical Physics81, 3188–3193 (1984). doi:10.1063/1.448024

  37. [37]

    Nonequilibrium free-energy calculation of solids using LAMMPS,

    R. J. R. X. Freitas, M. Asta, and M. de Koning, “Nonequilibrium free-energy calculation of solids using LAMMPS,”Computational Materials Science112, 333–341 (2016). doi:10.1016/j.commatsci.2015.10.050

  38. [38]

    High-temperature structural evolution of caesium and rubidium triiodoplumbates,

    D. M. Trots and S. V. Myagkota, “High-temperature structural evolution of caesium and rubidium triiodoplumbates,”Journal of Physics and Chemistry of Solids69, 2520–2526 (2008). doi:10.1016/j.jpcs.2008.05.007

  39. [39]

    PLUMED 2: New feathers for an old bird,

    G. A. Tribello, M. Bonomi, D. Branduardi, C. Camilloni, and G. Bussi, “PLUMED 2: New feathers for an old bird,”Computer Physics Communications185, 604–613 (2014). doi:10.1016/j.cpc.2013.09.018

  40. [40]

    Rethinking Metadynamics: From Bias Potentials to Probability Distributions,

    M. Invernizzi and M. Parrinello, “Rethinking Metadynamics: From Bias Potentials to Probability Distributions,”Journal of Physical Chemistry Letters11, 2731–2736 (2020). doi:10.1021/acs.jpclett.0c00497

  41. [41]

    Nonphysical sampling distributions in Monte Carlo free-energy estimation: Umbrella sampling,

    G. M. Torrie and J. P. Valleau, “Nonphysical sampling distributions in Monte Carlo free-energy estimation: Umbrella sampling,”Journal of Computational Physics23, 187–199 (1977). doi:10.1016/0021-9991(77)90121-8

  42. [42]

    The weighted histogram analysis method for free-energy calculations on biomolecules. I. The method,

    S. Kumar, J. M. Rosenberg, D. Bouzida, R. H. Swendsen, and P. A. Kollman, “The weighted histogram analysis method for free-energy calculations on biomolecules. I. The method,” Journal of Computational Chemistry13, 1011–1021 (1992). doi:10.1002/jcc.540130812

  43. [43]

    Thermal [1,5] sigmatropic alkyl shifts of isoindenes,

    W. R. Dolbier, Jr., K. E. Anapolle, L. McCullagh, K. Matsui, J. M. Riemann, and D. Rolison, “Thermal [1,5] sigmatropic alkyl shifts of isoindenes,”Journal of Organic Chemistry 44, 2845–2849 (1979). doi:10.1021/jo01330a006

  44. [44]

    Sigmatropic rearrangements of 1,1-diarylindenes: Migratory aptitudes of aryl migration in the ground and electronically 33 excited states,

    C. Manning, M. R. McClory, and J. J. McCullough, “Sigmatropic rearrangements of 1,1-diarylindenes: Migratory aptitudes of aryl migration in the ground and electronically 33 excited states,”Journal of Organic Chemistry46, 919–930 (1981). doi:10.1021/jo00318a018

  45. [45]

    Active learning meets metadynamics: Automated workflow for reactive machine learning interatomic potentials,

    V. Vitartas, H. Zhang, V. Juraskova, T. Johnston-Wood, and F. Duarte, “Active learning meets metadynamics: Automated workflow for reactive machine learning interatomic potentials,”Digital Discovery5, 108–122 (2026). doi:10.1039/D5DD00261C

  46. [46]

    NEP-MB-pol: A unified machine-learned framework for fast and accurate prediction of water’s thermodynamic and transport properties,

    K. Xu, T. Liang, N. Xu, P. Ying, S. Chen, N. Wei, J. Xu, and Z. Fan, “NEP-MB-pol: A unified machine-learned framework for fast and accurate prediction of water’s thermodynamic and transport properties,”npj Computational Materials11, 279 (2025). doi:10.1038/s41524-025-01777-1

  47. [47]

    Benchmark oxygen–oxygen pair-distribution function of ambient water from x-ray diffraction measurements with a wideQrange,

    L. B. Skinner, C. J. Benmore, J. K. R. Weber, J. B. Parise, and T. R. Hart, “Benchmark oxygen–oxygen pair-distribution function of ambient water from x-ray diffraction measurements with a wideQrange,”Journal of Chemical Physics138, 074506 (2013). doi:10.1063/1.4790861

  48. [48]

    The radial distribution functions of water and ice from 220 to 673 K and at pressures up to 400 MPa,

    A. K. Soper, “The radial distribution functions of water and ice from 220 to 673 K and at pressures up to 400 MPa,”Chemical Physics258, 121–137 (2000). doi:10.1016/S0301-0104(00)00179-8. 34 Appendix A: Fixed-environment tensorial cross-attention This Appendix states the TRACE attention algorithm at the level of its tensor depen- dencies and sparse impleme...

  49. [49]

    For each edgee,s(e) =jdenotes the sender andr(e) =ithe receiver or central atom

    Center states and fixed directed-edge tokens LetNbe the number of atoms and letEbe the image-resolved directed neighbor list, E={e= (j→i,S ij) :d e =∥r j −r i +S ijh∥< rc}.(A1) Its size isN e =|E|. For each edgee,s(e) =jdenotes the sender andr(e) =ithe receiver or central atom. Periodic images are distinct entries when they lie inside the cutoff. The ACE ...

  50. [50]

    Queries, keys, and equivariant values At blockt, the even scalar channels of the center state provideHqueries, q(t) ip =W Q,t p LNh h(t) i,ℓ=0 ∈R dk, p= 1, . . . , H.(A9) Only the invariant scalar part of each fixed edge token provides its keys, k(t) ep =W K,t p LNa(ae,ℓ=0)∈R dk.(A10) The complete edge tensor, including nonscalar irreps, provides one equi...

  51. [51]

    This product is square only for self-attention when the query and key sequences have the same length

    Why the score array is[N e, H], not[Ne, Ne, H] Scaled dot-product attention was introduced in the Transformer asQK T/√dk [16]. This product is square only for self-attention when the query and key sequences have the same length. For a particular TRACE centeriand one attention head, collect itsn i incoming edge keys and values into Ki ∈R ni×dk,V i ∈R ni×Dh...

  52. [52]

    The implemented radial-bias network is Linear(N r,max(16,4H))–SiLU– Linear(max(16,4H), H); unlike the descriptor radial network, these two linear maps include biases

    Radial logits and cutoff-preserving segment softmax The dot product is augmented by invariant radial terms, s(t) ep =η (t) ep +b (t) p (B(de))−softplus(λ (t) p )de,(A20) where the final term is present when the distance penalty is enabled, as it is for the re- ported models. The implemented radial-bias network is Linear(N r,max(16,4H))–SiLU– Linear(max(16...

  53. [53]

    (A28); vali- dation and inference use the undropped weights

    Equivariant aggregation, residual update, and scalar feed-forward map For each head, the weighted values are accumulated into their receivers, z(t) ip = X e:r(e)=i α(t) ep v(t) ep .(A28) During training, dropout is applied toα ep after normalization and before Eq. (A28); vali- dation and inference use the undropped weights. The reported configurations use...

  54. [54]

    Algorithmic sequence For one block, the implemented calculation can be summarized as follows:

  55. [55]

    Normalize the current center representation and project its even scalar channels to Q[N, H, dk]

  56. [56]

    Normalize the scalar part of the fixed edge tensor and project it toK[N e, H, dk]

  57. [57]

    GatherQ[receiver], contract the aligned query–key pairs overd k, and add the radial bias and nonnegative distance penalty

  58. [58]

    (A27); during training, apply post-softmax dropout without renormalization

    Apply the lower-bounded attention-temperature factor and the cutoff-preserving seg- ment softmax of Eq. (A27); during training, apply post-softmax dropout without renormalization

  59. [59]

    For each head, map the complete fixed edge tensor equivariantly to its value, multiply by the scalar attention weight, and scatter-add the result to the receiver. 42

  60. [60]

    Average the heads, apply the equivariant output projection and tied irrep-wise layer scale, and add the attention residual

  61. [61]

    At no step ish (t) s(e) read to construct the key or value

    Form the invariant tensor norms, normalize the combined scalar-and-norm vector, apply linear–SiLU–dropout–linear–dropout, multiply the scalar residual by its layer- scale vector, and add it only to the scalar channels. At no step ish (t) s(e) read to construct the key or value. Thus increasing the number of blocks changes the nonlinear interrogation of on...

  62. [62]

    Relabeling equiv- alent atoms correspondingly relabels the center outputs, while their energy sum restores permutation invariance

    Permutation symmetry , locality , and linear scaling Permuting edge storage leaves receiver-indexed reductions unchanged. Relabeling equiv- alent atoms correspondingly relabels the center outputs, while their energy sum restores permutation invariance. Because the attention weights areO(3)-invariant scalars, the val- ues are equivariant, parities are expl...