REVIEW 3 major objections 5 minor 92 references
This paper claims that machine-learning force fields for molecular dynamics can be reformulated as implicit fixed-point models, and that warm-starting the solver from previous timesteps cuts compute and memory by two- to five-fold while mat
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 12:32 UTC pith:K7CTXREL
load-bearing objection Solid DEQ-force-field paper with a genuinely useful warm-starting contribution, but the headline 2–5x speedup is measured in interaction-layer calls, not wall-clock, so the strong claim needs runtime validation. the 3 major comments →
Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that implicit modeling makes force-field inference amortizable along a molecular dynamics trajectory. Concretely, the paper shows that any explicit graph-neural-network force field can be converted into an implicit model by iterating one learned interaction layer, with input injection and normalization, until it satisfies h* = f(h*, x). The forces are then obtained by implicit differentiation, solving a second fixed-point equation for the adjoint u*, which requires storing only the final application of f. The authors demonstrate that when the solver is warm-started with linear extrapolation h(0)(t+Δt) = 2h*(t) - h*(t-Δt), the fixed point and adjoint converge to a residua
What carries the argument
The load-bearing mechanism is the fixed-point formulation of the latent representation, h* = f(h*, x), together with its implicit derivative. Because the mapping from positions to fixed points is differentiable (by the implicit function theorem), forces can be computed without unrolling the solver, using the adjoint fixed-point equation u* = (∂f/∂h*)ᵀ u* + ∂f_E/∂h*. Temporal amortization is achieved by warm-starting both solves with a linear (Adams-Bashforth-style) extrapolation of the previous two fixed points; training-time Jacobian and iterate-correction regularization keep the solver contractive enough to converge in one to two iterations. The entire argument rests on the smooth evolutio
Load-bearing premise
The central claim collapses if the number of calls to the interaction layer is not a faithful proxy for total cost—i.e., if the extra overhead of the fixed-point solver, residual checks, vector-Jacobian products, and extrapolation steps is large enough to erase the measured 2-5x reduction in layer calls.
What would settle it
Measure end-to-end wall-clock time per MD step for the implicit model versus its explicit counterpart with 3-5 layers, including all solver overhead, on a large system such as the double-walled nanotube, at matched force accuracy. If the implicit model is not faster (or not faster by the claimed factor), the 'compute = layer calls' assumption fails. Alternatively, run the solver with linear warm-start at a deliberately large timestep (e.g., 10 fs) on a high-temperature trajectory: if iteration counts jump far above two, the temporal-continuity premise is violated.
If this is right
- I-MLFFs reduce the per-step cost of force evaluation to about one interaction-layer call, bringing machine-learned force fields closer to classical force fields in speed without coarse graining or increasing the integration timestep.
- Memory use for force inference becomes independent of network depth (only the last layer application is stored), enabling larger atomistic systems on fixed-memory GPUs.
- The efficiency gain is architecture-agnostic, applying to invariant, Cartesian-equivariant, and spherical-tensor equivariant models, and is additive to future architectural and implementation improvements.
- Implicit models adaptively spend more solver iterations on rare, high-energy, or out-of-distribution conformations while remaining at one iteration near equilibrium, which preserves stability in NVE and NVT simulations.
- The fixed-point formulation extends the effective range of message passing beyond explicit depth, improving long-range and extrapolative prediction (e.g., cumulenes).
Where Pith is reading between the lines
- The warm-starting principle is not limited to force fields: any physics-based ML model whose inference is an iterative solve over a smoothly varying input sequence (e.g., neural wavefunctions, learned self-consistent-field solvers, Hamiltonian networks) could inherit similar amortization, as the paper's I-HNN experiment hints.
- Because the paper measures compute in interaction-layer calls, the practical wall-clock gain depends on the fixed-point solver overhead being negligible; profiling on large systems would confirm whether the 2-5x claim survives to wall-clock time, especially on GPUs where small kernels have launch overhead.
- One testable prediction is that the speedup grows as the integration timestep shrinks (fixed points become more similar) and shrinks as temperature rises or timestep increases—a direct consequence of the smoothness assumption the authors make.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes implicit machine learning force fields (I-MLFFs), replacing explicit stacks of graph-neural-network interaction layers with a fixed-point equation h* = f(h*, x). Forces are obtained by implicit differentiation, avoiding unrolled backpropagation, and the fixed-point and adjoint solvers are warm-started from previous MD timesteps, with linear extrapolation from the two previous fixed points. The authors claim that this reduces per-step inference cost to roughly one to two interaction-layer calls, giving a two- to five-fold reduction in compute and memory relative to explicit baselines, while preserving energy conservation. The method is demonstrated on SchNet, PaiNN, and SO3net backbones on MD17 and MD22, with additional stability and generalization experiments.
Significance. If the efficiency claim survives direct runtime measurement, the paper makes a valuable contribution: it couples MLFF inference with temporal coherence of MD trajectories, retains conservative forces through implicit differentiation, and offers architecture-agnostic improvements in interaction-layer call counts and memory. The strengths include extensive benchmarks across three architecture families and two benchmark suites, a clear algorithmic description, memory measurements, NVE/NVT stability tests, and a code release. The main unresolved issue is that the headline '2-5x compute/speedup' claim is supported only by counting interaction-layer calls, not by wall-clock time or FLOPs, so the practical significance is not yet fully established.
major comments (3)
- [Sec. II.C & Abstract] The headline 'two- to five-fold reduction in compute' rests on the statement in Sec. II.C that cost is measured as 'the number of calls to the interaction layer f' and that 'other contributions ... are generally negligible and comparable.' This assumption is load-bearing. The implicit pipeline adds residual-norm checks after every iteration, warm-start extrapolation (Eq. 6), storage/retrieval of previous fixed points, and the backward fixed-point solve (Eq. 4), which requires VJPs through f. For GNN layers with edge filters, normalization, and SiLU nonlinearities, reverse-mode cost is not guaranteed to be comparable to forward cost despite the Griewank citation. Since no wall-clock or FLOP comparison is reported, the abstract's 'simulation speedup' claim is not directly supported. Please add GPU wall-clock per MD step and a breakdown of solver/VJP overhead, or temper the claim to 'intera
- [Sec. II.C & Fig. 3] The 'matched computational cost' comparison in Fig. 3 requires a precise cost unit. An explicit K-layer model performs K forward calls and K backward VJPs; an implicit model performs I forward iterations and J backward iterations. If one VJP is counted as one f-call, the reported ratios depend on that equivalence. If VJPs are actually two to three times the forward cost, the cost-matched crossover and the claimed speedup change materially. Please specify how VJPs are converted into f-call units for each architecture and show sensitivity of the Fig. 3 results to this multiplier.
- [Sec. IV.B & IV.A] The paper states in Sec. IV.B that forces are independent of the warm start because implicit differentiation uses only local derivative information at the fixed point. Strictly, this requires the fixed-point equation h* = f(h*, x) to have a unique solution in the relevant region. Brouwer's theorem (Sec. IV.A) gives existence only, not uniqueness. If multiple fixed points or near-singular I - ∂f/∂h exist, different warm starts can select different branches, and the energy/force may not be a well-defined function of R. Please provide evidence of contraction or uniqueness (e.g., spectral radius estimates) or an empirical test comparing energies and forces obtained from multiple solver initializations along closed conformational loops.
minor comments (5)
- [Fig. 3] The placeholder text 'Lorem ipsum' appears in the figure/legend and must be removed before publication.
- [SI Sec. S4 vs Fig. 4] SI Sec. S4 refers to 'Fig. 4d' for the fine-tolerance temperature plot, but main-text Fig. 4 has only panels (a)-(c); the relevant panel is (c).
- [Fig. 1(d) & Sec. IV.E] Fig. 1(d) says the 2-5x footprint is 'averaged across MD17 and MD22 datasets at 300 K,' while MD17 trajectories are at 500 K and MD22 at 400-500 K (Sec. IV.E). Please clarify the temperature/protocol used for this panel.
- [Sec. II.A vs Sec. IV.C] The iterate-correction loss coefficient is reported as 10^3 in Sec. II.A but 10^4 in Sec. IV.C and in the final loss expression. Please unify.
- [Algorithm 1] Algorithm 1 has no maximum iteration count. For robustness in production MD, add a cap and define the behavior when the residual is not met within the cap.
Circularity Check
No significant circularity: efficiency claims rest on measured iteration counts and external accuracy benchmarks, not on self-referential definitions.
full rationale
I find no circular step. The implicit-force derivation (Eqs. 2-4) is a standard implicit-function-theorem argument and does not assume the efficiency conclusion. Accuracy is benchmarked against external MD17/MD22 DFT references, and the implicit-vs-explicit force-MAE comparisons are independent of the paper's fitted parameters. Hyperparameters (jac coefficient 0.32, itc weight 10^4, solver tolerance 10^-2) are tuned on aspirin and then transferred to other systems; this is a model-selection choice, not a target-equivalent fit. The warm-start speedups (Fig. 2D: 1.12 iterations with linear extrapolation; Fig. 3A: ~1-2 layer calls) are measured iteration counts to a stated residual, not quantities forced by construction, even though Eq. (6) supplies a nearby initial guess. Self-citations are to datasets (MD17/MD22) and to a generalization protocol from Ref. [12] that the paper reproduces rather than assumes. The one caveat that could weaken the headline claim is the cost metric: Sec. II.C defines cost as number of interaction-layer calls and assumes other contributions are negligible, so the 2-5x compute and memory advantage may not translate exactly to wall-clock time if solver/VJP overhead dominates. But this is an external-validity threat, not circularity: the f-call count is a stated, measurable proxy, and no equation in the paper reduces to the fitted benchmarks.
Axiom & Free-Parameter Ledger
free parameters (4)
- Forward/backward solver tolerance (epsilon) =
1e-2 (3.2e-3 for drift-free NVE)
- Jacobian regularization coefficient =
0.32
- Iterate correction coefficient and decay factor =
1e4, gamma=0.4
- Force loss weight alpha =
0.95
axioms (5)
- standard math Brouwer fixed-point theorem: a continuous self-map on a compact convex set has a fixed point.
- standard math Implicit Function Theorem: the fixed-point equation h* = f(h*, x) locally defines a differentiable map x -> h*(x).
- domain assumption The fixed-point trajectory h*(t) is smooth enough along MD trajectories for linear extrapolation to be an accurate warm-start.
- ad hoc to paper The number of interaction-layer calls is a faithful proxy for computational cost.
- domain assumption Residual tolerance 10^-2 is sufficient for accurate force prediction.
read the original abstract
We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point equations. In molecular simulations, this formulation enables intermediate representations to be reused across successive timesteps, thereby warm-starting force evaluation. The resulting models effectively combine the computational footprint of a shallow, single-layer MLFF with the representational capacity and accuracy of a deep neural network. Our approach unlocks architecture-agnostic efficiency gains that are inaccessible when force prediction and trajectory integration are considered separately. We demonstrate this across three major classes of graph neural networks: invariant, equivariant Cartesian tensor, and SO(3)-equivariant spherical-tensor architectures. Each yields a two- to five-fold reduction in compute and memory footprint. Crucially, these gains are achieved while retaining full atomistic resolution and the original integration timestep, avoiding spatial or temporal coarse graining. Our contribution therefore advances the scaling frontier of quantum-mechanically faithful molecular simulation, enabling longer trajectories and larger atomistic systems within fixed GPU memory and compute budgets, and thereby opening access to new insights across biomolecular and material systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Karplus and J
M. Karplus and J. A. McCammon, Molecular dynam- ics simulations of biomolecules, Nat. Struct. Biol.9, 646 (2002)
2002
-
[2]
O. T. Unke, S. Chmiela, H. E. Sauceda, M. Gastegger, I. Poltavsky, K. T. Sch¨ utt, A. Tkatchenko, and K.-R. M¨ uller, Machine learning force fields, Chem. Rev.121, 10142 (2021)
2021
-
[3]
Gasteiger, J
J. Gasteiger, J. Groß, and S. G¨ unnemann, Directional message passing for molecular graphs, inICLR(2020)
2020
-
[4]
Sch¨ utt, O
K. Sch¨ utt, O. Unke, and M. Gastegger, Equivariant mes- sage passing for the prediction of tensorial properties and molecular spectra, inICML(PMLR, 2021) pp. 9377– 9388
2021
-
[5]
T. W. Ko, J. A. Finkler, S. Goedecker, and J. Behler, A fourth-generation high-dimensional neural network po- tential with accurate electrostatics including non-local charge transfer, Nat. commun.12, 398 (2021)
2021
-
[6]
Gasteiger, F
J. Gasteiger, F. Becker, and S. G¨ unnemann, Gem- Net: Universal directional graph neural networks for molecules, inNeurIPS, Vol. 34 (2021) pp. 6790–6802
2021
-
[7]
Liao and T
Y.-L. Liao and T. Smidt, Equiformer: Equivariant graph attention transformer for 3D atomistic graphs, inICLR (2023)
2023
-
[8]
Y. Wang, S. Li, X. He, M. Li, Z. Wang, N. Zheng, B. Shao, T.-Y. Liu, and T. Wang, ViSNet: an equivariant geometry-enhanced graph neural network with vector- scalar interactive message passing for molecules, arXiv preprint arXiv:2210.16518 (2023)
Pith/arXiv arXiv 2023
-
[9]
Batatia, D
I. Batatia, D. P. Kov´ acs, G. Simm, C. Ortner, and G. Cs´ anyi, MACE: Higher order equivariant message passing neural networks for fast and accurate force fields, inNeurIPS, Vol. 35 (2022) pp. 11423–11436
2022
-
[10]
Batzner, A
S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Nat. Commun.13, 2453 (2022)
2022
-
[11]
Musaelian, S
A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, Learning local equivariant representations for large-scale atomistic dy- namics, Nat. Commun.14, 579 (2023)
2023
-
[12]
J. T. Frank, O. T. Unke, K.-R. M¨ uller, and S. Chmiela, A Euclidean transformer for fast and stable machine learned force fields, Nat. Commun.15, 6539 (2024)
2024
-
[13]
Chmiela, H
S. Chmiela, H. E. Sauceda, K.-R. M¨ uller, and A. Tkatchenko, Towards exact molecular dynamics simu- lations with machine-learned force fields, Nat. Commun. 9, 3887 (2018)
2018
-
[14]
Kabylda, J
A. Kabylda, J. T. Frank, S. Su´ arez-Dou, A. Khabibrakhmanov, L. Medrano Sandonas, O. T. Unke, S. Chmiela, K.-R. M¨ uller, and A. Tkatchenko, Molecular simulations with a pretrained neural network and universal pairwise force fields, J. Am. Chem. Soc. 147, 33723 (2025)
2025
-
[15]
D. P. Kov´ acs, J. H. Moore, N. J. Browning, I. Batatia, J. T. Horton, V. Kapil, W. C. Witt, I.-B. Magd˘ au, D. J. Cole, and G. Cs´ anyi, MACE-OFF23: Transferable ma- chine learning force fields for organic molecules, arXiv preprint arXiv:2312.15211 (2023)
Pith/arXiv arXiv 2023
-
[16]
M. E. Tuckerman, Ab initio molecular dynamics: basic concepts, current trends and novel applications, J. Phys.: Condens. Matter14, R1297 (2002)
2002
-
[17]
W. D. Cornell, P. Cieplak, C. I. Bayly, I. R. Gould, K. M. Merz, D. M. Ferguson, D. C. Spellmeyer, T. Fox, J. W. Caldwell, and P. A. Kollman, A second generation force field for the simulation of proteins, nucleic acids, and organic molecules, J. Am. Chem. Soc.117, 5179 (1995)
1995
-
[18]
A. D. MacKerell, D. Bashford, M. Bellott, R. L. Dunbrack, J. D. Evanseck, M. J. Field, S. Fischer, J. Gao, H. Guo, S. Ha, D. Joseph-McCarthy, L. Kuch- nir, K. Kuczera, F. T. K. Lau, C. Mattos, S. Mich- nick, T. Ngo, D. T. Nguyen, B. Prodhom, W. E. Rei- her, B. Roux, M. Schlenkrich, J. C. Smith, R. Stote, J. Straub, M. Watanabe, J. Wiorkiewicz-Kuczera, D. ...
1998
-
[19]
J. Wang, S. Olsson, C. Wehmeyer, A. P´ erez, N. E. Charron, G. de Fabritiis, F. No´ e, and C. Clementi, Ma- chine learning of coarse-grained molecular dynamics force fields, ACS Cent. Sci.5, 755 (2019)
2019
-
[20]
B. E. Husic, N. E. Charron, D. Lemm, J. Wang, A. P´ erez, M. Majewski, A. Kr¨ amer, Y. Chen, S. Olsson, G. de Fab- ritiis, F. No´ e, and C. Clementi, Coarse graining molecu- lar dynamics with graph neural networks, J. Chem. Phys. 153, 194101 (2020)
2020
-
[21]
Majewski, A
M. Majewski, A. P´ erez, P. Th¨ olke, S. Doerr, N. E. Char- ron, T. Giorgino, B. E. Husic, C. Clementi, F. No´ e, and G. De Fabritiis, Machine learning coarse-grained poten- tials of protein thermodynamics, Nat. Commun.14, 5739 (2023)
2023
-
[22]
N. E. Charron, K. Bonneau, A. S. Pasos-Trejo, A. Gul- jas, Y. Chen, F. Musil, J. Venturin, D. Gusew, I. Za- porozhets, A. Kr¨ amer, C. Templeton, A. Kelkar, A. E. P. Durumeric, S. Olsson, A. P´ erez, M. Majewski, B. E. Hu- sic, A. Patel, G. De Fabritiis, F. No´ e, and C. Clementi, Navigating protein landscapes with a machine-learned transferable coarse-gr...
2025
-
[23]
A. E. P. Durumeric, Y. Chen, A. S. Pasos-Trejo, F. No´ e, and C. Clementi, Learning data-efficient coarse-grained molecular dynamics from forces and noise, Nat. Com- mun.17, 2493 (2026)
2026
-
[24]
T. J. Lane, D. Shukla, K. A. Beauchamp, and V. S. Pande, To milliseconds and beyond: challenges in the simulation of protein folding, Curr. Opin. Struct. Biol. 23, 58 (2013)
2013
-
[25]
Simard, M
P. Simard, M. Ottaway, and D. Ballard, Fixed point anal- ysis for recurrent networks, NeurIPS1, 149 (1988)
1988
-
[26]
Miller and M
J. Miller and M. Hardt, Stable recurrent models, inInter- national Conference on Learning Representations(2019)
2019
-
[27]
S. Bai, J. Z. Kolter, and V. Koltun, Deep equilibrium models, inNeurIPS, Vol. 32 (Curran Associates, Inc.,
-
[28]
Winston and J
E. Winston and J. Z. Kolter, Monotone operator equilib- rium networks, inNeurIPS, Vol. 33 (Curran Associates, Inc., 2020) pp. 10718–10728
2020
-
[29]
Y. Lu, A. Zhong, Q. Li, and B. Dong, Beyond finite layer neural networks: Bridging deep architectures and numer- ical differential equations, inICML(PMLR, 2018) pp. 3276–3285. 12
2018
-
[30]
R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud, Neural ordinary differential equations, inNeurIPS, Vol. 31 (Curran Associates, Inc., 2018) pp. 6571–6583
2018
-
[31]
Haber and L
E. Haber and L. Ruthotto, Stable architectures for deep neural networks, Inverse Probl.34, 014004 (2017)
2017
-
[32]
Ruthotto and E
L. Ruthotto and E. Haber, Deep neural networks moti- vated by partial differential equations, J. Math. Imaging. Vis.62, 352 (2020)
2020
-
[33]
L. L. Schaaf, I. Batatia, J. Tilly, and T. D. Barrett, BoostMD: Accelerated molecular sampling leveraging ml force field features, inNeurIPS 2024 Workshop on Data- driven and Differentiable Simulations, Surrogates, and Solvers(2024)
2024
-
[34]
Burger, L
A. Burger, L. Thiede, A. Aspuru-Guzik, and N. Vijayku- mar, DEQuify your force field: Towards efficient simula- tions using deep equilibrium models, inAI for Accelerated Materials Design Workshop, ICLR(2025)
2025
-
[35]
F. L. Thiemann, T. Resch¨ utzegger, M. Esposito, T. Tad- dese, J. D. Olarte-Plata, and F. Martelli, Force-free molecular dynamics through autoregressive equivariant networks, arXiv preprint arXiv:2503.23794 (2025)
Pith/arXiv arXiv 2025
-
[36]
F. Bigi, S. Chong, A. Kristiadi, and M. Ceriotti, FlashMD: long-stride, universal prediction of molecular dynamics, inNeurIPS(2026)
2026
-
[37]
W. Ripken, M. Plainer, G. Lied, T. Frank, O. T. Unke, S. Chmiela, F. No´ e, and K.-R. M¨ uller, Learning hamilto- nian flow maps: Mean flow consistency for large-timestep molecular dynamics, arXiv preprint arXiv:2601.22123 (2026)
Pith/arXiv arXiv 2026
-
[38]
Chmiela, A
S. Chmiela, A. Tkatchenko, H. E. Sauceda, I. Poltavsky, K. T. Sch¨ utt, and K.-R. M¨ uller, Machine learning of ac- curate energy-conserving molecular force fields, Sci. Adv. 3, e1603015 (2017)
2017
-
[39]
W. Hu, M. Shuaibi, A. Das, S. Goyal, A. Sriram, J. Leskovec, D. Parikh, and C. L. Zitnick, ForceNet: A graph neural network for large-scale quantum calcula- tions, arXiv preprint arXiv:2103.01436 (2021)
Pith/arXiv arXiv 2021
-
[40]
C. L. Zitnick, A. Das, A. Kolluru, J. Lan, M. Shuaibi, A. Sriram, Z. Ulissi, and B. Wood, Spherical channels for modeling atomic interactions, inNeurIPS, Vol. 35 (2022)
2022
-
[41]
J. Gasteiger, M. Shuaibi, A. Sriram, S. G¨ unnemann, Z. Ulissi, C. L. Zitnick, and A. Das, GemNet-OC: Devel- oping graph neural networks for large and diverse molec- ular simulation datasets, Transactions on Machine Learn- ing Research 10.48550/arXiv.2204.02782 (2022)
-
[42]
Passaro and C
S. Passaro and C. L. Zitnick, Reducing SO(3) convolu- tions to SO(2) for efficient equivariant GNNs, inICML, Vol. 202 (PMLR, 2023) pp. 27420–27438
2023
-
[43]
Y.-L. Liao, B. M. Wood, A. Das, and T. Smidt, EquiformerV2: Improved equivariant transformer for scaling to higher-degree representations, inICLR(2024)
2024
-
[44]
M. Neumann, J. Gin, B. Rhodes, S. Bennett, Z. Li, H. Choubisa, A. Hussey, and J. Godwin, Orb: A fast, scalable neural network potential, arXiv preprint arXiv:2410.22570 (2024)
Pith/arXiv arXiv 2024
-
[45]
Eissler, T
M. Eissler, T. Korjakow, S. Ganscha, O. T. Unke, K.-R. M¨ uller, and S. Gugler, How simple can you go? an off- the-shelf transformer approach to molecular dynamics, J. Chem. Phys.164(2026)
2026
-
[46]
X. Fu, Z. Wu, W. Wang, T. Xie, S. Keten, R. Gomez- Bombarelli, and T. Jaakkola, Forces are not enough: Benchmark and critical evaluation for machine learn- ing force fields with molecular simulations, Trans. Mach. Learn. Res. (2023)
2023
-
[47]
F. Bigi, M. F. Langer, and M. Ceriotti, The dark side of the forces: assessing non-conservative force models for atomistic machine learning, inICML, Proceedings of Ma- chine Learning Research, Vol. 267, edited by A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (PMLR, 2025) pp. 4384–4414
2025
-
[48]
J. A. Keith, V. Vassilev-Galindo, B. Cheng, S. Chmiela, M. Gastegger, K.-R. M¨ uller, and A. Tkatchenko, Com- bining machine learning and computational chemistry for predictive insights into chemical systems, Chem. Rev. 121, 9816 (2021)
2021
-
[49]
K. T. Sch¨ utt, H. E. Sauceda, P.-J. Kindermans, A. Tkatchenko, and K.-R. M¨ uller, SchNet–a deep learn- ing architecture for molecules and materials, J. Chem. Phys.148(2018)
2018
-
[50]
N. Thomas, T. Smidt, S. Kearnes, L. Yang, L. Li, K. Kohlhoff, and P. Riley, Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds, arXiv preprint arXiv:1802.08219 (2018)
Pith/arXiv arXiv 2018
-
[51]
O. T. Unke, S. Chmiela, M. Gastegger, K. T. Sch¨ utt, H. E. Sauceda, and K.-R. M¨ uller, SpookyNet: Learning force fields with electronic degrees of freedom and nonlo- cal effects, Nat. Commun.12, 7273 (2021)
2021
-
[52]
X. Fu, B. M. Wood, L. Barroso-Luque, D. S. Levine, M. Gao, M. Dzamba, and C. L. Zitnick, Learning smooth and expressive interatomic potentials for physical prop- erty prediction, inICML, Vol. 267 (2025) pp. 17875– 17893
2025
-
[53]
B. M. Wood, M. Dzamba, X. Fu, M. Gao, M. Shuaibi, L. Barroso-Luque, K. Abdelmaqsoud, V. Gharakhanyan, J. R. Kitchin, D. S. Levine, K. Michel, A. Sriram, T. Co- hen, A. Das, A. Rizvi, S. J. Sahoo, Z. W. Ulissi, and C. L. Zitnick, UMA: A family of universal models for atoms, inNeurIPS(2026)
2026
-
[54]
L. E. J. Brouwer, ¨Uber abbildung von mannigfaltigkeiten, Mathematische Annalen71, 97 (1911)
1911
-
[55]
S. G. Krantz and H. R. Parks,The implicit function the- orem: history, theory, and applications(Springer Science & Business Media, 2002)
2002
-
[56]
Sch¨ utt, P.-J
K. Sch¨ utt, P.-J. Kindermans, H. E. Sauceda Felix, S. Chmiela, A. Tkatchenko, and K.-R. M¨ uller, Schnet: A continuous-filter convolutional neural network for mod- eling quantum interactions, inNeurIPS, Vol. 30 (Curran Associates, Inc., 2017) pp. 991–1001
2017
-
[57]
Grisafi, A
A. Grisafi, A. Fabrizio, B. Meyer, D. M. Wilkins, C. Corminboeuf, and M. Ceriotti, Equivariant graph neural networks for fast electron density estimation of molecules, liquids, and solids, npj Comput. Mater.8, 183 (2022)
2022
-
[58]
Esders, T
M. Esders, T. Schnake, J. Lederer, A. Kabylda, G. Mon- tavon, A. Tkatchenko, and K.-R. M¨ uller, Analyzing atomic interactions in molecules as learned by neural net- works, J. Chem. Theory Comput.21, 714 (2025)
2025
-
[59]
Chmiela, V
S. Chmiela, V. Vassilev-Galindo, O. T. Unke, A. Kabylda, H. E. Sauceda, A. Tkatchenko, and K.-R. M¨ uller, Accurate global machine learning force fields for molecules with hundreds of atoms, Sci. Adv.9, eadf0873 (2023)
2023
-
[60]
Kawaguchi, On the theory of implicit deep learning: Global convergence with implicit layers, inICML(2021)
K. Kawaguchi, On the theory of implicit deep learning: Global convergence with implicit layers, inICML(2021). 13
2021
-
[61]
S. Bai, V. Koltun, and Z. Kolter, Stabilizing equilib- rium models by jacobian regularization, inICML, Vol. 139 (PMLR, 2021) pp. 554–565
2021
-
[62]
S. Bai, Z. Geng, Y. Savani, and J. Z. Kolter, Deep equi- librium optical flow estimation, inProceedings of the IEEE/CVF conference on computer vision and pattern recognition(2022) pp. 610–620
2022
-
[63]
J. S. Spencer, D. Pfau, A. Botev, and W. M. C. Foulkes, Better, faster fermionic neural networks, arXiv preprint arXiv:2011.07125 (2020)
Pith/arXiv arXiv 2011
-
[64]
Hermann, Z
J. Hermann, Z. Sch¨ atzle, and F. No´ e, Deep-neural- network solution of the electronic Schr¨ odinger equation, Nat. Chem.12, 891 (2020)
2020
-
[65]
K. T. Sch¨ utt, M. Gastegger, A. Tkatchenko, K.-R. M¨ uller, and R. J. Maurer, Unifying machine learning and quantum chemistry with a deep neural network for molec- ular wavefunctions, Nat. Commun.10, 5024 (2019)
2019
-
[66]
F. Song and J. Feng, Neural network self-consistent fields for density functional theory, npj Comput. Mater. 10.1038/s41524-026-02110-0 (2026)
-
[67]
Zhang, C
H. Zhang, C. Liu, Z. Wang, X. Wei, S. Liu, N. Zheng, B. Shao, and T.-Y. Liu, Self-consistency training for density-functional-theory Hamiltonian prediction, in ICML, Proceedings of Machine Learning Research, Vol. 235, edited by R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (PMLR, 2024) pp. 59329–59357
2024
-
[68]
Z. Wang, C. Liu, N. Zou, H. Zhang, X. Wei, L. Huang, L. Wu, and B. Shao, Infusing self-consistency into density functional theory hamiltonian prediction via deep equi- librium models, inNeurIPS, NIPS ’24 (Curran Associates Inc., Red Hook, NY, USA, 2024)
2024
-
[69]
M. Cranmer, S. Greydanus, S. Hoyer, P. Battaglia, D. Spergel, and S. Ho, Lagrangian neural networks, arXiv preprint arXiv:2003.04630 (2020)
Pith/arXiv arXiv 2003
-
[70]
Greydanus, M
S. Greydanus, M. Dzamba, and J. Yosinski, Hamiltonian neural networks, NeurIPS32(2019)
2019
-
[71]
Sohl-Dickstein, E
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, Deep unsupervised learning using nonequi- librium thermodynamics, inICML(PMLR, 2015) pp. 2256–2265
2015
-
[72]
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, Score-based generative modeling through stochastic differential equations, inICLR(2021)
2021
-
[73]
A. Q. Nichol and P. Dhariwal, Improved denoising dif- fusion probabilistic models, inICML(PMLR, 2021) pp. 8162–8171
2021
-
[74]
Grathwohl, R
W. Grathwohl, R. T. Chen, J. Bettencourt, I. Sutskever, and D. Duvenaud, FFJORD: Free-form continuous dy- namics for scalable reversible generative models, inICLR (2019)
2019
-
[75]
Papamakarios, E
G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mo- hamed, and B. Lakshminarayanan, Normalizing flows for probabilistic modeling and inference, J. Mach. Learn. Res.22, 1 (2021)
2021
-
[76]
Zhang and R
B. Zhang and R. Sennrich, Root mean square layer nor- malization, inNeurIPS, Vol. 32 (2019)
2019
-
[77]
Y.-L. Liao, A. J. Hoffman, S. C. Shen, A. Duval, S. W. Norwood, and T. Smidt, EquiformerV3: Scaling effi- cient, expressive, and general SE(3)-equivariant graph attention transformers, arXiv preprint arXiv:2604.09130 (2026)
Pith/arXiv arXiv 2026
-
[78]
Girard, A fast ‘Monte-Carlo cross-validation’ proce- dure for large least squares problems with noisy data, Numer
A. Girard, A fast ‘Monte-Carlo cross-validation’ proce- dure for large least squares problems with noisy data, Numer. Math.56, 1–23 (1989)
1989
-
[79]
Loshchilov and F
I. Loshchilov and F. Hutter, Decoupled weight decay reg- ularization, inICLR(2019)
2019
-
[80]
Sch¨ utt, P
K. Sch¨ utt, P. Kessel, M. Gastegger, K. A. Nicoli, A. Tkatchenko, and K.-R. M¨ uller, SchNetPack: A deep learning toolbox for atomistic systems, J. Chem. Theory Comput.15, 448 (2018)
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.