Pith. sign in

REVIEW 3 major objections 6 minor 117 references

OpenQDC: Open Quantum Data Commons

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read OpenQDC collects 37 quantum-mechanical datasets, standardizes them into one Python-accessible format for training machine-learning interatomic potentials, and reports baseline results showing that established architectures often miss…

desk verdict A genuinely useful dataset aggregation effort whose benchmark section is currently invalid and whose isolated-atom energy normalization lacks the documentation and validation needed to trust it. read the letter →

arxiv 2411.19629 v1 pith:DEIIPAMO submitted 2024-11-29 physics.chem-ph cs.LG

classification physics.chem-phcs.LG
keywords quantumchemistrydatasetsmachinelearninginteratomicpotentialsdatasetstandardizationenergynormalizationmulti-fidelitymethodsmemory-mappedstoragemoleculardynamics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims to consolidate 37 published quantum-mechanical datasets, covering over 400 million molecular geometries and more than 250 distinct QM method/basis-set combinations, into one programmatically accessible, ML-ready repository. It argues this removes a major barrier for machine-learning interatomic potential (MLIP) development, since QM data were previously scattered across repositories with inconsistent formats, units, and missing metadata. The work supports the repository with a Python library that standardizes loading, unit conversion, QM-method naming, and energy normalization, and with baseline experiments on SchNet, DimeNet, and TorchMD-Net. A careful reader should care because the claim is that a single library now gives the MLIP community the kind of standardized data access that has driven progress in other machine-learning fields.

What carries the argument

The machinery is a memory-mapped storage layout plus a QM-method registry. Geometries, atomic charges, positions, per-geometry index ranges, energy labels, and force labels are stored in memory-mapped arrays that allow indexing and batching without loading data into RAM. QM methods are normalized into Python enums composed of a functional, a basis set, and a correction method, which standardizes naming across datasets and enables studying basis-set or correction effects. For energy normalization, the library pre-computes per-atom-type mean energies via linear or ridge regression and a per-atom residual scale, and it stores computed isolated-atom energies so users can convert potential energies to atomization energies while preserving extensivity.

What would settle it

Recompute isolated-atom energies for a sample of the QM methods in the repository using an independent electronic-structure code and compare them against the stored values; a disagreement beyond numerical tolerance, say greater than 1 meV per atom, would propagate directly into the energy labels users train on. Checking a dataset's distance-unit metadata against its original publication would similarly confirm whether the conversions that feed radial cutoffs are trustworthy.

Watch

Extended reading notes

Core claim

The central discovery is that the fragmented landscape of QM datasets can be unified at scale: 37 datasets, hundreds of millions of geometries, 70 atom types, and 250+ quantum methods are re-exposed through one interface with harmonized metadata. The paper further claims that normalization matters physically: converting potential energies to atomization energies by subtracting computed isolated-atom energies preserves size extensivity, while naive Z-scoring violates it. On the benchmark side, the paper reports that well-known architectures such as SchNet, DimeNet, and TorchMD-Net frequently fail to reach chemical accuracy of 1 kcal/mol on these datasets, indicating headroom for new architectures.

Load-bearing premise

The computed isolated-atom energies and unit metadata for QM methods whose original datasets lacked them are correct; if they are wrong, the atomization-energy normalization and every energy benchmark built on it are wrong.

Editorial extensions

If this is right

  • Any of the 37 datasets can be loaded in one line of Python with user-specified energy, distance, and array formats, while original units are preserved for conversion.
  • Multi-fidelity training is supported because the same molecular families appear at multiple levels of theory, such as QM9 and MultiXCQM9 with its 228 DFT variants.
  • Energy normalization via isolated-atom subtraction lets models trained on one dataset be applied to different system sizes without breaking size extensivity.
  • Interaction-energy datasets such as DES370K, DES5M, Splinter, Metcalf, X40, and L7 provide a dedicated track for training models of drug-target, drug-drug, and protein-protein interactions.
  • Baseline results show SchNet consistently behind the other two architectures on potential-energy datasets while being competitive on interaction datasets, suggesting different architectural biases for the two task families.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own admission that a bug affected dataset splits, making some random rather than molecule-based, implies the reported benchmark numbers should be treated as provisional until corrected re-runs appear.
  • If the computed isolated-atom energies are independently validated, the resulting harmonized multi-method repository would enable principled delta-learning and transfer-learning studies across levels of theory, which the paper only gestures toward.
  • The standardized method enum of functional, basis set, and correction opens a direct route to systematic studies of basis-set convergence and functional-family effects on learned potentials, going beyond what any single dataset paper provides.
  • One testable extension is to train two MLIPs on the same dataset with and without atomization-energy normalization and measure transfer error on larger molecules; the extensivity argument predicts the normalized model degrades less.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces OpenQDC, a Python library and dataset repository that consolidates 37 existing quantum-mechanical datasets, spanning over 250 QM methods and roughly 400 million geometries, into a standardized, programmatically accessible format for machine-learned interatomic potential (MLIP) training. The authors describe the data storage and loading infrastructure, unit handling, QM-method notation standardization, and energy normalization based on isolated-atom reference energies. They also report baseline MLIP experiments with SchNet, DimeNet, and TorchMD-Net on a subset of these datasets, presenting a leaderboard of mean absolute errors. The paper explicitly acknowledges a bug in the dataset splitting that contaminated some baseline results, and notes that corrected results are forthcoming.

Significance. If the repository and preprocessing are correct, OpenQDC would be a valuable community resource: it offers a single, large, multi-fidelity corpus of QM data with a clean Python interface, which could lower the barrier for MLIP research and support multi-task and transfer-learning studies. The paper is also honest about its limitations, including the known split bug. However, the scientific value of the benchmark results is currently weakened by that bug, and the central preprocessing step of computing isolated-atom energies for all QM methods is undocumented and unvalidated. The manuscript therefore has a plausible and potentially impactful core, but the evidence presented is not yet sufficient for its claims.

major comments (3)
  1. [Section 4 (The OpenQDC Library), paragraph on metadata and isolated atom energies] The manuscript states that most datasets lacked isolated-atom energies and that the authors 'computed these values for all QM methods in the repository.' These values underpin the atomization-energy normalization that the paper advertises as a physics-motivated, size-extensive approach, so an error in any one of them propagates into every energy label and derived benchmark. The paper does not report the computational protocol (software, basis set, spin-state treatment, charge handling, convergence criteria), does not list which methods were computed versus taken from the original sources, and provides no validation against known reference values (e.g., QM7/QM9 atomization energies or original dataset metadata). Please add this protocol and a validation table, or state explicitly which datasets rely on externally provided reference energies.
  2. [Section 5.2 (Results) and Table 2] The paper admits that 'a bug affected the dataset splits: some were random rather than molecule-based as intended,' yet Table 2 reports numerical MAE values and the surrounding text draws comparative conclusions, such as SchNet being 'consistently and significantly outperformed' by TorchMDNet and DimeNet and SchNet being more competitive on interaction datasets. These conclusions are not supported by the current, contaminated results. The authors should either remove the benchmark table and defer all architectural comparisons until the corrected splits are available, or clearly label Table 2 as preliminary and confine the discussion to the corrected experiments.
  3. [Section 4 (Dataset Storage) and Section 3 (Table 1)] The abstract and Table 1 present exact counts of geometries, energy labels, and force labels, but the paper provides no checksums, hash-based integrity verification, or scripts that reproduce these statistics from the downloaded source archives. Given that the central claim is the creation of a trustworthy, standardized data commons, adding a reproducibility appendix with dataset-manifest hashes and a count-validation script would materially strengthen the paper; currently these counts are unverifiable from the manuscript alone.
minor comments (6)
  1. [Section 3.1 and Section 3.2] The text contains unresolved placeholder references 'Appendix ??' for dataset visualizations; these should be fixed or removed before publication.
  2. [Section 3, Table 1] The table heading contains a typo, 'min-manximum', and the column header 'Atom Min/-Max' is unclear; please use 'Atom Min/Max' and correct the typo.
  3. [Section 5.2] The word 'significently' should be 'significantly'.
  4. [Section 2 (Related works)] The phrase 'website (to be release upon acceptance)' is grammatically incorrect; it should read 'to be released upon acceptance.'
  5. [Section 5.1 and Section 5.2] The statement in Section 5.1 that 'For all datasets, we performed a molecule/system-based split' is directly contradicted by the split-bug disclosure in Section 5.2; please reconcile these statements in the revised manuscript.
  6. [Section 5.2] The sentence 'Accordingly, we will only compare different architectures and discuss differences based on chemical space only after correcting the splits' suggests the results in Table 2 should be explicitly labeled as provisional; currently the text draws preliminary conclusions despite this caveat.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper aggregates external datasets and reports empirical benchmark errors; its normalization constants are preprocessing, not predictions.

full rationale

This is a dataset aggregation and benchmarking paper rather than a derivation chain. The only constructed numerical quantities are the isolated-atom reference energies and the per-atom shift/scale statistics used in Eq. (1). Neither is presented as a prediction of the benchmark results, and the reported MAEs are measured on held-out molecular splits using labels transformed by these preprocessing constants, which is standard normalization rather than a restatement of the input data. The computed isolated-atom energies are indeed a load-bearing and underdocumented assumption, as the skeptic notes, but missing validation is a correctness risk, not circularity: nothing in the paper defines those energies in terms of the benchmark outputs, nor are they fitted to the test labels. References to prior work, including the coauthor citation [33], are contextual and are not load-bearing for the central aggregation or benchmark claims; no self-citation chain forces the results. The acknowledged split bug in Table 2 is a self-reported benchmarking flaw with no circularity implications. Accordingly, no specific circular step can be quoted or exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claim is an infrastructure claim, so it rests on the correctness of the ingest/conversion process and the computed normalization constants rather than on a derived law. No free parameters are introduced for a scientific conclusion; the per-atom shifts and scales in Eq. 1 are data-fitting tools within the library.

free parameters (1)
  • per-atom energy shift mu_j,k and scale sigma_j per dataset/QM method = not reported; computed via linear/ridge regression
    Used in Eq. 1 to normalize energy labels before training; because these values are fit to the same data being predicted, benchmark MAEs depend on them. They are part of the library's normalization tooling, not of a scientific derivation.
assumptions (3)
  • domain assumption The original third-party QM datasets are correctly fetched, parsed, and converted to the unified memory-mapped format.
    Section 4 describes fetching and preprocessing but provides no verification (checksums, audit). Any conversion error would silently corrupt downstream MLIP training.
  • domain assumption The isolated atom energies computed by the authors for each QM method are correct.
    Section 4: 'Most datasets lacked this information, so we computed these values for all QM methods.' These values are essential for atomization-energy normalization and Eq. 1.
  • domain assumption The energy/force labels in the original datasets are accurate as published.
    The paper does not re-derive or check any QM labels; it trusts the source publications, which is standard for dataset aggregation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OpenQDC: Open Quantum Data Commons." pith.science (2026). https://pith.science/paper/DEIIPAMO

@misc{pith2026241119629,
  author       = {Pith},
  title        = {Pith review of: OpenQDC: Open Quantum Data Commons},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DEIIPAMO}},
  note         = {Machine review of arXiv:2411.19629}
}
read the original abstract

Machine Learning Interatomic Potentials (MLIPs) are a highly promising alternative to force-fields for molecular dynamics (MD) simulations, offering precise and rapid energy and force calculations. However, Quantum-Mechanical (QM) datasets, crucial for MLIPs, are fragmented across various repositories, hindering accessibility and model development. We introduce the openQDC package, consolidating 37 QM datasets from over 250 quantum methods and 400 million geometries into a single, accessible resource. These datasets are meticulously preprocessed, and standardized for MLIP training, covering a wide range of chemical elements and interactions relevant in organic chemistry. OpenQDC includes tools for normalization and integration, easily accessible via Python. Experiments with well-known architectures like SchNet, TorchMD-Net, and DimeNet reveal challenges for those architectures and constitute a leaderboard to accelerate benchmarking and guide novel algorithms development. Continuously adding datasets to OpenQDC will democratize QM dataset access, foster more collaboration and innovation, enhance MLIP development, and support their adoption in the MD field.

Figures

Figures reproduced from arXiv: 2411.19629 by the authors.

Figure 1
Figure 1. Overview of a dataset and its structure. [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Inference time over the potential datasets. [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

117 extracted references · 51 canonical work pages

  1. [1]

    Computational approaches for studying drug binding kinetics

    Julia Romanowska, Daria B Kokh, Jonathan C Fuller, and Rebecca C Wade. Computational approaches for studying drug binding kinetics. Thermodynamics and kinetics of drug binding, pages 211–235, 2015

  2. [2]

    Protein–ligand (un) binding kinetics as a new paradigm for drug discovery at the crossroad between experiments and modelling

    M Bernetti, Andrea Cavalli, and L Mollica. Protein–ligand (un) binding kinetics as a new paradigm for drug discovery at the crossroad between experiments and modelling. MedChem- Comm, 8(3):534–550, 2017

  3. [3]

    New approaches for computing ligand–receptor binding kinetics

    Neil J Bruce, Gaurav K Ganotra, Daria B Kokh, S Kashif Sadiq, and Rebecca C Wade. New approaches for computing ligand–receptor binding kinetics. Current opinion in structural biology, 49:1–10, 2018

  4. [4]

    Kinetics of ligand–protein dissociation from all-atom simulations: Are we there yet? Biochemistry, 58(3):156–165, 2018

    Joao Marcelo Lamim Ribeiro, Sun-Ting Tsai, Debabrata Pramanik, Yihang Wang, and Pratyush Tiwary. Kinetics of ligand–protein dissociation from all-atom simulations: Are we there yet? Biochemistry, 58(3):156–165, 2018

  5. [5]

    Molecular dynamics simulations in drug design

    John E Kerrigan. Molecular dynamics simulations in drug design. In silico models for drug discovery, pages 95–113, 2013

  6. [6]

    Molecular dynamics and protein function

    Martin Karplus and John Kuriyan. Molecular dynamics and protein function. Proceedings of the National Academy of Sciences, 102(19):6679–6685, 2005

  7. [7]

    Applications of molecular dynamics simulation in protein study

    Siddharth Sinha, Benjamin Tam, and San Ming Wang. Applications of molecular dynamics simulation in protein study. Membranes, 12(9):844, 2022

  8. [8]

    Gfn2-xtb—an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions

    Christoph Bannwarth, Sebastian Ehlert, and Stefan Grimme. Gfn2-xtb—an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions. Journal of Chemical Theory and Computation, 15(3):1652–1671, 2019

Show all 117 references
  1. [9]

    Dftb3: Extension of the self-consistent- charge density-functional tight-binding method (scc-dftb)

    Michael Gaus, Qiang Cui, and Marcus Elstner. Dftb3: Extension of the self-consistent- charge density-functional tight-binding method (scc-dftb). Journal of Chemical Theory and Computation, 7(4):931–948, 2011

  2. [10]

    Optimization of parameters for semiempirical methods v: Modification of nddo approximations and application to 70 elements

    James JP Stewart. Optimization of parameters for semiempirical methods v: Modification of nddo approximations and application to 70 elements. Journal of Molecular modeling, 13: 1173–1213, 2007

  3. [11]

    Generalized neural-network representation of high- dimensional potential-energy surfaces

    Jörg Behler and Michele Parrinello. Generalized neural-network representation of high- dimensional potential-energy surfaces. Physical review letters, 98(14):146401, 2007

  4. [12]

    Ani-1, a data set of 20 million calculated off-equilibrium conformations for organic molecules

    Justin S Smith, Olexandr Isayev, and Adrian E Roitberg. Ani-1, a data set of 20 million calculated off-equilibrium conformations for organic molecules. Scientific data, 4(1):1–8, 2017

  5. [13]

    Amp: A modular approach to machine learning in atomistic simulations

    Alireza Khorshidi and Andrew A Peterson. Amp: A modular approach to machine learning in atomistic simulations. Computer Physics Communications, 207:310–324, 2016

  6. [14]

    A reactive, scalable, and transferable model for molecular energies from a neural network approach based on local information

    Oliver T Unke and Markus Meuwly. A reactive, scalable, and transferable model for molecular energies from a neural network approach based on local information. The Journal of chemical physics, 148(24), 2018. 10

  7. [15]

    Equivariant transformers for neural network based molecular potentials

    Philipp Thölke and Gianni De Fabritiis. Equivariant transformers for neural network based molecular potentials. In International Conference on Learning Representations, 2021

  8. [16]

    Linear atomic cluster expansion force fields for organic molecules: beyond rmse

    Dávid Péter Kovács, Cas van der Oord, Jiri Kucera, Alice EA Allen, Daniel J Cole, Christoph Ortner, and Gábor Csányi. Linear atomic cluster expansion force fields for organic molecules: beyond rmse. Journal of chemical theory and computation, 17(12):7696–7711, 2021

  9. [17]

    Efficient and accurate machine- learning interpolation of atomic energies in compositions with many species

    Nongnuch Artrith, Alexander Urban, and Gerbrand Ceder. Efficient and accurate machine- learning interpolation of atomic energies in compositions with many species. Physical Review B, 96(1):014112, 2017

  10. [18]

    Towards universal neural network potential for material discovery applicable to arbitrary combination of 45 elements

    So Takamoto, Chikashi Shinagawa, Daisuke Motoki, Kosuke Nakago, Wenwen Li, Iori Kurata, Taku Watanabe, Yoshihiro Yayama, Hiroki Iriguchi, Yusuke Asano, et al. Towards universal neural network potential for material discovery applicable to arbitrary combination of 45 elements. ...

  11. [19]

    Directional message passing for molecular graphs

    Johannes Gasteiger, Janek Groß, and Stephan Günnemann. Directional message passing for molecular graphs. arXiv preprint arXiv:2003.03123, 2020

  12. [20]

    Gemnet: Universal directional graph neural networks for molecules

    Johannes Gasteiger, Florian Becker, and Stephan Günnemann. Gemnet: Universal directional graph neural networks for molecules. Advances in Neural Information Processing Systems, 34: 6790–6802, 2021

  13. [21]

    Spherical message passing for 3d molecular graphs

    Yi Liu, Limei Wang, Meng Liu, Yuchao Lin, Xuan Zhang, Bora Oztekin, and Shuiwang Ji. Spherical message passing for 3d molecular graphs. In International Conference on Learning Representations, 2021

  14. [22]

    E (n) equivariant graph neural networks

    Vıctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E (n) equivariant graph neural networks. In International conference on machine learning, pages 9323–9332. PMLR, 2021

  15. [23]

    A universal graph deep learning interatomic potential for the periodic table

    Chi Chen and Shyue Ping Ong. A universal graph deep learning interatomic potential for the periodic table. Nature Computational Science, 2(11):718–728, 2022

  16. [24]

    Equivariant message passing for the prediction of tensorial properties and molecular spectra

    Kristof Schütt, Oliver Unke, and Michael Gastegger. Equivariant message passing for the prediction of tensorial properties and molecular spectra. In International Conference on Machine Learning, pages 9377–9388. PMLR, 2021

  17. [25]

    Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds

    Nathaniel Thomas, Tess Smidt, Steven Kearnes, Lusann Yang, Li Li, Kai Kohlhoff, and Patrick Riley. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds. arXiv preprint arXiv:1802.08219, 2018

  18. [26]

    Unite: Unitary n-body tensor equivariant network with applications to quantum chemistry

    Zhuoran Qiao, Anders S Christensen, Matthew Welborn, Frederick R Manby, Anima Anand- kumar, and Thomas F Miller III. Unite: Unitary n-body tensor equivariant network with applications to quantum chemistry. arXiv preprint arXiv:2105.14655, 3, 2021

  19. [27]

    Ani-1: an extensible neural network potential with dft accuracy at force field computational cost

    Justin S Smith, Olexandr Isayev, and Adrian E Roitberg. Ani-1: an extensible neural network potential with dft accuracy at force field computational cost. Chemical science, 8(4):3192– 3203, 2017

  20. [28]

    Torchani: A free and open source pytorch-based deep learning implementation of the ani neural network potentials

    Xiang Gao, Farhad Ramezanghorbani, Olexandr Isayev, Justin S Smith, and Adrian E Roitberg. Torchani: A free and open source pytorch-based deep learning implementation of the ani neural network potentials. Journal of chemical information and modeling, 60(7):3408–3415, 2020

  21. [29]

    Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network

    Roman Zubatyuk, Justin S Smith, Jerzy Leszczynski, and Olexandr Isayev. Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network. Science advances, 5(8):eaav6490, 2019

  22. [30]

    Aimnet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs

    Dylan Anstine, Roman Zubatyuk, and Olexandr Isayev. Aimnet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs. arxiv, 2024

  23. [31]

    Mace-off23: Transfer- able machine learning force fields for organic molecules

    Dávid Péter Kovács, J Harry Moore, Nicholas J Browning, Ilyes Batatia, Joshua T Horton, Venkat Kapil, Ioan-Bogdan Magd˘au, Daniel J Cole, and Gábor Csányi. Mace-off23: Transfer- able machine learning force fields for organic molecules. arXiv preprint arXiv:2312.15211, 2023. 11

  24. [32]

    Forces are not enough: Benchmark and critical evaluation for machine learning force fields with molecular simulations

    Xiang Fu, Zhenghao Wu, Wujie Wang, Tian Xie, Sinan Keten, Rafael Gomez-Bombarelli, and Tommi Jaakkola. Forces are not enough: Benchmark and critical evaluation for machine learning force fields with molecular simulations. In AI for Science: Progress and Promises Workshop at Ne...

  25. [33]

    Scalable bayesian uncertainty quantifi- cation for neural network potentials: Promise and pitfalls

    Stephan Thaler, Gregor Doehner, and Julija Zavadlav. Scalable bayesian uncertainty quantifi- cation for neural network potentials: Promise and pitfalls. Journal of Chemical Theory and Computation, 19(14):4520–4532, 2023

  26. [34]

    Molecular dynamics simulations with quantum mechanics/molec- ular mechanics and adaptive neural networks

    Lin Shen and Weitao Yang. Molecular dynamics simulations with quantum mechanics/molec- ular mechanics and adaptive neural networks. Journal of Chemical Theory and Computation, 14(3):1442–1455, 2018

  27. [35]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09, 2009

  28. [36]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. arxiv, 2009

  29. [37]

    The cityscapes dataset for semantic urban scene understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognitio...

  30. [38]

    Pointer sentinel mixture models, 2016

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016

  31. [39]

    CCNet: Extracting high quality monolingual datasets from web crawl data

    Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. CCNet: Extracting high quality monolingual datasets from web crawl data. In Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Chris...

  32. [40]

    Unsupervised cross-lingual representation learning at scale

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wen- zek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoy- anov. Unsupervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Natalie Schlu...

  33. [41]

    Datasets: A community library for natural language processing

    Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Ca...

  34. [42]

    Torchvision: Pytorch’s computer vision library

    TorchVision maintainers and contributors. Torchvision: Pytorch’s computer vision library. https://github.com/pytorch/vision, 2016

  35. [43]

    Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development

    Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik. Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development. Proceedings of Neural Information Processin...

  36. [44]

    URL http://quantum-machine.org/datasets/

  37. [45]

    The molssi qcarchive project: An open-source platform to compute, organize, and share quantum chemistry data

    Daniel GA Smith, Doaa Altarawy, Lori A Burns, Matthew Welborn, Levi N Naden, Logan Ward, Sam Ellis, Benjamin P Pritchard, and T Daniel Crawford. The molssi qcarchive project: An open-source platform to compute, organize, and share quantum chemistry data. Wiley Interdisciplinar...

  38. [46]

    Colabfit exchange: Open-access datasets for data-driven interatomic potentials

    Joshua A Vita, Eric G Fuemmeler, Amit Gupta, Gregory P Wolfe, Alexander Quanming Tao, Ryan S Elliott, Stefano Martiniani, and Ellad B Tadmor. Colabfit exchange: Open-access datasets for data-driven interatomic potentials. The Journal of Chemical Physics, 159(15), 2023

  39. [47]

    Quantum chemical bench- mark energy and geometry database for molecular clusters and complex molecular systems (www

    Jan ˇRezáˇc, Petr Jure ˇcka, Kevin E Riley, Ji ˇrí ˇCern`y, Haydee Valdes, Krist `yna Pluháˇcková, Karel Berka, Tomáš ˇRezáˇc, Michal Pitoˇnák, Jiˇrí V ondrášek, et al. Quantum chemical bench- mark energy and geometry database for molecular clusters and complex molecular syste...

  40. [48]

    Cuby: An integrative framework for computational chemistry, 2016

    Jan ˇRezáˇc. Cuby: An integrative framework for computational chemistry, 2016

  41. [49]

    Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik

    Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W. Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik. Therapeutics data commons: Machine learning datasets and tasks for therapeutics. CoRR, abs/2102.09548, 2021. URL https: //arxiv.org/abs/2102.09548

  42. [50]

    Moleculenet: a benchmark for molecular machine learning

    Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530, 2018

  43. [51]

    Matthias Fey and Jan E. Lenssen. Fast graph representation learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019

  44. [52]

    Deep graph library: A graph-centric, highly-performant package for graph neural networks

    Minjie Wang, Da Zheng, Zihao Ye, Quan Gan, Mufei Li, Xiang Song, Jinjing Zhou, Chao Ma, Lingfan Yu, Yu Gai, et al. Deep graph library: A graph-centric, highly-performant package for graph neural networks. arXiv preprint arXiv:1909.01315, 2019

  45. [53]

    Torchdrug: A powerful and flexible machine learning platform for drug discovery, 2022

    Zhaocheng Zhu, Chence Shi, Zuobai Zhang, Shengchao Liu, Minghao Xu, Xinyu Yuan, Yangtian Zhang, Junkun Chen, Huiyu Cai, Jiarui Lu, Chang Ma, Runcheng Liu, Louis-Pascal Xhonneux, Meng Qu, and Jian Tang. Torchdrug: A powerful and flexible machine learning platform for drug disco...

  46. [54]

    M. Rupp, A. Tkatchenko, K.-R. Müller, and O. A. von Lilienfeld. Fast and accurate modeling of molecular atomization energies with machine learning. Physical Review Letters , 108: 058301, 2012

  47. [55]

    Electronic spectra from tddft and machine learning in chemical space

    Raghunathan Ramakrishnan, Mia Hartmann, Enrico Tapavicza, and O Anatole V on Lilienfeld. Electronic spectra from tddft and machine learning in chemical space. The Journal of chemical physics, 143(8), 2015

  48. [56]

    Quantum chemistry structures and properties of 134 kilo molecules

    Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole von Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules. Scientific Data, 1, 2014

  49. [58]

    Pubchemqc pm6: Data sets of 221 million molecules with optimized molecular geometries and electronic properties

    Maho Nakata, Tomomi Shimazaki, Masatomo Hashimoto, and Toshiyuki Maeda. Pubchemqc pm6: Data sets of 221 million molecules with optimized molecular geometries and electronic properties. Journal of Chemical Information and Modeling, 60(12):5891–5899, 2020

  50. [59]

    OGB-LSC: A large-scale challenge for machine learning on graphs

    Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. OGB-LSC: A large-scale challenge for machine learning on graphs. In Joaquin Vanschoren and Sai-Kit Yeung, editors, Proceedings of the Neural Information Processing Systems Track 13 on Datasets an...

  51. [60]

    Deepmd-kit: A deep learning package for many-body potential energy representation and molecular dynamics

    Han Wang, Linfeng Zhang, Jiequn Han, and E Weinan. Deepmd-kit: A deep learning package for many-body potential energy representation and molecular dynamics. Computer Physics Communications, 228:178–184, 2018

  52. [61]

    Dp-gen: A concurrent learning platform for the generation of reliable deep learning based potential energy models

    Yuzhi Zhang, Haidi Wang, Weijie Chen, Jinzhe Zeng, Linfeng Zhang, Han Wang, and E Weinan. Dp-gen: A concurrent learning platform for the generation of reliable deep learning based potential energy models. Computer Physics Communications, 253:107206, 2020

  53. [62]

    Pinn: A python library for building atomic neural networks of molecules and materials

    Yunqi Shao, Matti Hellstrom, Pavlin D Mitev, Lisanne Knijff, and Chao Zhang. Pinn: A python library for building atomic neural networks of molecules and materials. Journal of chemical information and modeling, 60(3):1184–1193, 2020

  54. [63]

    Elliott, and Ellad B

    Mingjian Wen, Yaser Afshar, Ryan S. Elliott, and Ellad B. Tadmor. KLIFF: A framework to develop physics-based and machine learning interatomic potentials. Computer Physics Communications, 272:108218, 2022. doi:10.1016/j.cpc.2021.108218

  55. [64]

    Jax md: a framework for differentiable physics

    Samuel Schoenholz and Ekin Dogus Cubuk. Jax md: a framework for differentiable physics. Advances in Neural Information Processing Systems, 33:11428–11441, 2020

  56. [65]

    Torchmd: A deep learning framework for molecular simulations

    Stefan Doerr, Maciej Majewski, Adrià Pérez, Andreas Kramer, Cecilia Clementi, Frank Noe, Toni Giorgino, and Gianni De Fabritiis. Torchmd: A deep learning framework for molecular simulations. Journal of chemical theory and computation, 17(4):2355–2363, 2021

  57. [66]

    Pelaez, Guillem Simeon, Raimondas Galvelis, Antonio Mirarchi, Peter Eastman, Stefan Doerr, Philipp Thölke, Thomas E

    Raul P. Pelaez, Guillem Simeon, Raimondas Galvelis, Antonio Mirarchi, Peter Eastman, Stefan Doerr, Philipp Thölke, Thomas E. Markland, and Gianni De Fabritiis. Torchmd-net 2.0: Fast neural network potentials for molecular simulations, 2024

  58. [67]

    Tensornet: Cartesian tensor representations for efficient learning of molecular potentials

    Guillem Simeon and Gianni De Fabritiis. Tensornet: Cartesian tensor representations for efficient learning of molecular potentials. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=BEHlPdBZ2e

  59. [68]

    Charron, Toni Giorgino, Brooke E

    Maciej Majewski, Adrià Pérez, Philipp Thölke, Stefan Doerr, Nicholas E. Charron, Toni Giorgino, Brooke E. Husic, Cecilia Clementi, Frank Noé, and Gianni De Fabritiis. Machine learning coarse-grained potentials of protein thermodynamics. Nature Communications, 14 (1), September...

  60. [69]

    Quantum-chemical insights from deep tensor neural networks

    Kristof T Schütt, Farhad Arbabzadah, Stefan Chmiela, Klaus R Müller, and Alexandre Tkatchenko. Quantum-chemical insights from deep tensor neural networks. Nature communi- cations, 8(1):13890, 2017

  61. [71]

    Schnet–a deep learning architecture for molecules and materials

    Kristof T Schütt, Huziel E Sauceda, P-J Kindermans, Alexandre Tkatchenko, and K-R Müller. Schnet–a deep learning architecture for molecules and materials. The Journal of Chemical Physics, 148(24), 2018

  62. [72]

    Schnetpack: A deep learning toolbox for atomistic systems

    KT Schutt, Pan Kessel, Michael Gastegger, KA Nicoli, Alexandre Tkatchenko, and K-R Muller. Schnetpack: A deep learning toolbox for atomistic systems. Journal of chemical theory and computation, 15(1):448–455, 2018

  63. [73]

    Thermal half-lives of azobenzene derivatives: Virtual screening based on intersystem crossing using a machine learning potential

    Simon Axelrod, Eugene Shakhnovich, and Rafael Gómez-Bombarelli. Thermal half-lives of azobenzene derivatives: Virtual screening based on intersystem crossing using a machine learning potential. ACS Central Science, 9(2):166–176, 2023. 14

  64. [74]

    Excited state non- adiabatic dynamics of large photoswitchable molecules using a chemically transferable ma- chine learning potential

    Simon Axelrod, Eugene Shakhnovich, and Rafael Gómez-Bombarelli. Excited state non- adiabatic dynamics of large photoswitchable molecules using a chemically transferable ma- chine learning potential. Nature communications, 13(1):3440, 2022

  65. [75]

    Mace: Higher order equivariant message passing neural networks for fast and accurate force fields

    Ilyes Batatia, David P Kovacs, Gregor Simm, Christoph Ortner, and Gábor Csányi. Mace: Higher order equivariant message passing neural networks for fast and accurate force fields. Advances in Neural Information Processing Systems, 35:11423–11436, 2022

  66. [76]

    The design space of e (3)-equivariant atom-centered interatomic potentials

    Ilyes Batatia, Simon Batzner, Dávid Péter Kovács, Albert Musaelian, Gregor NC Simm, Ralf Drautz, Christoph Ortner, Boris Kozinsky, and Gábor Csányi. The design space of e (3)-equivariant atom-centered interatomic potentials. arXiv preprint arXiv:2205.06643, 2022

  67. [77]

    A foundation model for atomistic materials chemistry

    Ilyes Batatia, Philipp Benner, Yuan Chiang, Alin M Elena, Dávid P Kovács, Janosh Riebesell, Xavier R Advincula, Mark Asta, William J Baldwin, Noam Bernstein, et al. A foundation model for atomistic materials chemistry. arXiv preprint arXiv:2401.00096, 2023

  68. [78]

    e3nn: Euclidean neural networks, 2022

    Mario Geiger and Tess Smidt. e3nn: Euclidean neural networks, 2022. URL https://arxiv. org/abs/2207.09453

  69. [79]

    Tobias Fink and Jean-Louis Reymond. Virtual exploration of the chemical universe up to 11 atoms of c, n, o, f: assembly of 26.4 million structures (110.9 million stereoisomers) and analysis for new ring systems, stereochemistry, physicochemical properties, compound classes, an...

  70. [80]

    Less is more: Sampling chemical space with active learning

    Justin S Smith, Ben Nebgen, Nicholas Lubbers, Olexandr Isayev, and Adrian E Roitberg. Less is more: Sampling chemical space with active learning. The Journal of chemical physics, 148 (24), 2018

  71. [81]

    The ani-1ccx and ani-1x data sets, coupled-cluster and density functional theory properties for molecules

    Justin S Smith, Roman Zubatyuk, Benjamin Nebgen, Nicholas Lubbers, Kipton Barros, Adrian E Roitberg, Olexandr Isayev, and Sergei Tretiak. The ani-1ccx and ani-1x data sets, coupled-cluster and density functional theory properties for molecules. Scientific data, 7(1): 134, 2020

  72. [82]

    Approaching coupled cluster accuracy with a general-purpose neural network potential through transfer learning

    Justin S Smith, Benjamin T Nebgen, Roman Zubatyuk, Nicholas Lubbers, Christian Devereux, Kipton Barros, Sergei Tretiak, Olexandr Isayev, and Adrian E Roitberg. Approaching coupled cluster accuracy with a general-purpose neural network potential through transfer learning. Natur...

  73. [83]

    Chembl: towards direct deposition of bioassay data

    David Mendez, Anna Gaulton, A Patrícia Bento, Jon Chambers, Marleen De Veij, Eloy Félix, María Paula Magariños, Juan F Mosquera, Prudence Mutowo, Michał Nowotka, et al. Chembl: towards direct deposition of bioassay data. Nucleic acids research, 47(D1):D930–D940, 2019

  74. [84]

    The s66x8 benchmark for noncovalent interactions revisited: explicitly correlated ab initio methods and density functional theory

    Brina Brauer, Manoj K Kesharwani, Sebastian Kozuch, and Jan ML Martin. The s66x8 benchmark for noncovalent interactions revisited: explicitly correlated ab initio methods and density functional theory. Physical Chemistry Chemical Physics, 18(31):20905–20925, 2016

  75. [85]

    Geom, energy-annotated molecular conforma- tions for property prediction and molecular generation

    Simon Axelrod and Rafael Gomez-Bombarelli. Geom, energy-annotated molecular conforma- tions for property prediction and molecular generation. Scientific Data, 9(1):185, 2022

  76. [86]

    Gfn2-xtb—an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions

    Christoph Bannwarth, Sebastian Ehlert, and Stefan Grimme. Gfn2-xtb—an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions. Journal of chemical theory and computatio...

  77. [87]

    Schnet: A continuous-filter convolutional neural network for modeling quantum interactions

    Kristof Schütt, Pieter-Jan Kindermans, Huziel Enoc Sauceda Felix, Stefan Chmiela, Alexandre Tkatchenko, and Klaus-Robert Müller. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. Advances in neural information processing systems, 30, 2017

  78. [88]

    Accurate global machine learning force fields for molecules with hundreds of atoms

    Stefan Chmiela, Valentin Vassilev-Galindo, Oliver T Unke, Adil Kabylda, Huziel E Sauceda, Alexandre Tkatchenko, and Klaus-Robert Müller. Accurate global machine learning force fields for molecules with hundreds of atoms. Science Advances, 9(2):eadf0873, 2023. 15

  79. [89]

    Molecule3d: A benchmark for predicting 3d geometries from molecular graphs

    Zhao Xu, Youzhi Luo, Xuan Zhang, Xinyi Xu, Yaochen Xie, Meng Liu, Kaleb Dickerson, Cheng Deng, Maho Nakata, and Shuiwang Ji. Molecule3d: A benchmark for predicting 3d geometries from molecular graphs. arXiv preprint arXiv:2110.01717, 2021

  80. [90]

    Multixc-qm9: Large dataset of molecular and reaction energies from multi-level quantum chemical methods

    Surajit Nandi, Tejs Vegge, and Arghya Bhowmik. Multixc-qm9: Large dataset of molecular and reaction energies from multi-level quantum chemical methods. Scientific Data, 10(1):783, 2023

  81. [91]

    Quantum chemistry structures and properties of 134 kilo molecules

    Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole V on Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules. Scientific data, 1(1):1–7, 2014

  82. [92]

    nabladft: Large-scale conformational energy and hamiltonian prediction benchmark and dataset

    Kuzma Khrabrov, Ilya Shenbin, Alexander Ryabov, Artem Tsypin, Alexander Telepov, Anton Alekseev, Alexander Grishin, Pavel Strashnov, Petr Zhilyaev, Sergey Nikolenko, et al. nabladft: Large-scale conformational energy and hamiltonian prediction benchmark and dataset. Physi- cal...

  83. [93]

    Orbnet denali: A machine learning potential for biological and organic chemistry with semi-empirical cost and dft accuracy

    Anders S Christensen, Sai Krishna Sirumalla, Zhuoran Qiao, Michael B O’Connor, Daniel GA Smith, Feizhi Ding, Peter J Bygrave, Animashree Anandkumar, Matthew Welborn, Frederick R Manby, et al. Orbnet denali: A machine learning potential for biological and organic chemistry with...

  84. [94]

    Pubchem 2023 update

    Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. Pubchem 2023 update. Nucleic acids research, 51(D1):D1373–D1380, 2023

  85. [95]

    Pubchemqc b3lyp/6-31g*//pm6 data set: The electronic structures of 86 million molecules using b3lyp/6-31g* calculations

    Maho Nakata and Toshiyuki Maeda. Pubchemqc b3lyp/6-31g*//pm6 data set: The electronic structures of 86 million molecules using b3lyp/6-31g* calculations. Journal of Chemical Information and Modeling, 63(18):5734–5754, 2023

  86. [96]

    Qm7-x, a comprehensive dataset of quantum- mechanical properties spanning the chemical space of small organic molecules

    Johannes Hoja, Leonardo Medrano Sandonas, Brian G Ernst, Alvaro Vazquez-Mayagoitia, Robert A DiStasio Jr, and Alexandre Tkatchenko. Qm7-x, a comprehensive dataset of quantum- mechanical properties spanning the chemical space of small organic molecules. Scientific data, 8(1):43, 2021

  87. [97]

    Qmugs, quantum mechanical properties of drug-like molecules

    Clemens Isert, Kenneth Atz, José Jiménez-Luna, and Gisbert Schneider. Qmugs, quantum mechanical properties of drug-like molecules. Scientific Data, 9(1):273, 2022

  88. [98]

    On the role of gradients for machine learning of molecular energies and forces

    Anders S Christensen and O Anatole V on Lilienfeld. On the role of gradients for machine learning of molecular energies and forces. Machine Learning: Science and Technology, 1(4): 045018, 2020

  89. [99]

    Machine learning of accurate energy-conserving molecular force fields

    Stefan Chmiela, Alexandre Tkatchenko, Huziel E Sauceda, Igor Poltavsky, Kristof T Schütt, and Klaus-Robert Müller. Machine learning of accurate energy-conserving molecular force fields. Science advances, 3(5):e1603015, 2017

  90. [100]

    Physnet: A neural network for predicting energies, forces, dipole moments, and partial charges

    Oliver T Unke and Markus Meuwly. Physnet: A neural network for predicting energies, forces, dipole moments, and partial charges. Journal of chemical theory and computation , 15(6): 3678–3693, 2019

  91. [101]

    Spice, a dataset of drug-like molecules and peptides for training machine learning potentials

    Peter Eastman, Pavan Kumar Behara, David L Dotson, Raimondas Galvelis, John E Herr, Josh T Horton, Yuezhi Mao, John D Chodera, Benjamin P Pritchard, Yuanqing Wang, et al. Spice, a dataset of drug-like molecules and peptides for training machine learning potentials. Scientific ...

  92. [102]

    Spice 2.0.1, April 2024

    Peter Eastman, Pavan Kumar Behara, David Dotson, Raimondas Galvelis, John Herr, Josh Horton, Yuezhi Mao, John Chodera, Benjamin Pritchard, Yuanqing Wang, Gianni De Fabritiis, and Thomas Markland. Spice 2.0.1, April 2024. URL https://doi.org/10.5281/zenodo. 10975225

  93. [103]

    tmqm dataset—quantum geometries and properties of 86k transition metal complexes

    David Balcells and Bastian Bjerkem Skjelstad. tmqm dataset—quantum geometries and properties of 86k transition metal complexes. Journal of chemical information and modeling, 60(12):6135–6146, 2020. 16

  94. [104]

    The cambridge structural database

    Colin R Groom, Ian J Bruno, Matthew P Lightfoot, and Suzanna C Ward. The cambridge structural database. Acta Crystallographica Section B: Structural Science, Crystal Engineering and Materials, 72(2):171–179, 2016

  95. [105]

    Transition1x- a dataset for building generalizable reactive machine learning potentials

    Mathias Schreiner, Arghya Bhowmik, Tejs Vegge, Jonas Busk, and Ole Winther. Transition1x- a dataset for building generalizable reactive machine learning potentials. Scientific Data, 9(1): 779, 2022

  96. [106]

    Systematic optimization of long-range corrected hybrid density functionals

    Jeng-Da Chai and Martin Head-Gordon. Systematic optimization of long-range corrected hybrid density functionals. The Journal of chemical physics, 128(8), 2008

  97. [107]

    Self-consistent molecular-orbital methods

    RHWJ Ditchfield, Warren J Hehre, and John A Pople. Self-consistent molecular-orbital methods. ix. an extended gaussian-type basis for molecular-orbital studies of organic molecules. The Journal of Chemical Physics, 54(2):724–728, 1971

  98. [108]

    Optimization methods for finding minimum energy paths

    Daniel Sheppard, Rye Terrell, and Graeme Henkelman. Optimization methods for finding minimum energy paths. The Journal of chemical physics, 128(13), 2008

  99. [109]

    Atlas of putative minima and low-lying energy networks of water clusters n= 3–25

    Avijit Rakshit, Pradipta Bandyopadhyay, Joseph P Heindel, and Sotiris S Xantheas. Atlas of putative minima and low-lying energy networks of water clusters n= 3–25. The Journal of chemical physics, 151(21), 2019

  100. [110]

    The flexible, polarizable, thole-type interaction potential for water (ttm2-f) revisited.The Journal of Physical Chemistry A, 110(11):4100–4106, 2006

    George S Fanourgakis and Sotiris S Xantheas. The flexible, polarizable, thole-type interaction potential for water (ttm2-f) revisited.The Journal of Physical Chemistry A, 110(11):4100–4106, 2006

  101. [111]

    Quantum chemical benchmark databases of gold-standard dimer interaction energies.Scientific data, 8(1):55, 2021

    Alexander G Donchev, Andrew G Taube, Elizabeth Decolvenaere, Cory Hargus, Robert T McGibbon, Ka-Hei Law, Brent A Gregersen, Je-Luen Li, Kim Palmo, Karthik Siva, et al. Quantum chemical benchmark databases of gold-standard dimer interaction energies.Scientific data, 8(1):55, 2021

  102. [112]

    S66: A well-balanced database of benchmark interaction energies relevant to biomolecular structures

    Jan Rezác, Kevin E Riley, and Pavel Hobza. S66: A well-balanced database of benchmark interaction energies relevant to biomolecular structures. Journal of chemical theory and computation, 7(8):2427–2438, 2011

  103. [113]

    Approaches for machine learning intermolecular interaction energies and application to energy components from symmetry adapted perturbation theory

    Derek P Metcalf, Alexios Koutsoukas, Steven A Spronk, Brian L Claus, Deborah A Loughney, Stephen R Johnson, Daniel L Cheney, and C David Sherrill. Approaches for machine learning intermolecular interaction energies and application to energy components from symmetry adapted per...

  104. [114]

    A quantum chemical interaction energy dataset for accurately modeling protein-ligand interactions

    Steven A Spronk, Zachary L Glick, Derek P Metcalf, C David Sherrill, and Daniel L Ch- eney. A quantum chemical interaction energy dataset for accurately modeling protein-ligand interactions. Scientific Data, 10(1):619, 2023

  105. [115]

    Learning local equivariant representations for large-scale atomistic dynamics

    Albert Musaelian, Simon Batzner, Anders Johansson, Lixin Sun, Cameron J Owen, Mordechai Kornbluth, and Boris Kozinsky. Learning local equivariant representations for large-scale atomistic dynamics. Nature Communications, 14(1):579, 2023

  106. [116]

    Torchmd-net: Equivariant transformers for neural network based molecular potentials

    Philipp Thölke and Gianni De Fabritiis. Torchmd-net: Equivariant transformers for neural network based molecular potentials. arXiv preprint arXiv:2202.02541, 2022

  107. [117]

    Big data meets quantum chemistry approximations: the δ-machine learning approach

    Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole V on Lilienfeld. Big data meets quantum chemistry approximations: the δ-machine learning approach. Journal of Chemical Theory and Computation, 11(5):2087–2096, 2015

  108. [118]

    Kingma and Jimmy Lei Ba

    Diederik P. Kingma and Jimmy Lei Ba. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, San Diego, CA, USA, May 7-9, 2015

  109. [119]

    Searching for activation functions

    Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017. 17 A Appendix A.1 Hyperparameters We employ the same hyperparameters for all models (summarized in table 3) and for all datasets throughout this work. Al...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.