REVIEW 3 major objections 5 minor 104 references
Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A kernel ridge regression protocol rewritten in terms of Structured Orthogonal Random Features turns Boltzmann conformer ensembles into trigonometric neural networks that predict experimental oxidation potentials and hydration energies to…
desk verdict A useful, carefully benchmarked SORF-based method for conformer-ensemble KRR with shipped code, but a stated Nf/Ninit inconsistency should be resolved before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Structured Orthogonal Random Features (SORF) approximation, which replaces the infinite-dimensional feature map of a Gaussian kernel with a finite feature vector whose entries are cosines of random projections computed through fast Walsh-Hadamard transforms. Because SORF features are additive, the paper composes them into layers: F for a single vector, S for sums over atoms or conformers, T to switch by nuclear charge, W for weighted sums over localized orbitals, N for normalization, and E for a 'mixed-extensive' layer that scales linearly with molecular size. Dot products of these layered features reproduce the kernel functions of Section 2.1, so the KRR prediction becomes $\vec{\alpha}\cdot\vec{z}(B)$, with coefficients found by an SVD-based solve of $(Z^T Z+\lambda I)\vec{\alpha}=Z^T\vec{q}$. The practical settings are $N_f=32768$ per F layer, $N_{\rm tr}=3$, and $N_{\rm conf}=32$ conformers per molecule, the latter two chosen by monitoring the convergence of the $Z^T Z$ matrix.
What would settle it
Train the best MSORF models (FCHL19, aSLATM) on LES and FreeSolv with $N_f$ doubled from 32768 to 65536 and compare test MAEs; if MAEs drop by more than the reported statistical error, the published learning curves are biased by the SORF approximation. Alternatively, on a small training subset, compare SORF-based predictions with exact Gaussian-kernel KRR on the same ensemble kernel; agreement to within the target accuracy would confirm convergence.
Extended reading notes
Core claim
The central claim is that a KRR model whose kernel acts on ensembles of conformers—a construction that would normally cost $O(N_{\rm conf}^2 N_a^2)$ per kernel evaluation—can be rewritten as a linear regression on layered SORF feature vectors, making both feature construction and prediction cost linear in molecular size. Applied to seven molecular representations, the resulting MSORF models reach mean absolute errors of 0.17–0.20 eV on the LES oxidation-potential dataset and 0.49–0.50 kcal/mol on FreeSolv hydration energies, which the author identifies as experimental accuracy and as comparable to state-of-the-art ML results, despite using conformers generated by a cheap force field. The paper further claims that the layered construction is general: global, local, element-switched, and orbital-based representations all fit into the same F/S/T/W/N/E layer scheme, with layer choices encoding whether the target property is intensive or extensive.
Load-bearing premise
The method's accuracy rests on the finite SORF approximation being converged: the feature count and other hyperparameters were set by checking the $Z^T Z$ matrix, not by checking final prediction errors, so a representation/property combination where the approximation is not converged would inherit an unquantified bias.
Editorial extensions
If this is right
- Learning curves for all tested representations drop systematically with training set size, reaching the 0.2 eV experimental-error threshold for oxidation potentials after a few hundred training molecules.
- The same trained protocol handles both intensive (oxidation potential) and extensive (hydration energy) quantities by swapping the layer combination, so the architecture encodes physical extensivity rather than relying on ad hoc feature engineering.
- Prediction cost is linear in molecule size rather than quadratic in atom counts, which matters when screening large electrolyte candidate molecules.
- Using only the lowest-energy conformer instead of a Boltzmann ensemble leaves MAE unchanged within statistical error for aSLATM, FCHL19, and CM, indicating that representation choice dominates over ensemble size for these datasets.
- The new self-consistent loss functions provide a principled way to optimize the many hyperparameters introduced by layered SORF features.
Reading between the lines
- If the finite SORF approximation is the bottleneck, doubling $N_f$ for the best representations should further lower or leave unchanged the reported MAEs; a scan of $N_f$ against final test MAE would quantify the residual bias that $Z^T Z$-convergence checks do not.
- Because the ensemble-vs-single-conformer difference is within noise, a simpler single-conformer MSORF could serve as a fast screening surrogate, while the ensemble form may become more valuable for floppy molecules or properties sensitive to population weights.
- The layered feature construction suggests a direct route to mixture properties by adding an extra summation layer over molecules, an extension the paper names as future work and that would connect to electrolyte-solvent mixture screening.
- The stoichiometry-dependent shift procedure, numerically unstable here for small training sets, may become useful with larger datasets, since it makes extensive-property predictions invariant to reference states.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MSORF, a protocol that rewrites kernel ridge regression for Boltzmann ensembles of conformers in terms of Structured Orthogonal Random Features (SORF), yielding trigonometric neural networks with predictions that scale linearly with molecular size. The method is tested on two experimental datasets relevant to battery electrolytes: oxidation potentials in acetonitrile (LES) and hydration free energies (FreeSolv), using a variety of molecular representations. The authors report MAEs close to estimated experimental errors after training on a few hundred molecules, and they also present self-consistent versions of the Huber and LogCosh loss functions used for hyperparameter optimization.
Significance. If the results hold, the paper provides a computationally efficient and representation-flexible framework for conformer-ensemble machine learning, with open-source code released in QML2. The empirical protocol is careful: four repeated random data splits, held-out test sets, and explicit checks of convergence with respect to some hyperparameters. The theoretical core is standard SORF kernel approximation and is sound. However, the significance is tempered by the paper's own finding that using only the minimum-energy conformer yields the same MAE within statistical error as the full ensemble, so the central 'conformer ensemble' motivation is not strongly supported by the numerical evidence; the main contribution is the layer construction and the loss functions.
major comments (3)
- [Section 3 and Eq. (13)] The statement 'By default we ran calculations with Nf = 32768 at all F layers' is inconsistent with the reported Ninit values for aSLATM (131072) and SOAP for LES (65536), because Eq. (13) defines Nf = Ninit * Nstack. The paper does not state Nstack for these cases. If Nstack = 1, the effective feature dimension is 131072 for aSLATM and 65536 for SOAP-LES, so the SORF approximation dimension is not held constant across representations. This confounds the cross-representation comparisons in Tables 2 and 3 and the conclusions about relative representation quality, and it makes the experimental setup not reproducible as written. Please specify Nstack and the actual Nf for every representation, and either match Nf across representations or analyze the impact of different Nf on the comparison.
- [Section 3] The convergence of Nconf (Morfeus conformer attempts) and Ntr (number of Hadamard products) is checked via the ZZT matrix, not via the final prediction MAEs. This leaves the SORF approximation error for the chosen Nf and Ninit unquantified for each representation and property. If the random-feature approximation is not converged for some combination, the reported learning curves inherit a bias that is not discussed; this is load-bearing for the claim that experimental accuracy is reached. Please provide at least one direct check of MAE versus Nf (or versus Nconf/Ntr) for a representative representation, or otherwise justify that the ZZT criterion is sufficient.
- [Eq. (8)] The definition of klin_FJK in Eq. (8) is missing a minus sign in the exponential. As written, it reads exp[ |v-v'|^2 / (2 sigma⊙sigma) ], which is not a kernel and would grow without bound for distant vectors. The subsequent Gaussian construction in Eq. (7) relies on klin_FJK being a positive-semidefinite kernel inside the normalization. Please correct the sign, presumably to exp[ -|v-v'|^2 / (2 sigma⊙sigma) ].
minor comments (5)
- [Throughout] There are numerous typographical errors, including 'Botlzmann' (Section 1), 'chosed' (Section 2.2), 'stochiometry' (Section 3), and 'ensembes' (Section 5). A careful proofread is needed.
- [Eq. (22)] The notation {X}_s=1^Nstack for concatenation is introduced in Eq. (13) but should be explicitly recalled at Eq. (22), where Nstack copies of X are concatenated; a brief clarification would improve readability.
- [Table 1] The layer-combination notation in Table 1 is very dense and hard to parse, especially with the multiple footnote symbols. Consider presenting a few explicit examples (e.g., for aSLATM with an intensive property) in the table or in the caption.
- [Section 4.2] The phrase "FreeSolv's 'default' experimental error of 0.6 kcal/mol" is vague; please specify the source of this estimate (e.g., the FreeSolv text notes or a specific reference).
- [Abstract and Section 4.3] The abstract emphasizes the use of Boltzmann ensembles, but Section 4.3 shows that using only the minimum-energy conformer yields MAEs equal within statistical error. This tension should be addressed explicitly in the abstract or the conclusions, even though the conclusions do mention the finding.
Circularity Check
No significant circularity: MSORF is a kernel-approximation construction validated on external experimental benchmarks with held-out hyperparameter optimization.
full rationale
MSORF is built by explicit algebraic rewriting of KRR in terms of SORF: Eq. (13) defines random Fourier features whose dot product approximates the Gaussian kernel via Eq. (14), and Eqs. (15)-(20) compose layers that reproduce klocal, kdlocal, and kFJK by construction. These are standard kernel-approximation identities, not fitted outputs. The benchmark targets are external: experimental oxidation potentials from Ref. [10] and FreeSolv hydration energies, with experimental-error thresholds taken from those external sources. Hyperparameters (sigma, lambda, loss smoothness) are optimized on held-out training subsets via leave-one-out error in Appendix A, and the reported MAEs are computed on separate test-set molecules, so the 'experimental accuracy' claim is not a refit of the target. The paper contains self-citations (FJK representation, Ref. [44]; QML2 implementation, Ref. [57]) and carries defaults such as Nf = 32768 from Ref. [21], but none of these is load-bearing for the central result: the best-performing representations (aSLATM, FCHL19, cMBDF) are external, and the SORF approximation theorem is cited from external work [17]. The paper also reports null results (ensemble vs. minimum-conformer differences within statistical error), which are inconsistent with a narrative where the outcome is forced by construction. Remaining concerns, such as the Nf/Ninit/Nstack bookkeeping for aSLATM and SOAP and the projection claim for SLATM, are reproducibility or correctness issues, not circularity. No prediction reduces by definition to a fitted parameter or to a self-citation chain.
Assumptions & free parameters
free parameters (10)
- sigma (kernel width) per F layer =
optimized per training subset via leave-one-out loss (values not tabulated in main text)
- lambda regularization =
optimized via BOSS and SLSQP for each sigma set
- sigma_mix =
initialized to 1, optimized
- Nf (random feature dimension per F layer) =
32768
- Ninit =
equal to Nf for most cases; SLATM 4096, SOAP 65536/32768, aSLATM 131072
- Ntr (number of Hadamard products) =
3
- Nconf (Morfeus conformer attempts) =
32
- rho_cut =
0.05
- gamma schedule and gradient steps =
gamma 0.5, 0.25, 0.125; gradient steps 0.5, 0.25, 0.125
- SOAP parameters (rcut, nmax, lmax) =
9.53, 8, 8
assumptions (6)
- standard math Random feature dot products approximate the target Gaussian kernel (Eq 14), with error controlled by Nstack and Ninit.
- domain assumption MMFF94 conformer energies and geometries at 298.15 K give Boltzmann weights adequate for the target properties.
- domain assumption The LES and FreeSolv experimental datasets are reliable, and their quoted experimental errors (0.2 eV and 0.6 kcal/mol) are valid accuracy targets.
- standard math Separate random feature blocks for different nuclear charges have zero expected cross terms, so STF reproduces kdlocal.
- ad hoc to paper The mixed-extensive E layer defines a valid kernel even though no closed-form kernel expression is given.
- domain assumption Huckel/6-311G orbitals with Boys localization and SMD solvation capture the orbital changes relevant to oxidation and hydration.
Cite this review
Pith. "Pith review of Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features." pith.science (2026). https://pith.science/paper/DBC2U5QU
@misc{pith2026250521247,
author = {Pith},
title = {Pith review of: Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features},
year = {2026},
howpublished = {\url{https://pith.science/paper/DBC2U5QU}},
note = {Machine review of arXiv:2505.21247}
}
read the original abstract
A computationally efficient protocol for machine learning in chemical space using Boltzmann ensembles of conformers as input is proposed; the method is based on rewriting Kernel Ridge Regression expressions in terms of Structured Orthogonal Random Features, yielding physics-motivated trigonometric neural networks. To evaluate the method's utility for materials discovery, we test it on experimental datasets of two quantities related to battery electrolyte design, namely oxidation potentials in acetonitrile and hydration energies, using several popular molecular representations to demonstrate the method's flexibility. Despite only using computationally cheap forcefield calculations for conformer generation, we observe systematic decrease of machine learning error with increased training set size in all cases, with experimental accuracy reached after training on hundreds of molecules and prediction errors being comparable to state-of-the-art machine learning approaches. We also present novel versions of Huber and LogCosh loss functions that made hyperparameter optimization of the new approach more convenient.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[21]
J, Faber, F
Browning, N. J, Faber, F. A, & von Lilienfeld , O. A. (2022) Gpu-accelerated approximate kernel method for quantum machine learning. J. Chem. Phys. 157 , 214801
2022
-
[1]
E, Oliynyk, A
Saal, J. E, Oliynyk, A. O, & Meredig, B. (2020) Machine learning in materials discovery: Confirmed predictions and their underlying approaches. Annu. Rev. Mater. Res. 50 , 49--69
2020
-
[2]
S, Barnard, A
Malica, C, Novoselov, K. S, Barnard, A. S, Kalinin, S. V, Spurgeon, S. R, Reuter, K, Alducin, M, Deringer, V. L, Csányi, G, Marzari, N, Huang, S, Cuniberti, G, Deng, Q, Ordejón, P, Cole, I, Choudhary, K, Hippalgaonkar, K, Zhu, R, von Lilienfeld, O. A, Hibat-Allah, M, Carrasquilla, J, Cisotto, G, Zancanaro, A, Wenzel, W, Ferrari, A. C, Ustyuzhanin, A, & Ro...
2025
-
[3]
P, Cs\' a nyi, G, & Ceriotti, M
De, S, Bart\' o k, A. P, Cs\' a nyi, G, & Ceriotti, M. (2016) Comparing molecules and solids across structural and alchemical space. Phys. Chem. Chem. Phys. 18 , 13754--13769
2016
-
[4]
Huang, B & von Lilienfeld , O. A. (2020) Quantum machine learning using atom-in-molecule-based fragments selected on the fly. Nat. Chem. 12
2020
-
[5]
P, Kov\' a cs, D
Darby, J. P, Kov\' a cs, D. P, Batatia, I, Caro, M. A, Hart, G. L. W, Ortner, C, & Cs\' a nyi, G. (2023) Tensor-reduced atomic density representations. Phys. Rev. Lett. 131 , 028001
2023
-
[6]
Welborn, M, Cheng, L, & Miller III , T. F. (2018) Transferability in machine learning for electronic structure via the molecular orbital basis. J. Chem. Theory Comput. 14
2018
-
[7]
R, & Corminboeuf, C
Fabrizio, A, Briling, K. R, & Corminboeuf, C. (2022) Spahm: the spectrum of approximated hamiltonian matrices representations. Digit. Discov. 1 , 286--294
2022
Show all 104 references
-
[8]
(2023) Matrix of orthogonalized atomic orbital coefficients representation for radicals and ions
Llenga, S & Gryn’ova, G. (2023) Matrix of orthogonalized atomic orbital coefficients representation for radicals and ions. J. Chem. Phys. 158 , 214116
2023
-
[9]
A, Hutchison, L, Huang, B, Gilmer, J, Schoenholz, S
Faber, F. A, Hutchison, L, Huang, B, Gilmer, J, Schoenholz, S. S, Dahl, G. E, Vinyals, O, Kearnes, S, Riley, P. F, & von Lilienfeld, O. A. (2017) Prediction errors of molecular machine learning models lower than hybrid dft error. J. Chem. Theory Comput. 13 , 5255--5264. PMID: 28926232
2017
-
[10]
(2024) Autonomous data extraction from peer reviewed literature for training machine learning models of oxidation potentials
Lee, S, Heinen, S, Khan, D, & Anatole von Lilienfeld, O. (2024) Autonomous data extraction from peer reviewed literature for training machine learning models of oxidation potentials. Mach. Learn.: Sci. Technol. 5 , 015052
2024
-
[11]
Vapnik, V. N. (1998) Statistical Learning Theory . (Wiley-Interscience)
1998
-
[12]
F, & von Lilienfeld , O
Weinreich, J, Lemm, D, von Rudorff , G. F, & von Lilienfeld , O. A. (2022) Ab initio machine learning of phase space averages. J. Chem. Phys. 157 , 024303
2022
-
[13]
(2020) Hydration free energies from kernel-based machine learning: Compound-database bias
Rauer, C & Bereau, T. (2020) Hydration free energies from kernel-based machine learning: Compound-database bias. J. Chem. Phys. 153 , 014101
2020
-
[14]
J, & von Lilienfeld , O
Weinreich, J, Browning, N. J, & von Lilienfeld , O. A. (2021) Machine learning of free energies in chemical compound space using ensemble representations: Reaching experimental uncertainty for solvation. J. Chem. Phys. 154 , 134113
2021
-
[15]
(2017) Molecular dynamics fingerprints (mdfp): Machine learning from md data to predict free-energy differences
Riniker, S. (2017) Molecular dynamics fingerprints (mdfp): Machine learning from md data to predict free-energy differences. J. Chem. Inf. Model. 57 , 726--741. PMID: 28368113
2017
-
[16]
(2007) Random Features for Large-Scale Kernel Machines eds
Rahimi, A & Recht, B. (2007) Random Features for Large-Scale Kernel Machines eds. Platt, J, Koller, D, Singer, Y, & Roweis, S. (Curran Associates, Inc.), Vol. 20
2007
-
[17]
X, Suresh, A
Yu, F. X, Suresh, A. T, Choromanski, K, Holtmann-Rice, D. N, & Kumar, S. (2016) Orthogonal Random Features
2016
-
[18]
Liu, F, Huang, X, Chen, Y, & Suykens, J. A. K. (2022) Random features for kernel approximation: A survey on algorithms, theory, and beyond. IEEE Trans. Pattern Anal. Mach. Intell. 44 , 7128--7148
2022
-
[19]
(2023) Simplex random features , ICML'23
Reid, I, Choromanski, K, Likhosherstov, V, & Weller, A. (2023) Simplex random features , ICML'23. (JMLR.org)
2023
-
[20]
(2024) Trigonometric Quadrature F ourier Features for Scalable G aussian Process Regression , Proceedings of Machine Learning Research eds
Li, K, Balakirsky, M, & Mak, S. (2024) Trigonometric Quadrature F ourier Features for Scalable G aussian Process Regression , Proceedings of Machine Learning Research eds. Dasgupta, S, Mandt, S, & Li, Y. (PMLR), Vol. 238, pp. 3484--3492
2024
-
[22]
J & Algazi, V
Fino, B. J & Algazi, V. R. (1976) Unified matrix treatment of the fast walsh-hadamard transform. IEEE on Trans. Comput. C-25 , 1142--1146
1976
-
[23]
V, Michiardi, P, & Filippone, M
Cutajar, K, Bonilla, E. V, Michiardi, P, & Filippone, M. (2017) Random Feature Expansions for Deep G aussian Processes , Proceedings of Machine Learning Research eds. Precup, D & Teh, Y. W. (PMLR), Vol. 70, pp. 884--893
2017
-
[24]
(2019) Deep kernel learning via random fourier features
Xie, J, Liu, F, Wang, K, & Huang, X. (2019) Deep kernel learning via random fourier features
2019
-
[25]
(2020) Walsh-Hadamard Variational Inference for Bayesian Deep Learning eds
Rossi, S, Marmin, S, & Filippone, M. (2020) Walsh-Hadamard Variational Inference for Bayesian Deep Learning eds. Larochelle, H, Ranzato, M, Hadsell, R, Balcan, M, & Lin, H. (Curran Associates, Inc.), Vol. 33, pp. 9674--9686
2020
-
[26]
(2022) On connecting deep trigonometric networks with deep G aussian processes: Covariance, expressivity, and neural tangent kernel
Lu, C.-K & Shafto, P. (2022) On connecting deep trigonometric networks with deep G aussian processes: Covariance, expressivity, and neural tangent kernel
2022
-
[27]
Damianou, A & Lawrence, N. D. (2013) Deep G aussian Processes , Proceedings of Machine Learning Research eds. Carvalho, C. M & Ravikumar, P. (PMLR, Scottsdale, Arizona, USA), Vol. 31, pp. 207--215
2013
-
[28]
(2009) Kernel Methods for Deep Learning eds
Cho, Y & Saul, L. (2009) Kernel Methods for Deep Learning eds. Bengio, Y, Schuurmans, D, Lafferty, J, Williams, C, & Culotta, A. (Curran Associates, Inc.), Vol. 22
2009
-
[29]
E, Leiter, K
Borodin, O, Olguin, M, Spear, C. E, Leiter, K. W, & Knap, J. (2015) Towards high throughput screening of electrochemical stability of battery electrolytes. Nanotechnol. 26 , 354003
2015
-
[30]
S, Qu, X, Jain, A, Ong, S
Cheng, L, Assary, R. S, Qu, X, Jain, A, Ong, S. P, Rajput, N. N, Persson, K, & Curtiss, L. A. (2015) Accelerating electrolyte discovery for energy storage with high-throughput screening. J. Phys. Chem. Lett. 6 , 283--291
2015
-
[31]
L & Guthrie, J
Mobley, D. L & Guthrie, J. P. (2014) Freesolv: a database of experimental and calculated hydration free energies, with input files. J. Comput. Aided Mol. Des. pp. 711--720
2014
-
[32]
Y, Loeffler, H
Duarte Ramos Matos , G, Kyu, D. Y, Loeffler, H. H, Chodera, J. D, Shirts, M. R, & Mobley, D. L. (2017) Approaches for calculating solvation free energies and enthalpies demonstrated with an update of the freesolv database. J. Chem. Eng. Data 62 , 1559--1569
2017
-
[33]
L, Shirts, M, Lim, N, Chodera, J, Beauchamp, K, & Lee-Ping
Mobley, D. L, Shirts, M, Lim, N, Chodera, J, Beauchamp, K, & Lee-Ping. (2018) Mobleylab/freesolv: Version 0.52
2018
-
[34]
(2022) A comprehensive survey of loss functions in machine learning
Wang, Q, Ma, Y, Zhao, K, & Tian, Y. (2022) A comprehensive survey of loss functions in machine learning. Ann. Data Sci. 9 , 187--212
2022
-
[35]
Huber, P. J. (1964) Robust Estimation of a Location Parameter . Ann. Math. Stat 35 , 73--101
1964
-
[36]
A, Xu, Y, & Zhang, H
Micchelli, C. A, Xu, Y, & Zhang, H. (2006) Universal kernels. J. Mach. Learn. Res. 7 , 2651--2667
2006
-
[37]
Rupp, M, Tkatchenko, A, M\"uller, K.-R, & von Lilienfeld , O. A. (2012) Fast and accurate modeling of molecular atomization energies with machine learning. Phys. Rev. Lett. 108 , 058301
2012
-
[38]
Ramakrishnan, R & von Lilienfeld , O. A. (2015) Many molecular properties from one kernel in chemical space. CHIMIA 69 , 182
2015
-
[39]
P, Kondor, R, & Cs\'anyi, G
Bart\'ok, A. P, Kondor, R, & Cs\'anyi, G. (2013) On representing chemical environments. Phys. Rev. B 87 , 184115
2013
-
[40]
A, Christensen, A
Faber, F. A, Christensen, A. S, Huang, B, & von Lilienfeld , O. A. (2018) Alchemical and structural distribution based representation for universal quantum machine learning. J. Chem. Phys. 148 , 241717
2018
-
[41]
S, Bratholm, L
Christensen, A. S, Bratholm, L. A, Faber, F. A, & von Lilienfeld , O. A. (2020) Fchl revisited: Faster and more accurate quantum machine learning. J. Chem. Phys. 152
2020
-
[42]
Khan, D, Heinen, S, & von Lilienfeld , O. A. (2023) Kernel based quantum machine learning at record rate: Many-body distribution functionals as compact representations. J. Chem. Phys. 159 , 034106
2023
-
[43]
Khan, D & von Lilienfeld , O. A. (2024) Generalized convolutional many body distribution functional representations. arXiv:2409.20471
2024 arXiv
-
[44]
Karandashev, K & von Lilienfeld , O. A. (2022) An orbital-based representation for accurate quantum machine learning. J. Chem. Phys. 156 , 114101
2022
-
[45]
B & Pedersen, M
Petersen, K. B & Pedersen, M. S. (2008) T he M atrix C ookbook. Version 20121115
2008
-
[46]
M, & Deane, C
Ebejer, J.-P, Morris, G. M, & Deane, C. M. (2012) Freely available conformer generation methods: How good are they? J. Chem. Inf. Model. 52 , 1146–1158
2012
-
[47]
https://kjelljorner.github.io/morfeus
(year?). https://kjelljorner.github.io/morfeus
-
[48]
Halgren, T. A. (1996) Merck molecular force field. i. basis, form, scope, parameterization, and performance of mmff94. J. Comput. Chem. 17 , 490--519
1996
-
[49]
Halgren, T. A. (1996) Merck molecular force field. ii. mmff94 van der waals and electrostatic parameters for intermolecular interactions. J. Comput. Chem. 17 , 520--552
1996
-
[50]
Halgren, T. A. (1996) Merck molecular force field. iii. molecular geometries and vibrational frequencies for mmff94. J. Comput. Chem. 17 , 553--586
1996
-
[51]
A & Nachbar, R
Halgren, T. A & Nachbar, R. B. (1996) Merck molecular force field. iv. conformational energies and geometries for mmff94. J. Comput. Chem. 17 , 587--615
1996
-
[52]
Halgren, T. A. (1996) Merck molecular force field. v. extension of mmff94 using experimental data, additional computational data, and empirical rules. J. Comput. Chem. 17 , 616--641
1996
-
[53]
Halgren, T. A. (1999) Mmff vi. mmff94s option for energy minimization studies. J. Comput. Chem. 20 , 720--729
1999
-
[54]
Halgren, T. A. (1999) Mmff vii. characterization of mmff94, mmff94s, and other widely available force fields for conformational energies and for intermolecular-interaction energies and geometries. J. Comput. Chem. 20 , 730--748
1999
-
[55]
(2014) Bringing the mmff force field to the rdkit: implementation and validation
Tosco, P, Stiefl, N, & Landrum, G. (2014) Bringing the mmff force field to the rdkit: implementation and validation. J. Cheminform. 6 , 37
2014
-
[56]
RDKit: Open-source cheminformatics
(year?). RDKit: Open-source cheminformatics. https://www.rdkit.org
-
[57]
QML2: Procedures for machine learning in chemistry. https://github.com/qml2code/qml2
Karandashev, K, Heinen, S, Khan, D, & Weinrech, J. (2024). "QML2: Procedures for machine learning in chemistry. https://github.com/qml2code/qml2"
2024
-
[58]
Himanen, L, J \"a ger, M. O. J, Morooka, E. V, Federici Canova, F, Ranawat, Y. S, Gao, D. Z, Rinke, P, & Foster, A. S. (2020) DScribe: Library of descriptors for machine learning in materials science . Comput. Phys. Commun. 247 , 106949
2020
-
[59]
V, J \"a ger, M
Laakso, J, Himanen, L, Homm, H, Morooka, E. V, J \"a ger, M. O, Todorovi\' c , M, & Rinke, P. (2023) Updates to the dscribe library: New descriptors and derivatives. J. Chem. Phys. 158
2023
-
[60]
(2024) Transfer learning for molecular property predictions from small datasets
Kirschbaum, T & Bande, A. (2024) Transfer learning for molecular property predictions from small datasets. AIP Adv. 14 , 105119
2024
-
[61]
(2015) Libcint: An efficient general integral library for gaussian basis functions
Sun, Q. (2015) Libcint: An efficient general integral library for gaussian basis functions. J. Comput. Chem. 36 , 1664--1671
2015
-
[62]
C, Blunt, N
Sun, Q, Berkelbach, T. C, Blunt, N. S, Booth, G. H, Guo, S, Li, Z, Liu, J, McClain, J. D, Sayfutyarova, E. R, Sharma, S, Wouters, S, & Chan, G. K.-L. (2018) Pyscf: the python-based simulations of chemistry framework. Wiley Interdiscip. Rev. Comput. Mol. Sci. 8 , e1340
2018
-
[63]
S, Bogdanov, N
Sun, Q, Zhang, X, Banerjee, S, Bao, P, Barbry, M, Blunt, N. S, Bogdanov, N. A, Booth, G. H, Chen, J, Cui, Z.-H, Eriksen, J. J, Gao, Y, Guo, S, Hermann, J, Hermes, M. R, Koh, K, Koval, P, Lehtola, S, Li, Z, Liu, J, Mardirossian, N, McClain, J. D, Motta, M, Mussard, B, Pham, H. ...
2020
-
[64]
J, Stewart, R
Hehre, W. J, Stewart, R. F, & Pople, J. A. (1969) Self‐consistent molecular‐orbital methods. i. use of gaussian expansions of slater‐type atomic orbitals. J. Chem. Phys. 51
1969
-
[65]
J, Ditchfield, R, Stewart, R
Hehre, W. J, Ditchfield, R, Stewart, R. F, & Pople, J. A. (1970) Self‐consistent molecular orbital methods. iv. use of gaussian expansions of slater‐type orbitals. extension to second‐row molecules. J. Chem. Phys. 52
1970
-
[66]
(2013) Intrinsic atomic orbitals: An unbiased bridge between quantum theory and chemical concepts
Knizia, G. (2013) Intrinsic atomic orbitals: An unbiased bridge between quantum theory and chemical concepts. J. Chem. Theory Comput. 9
2013
-
[67]
R, Calvino Alonso, Y, Fabrizio, A, & Corminboeuf, C
Briling, K. R, Calvino Alonso, Y, Fabrizio, A, & Corminboeuf, C. (2024) Spahm(a,b): Encoding the density information from guess hamiltonian in quantum machine learning representations. J. Chem. Theory Comput. 20 , 1108--1117. PMID: 38227222
2024
-
[68]
S, Seeger, R, & Pople, J
Krishnan, R, Binkley, J. S, Seeger, R, & Pople, J. A. (1980) Self‐consistent molecular orbital methods. xx. a basis set for correlated wave functions. J. Chem. Phys. 72 , 650--654
1980
-
[69]
D & Chandler, G
McLean, A. D & Chandler, G. S. (1980) Contracted gaussian basis sets for molecular calculations. i. second row atoms, z=11–18. J. Chem. Phys. 72 , 5639--5648
1980
-
[70]
N, Pross, A, McGrath , M
Glukhovtsev, M. N, Pross, A, McGrath , M. P, & Radom, L. (1995) Extension of gaussian‐2 (g2) theory to bromine‐ and iodine‐containing molecules: Use of effective core potentials. J. Chem. Phys. 103 , 1878--1885
1995
-
[71]
A, McGrath, M
Curtiss, L. A, McGrath, M. P, Blaudeau, J, Davis, N. E, Binning, Jr. , R. C, & Radom, L. (1995) Extension of gaussian‐2 theory to molecules containing third‐row atoms ga–kr. J. Chem. Phys. 103 , 6104--6113
1995
-
[72]
M & Boys, S
Foster, J. M & Boys, S. F. (1960) Canonical configurational interaction procedure. Rev. Mod. Phys. 32 , 300--302
1960
-
[73]
V, Cramer, C
Marenich, A. V, Cramer, C. J, & Truhlar, D. G. (2009) Universal solvation model based on solute electron density and on a continuum model of the solvent defined by the bulk dielectric constant and atomic surface tensions. J. Phys. Chem. B 113 , 6378--6396. PMID: 19366259
2009
-
[74]
A, M \"u ller, K.-R, & Tkatchenko, A
Hansen, K, Biegler, F, Ramakrishnan, R, Pronobis, W, von Lilienfeld , O. A, M \"u ller, K.-R, & Tkatchenko, A. (2015) Machine learning predictions of molecular properties: Accurate many-body potentials and nonlocality in chemical space. J. Phys. Chem. Lett. 6 , 2326--2331. PMI...
2015
-
[75]
Huang, B & von Lilienfeld , O. A. (2021) Ab initio machine learning in chemical compound space. Chem. Rev. 121 , 10001--10036. PMID: 34387476
2021
-
[76]
(2019) Molecule property prediction based on spatial graph embedding
Wang, X, Li, Z, Jiang, M, Wang, S, Zhang, S, & Wei, Z. (2019) Molecule property prediction based on spatial graph embedding. J. Chem. Inf. Model. 59 , 3817--3828. PMID: 31438677
2019
-
[77]
(2021) Graphical gaussian process regression model for aqueous solvation free energy prediction of organic molecules in redox flow batteries
Gao, P, Yang, X, Tang, Y.-H, Zheng, M, Andersen, A, Murugesan, V, Hollas, A, & Wang, W. (2021) Graphical gaussian process regression model for aqueous solvation free energy prediction of organic molecules in redox flow batteries. Phys. Chem. Chem. Phys. 23 , 24892--24904
2021
-
[78]
(2022) Molecular contrastive learning with chemical element knowledge graph
Fang, Y, Zhang, Q, Yang, H, Zhuang, X, Deng, S, Zhang, W, Qin, M, Chen, Z, Fan, X, & Chen, H. (2022) Molecular contrastive learning with chemical element knowledge graph. Proceedings of the AAAI Conference on Artificial Intelligence 36 , 3968--3976
2022
-
[79]
(2022) Accurate prediction of aqueous free solvation energies using 3d atomic feature-based graph neural network with transfer learning
Zhang, D, Xia, S, & Zhang, Y. (2022) Accurate prediction of aqueous free solvation energies using 3d atomic feature-based graph neural network with transfer learning. J. Chem. Inf. Model. 62 , 1840--1848
2022
-
[80]
L, & Izgorodina, E
Low, K, Coote, M. L, & Izgorodina, E. I. (2022) Explainable solvation free energy prediction combining graph neural networks with chemical intuition. J. Chem. Inf. Model. 62 , 5457--5470. PMID: 36317829
2022
-
[81]
(2023) Multitask deep ensemble prediction of molecular energetics in solution: From quantum mechanics to experimental properties
Xia, S, Zhang, D, & Zhang, Y. (2023) Multitask deep ensemble prediction of molecular energetics in solution: From quantum mechanics to experimental properties. J. Chem. Theory Comput. 19 , 659--668. PMID: 36607141
2023
-
[82]
K, Prakash, M
Yadav, A. K, Prakash, M. V, & Bandyopadhyay, P. (2025) Physics-based machine learning to predict hydration free energies for small molecules with a minimal number of descriptors: Interpretable and accurate. J. Phys. Chem. B. 129 , 1640--1647. PMID: 39841935
2025
-
[83]
(2025) A self-conformation-aware pre-training framework for molecular property prediction with substructure interpretability
Qiao, J, Jin, J, Wang, D, Teng, S, Zhang, J, Yang, X, Liu, Y, Wang, Y, Cui, L, Zou, Q, Su, R, & Wei, L. (2025) A self-conformation-aware pre-training framework for molecular property prediction with substructure interpretability. Nat. Commun. 16 , 4382
2025
-
[84]
(2021) A comprehensive survey on transfer learning
Zhuang, F, Qi, Z, Duan, K, Xi, D, Zhu, Y, Zhu, H, Xiong, H, & He, Q. (2021) A comprehensive survey on transfer learning. Proc. IEEE 109 , 43--76
2021
-
[85]
H & Green, W
Vermeire, F. H & Green, W. H. (2021) Transfer learning for solvation free energies: From quantum chemistry to experiments. Chem. Eng. J. 418 , 129307
2021
-
[86]
(2022) Improving molecular property prediction through a task similarity enhanced transfer learning strategy
Li, H, Zhao, X, Li, S, Wan, F, Zhao, D, & Zeng, J. (2022) Improving molecular property prediction through a task similarity enhanced transfer learning strategy. iScience 25
2022
-
[87]
(2025) Analyzing atomic interactions in molecules as learned by neural networks
Esders, M, Schnake, T, Lederer, J, Kabylda, A, Montavon, G, Tkatchenko, A, & M \"u ller, K.-R. (2025) Analyzing atomic interactions in molecules as learned by neural networks. J. Chem. Theory Comput. 21 , 714--729. PMID: 39792788
2025
-
[88]
(2022) Learning the laws of lithium-ion transport in electrolytes using symbolic regression
Flores, E, W\" o lke, C, Yan, P, Winter, M, Vegge, T, Cekic-Laskovic, I, & Bhowmik, A. (2022) Learning the laws of lithium-ion transport in electrolytes using symbolic regression. Digit. Discov. 1 , 440--447
2022
-
[89]
T, Joshi, S
Sose, A. T, Joshi, S. Y, Kunche, L. K, Wang, F, & Deshmukh, S. A. (2023) A review of recent advances and applications of machine learning in tribology. Phys. Chem. Chem. Phys. 25 , 4408--4443
2023
-
[90]
(2024) Calisol-23: Experimental electrolyte conductivity data for various li-salts and solvent combinations
de Blasio , P, Elsborg, J, Vegge, T, Flores, E, & Bhowmik, A. (2024) Calisol-23: Experimental electrolyte conductivity data for various li-salts and solvent combinations. Digit. Discov. 1 , 440--447
2024
-
[91]
R, Millman, K
Harris, C. R, Millman, K. J, van der Walt, S. J, Gommers, R, Virtanen, P, Cournapeau, D, Wieser, E, Taylor, J, Berg, S, Smith, N. J, Kern, R, Picus, M, Hoyer, S, van Kerkwijk, M. H, Brett, M, Haldane, A, del R \' i o, J. F, Wiebe, M, Peterson, P, G \' e rard-Marchant, P, Shepp...
2020
-
[92]
K, Pitrou, A, & Seibert, S
Lam, S. K, Pitrou, A, & Seibert, S. (2015) Numba: A llvm-based python jit compiler . pp. 1--6
2015
-
[93]
E, Haberland, M, Reddy, T, Cournapeau, D, Burovski, E, Peterson, P, Weckesser, W, Bright, J, van der Walt , S
Virtanen, P, Gommers, R, Oliphant, T. E, Haberland, M, Reddy, T, Cournapeau, D, Burovski, E, Peterson, P, Weckesser, W, Bright, J, van der Walt , S. J, Brett, M, Wilson, J, Millman, K. J, Mayorov, N, Nelson, A. R. J, Jones, E, Kern, R, Larson, E, Carey, C. J, Polat, \.I , Feng...
2020
-
[94]
C & Talbot, N
Cawley, G. C & Talbot, N. L. (2003) Efficient leave-one-out cross-validation of kernel fisher discriminant classifiers. Pattern Recognit. 36 , 2585--2592
2003
-
[95]
(2004) Convex Optimization
Boyd, S & Vandenberghe, L. (2004) Convex Optimization . (Cambridge University Press)
2004
-
[96]
(1987) Practical Methods of Optimization
Fletcher, R. (1987) Practical Methods of Optimization . (New York: Wiley)
1987
-
[97]
C & Nocedal, J
Liu, D. C & Nocedal, J. (2022) On the limited memory BFGS method for large scale optimization. Math. Program. 45 , 503--528
2022
-
[98]
U, Corander, J, & Rinke, P
Todorovi\' c , M, Gutmann, M. U, Corander, J, & Rinke, P. (2019) Bayesian inference of atomistic structure in functional materials. Npj Comput. Mater. 5 , 35
2019
-
[99]
Johnson, S. G. (2007) The NLopt nonlinear-optimization package (https://github.com/stevengj/nlopt)
2007
-
[100]
(1994) Algorithm 733: TOMP–Fortran modules for optimal control calculations
Kraft, D. (1994) Algorithm 733: TOMP–Fortran modules for optimal control calculations. ACM Trans. Math. Softw. 20 , 262–281
1994
-
[101]
Lewis, J. P. (1969) Homogeneous Functions and Euler's Theorem . (Palgrave Macmillan UK, London), pp. 297--303
1969
-
[102]
Solov'ev, V. N. (1983) On a criterion for convexity of a positive-homogeneous function. Mat. Sb. 46 , 285
1983
-
[103]
(2025) The elements of differentiable programming
Blondel, M & Roulet, V. (2025) The elements of differentiable programming
2025
-
[104]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence '...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.