Pith. sign in

REVIEW 3 major objections 5 minor 104 references

Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A kernel ridge regression protocol rewritten in terms of Structured Orthogonal Random Features turns Boltzmann conformer ensembles into trigonometric neural networks that predict experimental oxidation potentials and hydration energies to…

desk verdict A useful, carefully benchmarked SORF-based method for conformer-ensemble KRR with shipped code, but a stated Nf/Ninit inconsistency should be resolved before publication. read the letter →

arxiv 2505.21247 v2 pith:DBC2U5QU submitted 2025-05-27 physics.chem-ph physics.comp-ph

classification physics.chem-phphysics.comp-ph
keywords kernelridgeregressionstructuredorthogonalrandomfeaturesconformerensemblesBoltzmannweightingmolecularrepresentationsoxidationpotentialshydrationfreeenergybatteryelectrolytedesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes rewriting Kernel Ridge Regression over Boltzmann ensembles of conformers in terms of Structured Orthogonal Random Features, turning the kernel model into a trigonometric neural network whose feature vectors can be summed over atoms, orbitals, and conformers. The author claims this Multilevel SORF (MSORF) protocol, using only cheap MMFF94 force-field conformers, reaches experimental accuracy for predicting experimental oxidation potentials in acetonitrile (MAE 0.17–0.20 eV) and hydration free energies in FreeSolv (MAE 0.49–0.50 kcal/mol) after training on hundreds of molecules, with errors comparable to state-of-the-art ML methods. A secondary empirical claim is that for these properties, representing a molecule by a conformer ensemble rather than its lowest-energy conformer changes results only within statistical error, so how conformers are represented matters more than how many are used. The paper also introduces self-consistent versions of Huber and LogCosh loss functions that make hyperparameter optimization of the layered architecture more convenient.

What carries the argument

The load-bearing object is the Structured Orthogonal Random Features (SORF) approximation, which replaces the infinite-dimensional feature map of a Gaussian kernel with a finite feature vector whose entries are cosines of random projections computed through fast Walsh-Hadamard transforms. Because SORF features are additive, the paper composes them into layers: F for a single vector, S for sums over atoms or conformers, T to switch by nuclear charge, W for weighted sums over localized orbitals, N for normalization, and E for a 'mixed-extensive' layer that scales linearly with molecular size. Dot products of these layered features reproduce the kernel functions of Section 2.1, so the KRR prediction becomes $\vec{\alpha}\cdot\vec{z}(B)$, with coefficients found by an SVD-based solve of $(Z^T Z+\lambda I)\vec{\alpha}=Z^T\vec{q}$. The practical settings are $N_f=32768$ per F layer, $N_{\rm tr}=3$, and $N_{\rm conf}=32$ conformers per molecule, the latter two chosen by monitoring the convergence of the $Z^T Z$ matrix.

What would settle it

Train the best MSORF models (FCHL19, aSLATM) on LES and FreeSolv with $N_f$ doubled from 32768 to 65536 and compare test MAEs; if MAEs drop by more than the reported statistical error, the published learning curves are biased by the SORF approximation. Alternatively, on a small training subset, compare SORF-based predictions with exact Gaussian-kernel KRR on the same ensemble kernel; agreement to within the target accuracy would confirm convergence.

Watch

Extended reading notes

Core claim

The central claim is that a KRR model whose kernel acts on ensembles of conformers—a construction that would normally cost $O(N_{\rm conf}^2 N_a^2)$ per kernel evaluation—can be rewritten as a linear regression on layered SORF feature vectors, making both feature construction and prediction cost linear in molecular size. Applied to seven molecular representations, the resulting MSORF models reach mean absolute errors of 0.17–0.20 eV on the LES oxidation-potential dataset and 0.49–0.50 kcal/mol on FreeSolv hydration energies, which the author identifies as experimental accuracy and as comparable to state-of-the-art ML results, despite using conformers generated by a cheap force field. The paper further claims that the layered construction is general: global, local, element-switched, and orbital-based representations all fit into the same F/S/T/W/N/E layer scheme, with layer choices encoding whether the target property is intensive or extensive.

Load-bearing premise

The method's accuracy rests on the finite SORF approximation being converged: the feature count and other hyperparameters were set by checking the $Z^T Z$ matrix, not by checking final prediction errors, so a representation/property combination where the approximation is not converged would inherit an unquantified bias.

Editorial extensions

If this is right

  • Learning curves for all tested representations drop systematically with training set size, reaching the 0.2 eV experimental-error threshold for oxidation potentials after a few hundred training molecules.
  • The same trained protocol handles both intensive (oxidation potential) and extensive (hydration energy) quantities by swapping the layer combination, so the architecture encodes physical extensivity rather than relying on ad hoc feature engineering.
  • Prediction cost is linear in molecule size rather than quadratic in atom counts, which matters when screening large electrolyte candidate molecules.
  • Using only the lowest-energy conformer instead of a Boltzmann ensemble leaves MAE unchanged within statistical error for aSLATM, FCHL19, and CM, indicating that representation choice dominates over ensemble size for these datasets.
  • The new self-consistent loss functions provide a principled way to optimize the many hyperparameters introduced by layered SORF features.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the finite SORF approximation is the bottleneck, doubling $N_f$ for the best representations should further lower or leave unchanged the reported MAEs; a scan of $N_f$ against final test MAE would quantify the residual bias that $Z^T Z$-convergence checks do not.
  • Because the ensemble-vs-single-conformer difference is within noise, a simpler single-conformer MSORF could serve as a fast screening surrogate, while the ensemble form may become more valuable for floppy molecules or properties sensitive to population weights.
  • The layered feature construction suggests a direct route to mixture properties by adding an extra summation layer over molecules, an extension the paper names as future work and that would connect to electrolyte-solvent mixture screening.
  • The stoichiometry-dependent shift procedure, numerically unstable here for small training sets, may become useful with larger datasets, since it makes extensive-property predictions invariant to reference states.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces MSORF, a protocol that rewrites kernel ridge regression for Boltzmann ensembles of conformers in terms of Structured Orthogonal Random Features (SORF), yielding trigonometric neural networks with predictions that scale linearly with molecular size. The method is tested on two experimental datasets relevant to battery electrolytes: oxidation potentials in acetonitrile (LES) and hydration free energies (FreeSolv), using a variety of molecular representations. The authors report MAEs close to estimated experimental errors after training on a few hundred molecules, and they also present self-consistent versions of the Huber and LogCosh loss functions used for hyperparameter optimization.

Significance. If the results hold, the paper provides a computationally efficient and representation-flexible framework for conformer-ensemble machine learning, with open-source code released in QML2. The empirical protocol is careful: four repeated random data splits, held-out test sets, and explicit checks of convergence with respect to some hyperparameters. The theoretical core is standard SORF kernel approximation and is sound. However, the significance is tempered by the paper's own finding that using only the minimum-energy conformer yields the same MAE within statistical error as the full ensemble, so the central 'conformer ensemble' motivation is not strongly supported by the numerical evidence; the main contribution is the layer construction and the loss functions.

major comments (3)
  1. [Section 3 and Eq. (13)] The statement 'By default we ran calculations with Nf = 32768 at all F layers' is inconsistent with the reported Ninit values for aSLATM (131072) and SOAP for LES (65536), because Eq. (13) defines Nf = Ninit * Nstack. The paper does not state Nstack for these cases. If Nstack = 1, the effective feature dimension is 131072 for aSLATM and 65536 for SOAP-LES, so the SORF approximation dimension is not held constant across representations. This confounds the cross-representation comparisons in Tables 2 and 3 and the conclusions about relative representation quality, and it makes the experimental setup not reproducible as written. Please specify Nstack and the actual Nf for every representation, and either match Nf across representations or analyze the impact of different Nf on the comparison.
  2. [Section 3] The convergence of Nconf (Morfeus conformer attempts) and Ntr (number of Hadamard products) is checked via the ZZT matrix, not via the final prediction MAEs. This leaves the SORF approximation error for the chosen Nf and Ninit unquantified for each representation and property. If the random-feature approximation is not converged for some combination, the reported learning curves inherit a bias that is not discussed; this is load-bearing for the claim that experimental accuracy is reached. Please provide at least one direct check of MAE versus Nf (or versus Nconf/Ntr) for a representative representation, or otherwise justify that the ZZT criterion is sufficient.
  3. [Eq. (8)] The definition of klin_FJK in Eq. (8) is missing a minus sign in the exponential. As written, it reads exp[ |v-v'|^2 / (2 sigma⊙sigma) ], which is not a kernel and would grow without bound for distant vectors. The subsequent Gaussian construction in Eq. (7) relies on klin_FJK being a positive-semidefinite kernel inside the normalization. Please correct the sign, presumably to exp[ -|v-v'|^2 / (2 sigma⊙sigma) ].
minor comments (5)
  1. [Throughout] There are numerous typographical errors, including 'Botlzmann' (Section 1), 'chosed' (Section 2.2), 'stochiometry' (Section 3), and 'ensembes' (Section 5). A careful proofread is needed.
  2. [Eq. (22)] The notation {X}_s=1^Nstack for concatenation is introduced in Eq. (13) but should be explicitly recalled at Eq. (22), where Nstack copies of X are concatenated; a brief clarification would improve readability.
  3. [Table 1] The layer-combination notation in Table 1 is very dense and hard to parse, especially with the multiple footnote symbols. Consider presenting a few explicit examples (e.g., for aSLATM with an intensive property) in the table or in the caption.
  4. [Section 4.2] The phrase "FreeSolv's 'default' experimental error of 0.6 kcal/mol" is vague; please specify the source of this estimate (e.g., the FreeSolv text notes or a specific reference).
  5. [Abstract and Section 4.3] The abstract emphasizes the use of Boltzmann ensembles, but Section 4.3 shows that using only the minimum-energy conformer yields MAEs equal within statistical error. This tension should be addressed explicitly in the abstract or the conclusions, even though the conclusions do mention the finding.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: MSORF is a kernel-approximation construction validated on external experimental benchmarks with held-out hyperparameter optimization.

full rationale

MSORF is built by explicit algebraic rewriting of KRR in terms of SORF: Eq. (13) defines random Fourier features whose dot product approximates the Gaussian kernel via Eq. (14), and Eqs. (15)-(20) compose layers that reproduce klocal, kdlocal, and kFJK by construction. These are standard kernel-approximation identities, not fitted outputs. The benchmark targets are external: experimental oxidation potentials from Ref. [10] and FreeSolv hydration energies, with experimental-error thresholds taken from those external sources. Hyperparameters (sigma, lambda, loss smoothness) are optimized on held-out training subsets via leave-one-out error in Appendix A, and the reported MAEs are computed on separate test-set molecules, so the 'experimental accuracy' claim is not a refit of the target. The paper contains self-citations (FJK representation, Ref. [44]; QML2 implementation, Ref. [57]) and carries defaults such as Nf = 32768 from Ref. [21], but none of these is load-bearing for the central result: the best-performing representations (aSLATM, FCHL19, cMBDF) are external, and the SORF approximation theorem is cited from external work [17]. The paper also reports null results (ensemble vs. minimum-conformer differences within statistical error), which are inconsistent with a narrative where the outcome is forced by construction. Remaining concerns, such as the Nf/Ninit/Nstack bookkeeping for aSLATM and SOAP and the projection claim for SLATM, are reproducibility or correctness issues, not circularity. No prediction reduces by definition to a fitted parameter or to a self-citation chain.

Assumptions & free parameters 10 free parameters · 6 assumptions · 0 invented entities

The central claim rests mainly on standard SORF kernel approximation, on the adequacy of MMFF94 conformer Boltzmann weights, and on the reliability of the LES and FreeSolv experimental data. No new physical entities are introduced; the E layer and self-consistent loss functions are computational constructions whose kernel interpretation is partly asserted rather than fully derived.

free parameters (10)
  • sigma (kernel width) per F layer = optimized per training subset via leave-one-out loss (values not tabulated in main text)
    Controls the Gaussian kernel width in Eq (13); initial guess from RMS distance, refined by Algorithm 1 in Appendix A.
  • lambda regularization = optimized via BOSS and SLSQP for each sigma set
    Regularization in Eq (12); optimized on the training subset only.
  • sigma_mix = initialized to 1, optimized
    Interpolation parameter in the E layer Eq (22) controlling the mixed-extensive behavior.
  • Nf (random feature dimension per F layer) = 32768
    Chosen from Ref [21] rather than re-optimized or validated against final prediction errors.
  • Ninit = equal to Nf for most cases; SLATM 4096, SOAP 65536/32768, aSLATM 131072
    Padding dimension controlling SORF systematic bias; chosen to minimize bias, with projections for SLATM.
  • Ntr (number of Hadamard products) = 3
    Chosen by ZZT convergence scan over Ntr=1..7.
  • Nconf (Morfeus conformer attempts) = 32
    Convergence of the ZZT matrix for cMBDF on one training set; used throughout.
  • rho_cut = 0.05
    Cutoff for negligible Boltzmann weights and negligible FJK orbital contributions.
  • gamma schedule and gradient steps = gamma 0.5, 0.25, 0.125; gradient steps 0.5, 0.25, 0.125
    Annealing schedule for the self-consistent LogCosh loss in hyperparameter optimization.
  • SOAP parameters (rcut, nmax, lmax) = 9.53, 8, 8
    Taken from Ref [60] prior SOAP-KRR work on FreeSolv, not re-optimized.
assumptions (6)
  • standard math Random feature dot products approximate the target Gaussian kernel (Eq 14), with error controlled by Nstack and Ninit.
    Foundation of SORF from Ref [17]; used without explicit error bars in the reported predictions.
  • domain assumption MMFF94 conformer energies and geometries at 298.15 K give Boltzmann weights adequate for the target properties.
    Conformer generation in Section 3; the paper later shows ensemble versus min-conformer differences are within noise.
  • domain assumption The LES and FreeSolv experimental datasets are reliable, and their quoted experimental errors (0.2 eV and 0.6 kcal/mol) are valid accuracy targets.
    Used to define 'experimental accuracy' in Tables 2 through 5.
  • standard math Separate random feature blocks for different nuclear charges have zero expected cross terms, so STF reproduces kdlocal.
    Random phases uniform on [0,2pi) give zero mean feature blocks; implicit in Eq (16).
  • ad hoc to paper The mixed-extensive E layer defines a valid kernel even though no closed-form kernel expression is given.
    Eq (22) is a construction; the paper asserts sigma_mix interpolates between kernels but does not derive the kernel.
  • domain assumption Huckel/6-311G orbitals with Boys localization and SMD solvation capture the orbital changes relevant to oxidation and hydration.
    Replacement of STO-3G/IBO from Ref [44] described in Section 3; pair-determinant FJK results are worse, consistent with this being fragile.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features." pith.science (2026). https://pith.science/paper/DBC2U5QU

@misc{pith2026250521247,
  author       = {Pith},
  title        = {Pith review of: Kernel Ridge Regression for conformer ensembles made easy with Structured Orthogonal Random Features},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DBC2U5QU}},
  note         = {Machine review of arXiv:2505.21247}
}
read the original abstract

A computationally efficient protocol for machine learning in chemical space using Boltzmann ensembles of conformers as input is proposed; the method is based on rewriting Kernel Ridge Regression expressions in terms of Structured Orthogonal Random Features, yielding physics-motivated trigonometric neural networks. To evaluate the method's utility for materials discovery, we test it on experimental datasets of two quantities related to battery electrolyte design, namely oxidation potentials in acetonitrile and hydration energies, using several popular molecular representations to demonstrate the method's flexibility. Despite only using computationally cheap forcefield calculations for conformer generation, we observe systematic decrease of machine learning error with increased training set size in all cases, with experimental accuracy reached after training on hundreds of molecules and prediction errors being comparable to state-of-the-art machine learning approaches. We also present novel versions of Huber and LogCosh loss functions that made hyperparameter optimization of the new approach more convenient.

Figures

Figures reproduced from arXiv: 2505.21247 by the authors.

Figure 1
Figure 1. Mean Absolute Errors (MAE) of predictions of oxidation potentials Eox in acetonitrile from Ref. [10] obtained by combining MSORF with several molecular representations for different training set sizes Ntrain. “FJK (s. det.)” indicates FJK was used with a single Slater determinant. geometry, which is qualitatively correct, significantly more efficient. The remaining geometric representations are similar in terms of M… view at source ↗
Figure 2
Figure 2. Mean Absolute Error (MAE) of predictions of Ehyd of the FreeSolv dataset obtained by combining MSORF with several molecular representations for different training set sizes Ntrain. “S. det.” next to FJK entries indicates single Slater determinant was used, “m.-ext.” indicates using the “mixed-extensive” layer E (22). between MAEs observed in this work with different conformer representations. This implies that, at l… view at source ↗
Figure 3
Figure 3. Comparison of MAEs obtained by representing molecules either with a conformer ensemble or the lowest energy conformer (“min. conf.”) for different representations and training set sizes Ntrain. Eox and Ehyd are obtained from Ref. [10] and FreeSolv dataset [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figures from the paper (3 more)
Figure 1
Figure 1. Figure 1: Mean Absolute Errors (MAE) of predictions of oxidation potentials Eox in acetonitrile from Ref. [1] obtained by combining MSORF with several molecular representations for different training set sizes Ntrain. “FJK (s. det.)” indicates FJK was used with a single Slater d…
Figure 2
Figure 2. Figure 2: Mean Absolute Error (MAE) of predictions of Ehyd of the FreeSolv dataset obtained by combining MSORF with several molecular representations for different training set sizes Ntrain. “S. det.” next to FJK entries indicates single Slater determinant was used, “m.-ext.” in…
Figure 3
Figure 3. Figure 3: Comparison of MAEs obtained by representing molecules either with a conformer ensemble or the lowest energy conformer (“min. conf.”) for different representations and training set sizes Ntrain. Eox and Ehyd are obtained from Ref. [1] and FreeSolv dataset. Unlike [PITH…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

104 extracted references · 80 canonical work pages

  1. [21]

    J, Faber, F

    Browning, N. J, Faber, F. A, & von Lilienfeld , O. A. (2022) Gpu-accelerated approximate kernel method for quantum machine learning. J. Chem. Phys. 157 , 214801

  2. [1]

    E, Oliynyk, A

    Saal, J. E, Oliynyk, A. O, & Meredig, B. (2020) Machine learning in materials discovery: Confirmed predictions and their underlying approaches. Annu. Rev. Mater. Res. 50 , 49--69

  3. [2]

    S, Barnard, A

    Malica, C, Novoselov, K. S, Barnard, A. S, Kalinin, S. V, Spurgeon, S. R, Reuter, K, Alducin, M, Deringer, V. L, Csányi, G, Marzari, N, Huang, S, Cuniberti, G, Deng, Q, Ordejón, P, Cole, I, Choudhary, K, Hippalgaonkar, K, Zhu, R, von Lilienfeld, O. A, Hibat-Allah, M, Carrasquilla, J, Cisotto, G, Zancanaro, A, Wenzel, W, Ferrari, A. C, Ustyuzhanin, A, & Ro...

  4. [3]

    P, Cs\' a nyi, G, & Ceriotti, M

    De, S, Bart\' o k, A. P, Cs\' a nyi, G, & Ceriotti, M. (2016) Comparing molecules and solids across structural and alchemical space. Phys. Chem. Chem. Phys. 18 , 13754--13769

  5. [4]

    Huang, B & von Lilienfeld , O. A. (2020) Quantum machine learning using atom-in-molecule-based fragments selected on the fly. Nat. Chem. 12

  6. [5]

    P, Kov\' a cs, D

    Darby, J. P, Kov\' a cs, D. P, Batatia, I, Caro, M. A, Hart, G. L. W, Ortner, C, & Cs\' a nyi, G. (2023) Tensor-reduced atomic density representations. Phys. Rev. Lett. 131 , 028001

  7. [6]

    Welborn, M, Cheng, L, & Miller III , T. F. (2018) Transferability in machine learning for electronic structure via the molecular orbital basis. J. Chem. Theory Comput. 14

  8. [7]

    R, & Corminboeuf, C

    Fabrizio, A, Briling, K. R, & Corminboeuf, C. (2022) Spahm: the spectrum of approximated hamiltonian matrices representations. Digit. Discov. 1 , 286--294

Show all 104 references
  1. [8]

    (2023) Matrix of orthogonalized atomic orbital coefficients representation for radicals and ions

    Llenga, S & Gryn’ova, G. (2023) Matrix of orthogonalized atomic orbital coefficients representation for radicals and ions. J. Chem. Phys. 158 , 214116

  2. [9]

    A, Hutchison, L, Huang, B, Gilmer, J, Schoenholz, S

    Faber, F. A, Hutchison, L, Huang, B, Gilmer, J, Schoenholz, S. S, Dahl, G. E, Vinyals, O, Kearnes, S, Riley, P. F, & von Lilienfeld, O. A. (2017) Prediction errors of molecular machine learning models lower than hybrid dft error. J. Chem. Theory Comput. 13 , 5255--5264. PMID: 28926232

  3. [10]

    (2024) Autonomous data extraction from peer reviewed literature for training machine learning models of oxidation potentials

    Lee, S, Heinen, S, Khan, D, & Anatole von Lilienfeld, O. (2024) Autonomous data extraction from peer reviewed literature for training machine learning models of oxidation potentials. Mach. Learn.: Sci. Technol. 5 , 015052

  4. [11]

    Vapnik, V. N. (1998) Statistical Learning Theory . (Wiley-Interscience)

  5. [12]

    F, & von Lilienfeld , O

    Weinreich, J, Lemm, D, von Rudorff , G. F, & von Lilienfeld , O. A. (2022) Ab initio machine learning of phase space averages. J. Chem. Phys. 157 , 024303

  6. [13]

    (2020) Hydration free energies from kernel-based machine learning: Compound-database bias

    Rauer, C & Bereau, T. (2020) Hydration free energies from kernel-based machine learning: Compound-database bias. J. Chem. Phys. 153 , 014101

  7. [14]

    J, & von Lilienfeld , O

    Weinreich, J, Browning, N. J, & von Lilienfeld , O. A. (2021) Machine learning of free energies in chemical compound space using ensemble representations: Reaching experimental uncertainty for solvation. J. Chem. Phys. 154 , 134113

  8. [15]

    (2017) Molecular dynamics fingerprints (mdfp): Machine learning from md data to predict free-energy differences

    Riniker, S. (2017) Molecular dynamics fingerprints (mdfp): Machine learning from md data to predict free-energy differences. J. Chem. Inf. Model. 57 , 726--741. PMID: 28368113

  9. [16]

    (2007) Random Features for Large-Scale Kernel Machines eds

    Rahimi, A & Recht, B. (2007) Random Features for Large-Scale Kernel Machines eds. Platt, J, Koller, D, Singer, Y, & Roweis, S. (Curran Associates, Inc.), Vol. 20

  10. [17]

    X, Suresh, A

    Yu, F. X, Suresh, A. T, Choromanski, K, Holtmann-Rice, D. N, & Kumar, S. (2016) Orthogonal Random Features

  11. [18]

    Liu, F, Huang, X, Chen, Y, & Suykens, J. A. K. (2022) Random features for kernel approximation: A survey on algorithms, theory, and beyond. IEEE Trans. Pattern Anal. Mach. Intell. 44 , 7128--7148

  12. [19]

    (2023) Simplex random features , ICML'23

    Reid, I, Choromanski, K, Likhosherstov, V, & Weller, A. (2023) Simplex random features , ICML'23. (JMLR.org)

  13. [20]

    (2024) Trigonometric Quadrature F ourier Features for Scalable G aussian Process Regression , Proceedings of Machine Learning Research eds

    Li, K, Balakirsky, M, & Mak, S. (2024) Trigonometric Quadrature F ourier Features for Scalable G aussian Process Regression , Proceedings of Machine Learning Research eds. Dasgupta, S, Mandt, S, & Li, Y. (PMLR), Vol. 238, pp. 3484--3492

  14. [22]

    J & Algazi, V

    Fino, B. J & Algazi, V. R. (1976) Unified matrix treatment of the fast walsh-hadamard transform. IEEE on Trans. Comput. C-25 , 1142--1146

  15. [23]

    V, Michiardi, P, & Filippone, M

    Cutajar, K, Bonilla, E. V, Michiardi, P, & Filippone, M. (2017) Random Feature Expansions for Deep G aussian Processes , Proceedings of Machine Learning Research eds. Precup, D & Teh, Y. W. (PMLR), Vol. 70, pp. 884--893

  16. [24]

    (2019) Deep kernel learning via random fourier features

    Xie, J, Liu, F, Wang, K, & Huang, X. (2019) Deep kernel learning via random fourier features

  17. [25]

    (2020) Walsh-Hadamard Variational Inference for Bayesian Deep Learning eds

    Rossi, S, Marmin, S, & Filippone, M. (2020) Walsh-Hadamard Variational Inference for Bayesian Deep Learning eds. Larochelle, H, Ranzato, M, Hadsell, R, Balcan, M, & Lin, H. (Curran Associates, Inc.), Vol. 33, pp. 9674--9686

  18. [26]

    (2022) On connecting deep trigonometric networks with deep G aussian processes: Covariance, expressivity, and neural tangent kernel

    Lu, C.-K & Shafto, P. (2022) On connecting deep trigonometric networks with deep G aussian processes: Covariance, expressivity, and neural tangent kernel

  19. [27]

    Damianou, A & Lawrence, N. D. (2013) Deep G aussian Processes , Proceedings of Machine Learning Research eds. Carvalho, C. M & Ravikumar, P. (PMLR, Scottsdale, Arizona, USA), Vol. 31, pp. 207--215

  20. [28]

    (2009) Kernel Methods for Deep Learning eds

    Cho, Y & Saul, L. (2009) Kernel Methods for Deep Learning eds. Bengio, Y, Schuurmans, D, Lafferty, J, Williams, C, & Culotta, A. (Curran Associates, Inc.), Vol. 22

  21. [29]

    E, Leiter, K

    Borodin, O, Olguin, M, Spear, C. E, Leiter, K. W, & Knap, J. (2015) Towards high throughput screening of electrochemical stability of battery electrolytes. Nanotechnol. 26 , 354003

  22. [30]

    S, Qu, X, Jain, A, Ong, S

    Cheng, L, Assary, R. S, Qu, X, Jain, A, Ong, S. P, Rajput, N. N, Persson, K, & Curtiss, L. A. (2015) Accelerating electrolyte discovery for energy storage with high-throughput screening. J. Phys. Chem. Lett. 6 , 283--291

  23. [31]

    L & Guthrie, J

    Mobley, D. L & Guthrie, J. P. (2014) Freesolv: a database of experimental and calculated hydration free energies, with input files. J. Comput. Aided Mol. Des. pp. 711--720

  24. [32]

    Y, Loeffler, H

    Duarte Ramos Matos , G, Kyu, D. Y, Loeffler, H. H, Chodera, J. D, Shirts, M. R, & Mobley, D. L. (2017) Approaches for calculating solvation free energies and enthalpies demonstrated with an update of the freesolv database. J. Chem. Eng. Data 62 , 1559--1569

  25. [33]

    L, Shirts, M, Lim, N, Chodera, J, Beauchamp, K, & Lee-Ping

    Mobley, D. L, Shirts, M, Lim, N, Chodera, J, Beauchamp, K, & Lee-Ping. (2018) Mobleylab/freesolv: Version 0.52

  26. [34]

    (2022) A comprehensive survey of loss functions in machine learning

    Wang, Q, Ma, Y, Zhao, K, & Tian, Y. (2022) A comprehensive survey of loss functions in machine learning. Ann. Data Sci. 9 , 187--212

  27. [35]

    Huber, P. J. (1964) Robust Estimation of a Location Parameter . Ann. Math. Stat 35 , 73--101

  28. [36]

    A, Xu, Y, & Zhang, H

    Micchelli, C. A, Xu, Y, & Zhang, H. (2006) Universal kernels. J. Mach. Learn. Res. 7 , 2651--2667

  29. [37]

    Rupp, M, Tkatchenko, A, M\"uller, K.-R, & von Lilienfeld , O. A. (2012) Fast and accurate modeling of molecular atomization energies with machine learning. Phys. Rev. Lett. 108 , 058301

  30. [38]

    Ramakrishnan, R & von Lilienfeld , O. A. (2015) Many molecular properties from one kernel in chemical space. CHIMIA 69 , 182

  31. [39]

    P, Kondor, R, & Cs\'anyi, G

    Bart\'ok, A. P, Kondor, R, & Cs\'anyi, G. (2013) On representing chemical environments. Phys. Rev. B 87 , 184115

  32. [40]

    A, Christensen, A

    Faber, F. A, Christensen, A. S, Huang, B, & von Lilienfeld , O. A. (2018) Alchemical and structural distribution based representation for universal quantum machine learning. J. Chem. Phys. 148 , 241717

  33. [41]

    S, Bratholm, L

    Christensen, A. S, Bratholm, L. A, Faber, F. A, & von Lilienfeld , O. A. (2020) Fchl revisited: Faster and more accurate quantum machine learning. J. Chem. Phys. 152

  34. [42]

    Khan, D, Heinen, S, & von Lilienfeld , O. A. (2023) Kernel based quantum machine learning at record rate: Many-body distribution functionals as compact representations. J. Chem. Phys. 159 , 034106

  35. [43]

    Khan, D & von Lilienfeld , O. A. (2024) Generalized convolutional many body distribution functional representations. arXiv:2409.20471

  36. [44]

    Karandashev, K & von Lilienfeld , O. A. (2022) An orbital-based representation for accurate quantum machine learning. J. Chem. Phys. 156 , 114101

  37. [45]

    B & Pedersen, M

    Petersen, K. B & Pedersen, M. S. (2008) T he M atrix C ookbook. Version 20121115

  38. [46]

    M, & Deane, C

    Ebejer, J.-P, Morris, G. M, & Deane, C. M. (2012) Freely available conformer generation methods: How good are they? J. Chem. Inf. Model. 52 , 1146–1158

  39. [47]

    https://kjelljorner.github.io/morfeus

    (year?). https://kjelljorner.github.io/morfeus

  40. [48]

    Halgren, T. A. (1996) Merck molecular force field. i. basis, form, scope, parameterization, and performance of mmff94. J. Comput. Chem. 17 , 490--519

  41. [49]

    Halgren, T. A. (1996) Merck molecular force field. ii. mmff94 van der waals and electrostatic parameters for intermolecular interactions. J. Comput. Chem. 17 , 520--552

  42. [50]

    Halgren, T. A. (1996) Merck molecular force field. iii. molecular geometries and vibrational frequencies for mmff94. J. Comput. Chem. 17 , 553--586

  43. [51]

    A & Nachbar, R

    Halgren, T. A & Nachbar, R. B. (1996) Merck molecular force field. iv. conformational energies and geometries for mmff94. J. Comput. Chem. 17 , 587--615

  44. [52]

    Halgren, T. A. (1996) Merck molecular force field. v. extension of mmff94 using experimental data, additional computational data, and empirical rules. J. Comput. Chem. 17 , 616--641

  45. [53]

    Halgren, T. A. (1999) Mmff vi. mmff94s option for energy minimization studies. J. Comput. Chem. 20 , 720--729

  46. [54]

    Halgren, T. A. (1999) Mmff vii. characterization of mmff94, mmff94s, and other widely available force fields for conformational energies and for intermolecular-interaction energies and geometries. J. Comput. Chem. 20 , 730--748

  47. [55]

    (2014) Bringing the mmff force field to the rdkit: implementation and validation

    Tosco, P, Stiefl, N, & Landrum, G. (2014) Bringing the mmff force field to the rdkit: implementation and validation. J. Cheminform. 6 , 37

  48. [56]

    RDKit: Open-source cheminformatics

    (year?). RDKit: Open-source cheminformatics. https://www.rdkit.org

  49. [57]

    QML2: Procedures for machine learning in chemistry. https://github.com/qml2code/qml2

    Karandashev, K, Heinen, S, Khan, D, & Weinrech, J. (2024). "QML2: Procedures for machine learning in chemistry. https://github.com/qml2code/qml2"

  50. [58]

    Himanen, L, J \"a ger, M. O. J, Morooka, E. V, Federici Canova, F, Ranawat, Y. S, Gao, D. Z, Rinke, P, & Foster, A. S. (2020) DScribe: Library of descriptors for machine learning in materials science . Comput. Phys. Commun. 247 , 106949

  51. [59]

    V, J \"a ger, M

    Laakso, J, Himanen, L, Homm, H, Morooka, E. V, J \"a ger, M. O, Todorovi\' c , M, & Rinke, P. (2023) Updates to the dscribe library: New descriptors and derivatives. J. Chem. Phys. 158

  52. [60]

    (2024) Transfer learning for molecular property predictions from small datasets

    Kirschbaum, T & Bande, A. (2024) Transfer learning for molecular property predictions from small datasets. AIP Adv. 14 , 105119

  53. [61]

    (2015) Libcint: An efficient general integral library for gaussian basis functions

    Sun, Q. (2015) Libcint: An efficient general integral library for gaussian basis functions. J. Comput. Chem. 36 , 1664--1671

  54. [62]

    C, Blunt, N

    Sun, Q, Berkelbach, T. C, Blunt, N. S, Booth, G. H, Guo, S, Li, Z, Liu, J, McClain, J. D, Sayfutyarova, E. R, Sharma, S, Wouters, S, & Chan, G. K.-L. (2018) Pyscf: the python-based simulations of chemistry framework. Wiley Interdiscip. Rev. Comput. Mol. Sci. 8 , e1340

  55. [63]

    S, Bogdanov, N

    Sun, Q, Zhang, X, Banerjee, S, Bao, P, Barbry, M, Blunt, N. S, Bogdanov, N. A, Booth, G. H, Chen, J, Cui, Z.-H, Eriksen, J. J, Gao, Y, Guo, S, Hermann, J, Hermes, M. R, Koh, K, Koval, P, Lehtola, S, Li, Z, Liu, J, Mardirossian, N, McClain, J. D, Motta, M, Mussard, B, Pham, H. ...

  56. [64]

    J, Stewart, R

    Hehre, W. J, Stewart, R. F, & Pople, J. A. (1969) Self‐consistent molecular‐orbital methods. i. use of gaussian expansions of slater‐type atomic orbitals. J. Chem. Phys. 51

  57. [65]

    J, Ditchfield, R, Stewart, R

    Hehre, W. J, Ditchfield, R, Stewart, R. F, & Pople, J. A. (1970) Self‐consistent molecular orbital methods. iv. use of gaussian expansions of slater‐type orbitals. extension to second‐row molecules. J. Chem. Phys. 52

  58. [66]

    (2013) Intrinsic atomic orbitals: An unbiased bridge between quantum theory and chemical concepts

    Knizia, G. (2013) Intrinsic atomic orbitals: An unbiased bridge between quantum theory and chemical concepts. J. Chem. Theory Comput. 9

  59. [67]

    R, Calvino Alonso, Y, Fabrizio, A, & Corminboeuf, C

    Briling, K. R, Calvino Alonso, Y, Fabrizio, A, & Corminboeuf, C. (2024) Spahm(a,b): Encoding the density information from guess hamiltonian in quantum machine learning representations. J. Chem. Theory Comput. 20 , 1108--1117. PMID: 38227222

  60. [68]

    S, Seeger, R, & Pople, J

    Krishnan, R, Binkley, J. S, Seeger, R, & Pople, J. A. (1980) Self‐consistent molecular orbital methods. xx. a basis set for correlated wave functions. J. Chem. Phys. 72 , 650--654

  61. [69]

    D & Chandler, G

    McLean, A. D & Chandler, G. S. (1980) Contracted gaussian basis sets for molecular calculations. i. second row atoms, z=11–18. J. Chem. Phys. 72 , 5639--5648

  62. [70]

    N, Pross, A, McGrath , M

    Glukhovtsev, M. N, Pross, A, McGrath , M. P, & Radom, L. (1995) Extension of gaussian‐2 (g2) theory to bromine‐ and iodine‐containing molecules: Use of effective core potentials. J. Chem. Phys. 103 , 1878--1885

  63. [71]

    A, McGrath, M

    Curtiss, L. A, McGrath, M. P, Blaudeau, J, Davis, N. E, Binning, Jr. , R. C, & Radom, L. (1995) Extension of gaussian‐2 theory to molecules containing third‐row atoms ga–kr. J. Chem. Phys. 103 , 6104--6113

  64. [72]

    M & Boys, S

    Foster, J. M & Boys, S. F. (1960) Canonical configurational interaction procedure. Rev. Mod. Phys. 32 , 300--302

  65. [73]

    V, Cramer, C

    Marenich, A. V, Cramer, C. J, & Truhlar, D. G. (2009) Universal solvation model based on solute electron density and on a continuum model of the solvent defined by the bulk dielectric constant and atomic surface tensions. J. Phys. Chem. B 113 , 6378--6396. PMID: 19366259

  66. [74]

    A, M \"u ller, K.-R, & Tkatchenko, A

    Hansen, K, Biegler, F, Ramakrishnan, R, Pronobis, W, von Lilienfeld , O. A, M \"u ller, K.-R, & Tkatchenko, A. (2015) Machine learning predictions of molecular properties: Accurate many-body potentials and nonlocality in chemical space. J. Phys. Chem. Lett. 6 , 2326--2331. PMI...

  67. [75]

    Huang, B & von Lilienfeld , O. A. (2021) Ab initio machine learning in chemical compound space. Chem. Rev. 121 , 10001--10036. PMID: 34387476

  68. [76]

    (2019) Molecule property prediction based on spatial graph embedding

    Wang, X, Li, Z, Jiang, M, Wang, S, Zhang, S, & Wei, Z. (2019) Molecule property prediction based on spatial graph embedding. J. Chem. Inf. Model. 59 , 3817--3828. PMID: 31438677

  69. [77]

    (2021) Graphical gaussian process regression model for aqueous solvation free energy prediction of organic molecules in redox flow batteries

    Gao, P, Yang, X, Tang, Y.-H, Zheng, M, Andersen, A, Murugesan, V, Hollas, A, & Wang, W. (2021) Graphical gaussian process regression model for aqueous solvation free energy prediction of organic molecules in redox flow batteries. Phys. Chem. Chem. Phys. 23 , 24892--24904

  70. [78]

    (2022) Molecular contrastive learning with chemical element knowledge graph

    Fang, Y, Zhang, Q, Yang, H, Zhuang, X, Deng, S, Zhang, W, Qin, M, Chen, Z, Fan, X, & Chen, H. (2022) Molecular contrastive learning with chemical element knowledge graph. Proceedings of the AAAI Conference on Artificial Intelligence 36 , 3968--3976

  71. [79]

    (2022) Accurate prediction of aqueous free solvation energies using 3d atomic feature-based graph neural network with transfer learning

    Zhang, D, Xia, S, & Zhang, Y. (2022) Accurate prediction of aqueous free solvation energies using 3d atomic feature-based graph neural network with transfer learning. J. Chem. Inf. Model. 62 , 1840--1848

  72. [80]

    L, & Izgorodina, E

    Low, K, Coote, M. L, & Izgorodina, E. I. (2022) Explainable solvation free energy prediction combining graph neural networks with chemical intuition. J. Chem. Inf. Model. 62 , 5457--5470. PMID: 36317829

  73. [81]

    (2023) Multitask deep ensemble prediction of molecular energetics in solution: From quantum mechanics to experimental properties

    Xia, S, Zhang, D, & Zhang, Y. (2023) Multitask deep ensemble prediction of molecular energetics in solution: From quantum mechanics to experimental properties. J. Chem. Theory Comput. 19 , 659--668. PMID: 36607141

  74. [82]

    K, Prakash, M

    Yadav, A. K, Prakash, M. V, & Bandyopadhyay, P. (2025) Physics-based machine learning to predict hydration free energies for small molecules with a minimal number of descriptors: Interpretable and accurate. J. Phys. Chem. B. 129 , 1640--1647. PMID: 39841935

  75. [83]

    (2025) A self-conformation-aware pre-training framework for molecular property prediction with substructure interpretability

    Qiao, J, Jin, J, Wang, D, Teng, S, Zhang, J, Yang, X, Liu, Y, Wang, Y, Cui, L, Zou, Q, Su, R, & Wei, L. (2025) A self-conformation-aware pre-training framework for molecular property prediction with substructure interpretability. Nat. Commun. 16 , 4382

  76. [84]

    (2021) A comprehensive survey on transfer learning

    Zhuang, F, Qi, Z, Duan, K, Xi, D, Zhu, Y, Zhu, H, Xiong, H, & He, Q. (2021) A comprehensive survey on transfer learning. Proc. IEEE 109 , 43--76

  77. [85]

    H & Green, W

    Vermeire, F. H & Green, W. H. (2021) Transfer learning for solvation free energies: From quantum chemistry to experiments. Chem. Eng. J. 418 , 129307

  78. [86]

    (2022) Improving molecular property prediction through a task similarity enhanced transfer learning strategy

    Li, H, Zhao, X, Li, S, Wan, F, Zhao, D, & Zeng, J. (2022) Improving molecular property prediction through a task similarity enhanced transfer learning strategy. iScience 25

  79. [87]

    (2025) Analyzing atomic interactions in molecules as learned by neural networks

    Esders, M, Schnake, T, Lederer, J, Kabylda, A, Montavon, G, Tkatchenko, A, & M \"u ller, K.-R. (2025) Analyzing atomic interactions in molecules as learned by neural networks. J. Chem. Theory Comput. 21 , 714--729. PMID: 39792788

  80. [88]

    (2022) Learning the laws of lithium-ion transport in electrolytes using symbolic regression

    Flores, E, W\" o lke, C, Yan, P, Winter, M, Vegge, T, Cekic-Laskovic, I, & Bhowmik, A. (2022) Learning the laws of lithium-ion transport in electrolytes using symbolic regression. Digit. Discov. 1 , 440--447

  81. [89]

    T, Joshi, S

    Sose, A. T, Joshi, S. Y, Kunche, L. K, Wang, F, & Deshmukh, S. A. (2023) A review of recent advances and applications of machine learning in tribology. Phys. Chem. Chem. Phys. 25 , 4408--4443

  82. [90]

    (2024) Calisol-23: Experimental electrolyte conductivity data for various li-salts and solvent combinations

    de Blasio , P, Elsborg, J, Vegge, T, Flores, E, & Bhowmik, A. (2024) Calisol-23: Experimental electrolyte conductivity data for various li-salts and solvent combinations. Digit. Discov. 1 , 440--447

  83. [91]

    R, Millman, K

    Harris, C. R, Millman, K. J, van der Walt, S. J, Gommers, R, Virtanen, P, Cournapeau, D, Wieser, E, Taylor, J, Berg, S, Smith, N. J, Kern, R, Picus, M, Hoyer, S, van Kerkwijk, M. H, Brett, M, Haldane, A, del R \' i o, J. F, Wiebe, M, Peterson, P, G \' e rard-Marchant, P, Shepp...

  84. [92]

    K, Pitrou, A, & Seibert, S

    Lam, S. K, Pitrou, A, & Seibert, S. (2015) Numba: A llvm-based python jit compiler . pp. 1--6

  85. [93]

    E, Haberland, M, Reddy, T, Cournapeau, D, Burovski, E, Peterson, P, Weckesser, W, Bright, J, van der Walt , S

    Virtanen, P, Gommers, R, Oliphant, T. E, Haberland, M, Reddy, T, Cournapeau, D, Burovski, E, Peterson, P, Weckesser, W, Bright, J, van der Walt , S. J, Brett, M, Wilson, J, Millman, K. J, Mayorov, N, Nelson, A. R. J, Jones, E, Kern, R, Larson, E, Carey, C. J, Polat, \.I , Feng...

  86. [94]

    C & Talbot, N

    Cawley, G. C & Talbot, N. L. (2003) Efficient leave-one-out cross-validation of kernel fisher discriminant classifiers. Pattern Recognit. 36 , 2585--2592

  87. [95]

    (2004) Convex Optimization

    Boyd, S & Vandenberghe, L. (2004) Convex Optimization . (Cambridge University Press)

  88. [96]

    (1987) Practical Methods of Optimization

    Fletcher, R. (1987) Practical Methods of Optimization . (New York: Wiley)

  89. [97]

    C & Nocedal, J

    Liu, D. C & Nocedal, J. (2022) On the limited memory BFGS method for large scale optimization. Math. Program. 45 , 503--528

  90. [98]

    U, Corander, J, & Rinke, P

    Todorovi\' c , M, Gutmann, M. U, Corander, J, & Rinke, P. (2019) Bayesian inference of atomistic structure in functional materials. Npj Comput. Mater. 5 , 35

  91. [99]

    Johnson, S. G. (2007) The NLopt nonlinear-optimization package (https://github.com/stevengj/nlopt)

  92. [100]

    (1994) Algorithm 733: TOMP–Fortran modules for optimal control calculations

    Kraft, D. (1994) Algorithm 733: TOMP–Fortran modules for optimal control calculations. ACM Trans. Math. Softw. 20 , 262–281

  93. [101]

    Lewis, J. P. (1969) Homogeneous Functions and Euler's Theorem . (Palgrave Macmillan UK, London), pp. 297--303

  94. [102]

    Solov'ev, V. N. (1983) On a criterion for convexity of a positive-homogeneous function. Mat. Sb. 46 , 285

  95. [103]

    (2025) The elements of differentiable programming

    Blondel, M & Roulet, V. (2025) The elements of differentiable programming

  96. [104]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence '...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.