REVIEW 2 major objections 6 minor 36 references
Machine learning configuration interaction for ab initio potential energy curves
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read MLCI, using one neural network for ranking and hashing, computes ab initio potential energy curves for N2, H2O, and CO to near full-CI accuracy with far fewer processor hours than stochastic selection.
desk verdict A worthwhile extension of the author's MLCI method—the hash trick is genuinely practical—but the missing seed-dependence analysis makes the headline accuracy claims provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the artificial neural network used in two roles at once. Its output for a configuration is the predicted transformed coefficient, trained against |~c_i| = (0.4|ci| + 0.6 - cmin)/(1 - cmin) for coefficients above the cutoff, and the same output is turned into a hash key floor(output * 2L) + 1 so duplicates can be detected against a hash table of size about 2L without sorting. This removes the memory bottleneck of storing all single and double substitutions. Two further mechanisms carry the accuracy: configuration state functions with approximate orthonormalization, which guarantee a pure spin state, and geometry-to-geometry transfer of the wavefunction, which gives the network more important configurations to learn from.
What would settle it
Run the hash-based streaming selection and the full-space quicksort selection at every geometry of a potential curve and compare the final energies and configuration sets; if any geometry gives different energies, or if any configuration in the sorted-selection wavefunction would have been rejected early by the streaming rule, the central claim would be contradicted.
Extended reading notes
Core claim
On its own terms, the paper claims that MLCI can be made scalable enough to compute potential energy curves of near full-configuration-interaction quality. Using the neural network output as a hash value, the algorithm generates single and double substitutions, keeps the best L it has encountered where L is the current wavefunction size, and never stores the entire singles and doubles space; a single-point test on N2 reaches the same final energy as the old quicksort-based duplicate removal in 1816 seconds versus 2409. With configuration state functions and a transferred wavefunction, the best standard deviations of the energy difference are 0.56 kcal/mol for N2, 0.76 kcal/mol for CO, and 0.70 kcal/mol for H2O at a lower cutoff, while the N2 and CO curves use less wall time and far fewer processor hours than stochastic configuration selection despite running serially. The paper also reports that transferring only the wavefunction between geometries is the most accurate protocol, whereas transferring the neural network weights alone offers no clear accuracy gain.
Load-bearing premise
The method assumes the network's guesses about which electron configurations matter are trustworthy enough that discarding low-ranked candidates on the fly loses none of the important ones.
Editorial extensions
If this is right
- The ANN-as-hash removes the storage barrier of the singles and doubles space, so MLCI can be applied to systems whose single and double substitution space is too large to hold in memory.
- For N2 and CO, MLCI achieves lower curve errors than stochastic configuration selection while using substantially fewer processor hours, running in serial where the stochastic comparison ran in parallel.
- Transferring the wavefunction from a nearby geometry is systematically the most accurate transfer protocol, improving the curve error to 0.56 kcal/mol for N2 and 0.76 kcal/mol for CO.
- Using configuration state functions means the computed wavefunctions are pure spin states, so the potential curves are not contaminated by spin contamination.
- The method handles geometries ranging from single-reference to strongly multireference without choosing an active space, with multireference character reaching about 0.93 for CO at the longest bond length considered.
Reading between the lines
- If the neural network ranking stays reliable as the space grows, the hash-based streaming selection should scale to basis sets and molecules where even storing the sorted singles and doubles list is impossible, because memory is no longer the limiting resource.
- The paper's finding that transferring the ANN alone does not help suggests the learned weights encode geometry-specific importance patterns more than general electronic-structure knowledge; a deeper network trained differently might behave otherwise, as the paper itself notes as future work.
- A natural extension would be to run the hash-based and full-sort versions at every point of a potential curve, not just one geometry, to measure how often early rejection changes the final configuration set.
- The success of wavefunction warm-starting hints that selected CI methods more broadly could benefit from geometry-continuity information, not only from better importance estimators.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript extends the machine learning configuration interaction (MLCI) method in three ways: the ANN's output is used as a hash function for on-the-fly duplicate removal so that the full singles/doubles list need not be stored; configuration state functions are introduced to guarantee pure spin states; and transfer protocols between geometries are tested. The method is applied to potential energy curves of N2, H2O, and CO in cc-pVDZ, benchmarked against FCI data and compared with previous MCCI results. For N2 and CO the best MLCI protocols achieve sigma_deltaE of 0.56 and 0.76 kcal/mol, respectively, with substantially lower processor-hour usage than MCCI; for water, MLCI is less accurate at a comparable cutoff and requires a lower cutoff to match MCCI accuracy.
Significance. The proposed ANN-as-hash modification addresses a real scalability bottleneck in MLCI, and the CSF formulation is a natural extension. The validation is non-circular: energies are checked against independent FCI references, and the paper reports honest cases where MLCI does not outperform MCCI (water). The order-of-magnitude reductions in processor hours for N2 and CO, if reproducible, make this a useful contribution to selected CI methods. However, the absence of any run-to-run variance analysis for the stochastic ANN training means the margins over MCCI are not yet quantitatively established.
major comments (2)
- [§2.1, §3, Tables 1, 3, 4] The main results are single runs of a stochastic algorithm, and no seed dependence is reported. The algorithm uses random initial weights on [-0.1,0.1], a random 50/50 training/verification split each iteration, and SGD with shuffled data; §2.3 also mentions 'randomly swapping spins' in CSF construction. The seed is fixed only for the single-geometry hash-versus-quicksort test. The claimed advantage over MCCI rests on differences as small as 0.13 kcal/mol (CO: 0.76 vs 0.89 kcal/mol; N2: 0.56 vs 1.06 kcal/mol), which could be within run-to-run noise. Please report the mean and standard deviation (or at least the range) of sigma_deltaE, NPE, and timings over several independent seeds for the leading protocol of each system, and state whether the quoted numbers are representative single runs. Without this, the central comparison to MCCI is not quantitative.
- [§3, Tables 1, 3, 4] The efficiency comparison against MCCI relies on timings taken from Ref. 57. For N2 the paper states these were on the same hardware, but for H2O and CO no hardware statement is made. To make the 'substantially less processor hours' claim robust, please specify the hardware and parallel setup for all MCCI reference timings, or rerun at least the CO benchmark under identical conditions. This is especially important because the CO wall-time margin is small (19.84 vs 20.72 hours).
minor comments (6)
- [Figure 2 caption] The caption 'with a stretched geometry of 2.2225 Å' appears to be a typo; the figure plots errors against bond length across the whole curve, consistent with the text referring to 'the 15 bond lengths'.
- [Equation (1)] Equation (1) should be written with parentheses as 0.4|c_i| + (0.6 - c_min)/(1 - c_min); the current typesetting '0.4|ci| + 0.6 − cmin 1 − cmin' is ambiguous.
- [§2.2] The hash formula '⌊Output2L⌋+1' should read 'floor(Output x 2L) + 1', and the hash table size (2L) should be stated explicitly in the text.
- [§2.3] Please clarify whether 'randomly swapping spins' in the genealogical CSF construction is a deterministic canonicalization or a stochastic step; this affects the reproducibility of the wavefunction and should be addressed in the seed-dependence analysis requested above.
- [§3 and Fig. 1] Since the streaming 'keep the best L' rule is exact for fixed ANN predictions, the single-geometry comparison with quicksort is an implementation check rather than a test of a geometry-dependent effect; the paper should state this explicitly to avoid the impression that the hash approach was validated at only one bond length.
- [General] The manuscript does not include a data or code availability statement; given the stochastic nature of the method, providing the random seed protocol or the MLCI code would substantially aid reproduction.
Circularity Check
No significant circularity: MLCI energies are checked against external FCI benchmarks, and the ANN selection is not fitted to the target energies.
full rationale
The paper's central claim is that MLCI can efficiently produce accurate ab initio potential energy curves. The accuracy numbers (sigma_deltaE = 0.56, 0.76, and 0.70 kcal/mol in the best cases) are obtained by diagonalizing the Hamiltonian in an ANN-selected CSF space and comparing the resulting energies against FCI values. For N2 the FCI benchmarks come from Refs. 58-60, while for H2O and CO they are taken from Ref. 57. The ANN learns on the fly from intermediate wavefunction coefficients and reject sets; no ANN parameter or other fitted quantity is tuned to the FCI energies or to sigma_deltaE. Thus the benchmark comparison is not forced by construction. The use of the ANN output as a hash function is validated at one fixed seed as an implementation check of duplicate removal, and the resulting final energies are identical to the quicksort version, which is a consistency test rather than a prediction. The self-citations, especially Ref. 21 for the original MLCI protocol and Ref. 57 for prior MCCI timings and some FCI benchmarks, provide background and comparison data, but the present derivation does not reduce to those citations: the MLCI energies are computed in this paper and the N2 reference data are external. The stochastic nature of ANN initialization and SGD is a reproducibility or statistical-robustness concern for the MLCI-versus-MCCI error margins, but it is not circularity, because every reported energy remains a variational eigenvalue of the Hamiltonian in the selected space rather than a fitted output. No equation in the paper defines the predicted result in terms of the input data in a way that makes the accuracy comparison tautological.
Assumptions & free parameters
free parameters (3)
- Selection cutoff cmin =
5e-4 for N2 and CO; 1e-3, 5e-4, 2e-4 for H2O
- Hidden nodes in ANN (nh) =
40 (30 also tested for N2)
- Hash array size factor =
2L, where L is the current wavefunction size
assumptions (4)
- domain assumption FCI reference energies from Refs 57-60 are correct and were computed with the same basis sets, frozen orbitals, and geometries as the MLCI calculations.
- domain assumption The MCCI program framework (Refs 40-42) provides correct Hamiltonian and overlap matrix elements for CSFs, including the genealogical spin-adaptation that makes CSFs linearly independent.
- domain assumption A single-hidden-layer neural network with a handful of nodes can learn to rank important configurations well enough for the on-the-fly selection to converge to accurate energies.
- domain assumption The convergence criterion of Ref 43 is a sufficient stopping rule for potential energy curves.
Cite this review
Pith. "Pith review of Machine learning configuration interaction for ab initio potential energy curves." pith.science (2026). https://pith.science/paper/TQYP3XZJ
@misc{pith2026190807430,
author = {Pith},
title = {Pith review of: Machine learning configuration interaction for ab initio potential energy curves},
year = {2026},
howpublished = {\url{https://pith.science/paper/TQYP3XZJ}},
note = {Machine review of arXiv:1908.07430}
}
abstract
The concept of machine learning configuration interaction (MLCI) [J. Chem. Theory Comput. 2018, 14, 5739], where an artificial neural network (ANN) learns on the fly to select important configurations, is further developed so that accurate ab initio potential energy curves can be efficiently calculated. This development includes employing the artificial neural network also as a hash function for the efficient deletion of duplicates on the fly so that the singles and doubles space does not need to be stored and this barrier to scalability is removed. In addition configuration state functions are introduced into the approach so that pure spin states are guaranteed, and the transferability of data between geometries is exploited. This improved approach is demonstrated on potential energy curves for the nitrogen molecule, water, and carbon monoxide. The results are compared with full configuration interaction values, when available, and different transfer protocols are investigated. It is shown that, for all of the considered systems, accurate potential energy curves can now be efficiently computed with MLCI. For the potential curves of N$_{2}$ and CO, MLCI can achieve lower errors than stochastically selecting configurations while also using substantially less processor hours.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
(1) Qu, C.; Yu, Q.; Van Hoozen Jr., B. L.; Bowman, J. M.; Vargas-Hern ´ andez, R. A. Assessing Gaussian Process Regression and Permutationally Invar iant Polynomial Ap- proaches To Represent High-Dimensional Potential Energy Surfa ces. J. Chem. Theory Comput. 2018, 14,
work page 2018
-
[2]
J. Chem. Phys. 2000, 113,
work page 2000
-
[7]
(2) Schmitz, G.; Christiansen, O. Gaussian process regression to a ccelerate geometry opti- mizations relying on numerical differentiation. J. Chem. Phys. 2018, 148, 241704. (3) Wiens, A. E.; Copan, A. V.; Schaefer, H. F. Multi-fidelity Gaussian process modeling for chemical energy surfaces. Chem. Phys. Lett.: X 2019, 3, 100022. (4) Gastegger, M.; Marquetan...
work page 2018
-
[46]
(62) Coe, J. P.; Paterson, M. J. Investigating Multireference Cha racter and Correlation in Quantum Chemistry. J. Chem. Theory Comput. 2015, 11,
work page 2015
-
[58]
(46) Szepieniec, M.; Yeriskin, I.; Greer, J. C. Quasiparticle energiesand lifetimes in a metallic chain model of a tunnel junction. J. Chem. Phys. 2013, 138, 144105. (47) Kelly, T. P.; Perera, A.; Bartlett, R. J.; Greer, J. C. Monte Car lo configuration in- teraction with perturbation corrections for dissociation energies of first row diatomic molecules: C ...
work page 2013
-
[64]
(52) Zimmerman, P. M. Incremental full configuration interaction . J. Chem. Phys. 2017, 146, 104102. (53) Eriksen, J. J.; Lipparini, F.; Gauss, J. Virtual Orbital Many-Bo dy Expansions: A Possible Route towards the Full Configuration Interaction Limit. J. Phys. Chem. Lett. 2017, 8,
work page 2017
-
[181]
(42) Tong, L.; Nolan, M.; Cheng, T.; Greer, J. C. A Monte Carlo config uration generation computer program for the calculation of electronic states of atoms, molecules, and quan- tum dots. Comp. Phys. Comm. 2000, 131, 142, see https://github.com/MCCI/mcci. (43) Gy˝ orffy, W.; Bartlett, R. J.; Greer, J. C. Monte Carlo configu ration interaction pre- dictions ...
work page 2000
-
[291]
O.; Rupp, M.; von Lilienfeld, O
26 (17) Ramakrishnan, R.; Dral, P. O.; Rupp, M.; von Lilienfeld, O. A. Big Da ta Meets Quan- tum Chemistry Approximations: The ∆-Machine Learning Approach. J. Chem. Theory Comput. 2015, 11, 2087–2096. (18) McGibbon, R. T.; Taube, A. G.; Donchev, A. G.; Siva, K.; Hern´ andez, F.; Hargus, C.; Law, K.-H.; Klepeis, J. L.; Shaw, D. E. Improving the accuracy of...
work page 2015
Show all 36 references
-
[319]
Coordinate Descent Full Configuration I nteraction
(49) Wang, Z.; Li, Y.; Lu, J. Coordinate Descent Full Configuration I nteraction. J. Chem. Theory Comput. 2019, 15,
2019
-
[359]
P.; Murphy, P.; Paterson, M
(61) Coe, J. P.; Murphy, P.; Paterson, M. J. Applying Monte Carlo co nfiguration interac- tion to transition metal dimers: Exploring the balance between stat ic and dynamic correlation. Chem. Phys. Lett. 2014, 604,
2014
-
[431]
H.; Thom, A
(29) Booth, G. H.; Thom, A. J. W.; Alavi, A. Fermion Monte Carlo withou t fixed nodes: A game of life, death, and annihilation in Slater determinant space. J. Chem. Phys. 2009, 131, 054106. (30) Ohtsuka, Y. Spin-symmetry adaptation to the Monte Carlo co rrection configuration in...
2009
-
[618]
D.; Bartlett, R
(15) Purvis, G. D.; Bartlett, R. J. A full coupled-cluster singles and doubles model: The inclusion of disconnected triples. J. Chem. Phys. 1982, 76, 1910–1918. (16) Bartlett, R. J.; Musia/suppress l, M. Coupled-cluster theory in quantumchemistry. Rev. Mod. Phys. 2007, 79,
1982
-
[872]
Sem i-local machine- learned kinetic energy density functional with third-order gradients of electron density
(12) Seino, J.; Kageyama, R.; Fujinami, M.; Ikabata, Y.; Nakai, H. Sem i-local machine- learned kinetic energy density functional with third-order gradients of electron density. J. Chem. Phys. 2018, 148, 241705. (13) Cust´ odio, C. A.; Filletti, E. R.; Fran¸ ca, V. V. Artificia...
2018
-
[1139]
E.; Burke, K.; M¨ uller, K.-R
(11) Brockherde, F.; Vogt, L.; Li, L.; Tuckerman, M. E.; Burke, K.; M¨ uller, K.-R. Bypassing the Kohn-Sham equations with machine learning. Nat. Comm. 2017, 8,
2017
-
[1395]
E xcitation energies from diffusion Monte Carlo using selected configuration interaction n odes
28 (34) Scemama, A.; Benali, A.; Jacquemin, D.; Caffarel, M.; Loos, P.-F. E xcitation energies from diffusion Monte Carlo using selected configuration interaction n odes. J. Chem. Phys. 2018, 149, 034108. (35) Applencourt, T.; Gasperich, K.; Scemama, A. Spin adaptation w ith dete...
2018 arXiv
-
[1772]
Comparison of the open-shell state-univers al and state-selective coupled-cluster theories: H 4 and H8 models
(64) Li, X.; Paldus, J. Comparison of the open-shell state-univers al and state-selective coupled-cluster theories: H 4 and H8 models. J. Chem. Phys. 1995, 103,
1995
-
[1821]
(41) Greer, J. C. Monte Carlo Configuration Interaction. J. Comp. Phys. 1998, 146,
1998
-
[1886]
(14) Møller, C.; Plesset, M. S. Note on an Approximation Treatment f or Many-Electron Systems. Phys. Rev. 1934, 46,
1934
-
[2187]
T.; Arbabzadah, F.; Chmiela, S.; M¨ uller, K
(5) Sch¨ utt, K. T.; Arbabzadah, F.; Chmiela, S.; M¨ uller, K. R.; Tkatchenko, A. Quantum- chemical insights from deep tensor neural networks. Nat. Comm. 2017, 8, 13890. (6) Janet, J. P.; Kulik, H. J. Predicting electronic structure prope rties of transition metal complexes wi...
2017
-
[3404]
A.; Christensen, A
(8) Faber, F. A.; Christensen, A. S.; Huang, B.; von Lilienfeld, O. A. A lchemical and structural distribution based representation for universal qua ntum machine learning. J. Chem. Phys. 2018, 148, 241717. (9) Gubaev, K.; Podryabinkin, E. V.; Shapeev, A. V. Machine learning o...
2018
-
[3558]
I.; Bartlett, R
(50) Lyakh, D. I.; Bartlett, R. J. An adaptive coupled-cluster the ory: @CC approach. J. Chem. Phys. 2010, 133, 244112. (51) Bytautas, L.; Ruedenberg, K. A priori identification of config urational deadwood. Chem. Phys. 2009, 356,
2010
-
[3591]
A Mountaineering Strategy to Excited States: Highly Accurate Refe rence Energies and Benchmarks
(32) Loos, P.-F.; Scemama, A.; Blondel, A.; Garniron, Y.; Caffarel, M.; J acquemin, D. A Mountaineering Strategy to Excited States: Highly Accurate Refe rence Energies and Benchmarks. J. Chem. Theory Comput. 2018, 14,
2018
-
[3674]
M.; Lee, J.; Takeshita, T
27 (26) Tubman, N. M.; Lee, J.; Takeshita, T. Y.; Head-Gordon, M.; Birg itta Whaley, K. A deterministic alternative to the full configuration interaction quan tum Monte Carlo method. J. Chem. Phys. 2016, 145, 044112. (27) Ohtsuka, Y.; Hasegawa, J. Selected configuration interact...
2016
-
[4129]
Machine Learning Adaptive Basis Sets for Efficient Large Scale Density Functional Theory Simulation.J
(23) Sch¨ utt, O.; VandeVondele, J. Machine Learning Adaptive Basis Sets for Efficient Large Scale Density Functional Theory Simulation.J. Chem. Theory Comput. 2018, 14,
2018
-
[4168]
P.; Rancurel, P
(24) Huron, B.; Malrieu, J. P.; Rancurel, P. Iterative perturbationcalculations of ground and excited state energies from multiconfigurational zeroth-order w avefunctions. J. Chem. Phys. 1973, 58,
1973
-
[4189]
P.; Paterson, M
(63) Coe, J. P.; Paterson, M. J. Open-shell systems investigated with Monte Carlo configu- ration interaction. Int. J. Quantum Chem. 2016, 116,
2016
-
[4360]
Determinist ic Construction of Nodal Surfaces within Quantum Monte Carlo: The Case of FeS
(33) Scemama, A.; Garniron, Y.; Caffarel, M.; Loos, P.-F. Determinist ic Construction of Nodal Surfaces within Quantum Monte Carlo: The Case of FeS. J. Chem. Theory Comput. 2018, 14,
2018
-
[4633]
J.; Gauss, J
30 (54) Eriksen, J. J.; Gauss, J. Many-Body Expanded Full Configuration Interaction. I. Weakly Correlated Regime. J. Chem. Theory Comput. 2018, 14,
2018
-
[4772]
(21) Coe, J. P. Machine learning configuration interaction. J. Chem. Theory Comput. 2018, 14,
2018
-
[5137]
A.; Tkatchenko, A.; M¨ uller, K
25 (7) Hansen, K.; Montavon, G.; Biegler, F.; Fazli, S.; Rupp, M.; Scheffler , M.; von Lilien- feld, O. A.; Tkatchenko, A.; M¨ uller, K. R. Assessment and Validation of Machine Learning Methods for Predicting Molecular Atomization Energies. J. Chem. Theory Comput. 2013, 9,
2013
-
[5180]
J.; Gauss, J
(55) Eriksen, J. J.; Gauss, J. Many-Body Expanded Full Configura tion Inter- action. II. Strongly Correlated Regime. J. Chem. Theory Comput. 2019, doi:10.1021/acs.jctc.9b00456. (56) H.-J.Werner,; Knowles, P. J.; Knizia, G.; Manby, F. R.; Sch¨ utz, M .; Celani, P.; Gy¨ orffy, W.;...
2019 doi
-
[5354]
(40) Greer, J. C. Estimating full configuration interaction limits from a Monte Carlo selection of the expansion space. J. Chem. Phys. 1995, 103,
1995
-
[5739]
(22) Townsend, J.; Vogiatzis, K. D. Data-Driven Acceleration of the Coupled-Cluster Singles and Doubles Iterative Solver. J. Phys. Chem. Lett. 2019, 10,
2019
-
[5745]
A.; Tubman, N
(25) Holmes, A. A.; Tubman, N. M.; Umrigar, C. J. Heat-Bath Configu ration Interaction: An Efficient Selected Configuration Interaction Algorithm Inspired by Heat-Bath Sam- pling. J. Chem. Theory Comput. 2016, 12,
2016
-
[6343]
(20) Welborn, M.; Cheng, L.; Miller, III, T. F. Transferability in Machin e Learning for Electronic Structure via the Molecular Orbital Basis. J. Chem. Theory Comput. 2018, 14,
2018
-
[6677]
K.-L.; K´ allay, M.; Gauss, J
(59) Chan, G. K.-L.; K´ allay, M.; Gauss, J. State-of-the-art density matrix renormalization group and coupled cluster theory studies of the nitrogen binding cu rve. J. Chem. Phys. 2004, 121, 6110–6116. (60) Gwaltney, S. R.; Byrd, E. F. C.; Voorhis, T. V.; Head-Gordon, M . A p...
2004
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.