{"id":"98cfa441-0080-4072-8176-86b31d90b628","arxiv_id":"1908.07430","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MLCI with the ANN used as a hash function, configuration state functions, and geometry-to-geometry wavefunction transfer gives near-FCI potential energy curves for N2 and CO more cheaply than stochastic Monte Carlo CI.","lead":"This paper upgrades a machine-learning selected configuration interaction method so it can draw smooth potential energy curves for small molecules without storing the full list of generated configurations. The new version matches full configuration interaction accuracy on nitrogen and carbon monoxide curves while using far fewer processor hours than stochastic selection.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MLCI is stochastic (random ANN initialization and SGD), but no seed-dependence or run-to-run variance is reported; the claimed MLCI-vs-MCCI accuracy margins for N2 and CO could be within run-to-run noise.","rationale":"The paper makes a plausible empirical case: the hash-based duplicate removal is shown to reproduce the quicksort result at one N2 geometry, the CSF formulation guarantees spin purity, and the FCI comparisons for N2 and CO show small sigma_deltaE values. The water results are reported honestly, and the paper does not overclaim that MLCI beats MCCI for water. However, the headline comparison with MCCI is a quantitative comparison between stochastic methods, and the paper reports no measure of stochastic uncertainty for MLCI. The ANN training explicitly uses random initialization, random data splits, and SGD, so different seeds will generally produce different selected configuration spaces and energies. The seed is fixed only in the Section 3 comparison of duplicate-removal strategies; Tables 1-4 contain single values. The reported accuracy margins over MCCI are not obviously larger than the run-to-run spread such a procedure can exhibit, so without a seed study the central comparative claim is underdetermined. The reader's identified weakest assumption about streaming selection is less severe than it appears: the on-the-fly keep-best-L procedure is an exact priority queue over ANN predictions, so its equivalence to sorting is not contingent on ANN accuracy; the real vulnerability is the ANN ranking's accuracy and stability, which is only indirectly validated through final energies. I therefore recommend keeping the CONDITIONAL verdict, with the condition made explicit: report seed dependence or error bars, and ideally release the code and data for independent verification.","tokens_in":17222,"tokens_out":9009,"duration_ms":100963,"concrete_test":"Recompute the N2 (Table 1, transfer-wavefunction) and CO (Table 4, transfer-wavefunction) potential curves with at least 10 independent random seeds, using the same protocol and cutoff, and record sigma_deltaE and NPE for each seed. Ideally, also rerun the corresponding MCCI calculations from Ref. 57 under the same conditions. If the standard deviation of sigma_deltaE across seeds is small (e.g., <= 0.05 kcal/mol) and the lower bound of the MLCI distribution still beats the MCCI value, the comparative claim is robust. If the seed-to-seed spread is comparable to the 0.3-0.5 kcal/mol margin, the claim should be conditioned on seed-averaged or error-barred results.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing gap is statistical rather than algorithmic. MLCI as described in Section 2.1 is stochastic: ANN weights are initialized randomly on [-0.1, 0.1], data are split randomly into training and verification sets, and training uses stochastic gradient descent. The only place a seed is mentioned is Section 3, where it is fixed solely to compare duplicate-removal strategies at one N2 geometry. For the potential-energy-curve calculations in Tables 1, 3, and 4, no run-to-run variance or seed dependence is reported. The central quantitative claim is that MLCI gives sigma_deltaE = 0.56 kcal/mol for N2 and 0.76 kcal/mol for CO, beating prior MCCI results by roughly 0.3-0.5 kcal/mol. If different random seeds shift these sigma_deltaE values by more than that margin, the 'lower errors' comparison is not established. This concern is more load-bearing than the reader's focus on streaming selection: the on-the-fly 'keep the best L' rule is an exact priority queue given fixed ANN outputs, so its equivalence to quicksort does not depend on ANN reliability; the unvalidated part is the ANN ranking itself, which is only indirectly assessed through final energies.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript extends the machine learning configuration interaction (MLCI) method in three ways: the ANN's output is used as a hash function for on-the-fly duplicate removal so that the full singles/doubles list need not be stored; configuration state functions are introduced to guarantee pure spin states; and transfer protocols between geometries are tested. The method is applied to potential energy curves of N2, H2O, and CO in cc-pVDZ, benchmarked against FCI data and compared with previous MCCI results. For N2 and CO the best MLCI protocols achieve sigma_deltaE of 0.56 and 0.76 kcal/mol, respectively, with substantially lower processor-hour usage than MCCI; for water, MLCI is less accurate at a comparable cutoff and requires a lower cutoff to match MCCI accuracy.","tokens_in":17464,"tokens_out":16818,"duration_ms":155721,"significance":"The proposed ANN-as-hash modification addresses a real scalability bottleneck in MLCI, and the CSF formulation is a natural extension. The validation is non-circular: energies are checked against independent FCI references, and the paper reports honest cases where MLCI does not outperform MCCI (water). The order-of-magnitude reductions in processor hours for N2 and CO, if reproducible, make this a useful contribution to selected CI methods. However, the absence of any run-to-run variance analysis for the stochastic ANN training means the margins over MCCI are not yet quantitatively established.","major_comments":[{"comment":"The main results are single runs of a stochastic algorithm, and no seed dependence is reported. The algorithm uses random initial weights on [-0.1,0.1], a random 50/50 training/verification split each iteration, and SGD with shuffled data; §2.3 also mentions 'randomly swapping spins' in CSF construction. The seed is fixed only for the single-geometry hash-versus-quicksort test. The claimed advantage over MCCI rests on differences as small as 0.13 kcal/mol (CO: 0.76 vs 0.89 kcal/mol; N2: 0.56 vs 1.06 kcal/mol), which could be within run-to-run noise. Please report the mean and standard deviation (or at least the range) of sigma_deltaE, NPE, and timings over several independent seeds for the leading protocol of each system, and state whether the quoted numbers are representative single runs. Without this, the central comparison to MCCI is not quantitative.","section":"§2.1, §3, Tables 1, 3, 4"},{"comment":"The efficiency comparison against MCCI relies on timings taken from Ref. 57. For N2 the paper states these were on the same hardware, but for H2O and CO no hardware statement is made. To make the 'substantially less processor hours' claim robust, please specify the hardware and parallel setup for all MCCI reference timings, or rerun at least the CO benchmark under identical conditions. This is especially important because the CO wall-time margin is small (19.84 vs 20.72 hours).","section":"§3, Tables 1, 3, 4"}],"minor_comments":[{"comment":"The caption 'with a stretched geometry of 2.2225 Å' appears to be a typo; the figure plots errors against bond length across the whole curve, consistent with the text referring to 'the 15 bond lengths'.","section":"Figure 2 caption"},{"comment":"Equation (1) should be written with parentheses as 0.4|c_i| + (0.6 - c_min)/(1 - c_min); the current typesetting '0.4|ci| + 0.6 − cmin 1 − cmin' is ambiguous.","section":"Equation (1)"},{"comment":"The hash formula '⌊Output2L⌋+1' should read 'floor(Output x 2L) + 1', and the hash table size (2L) should be stated explicitly in the text.","section":"§2.2"},{"comment":"Please clarify whether 'randomly swapping spins' in the genealogical CSF construction is a deterministic canonicalization or a stochastic step; this affects the reproducibility of the wavefunction and should be addressed in the seed-dependence analysis requested above.","section":"§2.3"},{"comment":"Since the streaming 'keep the best L' rule is exact for fixed ANN predictions, the single-geometry comparison with quicksort is an implementation check rather than a test of a geometry-dependent effect; the paper should state this explicitly to avoid the impression that the hash approach was validated at only one bond length.","section":"§3 and Fig. 1"},{"comment":"The manuscript does not include a data or code availability statement; given the stochastic nature of the method, providing the random seed protocol or the MLCI code would substantially aid reproduction.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central issue is statistical: the manuscript's headline comparisons rest on single stochastic runs, and the CO margin (0.13 kcal/mol) is small enough that a seed-dependence study could overturn it. I would like the revision to include such a study before publication; the methodological contribution is otherwise sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper extends Coe's own MLCI with three concrete improvements: using the ANN output as a hash function to prune duplicates on the fly, switching to CSFs for spin purity, and testing geometry-transfer protocols. The hash trick is the real novelty—it removes the memory bottleneck of storing the full singles/doubles space and looks sound as an exact priority queue given fixed ANN outputs. The CSF adaptation and the systematic transfer study are solid, incremental work. On N2 and CO, MLCI with wavefunction transfer reaches sigma_deltaE of 0.56 and 0.76 kcal/mol respectively, and uses far fewer processor-hours than the cited MCCI results. The water section is honestly reported: MLCI underperforms MCCI at the same cutoff and only beats it with a tighter cutoff and longer time, which correctly tempers the abstract's 'efficiently' claim.\n\nThe soft spot is statistical, not algorithmic. MLCI is stochastic—random ANN initialization, random train/verification split, stochastic gradient descent—but the paper reports no seed dependence or run-to-run variance for any PEC. The N2 margin over MCCI (0.56 vs 1.06 kcal/mol) and especially CO (0.76 vs 0.89) could easily be within run-to-run noise. This matters more than the single-geometry validation of the hash method, which is actually fine for fixed ANN outputs; the unvalidated part is whether different ANN training runs produce consistently similar selections. The paper also does not release code or data, and the free parameters (cmin, nh, hash array size) get only cursory treatment. These are fixable with a few seed repetitions and a public repository, but without them the quantitative comparisons are not yet reproducible.\n\nWho is this for? People building selected CI methods or using MLCI for strongly correlated small molecules. It deserves a serious referee—the hash trick is worth the field's attention—but the referee should ask for seed-sensitivity analysis and code release before acceptance. I would not cite it yet for the timing or accuracy numbers, though I might cite the hash idea if it survives scrutiny.\n\nSend it to peer review, but expect revision.","headline":"A worthwhile extension of the author's MLCI method—the hash trick is genuinely practical—but the missing seed-dependence analysis makes the headline accuracy claims provisional.","tokens_in":17998,"tokens_out":2037,"would_cite":false,"duration_ms":22956,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MLCI, using one neural network for ranking and hashing, computes ab initio potential energy curves for N2, H2O, and CO to near full-CI accuracy with far fewer processor hours than stochastic selection.","keywords":["machine learning configuration interaction","selected configuration interaction","artificial neural network","hash function","configuration state functions","potential energy curves","full configuration interaction","wavefunction transfer"],"falsifier":"Run the hash-based streaming selection and the full-space quicksort selection at every geometry of a potential curve and compare the final energies and configuration sets; if any geometry gives different energies, or if any configuration in the sorted-selection wavefunction would have been rejected early by the streaming rule, the central claim would be contradicted.","tokens_in":1565,"feed_emoji":"⚛️","tokens_out":3817,"duration_ms":93040,"temperature":0.7,"pith_summary":"This paper extends machine learning configuration interaction (MLCI) so that it can produce accurate ab initio potential energy curves for small molecules rather than just single-point energies. The key move is to let one artificial neural network do two jobs at once: rank the importance of newly generated electron configurations and act as a hash function that removes duplicates on the fly, so the full list of single and double substitutions never has to be stored. The method also uses configuration state functions to guarantee pure spin states and reuses the wavefunction from one bond length as a starting point for the next. On nitrogen, water, and carbon monoxide, the resulting curves agree with full configuration interaction to within roughly a kilocalorie per mole once shifted, and for N2 and CO the approach beats stochastic configuration selection in both accuracy and processor time.","feed_headline":"MLCI computes molecular curves to near-exact accuracy","feed_subtitle":"A neural network ranks and hashes electron configurations, beating stochastic selection on processor time.","key_machinery":"The load-bearing object is the artificial neural network used in two roles at once. Its output for a configuration is the predicted transformed coefficient, trained against |~c_i| = (0.4|ci| + 0.6 - cmin)/(1 - cmin) for coefficients above the cutoff, and the same output is turned into a hash key floor(output * 2L) + 1 so duplicates can be detected against a hash table of size about 2L without sorting. This removes the memory bottleneck of storing all single and double substitutions. Two further mechanisms carry the accuracy: configuration state functions with approximate orthonormalization, which guarantee a pure spin state, and geometry-to-geometry transfer of the wavefunction, which gives the network more important configurations to learn from.","core_discovery":"On its own terms, the paper claims that MLCI can be made scalable enough to compute potential energy curves of near full-configuration-interaction quality. Using the neural network output as a hash value, the algorithm generates single and double substitutions, keeps the best L it has encountered where L is the current wavefunction size, and never stores the entire singles and doubles space; a single-point test on N2 reaches the same final energy as the old quicksort-based duplicate removal in 1816 seconds versus 2409. With configuration state functions and a transferred wavefunction, the best standard deviations of the energy difference are 0.56 kcal/mol for N2, 0.76 kcal/mol for CO, and 0.70 kcal/mol for H2O at a lower cutoff, while the N2 and CO curves use less wall time and far fewer processor hours than stochastic configuration selection despite running serially. The paper also reports that transferring only the wavefunction between geometries is the most accurate protocol, whereas transferring the neural network weights alone offers no clear accuracy gain.","pith_inferences":["If the neural network ranking stays reliable as the space grows, the hash-based streaming selection should scale to basis sets and molecules where even storing the sorted singles and doubles list is impossible, because memory is no longer the limiting resource.","The paper's finding that transferring the ANN alone does not help suggests the learned weights encode geometry-specific importance patterns more than general electronic-structure knowledge; a deeper network trained differently might behave otherwise, as the paper itself notes as future work.","A natural extension would be to run the hash-based and full-sort versions at every point of a potential curve, not just one geometry, to measure how often early rejection changes the final configuration set.","The success of wavefunction warm-starting hints that selected CI methods more broadly could benefit from geometry-continuity information, not only from better importance estimators."],"forward_implications":["The ANN-as-hash removes the storage barrier of the singles and doubles space, so MLCI can be applied to systems whose single and double substitution space is too large to hold in memory.","For N2 and CO, MLCI achieves lower curve errors than stochastic configuration selection while using substantially fewer processor hours, running in serial where the stochastic comparison ran in parallel.","Transferring the wavefunction from a nearby geometry is systematically the most accurate transfer protocol, improving the curve error to 0.56 kcal/mol for N2 and 0.76 kcal/mol for CO.","Using configuration state functions means the computed wavefunctions are pure spin states, so the potential curves are not contaminated by spin contamination.","The method handles geometries ranging from single-reference to strongly multireference without choosing an active space, with multireference character reaching about 0.93 for CO at the longest bond length considered."],"supporting_citations":[{"why":"Supplies the original MLCI method of training an ANN on the fly to rank configurations, which this paper extends.","marker":"[21]"},{"why":"Provides the stochastic configuration selection baseline and the FCI reference data for the H2O and CO potential curves.","marker":"[57]"},{"why":"Supplies the MCCI program framework for CSF Hamiltonian and overlap matrix elements and the genealogical spin-adaptation procedure.","marker":"[40-42]"},{"why":"Provide the full configuration interaction reference energies for the N2 potential curve used as the accuracy target.","marker":"[58-60]"},{"why":"Supplies the convergence criterion and the coefficient scaling that makes the CSFs approximately orthonormal.","marker":"[43]"},{"why":"Defines the non-parallelity error used to quantify how well the shape of a shifted potential curve matches FCI.","marker":"[64]"},{"why":"Supplies the Hartree-Fock orbitals and one- and two-electron integrals used as input to the MLCI calculations.","marker":"[56]"}],"fun_headline_variants":["Neural network hash speeds quantum chemistry curves","MLCI: near-exact molecular curves with less processing","AI ranks electron configurations for accurate molecular curves","MLCI beats stochastic selection on processor hours","Machine learning selects configurations for efficient ab initio curves"],"cache_read_input_tokens":20096,"weakest_assumption_plain":"The method assumes the network's guesses about which electron configurations matter are trustworthy enough that discarding low-ranked candidates on the fly loses none of the important ones.","fun_headline_variants_meta":{"raw":{"variants":["Neural network hash speeds quantum chemistry curves","MLCI: near-exact molecular curves with less processing","AI ranks electron configurations for accurate molecular curves","MLCI beats stochastic selection on processor hours","Machine learning selects configurations for efficient ab initio curves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3232,"prompt_tokens":954,"completion_tokens":2278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":2217}},"tokens_in":570,"tokens_out":2278,"duration_ms":16641,"temperature":1.0,"reasoning_tokens":2217,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:19:21.106166+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the hash-based streaming selection and the full-space quicksort selection at every geometry of a potential curve and compare the final energies and configuration sets; if any geometry gives different energies, or if any configuration in the sorted-selection wavefunction would have been rejected early by the streaming rule, the central claim would be contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the non-parallelity error used to quantify how well the shape of a shifted potential curve matches FCI."}],"review_version":1}