{"id":"371c79dd-ff2d-4056-8091-0befddb14f25","arxiv_id":"2608.02119","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The authors use a genetic algorithm to reoptimize STO-kG minimal basis sets up to k=11, yielding MSTO-kG bases that improve atomic FCI energies for H, Li, Be and some molecules, with fewer qubits than larger bases.","lead":"This paper re-optimizes tiny 'minimal' basis sets used in quantum chemistry, getting better atomic and molecular energies than the standard 6-31G basis for some light atoms while using the same or fewer qubits. It matters because near-term quantum computers can only handle a small number of orbitals, so a basis set that gives more accuracy per qubit could make quantum chemistry calculations feasible sooner.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Li 'surpass cc-pVQZ' claim rests on cc-pVQZ table entries that are numerically identical to 6-31G, which is physically impossible; a correct cc-pVQZ FCI value would almost certainly be lower than MSTO-11G.","rationale":"The strongest claim in the paper is the atomic comparison, and the Li 'surpass cc-pVQZ' statement is its most eye-catching quantitative assertion. That assertion is directly tied to Table III, where the cc-pVQZ values are identical to the 6-31G values. Such an identity is physically impossible for a full-basis FCI calculation and is very likely a table-generation artifact. The reader's weakest assumption identified the broader problem of comparing at equal active-space size rather than full-basis FCI; this stress-test narrows that concern to a concrete, checkable data error that alone invalidates the flagship Dunning comparison. I therefore do not think the reader's CONDITIONAL verdict needs to change: the underlying basis-set optimization idea may still be useful, and the Li versus 6-31G resource comparison may survive, but the paper must correct the Dunning entries and rephrase the abstract's cc-pVQZ claim before it can be accepted. I also note that Appendix A lists only l=0 parameters for all atoms, even though the resource tables count 10 spin orbitals for Li, implying the p functions are inherited from STO-kG; if the repository omits them, that is a separate reproducibility issue, but the decisive flaw is the invalid Dunning comparison values.","tokens_in":28178,"tokens_out":14986,"duration_ms":137318,"concrete_test":"Recompute the Li atom FCI energy in the full cc-pVQZ basis with PySCF using the basis from Basis Set Exchange, and also compute FCI with the MSTO-11G basis from the paper's GitHub repository. If full cc-pVQZ FCI is below -7.452345 (expected, since the exact nonrelativistic limit is -7.478060), the abstract's 'surpass cc-pVQZ' claim is false. If the authors intended an equal-qubit comparison, repeat with a 10-spin-orbital active space chosen by the same orbital-selection rule for cc-pVQZ and MSTO, and verify that the resulting cc-pVQZ value differs from the 6-31G value in Table III. Also recompute the H, Be, B, C, N, O, and F cc-pVTZ/cc-pVQZ rows to confirm they do not equal the corresponding 6-31G entries.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Tables II–IX, every cc-pVTZ and cc-pVQZ HF/FCI entry is numerically identical to the corresponding 6-31G entry (e.g., Li: -7.431235/-7.431554; H: -0.498233; F: -99.360218/-99.447423). This cannot be a full-basis FCI result: cc-pVQZ for Li should be within a few mHa of the exact nonrelativistic energy (-7.478060 Ha), not 46 mHa above it. The abstract's flagship claim that MSTO bases 'surpass the performance of cc-pVQZ' for Li therefore compares MSTO-11G (-7.452345 Ha) against an invalid reference value. The paper's own benchmark, Hylleraas-infinity (-7.478060 Ha), shows MSTO-11G is 25.7 mHa above the exact result, so a genuine cc-pVQZ FCI would very likely be lower than MSTO-11G. The same copy artifact affects the 'comparable to 6-31G' comparisons: for C–F the full 6-31G FCI entries are 16–56 mHa below MSTO, contradicting the abstract's blanket claim. The central 'better accuracy with fewer qubits' story may survive for Li versus 6-31G, but the Dunning-based headline is unsupported as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes re-optimizing minimal STO-kG basis sets (k=2–11) for atoms H through F (excluding He) using a memetic algorithm, with the goal of improving quantum-chemistry energies on quantum computers while keeping the number of spin orbitals (and hence qubits) fixed. The authors report atomic FCI energies for MSTO-kG bases, compare them with STO-kG, 6-31G, and Dunning basis sets, present molecular potential energy curves for H2, Li2, C2, LiH, BeH, and BeH2, and estimate quantum resources for VQE-UCCSD, QPE-CASCI, and HHL-LCCSD using Li as a representative example. The central claim is that MSTO bases achieve better or comparable energies than larger standard bases with fewer qubits and gates.","tokens_in":28594,"tokens_out":5930,"duration_ms":53325,"significance":"The idea of one-time classical preprocessing to design qubit-efficient basis sets is genuinely useful for near-term quantum chemistry, and the open-source implementation (https://github.com/subimal/MSTO-kG) is a commendable contribution. The resource-scaling analysis for VQE, QPE, and HHL is transparent and includes explicit gate-count formulas. However, as presented, the headline numerical comparisons are not reliable: several table entries for the Dunning basis sets are physically implausible, and the abstract's blanket statement that MSTO FCI energies are 'comparable or sometimes even lower' than 6-31G is contradicted by the paper's own tables for B, C, N, O, and F. The underlying methodology is defensible, but the manuscript needs substantial corrections and clarifications before its central claims can be considered sound.","major_comments":[{"comment":"The cc-pVTZ and cc-pVQZ entries in Tables II–IX are numerically identical to the 6-31G entries for every atom (e.g., Li: -7.431235/-7.431554; F: -99.360218/-99.447423). A genuine full-basis FCI calculation with cc-pVQZ cannot produce exactly the same energy as 6-31G; for Li, the tabulated cc-pVQZ FCI value of -7.431554 Ha is 46.5 mHa above the Hylleraas-infinity value of -7.478060 Ha that the paper itself quotes in Table III. Therefore the statement in Section IV.A and in the abstract that MSTO-11G (-7.452345 Ha) 'surpasses the performance of cc-pVQZ' is not supported by the data as printed. The authors must recompute the Dunning-basis values, or explicitly state and justify if these are truncated-active-space results, and revise all claims based on them.","section":"Appendix B, Tables II–IX; Section IV.A"},{"comment":"The abstract states: 'The ground state energies of H through F using our MSTO bases at FCI level of theory yield ground state energies that are comparable or sometimes even lower than those obtained using 6-31G basis sets.' This is contradicted by the paper's own data: MSTO-11G is higher than the 6-31G FCI energy by 6.8 mHa for B, 16.3 mHa for C, 30.4 mHa for N, 49.9 mHa for O, and 56.0 mHa for F. The Conclusion (Section VII) already concedes the limitation for N, O, and F. The abstract and the corresponding sentences in the Introduction should be corrected to reflect the actual, element-dependent behavior.","section":"Abstract and Section IV.A, Tables V–IX"},{"comment":"The comparison between MSTO bases (10 spin orbitals) and the reference bases 6-31G, cc-pVDZ, cc-pVTZ, and cc-pVQZ is ambiguous: the text discusses a 10-spin-orbital active space for the reference bases in Table I, but the tables in Appendix B do not state whether the reported 6-31G and Dunning energies are full-basis FCI or truncated-active-space FCI. Because the central claim of 'better accuracy with fewer qubits' depends on the reference calculations being the standard full-basis results, the authors must specify the level of calculation for every table entry and ensure that the comparison is consistent.","section":"Section IV.A and Table I"},{"comment":"The resource analysis uses FCI energies in each basis as proxies for the energies that VQE-UCCSD, QPE-CASCI, and HHL-LCCSD would produce. The authors disclose this approximation, but the resulting claim that MSTO bases 'yield better energies than the competing basis sets while incurring fewer qubits' rests on the energy values that are in question due to the Dunning-table errors and on the assumption that the quantum algorithms recover the full FCI correlation energy. The authors should re-evaluate the resource comparison using corrected energies and should discuss the validity of the FCI proxy for each algorithm, particularly for UCCSD and LCCSD, which are approximate methods.","section":"Section V, resource estimation"},{"comment":"The basis parameters are optimized by minimizing CISD energies of the same atoms that are subsequently evaluated at the FCI level. This creates a circularity in the interpretation of the atomic energy improvements over the unoptimized STO bases: the improvement partly reflects the optimization target rather than an independent predictive gain. The authors should clarify how much of the reported atomic improvement is independent, for example by reporting results for molecules or for test atoms not used in the optimization. This does not invalidate the approach, but it is load-bearing for the atomic comparisons as currently presented.","section":"Section III and Section IV.A"}],"minor_comments":[{"comment":"The header contains a typo: 'Hyleraas' should be written as 'Hylleraas'.","section":"Table III header"},{"comment":"Equation (2) is presented without clear definitions of all symbols beyond the text; please add a sentence explaining that s_ij = alpha_i + alpha_j and that the normalization condition was used to determine d_i.","section":"Section II, Eq. (2)"},{"comment":"The notation is inconsistent: 'STO-KG', 'STO-kG', and 'MSTO-kG' are all used. Please standardize to 'STO-kG' and 'MSTO-kG'.","section":"Throughout"},{"comment":"The Pauli-term counts for STO-6G and MSTO-11G are reported as identical (156) because both have 10 spin orbitals; please mention that this is expected due to the same number of spatial orbitals, not because the integrals are identical.","section":"Section V, Table XI and XII"},{"comment":"Some reference entries are incomplete or inconsistent (e.g., Ref. [34] lacks full page range, Ref. [7] is an arXiv identifier without a title). Please harmonize the bibliography style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The identical cc-pVTZ/cc-pVQZ entries in Tables II–IX strongly suggest a data-handling or copy-paste error. The editor may wish to ask the authors to provide the calculation logs or scripts that generated these tables, and to confirm whether the reported Dunning-basis energies were ever actually computed. The open-source code is a positive feature, but the paper's consistency must be restored before it is suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core here is real: a memetic optimization of STO-kG exponents and contraction coefficients up to k=11 for H–F, with open-source code and basis parameters released. That is a concrete extension of Andrade et al. and a legitimate tool for qubit-limited quantum chemistry. But you can't trust the abstract as written.\n\nWhat is genuinely new: the systematic extension to k=11, the resource accounting for VQE, QPE, and HHL, and the molecular PECs. The H2 failure is honestly flagged and consistent with earlier work. The Li example—same qubits as STO-6G, better FCI energy than 6-31G in the 10-orbital space—is a meaningful demonstration of the core idea. The code being public is real evidence and deserves credit.\n\nThe soft spots are serious but fixable. First, the cc-pVTZ and cc-pVQZ columns in every table are numerically identical to the 6-31G entries. No physical basis gives Li cc-pVQZ FCI = -7.431554; the exact nonrelativistic energy is -7.478060, and the paper's own Hylleraas line shows MSTO-11G is 25.7 mHa above exact. So the abstract's claim that MSTO 'surpasses cc-pVQZ' for Li is unsupported by the paper's own data. Second, the abstract says FCI energies for H–F are comparable to or lower than 6-31G, but the tables show C, N, O, F are 16–56 mHa higher, and the conclusion later concedes N, O, F are higher. The abstract and conclusion contradict each other. Third, Appendix A lists only s-type exponents and contraction coefficients for every element. If p functions were left at their STO values, that should be stated; if they were optimized, the tables are missing data. Either way, reproducing the B–F results is not possible from the released basis parameters as presented. The active-space truncation issue in Table I also shows the 'fewer qubits' advantage is a trade-off for B–F, not a free lunch; it works for Li, and the paper should say so.\n\nThe circularity of optimizing against atomic CISD and then re-evaluating the same atoms at FCI is not a real flaw—it is a legitimate design objective, and the molecular tests are independent. The resource estimates are order-of-magnitude, clearly labeled, and adequate for comparing basis sets.\n\nWho is this for? People doing near-term quantum chemistry resource estimation and basis set design. It deserves serious peer review because the method is sound and the code is real, but the authors need to correct the Dunning tables, revise the abstract to match their own tables, and complete the basis parameter listings. The fixable errors are concentrated in presentation and honesty of the headline claims, not in the underlying approach.","headline":"Useful basis-optimization pipeline and open-source data, but the headline accuracy claims rest on Dunning table columns that are plain copy-paste errors.","tokens_in":29047,"tokens_out":7054,"would_cite":false,"duration_ms":60075,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that re-optimized minimal STO-kG basis sets reach or beat 6-31G-quality atomic energies with only 10 spin orbitals, and for lithium exceed cc-pVQZ.","keywords":["modified minimal basis sets","basis set optimization","quantum chemistry","quantum computing","variational quantum eigensolver","quantum phase estimation","contracted Gaussian basis","correlation energy"],"falsifier":"Repeat the Li FCI calculation with the full cc-pVQZ basis, with all spin orbitals active, and compare against the MSTO-11G result of $-7.452345\\,\\mathrm{Ha}$; if the full cc-pVQZ FCI energy is lower, the paper's claim that MSTO surpasses cc-pVQZ fails. Likewise, run full 6-31G FCI for B, C, N, O, and F and check whether MSTO energies remain below the untruncated 6-31G values.","tokens_in":27972,"feed_emoji":"⚛️","tokens_out":5254,"duration_ms":47527,"temperature":0.7,"pith_summary":"This paper argues that re-optimizing the exponents and contraction coefficients of minimal STO-kG basis sets gives atomic ground-state energies comparable to or better than the much larger 6-31G basis while using the same 10 spin orbitals as the minimal basis, and for lithium even surpassing cc-pVQZ. Because contracted minimal bases keep the spin-orbital count fixed as the number of Gaussians k grows, the improvement costs no extra qubits in a quantum algorithm. The authors generate MSTO-kG bases (k=2 through 11) for H through F with a genetic-algorithm-plus-refinement scheme and benchmark FCI energies, plus VQE, QPE, and HHL resource counts for lithium. If correct, the approach gives near-term quantum chemistry a better accuracy-per-qubit trade-off than choosing a standard larger basis and truncating its active space.","feed_headline":"Reoptimized minimal bases beat bigger sets with fewer qubits","feed_subtitle":"For H through F, contracted MSTO-kG energies match or beat 6-31G with 10 spin orbitals instead of 18.","key_machinery":"The load-bearing object is the contracted Gaussian minimal basis: each atomic orbital is a fixed linear combination of k zero-centered Gaussians, so the number of contracted functions, and hence the number of spin orbitals and qubits, is independent of k. The memetic algorithm—a genetic algorithm with elitist selection, interpolative mutation, and discrete-parameter crossover, followed by parallel aggressive refinement—tunes the k exponents and k contraction coefficients against the CISD ground-state energy. Because k can grow to 11 without enlarging the orbital space, the same qubit footprint yields progressively lower energies, which is the mechanism behind the claim of better accuracy at a fixed quantum-resource budget.","core_discovery":"The paper constructs modified minimal Slater-type-orbital basis sets, MSTO-kG for k=2 through 11, for the atoms H through F, excluding He, by memetic optimization of Gaussian exponents and contraction coefficients against CISD energies. At the FCI level these bases yield ground-state energies that are comparable to or lower than 6-31G for all the atoms studied, and for Li the MSTO-11G energy of $-7.452345\\,\\mathrm{Ha}$ falls below the truncated cc-pVQZ value of $-7.431554\\,\\mathrm{Ha}$, while using only 10 spin orbitals compared with 18 for 6-31G and 28 for cc-pVDZ. For molecules, the FCI results from MSTO bases are comparable to or better than 6-31G for Li$_2$, LiH, BeH, and BeH$_2$, and for C$_2$ at the CISD level, whereas H$_2$ is a known failure. Resource estimates for VQE-UCCSD, QPE-CASCI, and HHL-LCCSD on Li show far fewer qubits and CX gates with MSTO bases than with 6-31G or cc-pVDZ along with lower FCI energies.","pith_inferences":["Editorial inference: The same memetic recipe could be applied to polarization or split-valence bases to push the accuracy-per-qubit frontier further; the paper notes 6-31G is already near-optimal, so gains there would be smaller.","Editorial inference: The near-parity with cc-pVQZ for Li hints that re-optimized contracted bases may be systematically beneficial for one- and two-valence-electron atoms; extending to Na, K, or Mg would test this without new methodology.","Editorial inference: Because contracted basis size and qubit footprint are decoupled from k, basis-set optimization can be composed with qubit tapering and orbital-selection heuristics, potentially stacking resource savings.","Editorial inference: If the truncated-active-space benchmark is replaced by a full-basis comparison, MSTO's margin will shrink or invert; the practical claim is best read as 'best accuracy within a fixed small active space.'"],"forward_implications":["With MSTO bases, a quantum computation for these atoms uses 10 qubits instead of 18 or 28 while targeting lower energies, directly expanding what is feasible on near-term hardware.","The VQE-UCCSD gate count for Li drops from 17,420 CX gates with 6-31G to 2,246 with MSTO-11G, and QPE and HHL logical T-gate counts fall because gate estimates scale polynomially with spin-orbital count.","One-time classical basis-set optimization can be reused for any quantum-chemical calculation on molecules built from H through F atoms, including hybrid STO/MSTO combinations.","For N, O, and F the MSTO bases still lack sufficient virtual orbitals, so FCI correlation energy is near zero and a quantum computer gains nothing over Hartree-Fock there.","H$_2$ is a counterexample where MSTO underperforms both STO-6G and 6-31G, so the method is not universally better for molecules."],"supporting_citations":[{"why":"Supplies the original STO-kG basis sets, the starting point and benchmark that the memetic algorithm improves on.","marker":"[27]"},{"why":"Earlier simulated-annealing optimization of STO-3G and STO-6G bases; provides the comparison point and the known poor H2 behaviour.","marker":"[34]"},{"why":"Provides the electronic-structure package used for all HF, FCI, and CISD energy benchmarks in the paper.","marker":"[35]"},{"why":"Companion reference for the computational package, supporting the energy computations.","marker":"[36]"},{"why":"Supplies the high-accuracy Hylleraas benchmark for the Li atom used to contextualize the MSTO-11G result.","marker":"[37]"},{"why":"Introduces the VQE algorithm whose UCCSD resource counts are compared across basis sets.","marker":"[12]"},{"why":"Establishes quantum phase estimation for chemistry, the framework behind the QPE-CASCI resource estimates.","marker":"[2]"},{"why":"Gives the Pauli-gadget CX count used in the VQE and QPE gate estimates.","marker":"[40]"},{"why":"Provides the controlled-unitary decomposition count used for the QPE CX scaling expression.","marker":"[41]"}],"fun_headline_variants":["Optimized bases: same accuracy, fewer qubits","Fewer qubits, better energies via basis optimization","MSTO bases match big sets with fewer qubits","Basis tweaks cut qubit count for quantum chemistry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparisons to 6-31G and Dunning bases are made with those larger bases truncated to the same 10 spin orbitals; if one instead compares full active-space FCI in each basis, the larger bases recover much more correlation energy and the apparent advantage of MSTO bases may disappear.","fun_headline_variants_meta":{"raw":{"variants":["Optimized bases: same accuracy, fewer qubits","Fewer qubits, better energies via basis optimization","MSTO bases match big sets with fewer qubits","Basis tweaks cut qubit count for quantum chemistry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":3124,"prompt_tokens":1182,"completion_tokens":1942,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":798,"completion_tokens_details":{"reasoning_tokens":1877}},"tokens_in":798,"tokens_out":1942,"duration_ms":13705,"temperature":1.0,"reasoning_tokens":1877,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:01:18.377120+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the Li FCI calculation with the full cc-pVQZ basis, with all spin orbitals active, and compare against the MSTO-11G result of $-7.452345\\,\\mathrm{Ha}$; if the full cc-pVQZ FCI energy is lower, the paper's claim that MSTO surpasses cc-pVQZ fails. Likewise, run full 6-31G FCI for B, C, N, O, and F and check whether MSTO energies remain below the untruncated 6-31G values.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original STO-kG basis sets, the starting point and benchmark that the memetic algorithm improves on."},{"cited_title":"Nagy and F","cited_arxiv_id":null,"evidence_quote":"Earlier simulated-annealing optimization of STO-3G and STO-6G bases; provides the comparison point and the known poor H2 behaviour."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the electronic-structure package used for all HF, FCI, and CISD energy benchmarks in the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Companion reference for the computational package, supporting the energy computations."},{"cited_title":"Shavitt, in Methods in Computational Physics: Ad- vances in Research and Applications, edited by B","cited_arxiv_id":null,"evidence_quote":"Supplies the high-accuracy Hylleraas benchmark for the Li atom used to contextualize the MSTO-11G result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes quantum phase estimation for chemistry, the framework behind the QPE-CASCI resource estimates."},{"cited_title":"Walton, O","cited_arxiv_id":null,"evidence_quote":"Gives the Pauli-gadget CX count used in the VQE and QPE gate estimates."},{"cited_title":"Chinnasamy, M","cited_arxiv_id":null,"evidence_quote":"Provides the controlled-unitary decomposition count used for the QPE CX scaling expression."}],"review_version":2}