{"id":"d53ae784-b4f4-40d7-b096-e11b5369518d","arxiv_id":"2607.20585","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An RBM-guided selected-CI solver inside DMET reaches the DMET-CASCI energy within 1.6 mHa using ~4% of the symmetry-valid configuration subspace on an 11-fragment protein–ligand model.","lead":"Researchers combined a machine-learning model with quantum sampling to select a compact set of molecular configurations for computing ground-state energies inside a quantum embedding framework. On a SARS-CoV-2 protease–drug complex, the approach reached the target accuracy while using about 4% of the configuration space, but the comparison baseline was not allowed to converge.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central 'DMET-SQD failed' claim rests on a halted, non-converged baseline; convergence behavior at the same εspb is untested, so the compactness advantage may be overstated.","rationale":"The strongest claim and abstract hinge on the contrast between DMET-QSCI-RBM's ~4% subspace success and DMET-SQD's 'failure' at 19–20%. That failure is not established because the SQD run was stopped early; no evidence shows it would not have converged with further iterations. This is a load-bearing concern because it directly affects the central comparative statement, not just the internal mechanism. The reader's weakest assumption focuses on the RBM's equal-weighted training, which is also a valid concern about whether the learned proposal is essential, but the non-converged baseline is more immediately damaging to the headline empirical claim. The proposed test—continuing the SQD loop—would settle whether the comparison is fair; if SQD eventually converges, the paper's conclusion must be revised, while if it remains above chemical accuracy, the claim is strengthened. Thus the reader's CONDITIONAL verdict remains appropriate pending this test, and I partially agree with the reader's identified weakest assumption.","tokens_in":29542,"tokens_out":5081,"duration_ms":55111,"concrete_test":"Run the DMET-SQD (εspb = √|S|/2) self-consistency loop for at least 10–15 µ-iterations (or until |E − E_DMET−CASCI| stabilizes or crosses below 1.6 mHa) using the same fragment set, basis, and hardware/noise conditions (or a validated classical noise model calibrated to the reported April 2026 calibration data). Record the subspace fraction at the first iteration satisfying chemical accuracy, or report the plateau error if none. If a crossing occurs, the 'SQD failed' claim is refuted and the comparison must be restated as 'SQD requires X% vs RBM 4%.' If the error remains above 1.6 mHa for all iterations, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 reports that the truncated DMET-SQD baseline (εspb = √|S|/2) was halted after five µ-iterations 'owing to the prohibitive QPU time cost,' with |E − E_DMET−CASCI| stuck at ~10^-3–10^-2 Ha and non-monotonic in µ. The paper explicitly labels this run 'Not Converged.' The central comparative claim—that standard DMET-SQD 'failed to reach chemical accuracy' while accessing up to 19–20% of the subspace—therefore depends on the assumption that this run would never converge if continued. The observed decoupling between a small µ-residual and large energy error does not establish that; it only shows µ-convergence is not sufficient for energy convergence. If continued iterations brought the energy below 1.6 mHa at, say, 25% subspace, the headline advantage would shrink from 'failed vs 4%' to '25% vs 4%' and the abstract's 'failed' would be false. The comparison is also asymmetric: DMET-QSCI-RBM was itself halted after only three µ-iterations, so the claimed advantage is not demonstrated against a converged SQD baseline at comparable truncation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes QSCI-RBM, an iterative RBM-guided subspace-expansion protocol for sample-based quantum diagonalization (SQD/QSCI), and integrates it into a DMET embedding loop as the impurity solver (DMET-QSCI-RBM). The RBM is trained on the current determinant memory with equal per-determinant weights; Gibbs sampling proposes new determinants, which are symmetry-filtered, proliferated over spin strings, and selected into a persistent memory by CI-coefficient thresholding. The method is first validated on C2H4 and CH5NO against FCI, then applied to the Carmofur/SARS-CoV-2 Mpro complex fragmented into 11 DMET impurities with a (16e,16o) active space on IBM Heron hardware. The central reported result is that DMET-QSCI-RBM reaches chemical accuracy relative to DMET-CASCI in two independent runs (errors 9.325e-4 and 4.257e-4 Ha) while accessing ~3.9% of the symmetry space, whereas a truncated DMET-SQD baseline at εspb = sqrt(|S|)/2 accessed ~19% without reaching chemical accuracy and an effectively untruncated DMET-SQD baseline (εspb = 1e8) converged only at ~97% of the symmetry space.","tokens_in":29936,"tokens_out":6001,"duration_ms":58984,"significance":"If the compactness claim holds, the result is practically significant: it addresses the classical diagonalization cost that dominates SQD-based DMET simulations and demonstrates a hardware-scale application to a protein-ligand complex. The paper has several strengths: two independent hardware re-solves at the final chemical-potential point; comparison against a DMET-CASCI reference; small-molecule validation against FCI; transparent reporting of the halted runs and the reuse of shared hardware sampling data for the first two DMET chemical-potential iterations. However, the main comparative claim against standard DMET-SQD rests on a non-converged baseline, and the paper does not establish that the RBM itself, rather than the iterative CI-threshold memory update, is responsible for the improved compactness. Both of these points are load-bearing for the abstract's claims.","major_comments":[{"comment":"The central comparison against standard DMET-SQD is not established. The εspb = sqrt(|S|)/2 run is explicitly labeled 'Not Converged' and was halted after five µ-iterations 'owing to prohibitive QPU time cost.' Non-monotonic E(µ) over five iterations is not evidence that the run would never converge; a truncated SCI solver can oscillate before eventually converging. The abstract's statement that standard DMET-SQD 'failed to reach chemical accuracy' therefore overstates what the data show. The comparison is also asymmetric: DMET-QSCI-RBM itself was halted after three µ-iterations. Please either continue the truncated SQD run to convergence (or run it at a less restrictive cap that still limits the subspace) or soften all claims to 'did not reach chemical accuracy within the attempted iterations.' This is not a presentation issue; it underpins the headline compactness advantage.","section":"§4.3, Table 1, Fig. 8"},{"comment":"The paper's mechanism claim — that the RBM 'learns the underlying probability distribution of dominant determinants' — is not directly supported. Step 5 trains the RBM with equal weights wφ = 1 on the current memory, not on |cφ|², so the RBM is not learning the ground-state determinant distribution. The persistent memory is instead updated by CI-coefficient thresholding (Eq. (4)), and the paper does not provide an ablation separating the contribution of RBM-proposed determinants from that of the raw hardware samples plus spin-string proliferation. Section 4.7 explicitly leaves open whether the high-excitation hardware samples 'contribute meaningfully to the converged ground-state wavefunction.' Without an ablation (e.g., replacing the RBM proposals with random spin-adapted determinants or with the raw hardware singletons at fixed memory size), the observed compactness cannot be attribute","section":"§2.1, Eq. (4); §4.7"},{"comment":"The DMET-QSCI-RBM result is reported at a halted chemical-potential trajectory (residual ~1.4e-4, not self-consistent), and the paper draws a strong conclusion that µ-convergence is not necessary for energy accuracy. This conclusion is based on two runs at a single non-equilibrium µ. The observation is interesting, but the paper should either demonstrate stability of the energy under continuation of the µ loop or explicitly frame the result as a proof-of-principle at a non-equilibrium embedding potential. The abstract and conclusion should state symmetrically that both DMET-SQD (εspb = sqrt(|S|)/2) and DMET-QSCI-RBM were halted before self-consistency, and that the advantage is measured at this halting point.","section":"§4.4–4.5, Fig. 9"}],"minor_comments":[{"comment":"The molecule is called 'methanolamine' in the main text and 'methoxyamine' in the Supplementary Material; the chemical formula CH5NO corresponds to methoxyamine. Please unify the nomenclature.","section":"§3.3 and Supplementary A"},{"comment":"The text says all Carmofur experiments used 'IBM Boston Heron R3', but the captions of Figs. 8 and 9 say 'IBM Heron Fez'. The Supplementary calibration table also lists Boston. Please correct the captions.","section":"§4.1, §4.2, Figs. 8–9"},{"comment":"The notation S.S. is used interchangeably with |S| in several places, and the caption of Fig. 10 refers to 'S.S' without defining it. Please define the symmetry space once and use consistent notation.","section":"§2.1, Eq. (5)"},{"comment":"The non-converged DMET-SQD energy lies below the DMET-CASCI reference energy by 22 mHa. A one-sentence explanation (e.g., effect of the non-converged chemical potential or of the truncated subspace on the embedding energy) would prevent the apparent variational violation from confusing readers.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: the method seems to work, but the paper's flagship comparison—\"DMET-SQD failed\"—is built on a baseline the authors themselves label 'Not Converged.' The actual empirical contribution is narrower: on an 11-fragment Carmofur/Mpro model, two independent DMET-QSCI-RBM runs at a halted chemical potential land within 0.94 and 0.43 mHa of DMET-CASCI while using roughly 3.4–3.9% of the symmetry-valid subspace. That is a genuine hardware demonstration, and the small-molecule validation (C2H4, CH5NO) shows smooth, monotonic convergence to chemical accuracy. The paper also does the resource accounting properly: circuit depths, QPU times, calibration data, and the fact that the first two micro-points shared hardware data across solvers. It is honest about several limitations, including that it does not establish whether the high-excitation hardware samples actually matter for the converged state.\n\nThe soft spots are real. The central \"standard SQD failed\" claim rests on DMET-SQD run with eps_spb = sqrt(|S|)/2 that was halted after five micro-iterations \"owing to prohibitive QPU time cost\" and explicitly marked not converged, with non-monotonic energy error stuck at ~1e-2 to 1e-3 Ha. That shows micro-convergence is not sufficient for energy convergence; it does not show that continuing that baseline would never reach chemical accuracy. The QSCI-RBM run was itself halted after three micro-iterations, so the comparison is between two unfinished trajectories, one labeled failed, the other labeled successful. The abstract's \"failed to reach chemical accuracy\" is too strong. Also, the RBM is trained with equal weights on the current determinant memory, not on |c|^2; the phrase \"learn the underlying probability distribution\" overstates what the model does. No code or data are released, RBM hyperparameters are not reported, and there are no error bars beyond two runs.\n\nThat said, the core loop is not circular: final energies are checked against an external DMET-CASCI reference. And the converged eps_spb = 1e8 DMET-SQD run at ~97% subspace provides a useful reference point. The method is a legitimate extension of prior RBM/SQD/DMET work, not a breakthrough, and the central comparison needs to be reworked. But it is a coherent, honestly written paper with real hardware data, and the flaws are addressable. I would send it to a competent referee rather than desk-reject. A revised version that continues the SQD baseline, reports RBM settings, and softens the comparison would be worth citing.","headline":"The core method has real hardware evidence behind it, but the paper's central 'DMET-SQD failed' comparison is built on a baseline the authors themselves label 'Not Converged,' so the abstract overstates the case.","tokens_in":30415,"tokens_out":2944,"would_cite":true,"duration_ms":28421,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a Restricted Boltzmann Machine trained on quantum-sampled determinants can propose compact, physically relevant configuration subspaces for selected configuration interaction, letting a DMET-embedded calculation reach","keywords":["sample-based quantum diagonalization","quantum selected configuration interaction","density matrix embedding theory","restricted Boltzmann machine","configuration subspace compactness","chemical accuracy","protein-ligand complex","hybrid quantum-classical algorithm"],"falsifier":"Run QSCI-RBM side by side with a version where the RBM is replaced by uniform random selection of symmetry-valid determinants at the same memory size and shot budget; if the random version matches or beats the RBM version at equal subspace fraction on C2H4 or the protein fragments, the compactness claim fails. A second check: retrain the RBM on |c|²-weighted determinants instead of equal weights and see whether the subspace needed for chemical accuracy shrinks, grows, or stays the same.","tokens_in":29438,"feed_emoji":"🧪","tokens_out":6388,"duration_ms":57528,"temperature":0.7,"pith_summary":"The paper sets out to fix the main bottleneck in sample-based quantum diagonalization: the recovered configuration subspace is often larger than necessary and not optimally concentrated on the determinants that matter. It introduces QSCI-RBM, a loop in which a Restricted Boltzmann Machine learns from the determinants already accumulated from quantum sampling, Gibbs-samples new symmetry-valid candidates, and the enlarged set is diagonalized classically after a spin-string proliferation step. Embedded in DMET fragments of a protein-ligand complex, the method reaches chemical accuracy while accessing only about 3.9% of the symmetry-preserving space, whereas the standard DMET-SQD baseline fails at about 19% and only reaches chemical accuracy by accessing nearly the whole space. A sympathetic reader would care because the classical diagonalization cost in embedding workflows grows with subspace size, so a learned proposal that concentrates on dominant determinants attacks the scaling wall directly without deeper quantum circuits or perturbative seeding.","feed_headline":"Learned sampler hits chemical accuracy at 4% of configuration space","feed_subtitle":"RBM-guided determinant selection reaches chemical accuracy for a protein-ligand complex at a fraction of the subspace.","key_machinery":"The carrying mechanism is an iterative RBM-guided subspace expansion. A Restricted Boltzmann Machine—an energy-based generative model with visible and hidden binary units—is retrained each iteration on the current determinant memory (all retained determinants weighted equally, excluding the Hartree-Fock reference). Its Gibbs samples are filtered for electron-number and spin symmetry and for novelty, unioned with the memory, and diagonalized after spin-string proliferation, which forms the tensor product of the unique α- and β-spin strings to surface cross-paired determinants never explicitly sampled. The top determinants by squared CI coefficient are then appended persistently to the memory.","core_discovery":"The central discovery is that machine-learned configuration generation, trained with equal weights on the current determinant memory rather than on the CI coefficients themselves, produces a determinant subspace whose energy per accessed determinant is far higher than configuration-recovery-based selection. The paper demonstrates this on the 11-fragment Carmofur–SARS-CoV-2 Mpro complex, where QSCI-RBM reaches chemical accuracy at roughly 3.9% of the symmetry-preserving space, and on small molecules, where chemical accuracy is reached at 2.12% (C2H4) and 0.06% (CH5NO) of the full symmetry space. A second finding is that energy accuracy is governed by the quality and compactness of the accumul","pith_inferences":["The key untested assumption is that the RBM's proposals beat random symmetry-valid selection; a direct ablation replacing the RBM with uniform random proposals at equal memory size would settle whether the compactness comes from learning or from the iterative memory/proliferation machinery.","Because the hardware samples concentrate at excitation orders 4–7, the method may be most useful in multi-reference regimes beyond doubles-restricted perturbation theory; testing on a strongly correlated system with known higher-order dominance would sharpen that claim.","The paper's own admission that hyperparameters were chosen empirically suggests subspace size at chemical accuracy is an upper bound, not a tuned optimum; a systematic search could push compactness further or reveal sensitivity.","Cross-fragment transfer of the learned distribution is unexamined; if the RBM must be retrained from scratch per fragment and µ-iteration, the classical overhead of training could offset some diagonalization savings at scale."],"forward_implications":["If determinant-memory quality rather than µ self-consistency governs accuracy, DMET loops can be halted after a few chemical-potential values, cutting rounds of quantum sampling.","Classical diagonalization cost per fragment falls roughly in proportion to subspace size, giving about a 5× reduction versus the non-converged SQD baseline and about 25× versus the converged one.","Chemical accuracy can be reached without MP2 or other perturbative seeding, directly from hardware samples and their proliferated partners.","The chemical-potential residual is not a reliable convergence proxy for truncated-subspace solvers; subspace-quality metrics must accompany it.","On small molecules, the approach beats the CCSD reference error for C2H4 and reaches chemical accuracy for CH5NO while accessing only 0.06% of a 6×10^9-determinant space."],"fun_headline_variants":["RBM-trained sampler hits chemical accuracy at 4% of space","AI-driven determinant selection: 4% subspace, chemical accuracy","Quantum RBM reaches chemical accuracy at 4% configuration space","Learned subspace generation: 4% of space gets chemical accuracy"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that Gibbs-sampled determinants from an RBM trained with equal weights on the current memory, after spin-string proliferation, actually span the dominant correlation space; if the learned proposals are no better than random symmetry-valid determinants, the compactness advantage disappears and the method reduces to a random sampler—and the paper itself leaves open whether the high-excitation hardware samples contribute meaningfully to the converged","fun_headline_variants_meta":{"raw":{"variants":["RBM-trained sampler hits chemical accuracy at 4% of space","AI-driven determinant selection: 4% subspace, chemical accuracy","Quantum RBM reaches chemical accuracy at 4% configuration space","Learned subspace generation: 4% of space gets chemical accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1386,"prompt_tokens":826,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":487}},"tokens_in":570,"tokens_out":560,"duration_ms":5845,"temperature":1.0,"reasoning_tokens":487,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:45:43.907402+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run QSCI-RBM side by side with a version where the RBM is replaced by uniform random selection of symmetry-valid determinants at the same memory size and shot budget; if the random version matches or beats the RBM version at equal subspace fraction on C2H4 or the protein fragments, the compactness claim fails. A second check: retrain the RBM on |c|²-weighted determinants instead of equal weights and see whether the subspace needed for chemical accuracy shrinks, grows, or stays the same.","supporting_citations":[],"review_version":1}