{"id":"71f86890-6de9-4c1b-bce4-7516e09fce3d","arxiv_id":"2608.11569","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Uniform random classical sampling can match noisy SQD benchmarks when the diagonalization subspace grows unchecked, and multi-basis NOCI measurement improves sample efficiency under fixed classical budgets.","lead":"A quantum chemistry method called sample-based quantum diagonalization (SQD) can appear to work well because the classical computer is doing more work than reported, not because the quantum sampling is smart. The paper proposes measuring in several optimized orbital bases to make the quantum sampling find better configurations with the same classical budget.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"NOCI advantage is benchmarked at fixed diagonalization dimension d, but NOCI-SQD spends far more classical work per matrix element; the 'fixed classical resource budget' claim is not actually controlled.","rationale":"The reader's weakest assumption concerns the qiskit-sqd-addon default spin-product expansion and whether literature SQD experiments use that setting. That is a legitimate caveat, but it mainly weakens the interpretation of prior benchmarks; the normative claim that d must be controlled survives even if the default differs. The resource-control discrepancy, by contrast, affects the paper's principal positive contribution: the NOCI-SQD advantage is presented as persisting under fixed classical resource budgets, yet the benchmarks hold d fixed while Table I shows per-matrix-element costs differing by orders of magnitude. This is not a disagreement with the field's consensus; it is an internal inconsistency between the stated resource-fairness claim and the authors' own complexity accounting. The paper does honestly list the NOCI overhead in Table I and the Discussion, so the fix is straightforward: either relabel the claim as 'same diagonalization size' or provide FLOP-controlled comparisons. I would keep the reader's CONDITIONAL verdict but make the explicit condition that Figs. 4-5 be repeated with matched total classical cost, including NOCI basis construction, and that statistical error bars from the ten trials be reported. Without this, the central positive claim about higher-quality configurations is not fully supported.","tokens_in":20492,"tokens_out":10918,"duration_ms":132579,"concrete_test":"Recompute the H12 STO-6G benchmarks of Fig. 5 under an explicit total-FLOP budget for the classical post-processing. For each method, use Table I's cost model: standard SQD cost is roughly d^2 times the sparse Slater-Condon element cost, while NOCI-SQD cost is roughly d^2 times O(N_cdf * kappa^3) plus O(M^2 * N_cdf * kappa^3) for basis construction. Reduce d for NOCI-SQD until its classical cost matches the HF/LUCJ SQD cost at the plotted d values, then compare final energy errors. If the NOCI advantage vanishes or reverses, the 'fixed classical resource budget' claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The positive NOCI claim rests on resource control that does not control resources. The abstract says improvements persist under fixed classical resource budgets, and Section III says improvements persist after controlling for classical resources required for diagonalization. In Figs. 4 and 5, the only controlled quantity is d, the diagonalization subspace dimension. Table I, however, gives the per-matrix-element cost of NOCI-SQD as O(N_cdf * kappa^3), versus O(1) to O(N_e^2) for standard SQD, and notes that NOCI projected matrices are dense while standard SQD matrices are typically sparse, with NOCI memory cost 2O(d^2). Thus equal d does not imply equal classical cost: a NOCI run can spend orders of magnitude more classical work per retained configuration. The conclusion that NOCI configurations are 'higher quality rather than more numerous' is therefore not established by the presented benchmarks; the energy improvement could be purchased with extra classical work per configuration. If the intended claim is only 'same d,' it should be stated that way; as written, the resource-fairness claim is an internal inconsistency with the paper's own cost model.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses benchmarking and improvement of sample-based quantum diagonalization (SQD) for molecular ground-state energies. The authors simulate SQD with LUCJ ansatz states on N2, H8, and H12 and show that when the classical diagonalization subspace is allowed to grow through spin-product expansion and configuration recovery, uniform random sampling can match or exceed noisy SQD, with the apparent improvement tracking the growth of the diagonalization dimension d. They then propose NOCI-SQD, in which measurements are distributed over multiple classically optimized non-orthogonal orbital bases, and report lower energy errors at fixed d for H12 compared with Hartree-Fock-only measurement, while N2 shows only diversity gains without energy improvement. The paper concludes that fair SQD benchmarking must control d and that measurement-basis engineering is a promising direction for improving quantum sampling methods.","tokens_in":20779,"tokens_out":11064,"duration_ms":120726,"significance":"The negative benchmarking result is significant: it provides a concrete mechanism by which reported SQD gains can be classical post-processing artifacts, and it yields a simple, adoptable prescription that fair benchmarking must report and control the diagonalization dimension d. The d-growth analysis across N2, H8, and H12 is consistent, and the uniform-random baseline is an appropriate control. The NOCI measurement idea is original, and the H12 energy improvements at fixed d are encouraging, as is the honest reporting that N2 shows no energy benefit. However, the resource-fairness claim for NOCI is not supported by the benchmarks: the experiments fix only d, while the paper's own cost model indicates that NOCI-SQD spends substantially more classical work per retained configuration. The absence of code and data also limits independent verification, although the methods are described in enough detail to be reproducible in principle.","major_comments":[{"comment":"The abstract and conclusion claim that NOCI improvements 'persist even under fixed classical resource budgets' and 'after controlling for classical resources required for diagonalization.' The benchmarks in Figs. 4 and 5 fix only the diagonalization dimension d, not the total classical work. Table I and Appendix E3 show that a NOCI-SQD matrix element costs O(N_cdf kappa^3) versus O(1)-O(N_e^2) for standard SQD, and that NOCI-SQD stores dense 2O(d^2) matrices whereas standard SQD stores a typically sparse O(d^2) matrix. Equal d therefore does not imply equal classical cost; a NOCI run at the same d can expend orders of magnitude more classical effort per retained configuration. The conclusion that NOCI configurations are 'higher quality rather than more numerous' is not established by these experiments. Please either compare at equal estimated classical-work budgets (e.g., FLOPs or memory), or explicitly restate the claim as an improvement at fixed diagonalization dimension only.","section":"Abstract, Section III, Table I, Appendix E3"},{"comment":"The negative benchmarking claim is framed in the abstract as applying when 'classical resources are not explicitly constrained,' and the paper also asserts that Cartesian-product spin expansion is 'the default setting under which most SQD experiments in literature appear to have been performed using the software package in Ref. [24].' The only support for this literature-level claim is the qiskit-sqd-addon default; no survey of the cited SQD demonstrations is provided. If a substantial fraction of prior demonstrations already controlled d or used a restricted spin expansion, the statement that uniform random sampling reproduces SQD benchmarks would not generalize. Please provide explicit evidence for the default-setting claim, or soften the wording to state that uniform random sampling can reproduce SQD benchmarks when the Cartesian-product expansion is used and d is uncontrolled.","section":"Section II, Figs. 2-3, Section V.A, Appendix B"}],"minor_comments":[{"comment":"The text states that ten trials are used per data point, but the figures show no error bars or scatter. Please report standard errors or show individual trial outcomes, as the central quantitative claim is that uniform random sampling matches or exceeds noisy SQD.","section":"Section V.A, Figs. 2, 3, 5"},{"comment":"The phrase 'improving beyond the NOCI ground-state energy itself' is confusing: since sampled configurations are added to the NOCI subspace, a lower energy than the (M+1)-determinant NOCI energy is expected and is not an anomaly. Consider rewording.","section":"Section II, Fig. 5 caption"},{"comment":"The claim that Cartesian-product spin expansion is the default in 'most SQD experiments in literature' would be strengthened by a citation or a brief survey of the referenced SQD papers; currently only the software package default is cited.","section":"Section V.A"},{"comment":"The statement that code and data are 'available upon reasonable request' is weak for a computational benchmarking paper; a public repository or permanent DOI would materially aid reproducibility and independent verification.","section":"Data and code availability"},{"comment":"The noise model is restricted to local depolarizing noise on two-qubit gates. A sentence noting that readout errors and other hardware noise channels may alter the diversity-enhancement effect would help calibrate the claim's applicability to real devices.","section":"Section V.A, Fig. 2"}],"recommendation":"major_revision","confidential_remarks":"The resource-fairness concern raised during review is confirmed: the paper's own cost model in Table I contradicts the 'fixed classical resource budgets' wording in the abstract and conclusion. The negative benchmarking contribution is solid and likely publishable once the positive NOCI claim is restated to match what the experiments actually control. The required changes are primarily rewording and possibly adding resource-matched benchmarks, so I would not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: the negative result is real and worth taking seriously. The paper shows that under the default qiskit-sqd-addon settings—Cartesian spin-product expansion and sub-unity carryover—the classical diagonalization subspace grows with noise, and once that growth is accounted for, uniform random sampling reproduces or beats noisy LUCJ-based SQD. That is a genuinely important benchmarking critique, and it is well supported by the N2, H8, and H12 simulations. They set carryover to unity, vary the depolarizing rate, and plot d explicitly, so the mechanism is clear and the conclusion is not an artifact of one molecule. If this holds for the protein demonstrations in Refs. [15,16], those results need re-reading.\n\nThe NOCI-SQD protocol is a reasonable idea and the H12 results show a real sample-efficiency gain at fixed d, but the paper oversells it. The abstract and Section III claim improvements persist under fixed classical resource budgets, and the paper equates that with fixing d. Their own cost model in Table I contradicts that: NOCI-SQD costs O(N_cdf κ^3) per matrix element versus O(1)–O(N_e^2) for standard SQD, stores 2O(d^2) instead of O(d^2), and works with dense matrices rather than sparse ones. So equal d is not equal classical work. The claim that NOCI configurations are 'higher quality rather than more numerous' is not established by these benchmarks; the energy gain could simply be purchased with more classical effort per retained configuration. That is the load-bearing flaw in the positive claims. Fixing it is straightforward: either control for actual classical cost (flops or runtime), or rewrite the claim to say 'same diagonalization dimension' and explicitly discuss the trade-off. The N2 results in Fig. 8 also show no energy advantage, only diversity, which the authors acknowledge but should be more prominent in the main text.\n\nOther soft spots are minor: no code or data shipped (the paper says available on request), no visible error bars despite stochastic subsampling, and the simulations use a local depolarizing noise model rather than hardware noise. None of these undermine the negative critique.\n\nThis paper deserves a serious referee. The benchmarking result alone is enough to justify peer review, and the NOCI protocol is worth evaluating even if the resource-fairness claim needs correction. I would send it out, and I would tell the authors to fix the resource-budget language, report error bars, and release code and data before acceptance.","headline":"The negative benchmarking critique of SQD is solid and important, but the NOCI advantage is oversold: equal diagonalization dimension does not mean equal classical cost, so the 'fixed classical resource budget' claim needs a rewrite.","tokens_in":689,"tokens_out":1337,"would_cite":true,"duration_ms":29169,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sample-based quantum diagonalization can look successful for purely classical reasons: when the diagonalization subspace is allowed to grow, uniform random sampling matches or beats quantum-backed SQD benchmarks, so fair benchmarking must…","keywords":["sample-based quantum diagonalization","quantum-selected configuration interaction","non-orthogonal configuration interaction","electronic structure","quantum chemistry","benchmarking","diagonalization subspace","noise resilience"],"falsifier":"Take a published SQD hardware experiment, fix the diagonalization subspace size $d$ to the number of unique sampled configurations at each iteration, and compare the reported ground-state energy with uniform random sampling under the same $d$. If the reported energy error does not increase when local depolarizing noise is added and does not stay at or below the uniform-random value at matched $d$, the central claim is falsified; measuring how $d$ grows across iterations in the reference implementation would also directly test the proposed mechanism.","tokens_in":20311,"feed_emoji":"⚛️","tokens_out":7563,"duration_ms":76252,"temperature":0.7,"pith_summary":"The paper argues that a leading quantum-selected configuration interaction method, sample-based quantum diagonalization (SQD), can appear to perform well for reasons that have nothing to do with quantum sampling: uncontrolled growth of the classical diagonalization subspace. When the subspace size is not fixed, classical uniform random sampling reproduces or even outperforms SQD benchmarks, and adding local depolarizing noise improves apparent performance simply by generating more distinct configurations. The paper therefore proposes that any fair SQD benchmark must control the diagonalization size over unique samples. It then introduces a measurement protocol based on non-orthogonal configuration interaction (NOCI), spreading measurements across optimized orbital bases, and reports improved sample efficiency that persists under fixed classical budgets, indicating higher-quality configurations rather than merely more numerous ones.","feed_headline":"Classical random sampling can beat quantum SQD benchmarks","feed_subtitle":"Fixing the diagonalization subspace size removes noise's apparent boost and exposes a purely classical effect.","key_machinery":"The central object is the classical diagonalization subspace of dimension $d$: the number of unique Slater determinants on which the molecular Hamiltonian is projected. Two growth mechanisms drive the critique: spin-product expansion, which splits a measured bitstring into $\\alpha$- and $\\beta$-spin halfstrings and recombines them (often as a Cartesian product, the default in the reference implementation), and configuration carryover, which appends important samples from previous iterations. The constructive machinery is a NOCI measurement basis: a set of orbital rotations $\\{U_m\\}$ optimized classically so that each rotated determinant $|\\Phi_m\\rangle = U_m|\\Psi_0\\rangle$ lowers the NOCI ground-state energy, with measurements distributed across these bases. Because determinants from different bases are non-orthogonal, the method assembles overlap and Hamiltonian matrices into a generalized eigenvalue problem $\\tilde{H}\\mathbf{c} = E\\tilde{\\Gamma}\\mathbf{c}$, solved with a generalized Davidson iteration.","core_discovery":"The paper's central claim is that SQD's apparent robustness to noise is largely an artifact of classical post-processing. Reproducing standard N2 benchmarks, the authors find that replacing quantum measurements with uniformly random configurations matches or beats LUCJ-based SQD whenever the diagonalization subspace is allowed to grow through spin-product expansion and configuration recovery. Increasing local depolarizing noise helps only because it increases the diversity of sampled bitstrings, which after correction and expansion enlarges the subspace. Once the diagonalization size $d$ is fixed, noise no longer helps, and noisy LUCJ converges to uniform random sampling from below. The constructive claim is that measuring in a non-orthogonal basis, built classically from NOCI orbital rotations, discovers energy-lowering configurations more efficiently than measuring only in the Hartree–Fock basis, even with $d$ controlled.","pith_inferences":["As an extension, every QSCI benchmark could require reporting $d$ per iteration and matching a uniform-random-sampling control at equal $d$; that single change would reframe several published SQD demonstrations.","The NOCI advantage is largest when the underlying ansatz is weakest and at small diagonalization sizes, so measurement-basis engineering looks most useful for near-term noisy hardware rather than as an asymptotic quantum speedup.","Because uniform random sampling in the NOCI basis outperforms LUCJ there, a purely classical strongly correlated method suggests itself: classically optimize orbital rotations, sample configurations randomly, and diagonalize, which could be tested against selected CI.","The quantitative noise model is local depolarizing noise, so the thresholds would likely shift under readout errors or crosstalk; testing the NOCI protocol under those noise models is a concrete next step."],"forward_implications":["SQD results that do not report the diagonalization size $d$ over unique samples cannot be read as evidence of quantum sampling quality, because noise-driven improvements vanish once $d$ is controlled.","Uniform random sampling should be a standard control in SQD benchmarking; in a NOCI basis it can outperform LUCJ sampling for correlated systems, exposing ansatz limitations that Hartree–Fock-basis benchmarks hide.","Measuring in NOCI bases improves sample efficiency under fixed classical resource budgets, so the benefit is higher-quality configurations rather than a larger search space.","The NOCI rotations can be folded into the terminal orbital-rotation layer of LUCJ ansätze, so the multi-basis measurements require no additional quantum circuit depth.","The classical cost shifts to non-orthogonal matrix-element evaluation and overlap-matrix storage, with the paper's scaling estimates showing future work must target scalable NOCI basis construction and inter-basis matrix elements."],"supporting_citations":[{"why":"The SQD method whose benchmarks and post-processing pipeline (configuration recovery, carryover, spin-product expansion) are reproduced and critiqued.","marker":"[7]"},{"why":"Introduces the QSCI strategy of selecting configurations by quantum sampling and diagonalizing classically.","marker":"[6]"},{"why":"Supplies the prior concern about sampling inefficiencies in quantum-selected configuration interaction that motivates the analysis.","marker":"[17]"},{"why":"The reference SQD implementation whose default Cartesian-product spin expansion produces the uncontrolled growth in $d$.","marker":"[24]"},{"why":"Defines the local unitary cluster Jastrow (LUCJ) ansatz used for the simulated quantum states.","marker":"[28]"},{"why":"Provides the iterative Davidson method used for the large projected diagonalizations.","marker":"[25]"},{"why":"Supplies compressed double factorization, used to evaluate non-orthogonal Hamiltonian matrix elements efficiently.","marker":"[29]"},{"why":"Provides the generalized non-orthogonal matrix-element formalism underlying the NOCI diagonalization.","marker":"[22]"}],"fun_headline_variants":["SQD noise benefit vanishes when diagonalization size is fixed","Classical random sampling beats SQD without subspace control","NOCI measurement basis improves quantum sampling efficiency","Fixed subspace reveals classical parity in SQD benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulated SQD pipeline matches literature practice: the diagonalization subspace is built from the Cartesian product of alpha- and beta-spin halfstrings (the default in the commonly used reference implementation), and local depolarizing noise on two-qubit gates captures hardware behaviour; if actual demonstrations already control $d$ or use a different spin expansion, the critique may not apply.","fun_headline_variants_meta":{"raw":{"variants":["SQD noise benefit vanishes when diagonalization size is fixed","Classical random sampling beats SQD without subspace control","NOCI measurement basis improves quantum sampling efficiency","Fixed subspace reveals classical parity in SQD benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1316,"prompt_tokens":936,"completion_tokens":380,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":319}},"tokens_in":552,"tokens_out":380,"duration_ms":4427,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:35:00.499260+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a published SQD hardware experiment, fix the diagonalization subspace size $d$ to the number of unique sampled configurations at each iteration, and compare the reported ground-state energy with uniform random sampling under the same $d$. If the reported energy error does not increase when local depolarizing noise is added and does not stay at or below the uniform-random value at matched $d$, the central claim is falsified; measuring how $d$ grows across iterations in the reference implementation would also directly test the proposed mechanism.","supporting_citations":[{"cited_title":"Kanno, M","cited_arxiv_id":null,"evidence_quote":"Introduces the QSCI strategy of selecting configurations by quantum sampling and diagonalizing classically."},{"cited_title":"State-preparation and basis construction","cited_arxiv_id":null,"evidence_quote":"Supplies the prior concern about sampling inefficiencies in quantum-selected configuration interaction that motivates the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The reference SQD implementation whose default Cartesian-product spin expansion produces the uncontrolled growth in $d$."},{"cited_title":"Matsuzawa and Y","cited_arxiv_id":null,"evidence_quote":"Defines the local unitary cluster Jastrow (LUCJ) ansatz used for the simulated quantum states."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the iterative Davidson method used for the large projected diagonalizations."},{"cited_title":"Motta, K","cited_arxiv_id":null,"evidence_quote":"Supplies compressed double factorization, used to evaluate non-orthogonal Hamiltonian matrix elements efficiently."},{"cited_title":"A General Approach for Multireference Ground and Excited States using Non-Orthogonal Configuration Interaction","cited_arxiv_id":"1905.02626","evidence_quote":"Provides the generalized non-orthogonal matrix-element formalism underlying the NOCI diagonalization."}],"review_version":1}