Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Hardware Robustness of Sample-Based Quantum Diagonalization

T0 review · 2 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Sample-based Quantum Diagonalization is robust across CCSD inputs, qubit layouts, and shot budgets on real hardware.

desk verdict A genuinely useful first map of SQD's deployment robustness, but the headline magnitudes are single-run effects that need repeated trials before they can be taken to the bank. read the letter →

arxiv 2607.18196 v1 pith:AMAU3TTS submitted 2026-07-20 quant-ph cs.CR

classification quant-phcs.CR
keywords sample-basedquantumdiagonalizationhybridquantum-classicalalgorithmschemistryrobustnessCCSDamplitudesqubitlayoutnoisemitigationshotbudget
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports an empirical study of Sample-based Quantum Diagonalization (SQD), a hybrid quantum-classical method that replaces variational optimization with a self-consistent recovery loop over quantum-device samples. The authors test three deployment choices a practitioner controls—classical CCSD amplitudes that seed the ansatz, qubit layout and noise-mitigation settings, and the QPU shot budget—on a 156-qubit processor for BeH2, H2O, and N2. Across these probes, the recovery loop absorbs most differences: even zeroing all double-excitation amplitudes shifts converged energies by less than about 7 millihartree, layout and mitigation spreads narrow from about 40x at iteration 0 to about 1.4x after five iterations, and accuracy saturates at 10^3–10^4 shots, with 10^5 shots slightly worse. The upshot: SQD offers practical deployment robustness, and expensive choices like very large shot budgets may not buy accuracy.

What carries the argument

The recovery loop itself. At each iteration SQD checks measured bitstrings against particle-number constraints, repairs or filters them using average orbital occupancies, groups survivors into batches, and diagonalizes the molecular Hamiltonian within each batch; the lowest-energy batch updates occupancies for the next iteration. This mechanism rebuilds the diagonalization subspace from the Hamiltonian and occupancy feedback rather than from the quality of any individual sample, which is why perturbed classical inputs, poor layouts, and different mitigation settings lose most of their influence after a few iterations.

What would settle it

Repeat each key comparison on the same device many times—especially H2O with zeroed t2 and the 10^4-vs-10^5 shot case—and check whether the standard deviation across runs is smaller than the reported differences (roughly 7 mHa for the zeroing shift and 2.6 mHa for the shot-budget reversal). If repeated runs show spread comparable to or larger than these gaps, the robustness ordering collapses.

Watch

Extended reading notes

Core claim

The central claim is that the SQD recovery loop—the self-consistent cycle that filters or repairs sampled bitstrings, selects a working set of configurations, and diagonalizes the Hamiltonian in their span—makes the final energy insensitive to many deployment choices. Structured perturbations of the t2 CCSD amplitudes, including sign flips, HOMO scaling, and complete zeroing, move the converged H2O energy by at most a few millihartree relative to clean baseline; across BeH2, H2O, and N2, zeroing increases error by less than 7 mHa. Layout and noise-mitigation choices produce large iteration-0 differences (up to 40x) that shrink to about 1.4x after five recovery iterations. Shot-budget sweeps

Load-bearing premise

The load-bearing premise is that single, unrepeated hardware runs per configuration give reliable effect sizes for the reported robustness gaps; if run-to-run shot noise or calibration drift is comparable to those gaps, the claimed 'modest shifts' and 'slightly worse' ordering are not established.

Editorial extensions

If this is right

  • Practitioners can safely reuse or tolerate imperfect CCSD inputs: even complete zeroing of double-excitation amplitudes leaves converged SQD energies within roughly 7 mHa of the clean baseline.
  • Qubit layout and noise-mitigation choices matter mainly for the first recovery iteration; after five iterations the 40x spread collapses to about 1.4x, so a poor initial layout need not spoil the final answer.
  • A moderate shot budget of 10^3–10^4 shots is enough; allocating 10^5 shots does not improve and can slightly worsen recovered accuracy, so excess sampling is wasted QPU time.
  • The zigzag layout is a sensible default because it gives a much cleaner iteration-0 distribution, even though recovery later narrows the advantage.
  • The robustness pattern tracks molecular correlation, not qubit count: single-reference BeH2 is essentially unaffected by t2 zeroing, while correlated N2 shows the largest absolute error.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If this robustness transfers to other molecules and devices, SQD's practical advantage over variational methods widens because the method removes the optimizer and its hyperparameter tuning from the quantum pipeline.
  • The shot-budget plateau suggests sample diversity, not raw count, drives recovery; varying samples_per_batch while holding total shots fixed would be a direct test of the working-set explanation.
  • Because each configuration was measured once, the reported gaps (e.g., 1.06 vs 2.10 mHa, or 1.1 vs 3.7 mHa) are single-run differences; repeated trials could either tighten these orderings or reveal run-to-run noise large enough to invert them.
  • Layout and mitigation being washed out implies SQD may tolerate less aggressive noise mitigation, but the paper's data cannot yet support skipping mitigation entirely in noisier regimes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports an empirical robustness study of sample-based quantum diagonalization (SQD) on IBM Heron hardware (ibmq_fez). It probes three deployment choices: structured perturbations of CCSD t2 amplitudes (Q1), qubit layout and noise mitigation (Q2), and shot budget scaling (Q3), using BeH2, H2O, and N2 active spaces with CASCI references. The main claims are that t2 perturbations, including complete zeroing, shift converged SQD energies by less than ~7 mHa; layout and mitigation differences, although large at iteration 0, narrow to ~1.4x after five recovery iterations; and accuracy saturates at moderate shot budgets, with the largest budget (10^5) slightly worse than smaller budgets.

Significance. If the quantitative findings hold, the paper offers practically useful guidance for SQD users: moderate shot budgets are sufficient, imperfect CCSD inputs are tolerated, and layout/mitigation choices are largely washed out by the recovery loop. Strengths of the work include the use of external CASCI reference energies (Eq. 4), explicit perturbation transformations (Eqs. 1-3), per-iteration trajectories for every configuration, and a cross-molecule comparison. However, every quantitative comparison is based on a single unrepeated hardware run per configuration, and the paper itself concedes this limitation only in Section V-B when declining to interpret mitigation sign. Because the reported effect sizes in Q1 and Q3 are comparable to plausible run-to-run drift and shot noise, the magnitudes and even some directions of the central robustness claims are not yet established. The qualitative phenomenon — that the recovery loop reduces initial differences — is plausible, but the quantitative guidance requires repeated trials or propagated uncertainty.

major comments (2)
  1. [§V-B; Figs. 2–4] The manuscript states, 'we do not have repeated hardware trials' (Section V-B), but uses this caveat only for the mitigation sign. The same limitation applies to the load-bearing claims in Q1 and Q3. In Fig. 2, P1 and P2 yield 1.06 and 1.26 mHa versus a clean baseline of 2.10 mHa — a difference of ~0.8–1.0 mHa that could be calibration drift. In Fig. 4b, the H2O 10^4 vs 10^5 gap is 1.1 vs 3.7 mHa; in Fig. 4a, the N2 gap is 1.93 vs 2.2×10^-2 Ha. These gaps are comparable to the mitigation-induced shifts that the authors themselves decline to interpret. Without repeated runs or an uncertainty model, the claims of 'modest shifts' and 'slightly worse' at 10^5 shots are not established. Please add repeated trials, error bars, or explicitly rephrase all magnitudes as single-run point estimates.
  2. [§IV, Table I; §V-C, Fig. 4] For the H2O Q3 sweep, Table I sets samples per batch = 300 and K = 5, implying that 1,500 recovered bitstrings are needed per recovery iteration. The N_shots = 100 configuration cannot supply 1,500 unique QPU measurement outcomes. The manuscript does not specify how the parallel batches are formed in this regime (e.g., sampling with replacement, reusing bitstrings, or reducing effective batch size). This matters because the 'too sparse for effective recovery' conclusion for 100 shots may reflect this batching implementation rather than intrinsic SQD shot-budget behavior. Please clarify the batching procedure for N_shots < K × samples_per_batch.
minor comments (5)
  1. [Eq. (2)] The symbol H is used for the HOMO orbital index set, which is easy to confuse with the Hamiltonian. Please rename, e.g., 'HOMO' or S_h, and specify whether H contains one orbital or multiple degenerate frontier orbitals.
  2. [Eq. (2), §III-A] The 2π multiplier in P2 is arbitrary; this is acceptable as a stress test, but the paper should acknowledge that it samples only one point in a continuum of scaling factors.
  3. [Abstract; §V-C] The abstract and Section V-C attribute the saturation to 'working-set selection' without direct evidence. This is a plausible hypothesis; please label it as such or provide supporting data (e.g., working-set overlap statistics across shot budgets).
  4. [Figs. 2–4] Each figure would benefit from a statement in the caption that every point is a single unrepeated hardware run and that no error bars are shown. A tabulated final-iteration energy with run metadata would also improve reproducibility.
  5. [References] Reference [14] has a typo: 'Sigbahn' should be 'Siegbahn'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: SQD robustness claims are measured against external CASCI references with explicit perturbation equations and no fitted parameters.

full rationale

The paper's claims are empirical measurements of SQD on IBM hardware, not derivations from fitted inputs. The energy error is defined in Eq. (4) relative to externally computed CASCI references, and the Q1 perturbations are explicit parameter-free transformations (Eqs. 1-3). Q2 and Q3 vary layout/mitigation and shot budget while holding documented defaults from Table I; no parameter is fit to the reported errors, so no 'prediction' is forced by construction. The recovery loop is a self-consistent algorithm, but it is the object under study, not a hidden input to the analysis. The reference list contains no works by the present authors, so the argument does not depend on self-citation. The only substantive caveat—single unrepeated hardware runs and lack of propagated uncertainty—is a measurement-representativeness concern, not a circularity concern. Therefore no circular step is present and the score is 0.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claims depend on hand-chosen hyperparameters (T, K, samples per batch, the 2π probe) and on the SQD/LUCJ algorithm behaving exactly as its developers describe [4],[5],[8]; no new entities are introduced. The most fragile ledger items are the implicit single-run measurement model and the underspecified low-shot batch construction.

free parameters (4)
  • P2 HOMO-scaling multiplier 2π =
    Hand-chosen perturbation magnitude in Eq. (2) for the Q1 P2 stress test. Not fitted to data, but arbitrary and it defines the severity of the 'strong distortion' probed.
  • Samples per batch = 300 (H2O, BeH2); 500 (N2)
    Table I default that caps the recovered working set each iteration. The Q3 saturation conclusion ('additional shots do not enlarge the useful subspace') is contingent on this value being held fixed.
  • Recovery iterations T = 5
    Table I default. Every 'converged' error is read at iteration 4 (0-indexed), so all convergence and narrowing claims are claims about T=5.
  • Parallel batches K = 5
    Table I default. Together with samples-per-batch, K defines how many bitstrings enter subspace diagonalization per iteration.
assumptions (5)
  • domain assumption Bitstrings sampled from the CCSD-initialized LUCJ circuit on ibmq_fez contain the determinants needed for subspace recovery.
    Section II; this is the premise of SQD itself, inherited from Refs. [4],[5] and not tested here.
  • domain assumption A single unrepeated hardware run per configuration is representative of SQD performance on ibmq_fez.
    Section V relies on unrepeated runs; the paper flags this only for the Q2 mitigation sign (Section V-B: 'we do not have repeated hardware trials') and does not propagate it to Q1/Q3 magnitudes.
  • domain assumption CASCI active-space energies (Eq. 4) are the correct ground-truth reference for the three active spaces.
    Section IV; standard practice (Ref. [14]), but the active-space choice itself is an unexamined modeling layer.
  • domain assumption Five recovery iterations suffice, i.e., iteration-4 values are the converged energies.
    Section IV: 'using the final value after T=5 recovery iterations'; the Q2 narrowing claim is read at iteration 4.
  • ad hoc to paper K×samples-per-batch bitstrings can be formed even when N_shots is smaller than that product (e.g., N_shots=100 with 1500 needed).
    Table I vs. Section V-C; the paper never specifies resampling or batch shrinking when the shot pool is undersized, so the low-shot regime of Q3 rests on this unstated procedure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hardware Robustness of Sample-Based Quantum Diagonalization." pith.science (2026). https://pith.science/paper/AMAU3TTS

@misc{pith2026260718196,
  author       = {Pith},
  title        = {Pith review of: Hardware Robustness of Sample-Based Quantum Diagonalization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AMAU3TTS}},
  note         = {Machine review of arXiv:2607.18196}
}
read the original abstract

Sample-based Quantum Diagonalization (SQD) is a hybrid quantum-classical method that replaces variational optimization with a self-consistent recovery loop over QPU samples. Although SQD is considered robust to noisy samples and imperfect classical inputs, its robustness across practical deployment choices has not been systematically analyzed. As a result, shot budgets, qubit layouts, noise mitigation strategies, and the coupled-cluster singles and doubles (CCSD) amplitudes that initialize the ansatz are often chosen without clear empirical guidance. We analyze SQD robustness on IBM Heron hardware across these dimensions. Structured CCSD-amplitude perturbations, including complete zeroing, produce only modest energy shifts from the clean baseline. Differences across layouts and noise-mitigation settings are large in the first recovery iteration but narrow within a few iterations. Accuracy saturates at moderate shot budgets, while very large budgets slightly worsen recovered energies, likely because working-set selection limits the value of additional samples. These results identify where SQD provides genuine deployment robustness and where its limits remain.

Figures

Figures reproduced from arXiv: 2607.18196 by the authors.

Figure 1
Figure 1. SQD workflow and the three robustness probes (Q1–Q3) analyzed in this work. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. H2O SQD error under structured t2 perturbations. while N2 uses 32 qubits. Classical CASCI calculations provide the reference ground-state energies: ECASCI = −15.799481 Ha for BeH2, −76.119926 Ha for H2O, and −109.046672 Ha for N2 [14]. We use the standard chemical-accuracy threshold of 1.6 × 10−3 Ha (1 kcal/mol) as a visual reference in all error plots. Mean-field and CCSD calculations are performed in PySCF [6], pr… view at source ↗
Figure 3
Figure 3. N2 SQD error across layout and mitigation settings. 0 1 2 3 4 10−3 10−2 10−1 100 101 Iteration Energy error (Ha) (a) N2 0 1 2 3 4 10−3 10−2 10−1 Iteration (b) H2O [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: SQD energy error vs. recovery-loop iteration on [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Machine learning for sample-based quantum diagonalization: generative configuration recovery and the classical-simulability frontier

    quant-ph 2026-08 conditional novelty 6.0 of 10

    A critical review plus small exact-FCI experiments concludes that sample-based quantum diagonalization has not beaten classical selected CI and maps where, if anywhere, a quantum or generative advantage could survive.

Reference graph

Works this paper leans on

14 extracted references · cited by 1 Pith paper

  1. [1]

    Peruzzo, J

    A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’Brien, ”A variational eigenvalue solver on a photonic quantum processor,” Nature Communications, vol. 5, no. 1, p. 4213, 2014

  2. [2]

    J. R. McClean, S. Boixo, V . N. Smelyanskiy, R. Babbush, and H. Neven, ”Barren plateaus in quantum neural network training landscapes,” Nature Communications, vol. 9, no. 1, p. 4812, Nov. 2018

  3. [3]

    Scriva, N

    G. Scriva, N. Astrakhantsev, S. Pilati, and G. Mazzola, ”Challenges of variational quantum optimization with measurement shot noise,” Physical Review A, vol. 109, no. 3, p. 032408, Mar. 2024

  4. [4]

    Robledo-Moreno, M

    J. Robledo-Moreno, M. Motta, H. Haas, A. Javadi-Abhari, P. Jurcevic, W. Kirby, S. Martiel, K. Sharma, S. Sharma, T. Shirakawa, and I. Sitdikov, ”Chemistry beyond the scale of exact diagonalization on a quantum-centric supercomputer,” Science Advances, vol. 11, no. 25, eadu9991, Jun. 2025

  5. [5]

    Motta, K

    M. Motta, K. J. Sung, K. B. Whaley, M. Head-Gordon, and J. Shee, ”Bridging physical intuition and hardware efficiency for correlated electronic states: the local unitary cluster Jastrow ansatz for electronic structure,” Chemical Science, vol. 14, no. 40, pp. 11213–11227, 2023

  6. [6]

    Q. Sun, T. C. Berkelbach, N. S. Blunt, G. H. Booth, S. Guo, Z. Li, J. Liu, J. D. McClain, E. R. Sayfutyarova, S. Sharma, and S. Wouters, ”The Python-based simulations of chemistry framework (PySCF),” WIREs Computational Molecular Science, vol. 8, no. 1, e1340, 2018

  7. [7]

    R. J. Bartlett and M. Musiał, ”Coupled-cluster theory in quantum chemistry,” Reviews of Modern Physics, vol. 79, no. 1, pp. 291–352, Jan. 2007

  8. [8]

    Available: https://quantum.cloud.ibm.com/docs/en/guides/qiskit-addons-sqd

    IBM Quantum, ”Qiskit add-on: Sample-based quan- tum diagonalization (SQD),” [Online]. Available: https://quantum.cloud.ibm.com/docs/en/guides/qiskit-addons-sqd

Show all 14 references
  1. [9]

    F. A. Evangelista, G. K.-L. Chan, and G. E. Scuseria, ”Exact parameter- ization of fermionic wave functions via unitary coupled cluster theory,” The Journal of Chemical Physics, vol. 151, no. 24, p. 244112, Dec. 2019

  2. [10]

    G. Li, Y . Ding, and Y . Xie, ”Tackling the qubit mapping problem for NISQ-era quantum devices,” in Proc. 24th Int. Conf. Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2019, pp. 1001-1014

  3. [11]

    Chamberland, G

    C. Chamberland, G. Zhu, T. J. Yoder, J. B. Hertzberg, and A. W. Cross, ”Topological and subsystem codes on low-degree graphs with flag qubits,” Physical Review X, vol. 10, no. 1, p. 011022, Jan. 2020

  4. [12]

    Viola and S

    L. Viola and S. Lloyd, ”Dynamical suppression of decoherence in two- state quantum systems,” Physical Review A, vol. 58, no. 4, pp. 2733– 2744, Oct. 1998

  5. [13]

    J. J. Wallman and J. Emerson, ”Noise tailoring for scalable quantum computation via randomized compiling,” Physical Review A, vol. 94, no. 5, p. 052325, Nov. 2016

  6. [14]

    B. O. Roos, P. R. Taylor, and P. E. M. Sigbahn, ”A complete active space SCF method (CASSCF) using a density matrix formulated super- CI approach,” Chemical Physics, vol. 48, no. 2, pp. 157–173, May 1980

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.