REVIEW 2 major objections 5 minor 1 cited by
Hardware Robustness of Sample-Based Quantum Diagonalization
T0 review · 2 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Sample-based Quantum Diagonalization is robust across CCSD inputs, qubit layouts, and shot budgets on real hardware.
desk verdict A genuinely useful first map of SQD's deployment robustness, but the headline magnitudes are single-run effects that need repeated trials before they can be taken to the bank. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The recovery loop itself. At each iteration SQD checks measured bitstrings against particle-number constraints, repairs or filters them using average orbital occupancies, groups survivors into batches, and diagonalizes the molecular Hamiltonian within each batch; the lowest-energy batch updates occupancies for the next iteration. This mechanism rebuilds the diagonalization subspace from the Hamiltonian and occupancy feedback rather than from the quality of any individual sample, which is why perturbed classical inputs, poor layouts, and different mitigation settings lose most of their influence after a few iterations.
What would settle it
Repeat each key comparison on the same device many times—especially H2O with zeroed t2 and the 10^4-vs-10^5 shot case—and check whether the standard deviation across runs is smaller than the reported differences (roughly 7 mHa for the zeroing shift and 2.6 mHa for the shot-budget reversal). If repeated runs show spread comparable to or larger than these gaps, the robustness ordering collapses.
Extended reading notes
Core claim
The central claim is that the SQD recovery loop—the self-consistent cycle that filters or repairs sampled bitstrings, selects a working set of configurations, and diagonalizes the Hamiltonian in their span—makes the final energy insensitive to many deployment choices. Structured perturbations of the t2 CCSD amplitudes, including sign flips, HOMO scaling, and complete zeroing, move the converged H2O energy by at most a few millihartree relative to clean baseline; across BeH2, H2O, and N2, zeroing increases error by less than 7 mHa. Layout and noise-mitigation choices produce large iteration-0 differences (up to 40x) that shrink to about 1.4x after five recovery iterations. Shot-budget sweeps
Load-bearing premise
The load-bearing premise is that single, unrepeated hardware runs per configuration give reliable effect sizes for the reported robustness gaps; if run-to-run shot noise or calibration drift is comparable to those gaps, the claimed 'modest shifts' and 'slightly worse' ordering are not established.
Editorial extensions
If this is right
- Practitioners can safely reuse or tolerate imperfect CCSD inputs: even complete zeroing of double-excitation amplitudes leaves converged SQD energies within roughly 7 mHa of the clean baseline.
- Qubit layout and noise-mitigation choices matter mainly for the first recovery iteration; after five iterations the 40x spread collapses to about 1.4x, so a poor initial layout need not spoil the final answer.
- A moderate shot budget of 10^3–10^4 shots is enough; allocating 10^5 shots does not improve and can slightly worsen recovered accuracy, so excess sampling is wasted QPU time.
- The zigzag layout is a sensible default because it gives a much cleaner iteration-0 distribution, even though recovery later narrows the advantage.
- The robustness pattern tracks molecular correlation, not qubit count: single-reference BeH2 is essentially unaffected by t2 zeroing, while correlated N2 shows the largest absolute error.
Reading between the lines
- If this robustness transfers to other molecules and devices, SQD's practical advantage over variational methods widens because the method removes the optimizer and its hyperparameter tuning from the quantum pipeline.
- The shot-budget plateau suggests sample diversity, not raw count, drives recovery; varying samples_per_batch while holding total shots fixed would be a direct test of the working-set explanation.
- Because each configuration was measured once, the reported gaps (e.g., 1.06 vs 2.10 mHa, or 1.1 vs 3.7 mHa) are single-run differences; repeated trials could either tighten these orderings or reveal run-to-run noise large enough to invert them.
- Layout and mitigation being washed out implies SQD may tolerate less aggressive noise mitigation, but the paper's data cannot yet support skipping mitigation entirely in noisier regimes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an empirical robustness study of sample-based quantum diagonalization (SQD) on IBM Heron hardware (ibmq_fez). It probes three deployment choices: structured perturbations of CCSD t2 amplitudes (Q1), qubit layout and noise mitigation (Q2), and shot budget scaling (Q3), using BeH2, H2O, and N2 active spaces with CASCI references. The main claims are that t2 perturbations, including complete zeroing, shift converged SQD energies by less than ~7 mHa; layout and mitigation differences, although large at iteration 0, narrow to ~1.4x after five recovery iterations; and accuracy saturates at moderate shot budgets, with the largest budget (10^5) slightly worse than smaller budgets.
Significance. If the quantitative findings hold, the paper offers practically useful guidance for SQD users: moderate shot budgets are sufficient, imperfect CCSD inputs are tolerated, and layout/mitigation choices are largely washed out by the recovery loop. Strengths of the work include the use of external CASCI reference energies (Eq. 4), explicit perturbation transformations (Eqs. 1-3), per-iteration trajectories for every configuration, and a cross-molecule comparison. However, every quantitative comparison is based on a single unrepeated hardware run per configuration, and the paper itself concedes this limitation only in Section V-B when declining to interpret mitigation sign. Because the reported effect sizes in Q1 and Q3 are comparable to plausible run-to-run drift and shot noise, the magnitudes and even some directions of the central robustness claims are not yet established. The qualitative phenomenon — that the recovery loop reduces initial differences — is plausible, but the quantitative guidance requires repeated trials or propagated uncertainty.
major comments (2)
- [§V-B; Figs. 2–4] The manuscript states, 'we do not have repeated hardware trials' (Section V-B), but uses this caveat only for the mitigation sign. The same limitation applies to the load-bearing claims in Q1 and Q3. In Fig. 2, P1 and P2 yield 1.06 and 1.26 mHa versus a clean baseline of 2.10 mHa — a difference of ~0.8–1.0 mHa that could be calibration drift. In Fig. 4b, the H2O 10^4 vs 10^5 gap is 1.1 vs 3.7 mHa; in Fig. 4a, the N2 gap is 1.93 vs 2.2×10^-2 Ha. These gaps are comparable to the mitigation-induced shifts that the authors themselves decline to interpret. Without repeated runs or an uncertainty model, the claims of 'modest shifts' and 'slightly worse' at 10^5 shots are not established. Please add repeated trials, error bars, or explicitly rephrase all magnitudes as single-run point estimates.
- [§IV, Table I; §V-C, Fig. 4] For the H2O Q3 sweep, Table I sets samples per batch = 300 and K = 5, implying that 1,500 recovered bitstrings are needed per recovery iteration. The N_shots = 100 configuration cannot supply 1,500 unique QPU measurement outcomes. The manuscript does not specify how the parallel batches are formed in this regime (e.g., sampling with replacement, reusing bitstrings, or reducing effective batch size). This matters because the 'too sparse for effective recovery' conclusion for 100 shots may reflect this batching implementation rather than intrinsic SQD shot-budget behavior. Please clarify the batching procedure for N_shots < K × samples_per_batch.
minor comments (5)
- [Eq. (2)] The symbol H is used for the HOMO orbital index set, which is easy to confuse with the Hamiltonian. Please rename, e.g., 'HOMO' or S_h, and specify whether H contains one orbital or multiple degenerate frontier orbitals.
- [Eq. (2), §III-A] The 2π multiplier in P2 is arbitrary; this is acceptable as a stress test, but the paper should acknowledge that it samples only one point in a continuum of scaling factors.
- [Abstract; §V-C] The abstract and Section V-C attribute the saturation to 'working-set selection' without direct evidence. This is a plausible hypothesis; please label it as such or provide supporting data (e.g., working-set overlap statistics across shot budgets).
- [Figs. 2–4] Each figure would benefit from a statement in the caption that every point is a single unrepeated hardware run and that no error bars are shown. A tabulated final-iteration energy with run metadata would also improve reproducibility.
- [References] Reference [14] has a typo: 'Sigbahn' should be 'Siegbahn'.
Circularity Check
No circularity: SQD robustness claims are measured against external CASCI references with explicit perturbation equations and no fitted parameters.
full rationale
The paper's claims are empirical measurements of SQD on IBM hardware, not derivations from fitted inputs. The energy error is defined in Eq. (4) relative to externally computed CASCI references, and the Q1 perturbations are explicit parameter-free transformations (Eqs. 1-3). Q2 and Q3 vary layout/mitigation and shot budget while holding documented defaults from Table I; no parameter is fit to the reported errors, so no 'prediction' is forced by construction. The recovery loop is a self-consistent algorithm, but it is the object under study, not a hidden input to the analysis. The reference list contains no works by the present authors, so the argument does not depend on self-citation. The only substantive caveat—single unrepeated hardware runs and lack of propagated uncertainty—is a measurement-representativeness concern, not a circularity concern. Therefore no circular step is present and the score is 0.
Assumptions & free parameters
free parameters (4)
- P2 HOMO-scaling multiplier 2π =
2π
- Samples per batch =
300 (H2O, BeH2); 500 (N2)
- Recovery iterations T =
5
- Parallel batches K =
5
assumptions (5)
- domain assumption Bitstrings sampled from the CCSD-initialized LUCJ circuit on ibmq_fez contain the determinants needed for subspace recovery.
- domain assumption A single unrepeated hardware run per configuration is representative of SQD performance on ibmq_fez.
- domain assumption CASCI active-space energies (Eq. 4) are the correct ground-truth reference for the three active spaces.
- domain assumption Five recovery iterations suffice, i.e., iteration-4 values are the converged energies.
- ad hoc to paper K×samples-per-batch bitstrings can be formed even when N_shots is smaller than that product (e.g., N_shots=100 with 1500 needed).
Cite this review
Pith. "Pith review of Hardware Robustness of Sample-Based Quantum Diagonalization." pith.science (2026). https://pith.science/paper/AMAU3TTS
@misc{pith2026260718196,
author = {Pith},
title = {Pith review of: Hardware Robustness of Sample-Based Quantum Diagonalization},
year = {2026},
howpublished = {\url{https://pith.science/paper/AMAU3TTS}},
note = {Machine review of arXiv:2607.18196}
}
read the original abstract
Sample-based Quantum Diagonalization (SQD) is a hybrid quantum-classical method that replaces variational optimization with a self-consistent recovery loop over QPU samples. Although SQD is considered robust to noisy samples and imperfect classical inputs, its robustness across practical deployment choices has not been systematically analyzed. As a result, shot budgets, qubit layouts, noise mitigation strategies, and the coupled-cluster singles and doubles (CCSD) amplitudes that initialize the ansatz are often chosen without clear empirical guidance. We analyze SQD robustness on IBM Heron hardware across these dimensions. Structured CCSD-amplitude perturbations, including complete zeroing, produce only modest energy shifts from the clean baseline. Differences across layouts and noise-mitigation settings are large in the first recovery iteration but narrow within a few iterations. Accuracy saturates at moderate shot budgets, while very large budgets slightly worsen recovered energies, likely because working-set selection limits the value of additional samples. These results identify where SQD provides genuine deployment robustness and where its limits remain.
Figures
Forward citations
Cited by 1 Pith paper
-
Machine learning for sample-based quantum diagonalization: generative configuration recovery and the classical-simulability frontier
A critical review plus small exact-FCI experiments concludes that sample-based quantum diagonalization has not beaten classical selected CI and maps where, if anywhere, a quantum or generative advantage could survive.
Reference graph
Works this paper leans on
-
[1]
Peruzzo, J
A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’Brien, ”A variational eigenvalue solver on a photonic quantum processor,” Nature Communications, vol. 5, no. 1, p. 4213, 2014
2014
-
[2]
J. R. McClean, S. Boixo, V . N. Smelyanskiy, R. Babbush, and H. Neven, ”Barren plateaus in quantum neural network training landscapes,” Nature Communications, vol. 9, no. 1, p. 4812, Nov. 2018
2018
-
[3]
Scriva, N
G. Scriva, N. Astrakhantsev, S. Pilati, and G. Mazzola, ”Challenges of variational quantum optimization with measurement shot noise,” Physical Review A, vol. 109, no. 3, p. 032408, Mar. 2024
2024
-
[4]
Robledo-Moreno, M
J. Robledo-Moreno, M. Motta, H. Haas, A. Javadi-Abhari, P. Jurcevic, W. Kirby, S. Martiel, K. Sharma, S. Sharma, T. Shirakawa, and I. Sitdikov, ”Chemistry beyond the scale of exact diagonalization on a quantum-centric supercomputer,” Science Advances, vol. 11, no. 25, eadu9991, Jun. 2025
2025
-
[5]
Motta, K
M. Motta, K. J. Sung, K. B. Whaley, M. Head-Gordon, and J. Shee, ”Bridging physical intuition and hardware efficiency for correlated electronic states: the local unitary cluster Jastrow ansatz for electronic structure,” Chemical Science, vol. 14, no. 40, pp. 11213–11227, 2023
2023
-
[6]
Q. Sun, T. C. Berkelbach, N. S. Blunt, G. H. Booth, S. Guo, Z. Li, J. Liu, J. D. McClain, E. R. Sayfutyarova, S. Sharma, and S. Wouters, ”The Python-based simulations of chemistry framework (PySCF),” WIREs Computational Molecular Science, vol. 8, no. 1, e1340, 2018
2018
-
[7]
R. J. Bartlett and M. Musiał, ”Coupled-cluster theory in quantum chemistry,” Reviews of Modern Physics, vol. 79, no. 1, pp. 291–352, Jan. 2007
2007
-
[8]
Available: https://quantum.cloud.ibm.com/docs/en/guides/qiskit-addons-sqd
IBM Quantum, ”Qiskit add-on: Sample-based quan- tum diagonalization (SQD),” [Online]. Available: https://quantum.cloud.ibm.com/docs/en/guides/qiskit-addons-sqd
Show all 14 references
-
[9]
F. A. Evangelista, G. K.-L. Chan, and G. E. Scuseria, ”Exact parameter- ization of fermionic wave functions via unitary coupled cluster theory,” The Journal of Chemical Physics, vol. 151, no. 24, p. 244112, Dec. 2019
2019
-
[10]
G. Li, Y . Ding, and Y . Xie, ”Tackling the qubit mapping problem for NISQ-era quantum devices,” in Proc. 24th Int. Conf. Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2019, pp. 1001-1014
2019
-
[11]
Chamberland, G
C. Chamberland, G. Zhu, T. J. Yoder, J. B. Hertzberg, and A. W. Cross, ”Topological and subsystem codes on low-degree graphs with flag qubits,” Physical Review X, vol. 10, no. 1, p. 011022, Jan. 2020
2020
-
[12]
Viola and S
L. Viola and S. Lloyd, ”Dynamical suppression of decoherence in two- state quantum systems,” Physical Review A, vol. 58, no. 4, pp. 2733– 2744, Oct. 1998
1998
-
[13]
J. J. Wallman and J. Emerson, ”Noise tailoring for scalable quantum computation via randomized compiling,” Physical Review A, vol. 94, no. 5, p. 052325, Nov. 2016
2016
-
[14]
B. O. Roos, P. R. Taylor, and P. E. M. Sigbahn, ”A complete active space SCF method (CASSCF) using a density matrix formulated super- CI approach,” Chemical Physics, vol. 48, no. 2, pp. 157–173, May 1980
1980
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.