REVIEW 4 major objections 4 minor 22 references
The Stabilizer Bootstrap of Quantum Machine Learning with up to 10000 qubits
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Whether stabilizer bootstrap can pre-optimize a variational quantum circuit classically is set by the observable and entanglement structure, with constant success in some cases and exponential decay in others.
desk verdict The four exact sampling probabilities are a real, useful result, but the strong/weak taxonomy overreaches: it is built on an unvalidated proxy and fitted exponents, so the quantum-advantage boundary claims need a rewrite. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the stabilizer bootstrap itself: a restricted variational circuit made entirely of Clifford operations by forcing every $R_y$ angle to lie in $\{0, \pi/2, \pi, -\pi/2\}$, so the state can be simulated classically through the Gottesman-Knill theorem. The bootstrap samples these discrete angles, evaluates a mean-squared-error loss from measurements of a chosen Pauli-string observable, and hands the best points to a Bayesian optimizer (random-forest surrogate with a greedy acquisition function) that refines the angles before any quantum execution. The workhorse quantity is the probability of nontrivial sampling—the chance that a random Clifford angle assignment yields measurement outcome $\pm 1$ instead of $0$. If that probability is high, the optimizer has signal to work with; if all samples are trivial, the acquisition function degenerates to random exploration (the paper shows the expected-improvement score becomes proportional to the predictive variance alone). The proofs proceed by mathematical induction on the qubit count, tracking how $1$-, $-1$-, and $0$-states of the $Z$-string observable evolve through the CNOT layer.
What would settle it
One concrete check: for a weak-enhancement configuration (e.g., linear CNOTs with a $Z$-string observable), enumerate or sample the Clifford angle assignments at $n=30$ and count the fraction with nonzero expectation; if it is not close to $1/2^{16}$, the exponential formula fails. A second check targets the proxy itself: train the same variational circuit after bootstrap initialization and after random initialization on several datasets of growing size; if a strong-enhancement configuration yields no persistent improvement in final loss or convergence speed, then the sampling probability is not the right measure of enhancement.
Extended reading notes
Core claim
The paper's core claim is that the stabilizer bootstrap's effectiveness is governed by one number: the probability, over uniformly random choices of the discrete Clifford angles $\{0, \pi/2, \pi, -\pi/2\}$ for the $R_y$ layers, that a chosen Pauli-string observable has a nonzero expectation value on the prepared state. For single-layer ansatze with linear or reverse-linear CNOT entanglement, this probability is exactly $1/4$ for the strong pairings—linear CNOTs with an $X$-string, or reverse-linear CNOTs with a $Z$-string—and is $1/2^{\lceil n/2+1\rceil}$ for the weak pairings, namely linear CNOTs with a $Z$-string and reverse-linear CNOTs with an $X$-string. The strong pairings give constant probability independent of qubit number, so the bootstrap can classically find good initial points at any scale; the weak pairings lose this handle exponentially. For generic observables that mix $X$ and $Z$ factors (domain-wall strings), numerical experiments find the probability follows $p(r,n)=1/(4n^\nu)$, with the decay exponent $\nu$ interpolating between $0$ and $(n/2-1)\log 2/\log n$; most real cases lie between the two extremes. The paper also reports that as the number of training samples grows, the minimum loss achieved by sampling rises and the variance of sampled losses falls, indicating exponentially decreasing search efficiency with dataset size.
Load-bearing premise
The load-bearing premise is that the probability of getting a nonzero ($\pm 1$) measurement from a random Clifford-angle sample is a faithful measure of how much the stabilizer bootstrap improves the final trained circuit; this proxy is used throughout the paper but is never validated against end-to-end training outcomes such as final loss or accuracy.
Editorial extensions
If this is right
- Variational circuits in the strong-enhancement regime can have their initial angles found classically with constant probability at essentially any qubit count, so the bootstrap is a reliable polynomial-time pre-optimizer for those ansatz/observable pairs.
- Circuits in the weak-enhancement regime require exponentially many Clifford samples (order $2^{n/2}$) to find a nonzero starting point, so the bootstrap cannot pre-train them at large scale and quantum advantage is not undercut by this method.
- For mixed $X/Z$ observables the decay is polynomial in $n$ with an exponent $\nu$ that can be measured; comparing $\nu$ across candidate ansatze gives a quantitative design criterion for choosing circuits that retain a quantum advantage.
- Dataset size acts as an independent resource: increasing the number of training samples from 100 to 1000 lowers minimum loss quality and sampling variance roughly exponentially, so classical pre-optimization becomes harder as the learning task grows.
- Because the whole pipeline (encoding, ansatz, measurement) is Clifford when restricted to these angles, the bootstrap connects fault-tolerant-era Clifford simulation techniques directly to near-term variational quantum machine learning.
Reading between the lines
- Going beyond the paper's stated conclusions: the strong/weak split gives a practical classical-simulatability test—a configuration with constant bootstrap success is one where a classical observer can, with constant probability, find a provably good initial state, which makes claims of quantum advantage for that configuration harder to defend.
- A natural testable extension is to replace the paper's proxy (probability of nonzero sampling) with direct end-to-end metrics such as final test accuracy or gradient variance, and check whether the strong/weak boundary still holds.
- The observed $X/Z$ duality between linear and reverse-linear entanglement points to an underlying Clifford-group symmetry; proving the intermediate exponents $\nu(r)$ from stabilizer theory would turn the numerical interpolation into a phase diagram rather than a fit.
- The dataset-size dependence suggests that data encoding is part of the classical-simulability boundary; testing different encoding maps (not just $\{0,1,2,3\}\to\{0,\pi,\pi/2,-\pi/2\}$) would reveal whether the exponential decline with dataset size is universal.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the stabilizer bootstrap, a method to pre-optimize variational quantum machine learning circuits by restricting rotation angles to Clifford values and simulating the resulting circuits classically via the Gottesman-Knill theorem. The central quantity is the probability p(±1) that a uniformly random Clifford angle assignment yields a nonzero (sign ±1) expectation value for a given observable and entanglement structure. The authors prove exact values of this probability for single-layer circuits with linear or reverse-linear CNOT entanglement and Pauli-X or Pauli-Z string observables (Table I and Appendix B), use these to define 'strong' versus 'weak' stabilizer enhancement, and fit an intermediate decay law p(r,n)=1/(4 n^ν) for mixed X/Z observables. They report high-performance simulations with up to 10000 qubits and datasets up to 1000 points, and from these they infer that the possibility of improvements from the stabilizer bootstrap depends on observable structure and dataset size, with implications for the boundary of classical simulability and quantum advantage.
Significance. If the exact counting results are correct, the paper provides a clean theoretical classification for a simple but nontrivial family of Clifford-constrained variational circuits, which is a useful step toward understanding when classical pre-optimization of QML circuits remains feasible. The demonstration of classical simulation at 10000 qubits using stabilizer techniques is technically impressive and illustrates the practical reach of Gottesman-Knill simulation. The strong/weak enhancement taxonomy, however, is only as valuable as the proxy on which it is built; the paper does not yet show that the nontrivial-sampling probability predicts actual improvement in downstream variational training. The fitted intermediate exponents and dataset-size scaling claims also lack statistical support. These gaps currently prevent the paper from fully supporting its broader claims about the boundary toward quantum advantages.
major comments (4)
- [Appendix B, Theorem 1 proof] The induction proof contains an unjustified probability identity: the text states 'p(the first n qubits are in the 1-state) = p(the first n qubits are in the 1-state | the n-th qubit is |0⟩ or |1⟩)' and justifies it by saying the condition is necessary. Necessity does not equate unconditional and conditional probabilities; a factor of p(the n-th qubit is |0⟩ or |1⟩) is missing. The proof also mixes the induction index k with the summation index k and uses undefined bit-string notation. Since Theorem 1 is the basis for the constant probability 1/4 in Table I and for the strong-enhancement classification, this gap must be fixed before the central classification is rigorous.
- [Appendix B, Theorem 2 proof] The induction proof for the exponential-decay case has multiple typographical and logical issues: the reference 'additional table ??' is left unresolved in both the odd and even parts; factors such as 1/8 · 1/2^{k+2} appear without a clear derivation of where the 1/8 comes from; and the bit indices in expressions like |00x_{i3}...x_{in+2}⟩ are inconsistent with the summation index. Because Theorem 2 supplies the exponential decay that defines weak stabilizer enhancement, the proof needs to be rewritten carefully with correct notation and complete case analysis before it can be considered verified.
- [Section II, Eq. (1) and FIG. 6] The decay law p(r,n)=1/(4 n^ν) is introduced as a hypothesis and the exponent ν is then fitted to the very data that the formula is intended to characterize. No error bars, confidence intervals, or goodness-of-fit measures are reported, and Appendix C explicitly states that the maximum probability 1/4 is 'not a mathematical conclusion but derived from various experiments.' Thus the intermediate 'most situations fall in between' claim is an empirical extrapolation, not a derived result, and it cannot support a rigorous strong/weak dichotomy as presented.
- [Section II and FIG. 7] The claim that 'as the dataset size increases, the search efficiency declines exponentially' is based on plots of minimum loss and variance with no error bars, no repeated trials, and no fitted decay function. More importantly, the paper never validates the central proxy: a nontrivial sample (expectation ±1) is asserted to be the key metric for sampling efficiency, but no experiment demonstrates that higher p(±1) correlates with lower loss, faster convergence, or better final accuracy in the dataset experiments. Without this validation, the strong/weak enhancement taxonomy may not reflect actual end-to-end QML performance.
minor comments (4)
- [Section II, paragraph on general observables] The sentence 'The experimental results for the observables consisting of X operators and Z operators and Clifford stabilizer are shown in FIG.2' appears to reference the wrong figure; the relevant results are in FIG.4 and FIG.5, not the ansatz diagram in FIG.2.
- [Section II, FIG. 6 caption] The caption contains a typo: 'exponent carve' should be 'exponent curve', and 'Equation II' should refer to a numbered equation rather than a section.
- [Appendix B, Theorem 2 proof] The proof would benefit from a preliminary lemma that tracks the probability distribution over the first two qubits, since the current text refers to 'TABLE II' for the two-qubit case but leaves the table reference incomplete in the even case.
- [Appendix A, GP/EI derivation] The statement that setting the noise parameter σ_n^2 to 0 yields μ(x)=0 is correct only when all observed values y are zero; this assumption should be stated explicitly in the derivation.
Circularity Check
No significant circularity: the main theorems are self-contained induction proofs over Clifford circuits, and the fitted exponent is presented as an empirical characterization rather than as a prediction.
full rationale
The paper's central theoretical results (Theorems 1-4 in Appendix B) are derived by mathematical induction from the explicit Clifford ansatz with Ry angles restricted to {0, pi/2, pi, -pi/2} and the observable being a Pauli-X or Pauli-Z string. These proofs do not invoke fitted parameters, empirical data, or prior work; they are parameter-free counting arguments. The probability of measuring 1 or -1 is computed directly from the stabilizer circuit structure, so the strong/weak classification for the four extreme cases is not circular. The intermediate-regime claim ('most situations fall in between') is supported by the fitted exponent nu in Eq. (II), but the paper explicitly describes this as an approximation of decay curves ('we try to use the following function to approximate the decay curves'), and the strong and weak extremes are proven independently. The CAFQA citation [8] is motivational background for the stabilizer bootstrap method, not a load-bearing premise for the new probability theorems; the paper also acknowledges that the maximal probability bound 1/4 is 'not a mathematical conclusion but derived from various experiments.' No step was found where the derivation reduces by construction to its own inputs or where a fitted parameter is relabeled as a prediction. The p(1) metric is a chosen proxy for sampling efficiency, but the paper does not define 'improvement' as p(1); it proves the metric's values and separately argues the importance of nontrivial sampling in Appendix A. Thus no circularity is established.
Assumptions & free parameters
free parameters (2)
- critical exponent ν =
varies with r and n; e.g., ν=0 for strong, ν=(n/2−1)log2/log n for weak extremes
- dataset feature-to-angle mapping =
features mapped to {0,1,2,3} -> angles {0,π,π/2,−π/2}
assumptions (4)
- standard math Gottesman-Knill theorem: Clifford circuits can be classically simulated in polynomial time
- domain assumption The measurement 'outcome' is defined as the sign of the Pauli expectation value, with 0 for states not in a ±1 eigenspace
- domain assumption Uniform distribution over the four angles {0,π/2,π,−π/2} when computing sampling probability
- ad hoc to paper The decay law p(r,n)=1/(4 n^ν) for intermediate observables is extrapolated from experiments
invented entities (3)
-
Strong stabilizer enhancement
-
Weak stabilizer enhancement
-
Critical exponent ν
Cite this review
Pith. "Pith review of The Stabilizer Bootstrap of Quantum Machine Learning with up to 10000 qubits." pith.science (2026). https://pith.science/paper/BDJAJE7D
@misc{pith2026241211356,
author = {Pith},
title = {Pith review of: The Stabilizer Bootstrap of Quantum Machine Learning with up to 10000 qubits},
year = {2026},
howpublished = {\url{https://pith.science/paper/BDJAJE7D}},
note = {Machine review of arXiv:2412.11356}
}
read the original abstract
Quantum machine learning is considered one of the flagship applications of quantum computers, where variational quantum circuits could be the leading paradigm both in the near-term quantum devices and the early fault-tolerant quantum computers. However, it is not clear how to identify the regime of quantum advantages from these circuits, and there is no explicit theory to guide the practical design of variational ansatze to achieve better performance. We address these challenges with the stabilizer bootstrap, a method that uses stabilizer-based techniques to optimize quantum neural networks before their quantum execution, together with theoretical proofs and high-performance computing with 10000 qubits or random datasets up to 1000 data. We find that, in a general setup of variational ansatze, the possibility of improvements from the stabilizer bootstrap depends on the structure of the observables and the size of the datasets. The results reveal that configurations exhibit two distinct behaviors: some maintain a constant probability of circuit improvement, while others show an exponential decay in improvement probability as qubit numbers increase. These patterns are termed strong stabilizer enhancement and weak stabilizer enhancement, respectively, with most situations falling in between. Our work seamlessly bridges techniques from fault-tolerant quantum computing with applications of variational quantum algorithms. Not only does it offer practical insights for designing variational circuits tailored to large-scale machine learning challenges, but it also maps out a clear trajectory for defining the boundaries of feasible and practical quantum advantages.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
A comprehensive review of quantum machine learning: from nisq to fault tolerance
Yunfei Wang and Junyu Liu. A comprehensive review of quantum machine learning: from nisq to fault tolerance. Reports on Progress in Physics, 2024
work page 2024
-
[2]
Quantum machine learning algorithms for drug discovery applications
Kushal Batra, Kimberley M Zorn, Daniel H Foil, Eni Minerali, Victor O Gawriljuk, Thomas R Lane, and Sean Ekins. Quantum machine learning algorithms for drug discovery applications. Journal of chemical information and modeling, 61(6):2641–2647, 2021
work page 2021
-
[3]
Y. Cao, J. Romero, and A. Aspuru-Guzik. Potential of quantum computing for drug discovery. IBM Journal of Research and Development, 62(6):6:1–6:20, 2018
2018
-
[4]
A variational eigenvalue solver on a photonic quantum processor
Alberto Peruzzo, Jarrod McClean, Peter Shadbolt, Man- Hong Yung, Xiao-Qi Zhou, Peter J Love, Al´ an Aspuru- Guzik, and Jeremy L O’Brien. A variational eigenvalue solver on a photonic quantum processor. Nature Com- munications, 5:4213, 2014
work page 2014
-
[5]
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. pages 249–256, 2010
work page 2010
-
[6]
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
2015
-
[7]
Solving the 3d ising model with the conformal bootstrap ii
Sheer El-Showk, Miguel F Paulos, David Poland, Slava Rychkov, David Simmons-Duffin, and Alessandro Vichi. Solving the 3d ising model with the conformal bootstrap ii. c c-minimization and precise critical exponents. Jour- nal of Statistical Physics, 157:869–914, 2014
work page 2014
-
[8]
Gokul Subramanian Ravi, Pranav Gokhale, Yi Ding, William Kirby, Kaitlin Smith, Jonathan M. Baker, Pe- ter J. Love, Henry Hoffmann, Kenneth R. Brown, and Frederic T. Chong. Cafqa: A classical simulation boot- strap for variational quantum algorithms. ASPLOS 2023, page 15–29, New York, NY, USA, 2022. Association for Computing Machinery
work page 2023
Show all 22 references
-
[9]
M. S. Rudolph, J. Miller, D. Motlagh, P. J. Coles, and Z. Wang. Synergistic pretraining of parametrized quan- tum circuits via tensor networks. Nature Communica- tions, 14(1):8367, 2023
2023
-
[10]
Escaping from the barren plateau via gaussian ini- tializations in deep variational quantum circuits
Kaining Zhang, Liu Liu, Min-Hsiu Hsieh, and Dacheng Tao. Escaping from the barren plateau via gaussian ini- tializations in deep variational quantum circuits. 2024
2024
-
[11]
James Dborin, Fergus Barratt, Vinul Wimalaweera, Lewis Wright, and Andrew G. Green. Matrix product state pre-training for quantum machine learning, 2021
2021
-
[12]
Avoiding barren plateaus via Gaussian Mixture Model
Xiao Shi and Yun Shang. Avoiding barren plateaus via Gaussian Mixture Model. 2 2024
2024
-
[13]
While these advances show important progress, current research remains lim- ited to small-scale systems and has no discussion of the dependence on involved dataset
initialization and even could show 2.5× faster conver- gence than HF for small molecules. While these advances show important progress, current research remains lim- ited to small-scale systems and has no discussion of the dependence on involved dataset. In fact, since the op-...
2024 arXiv
-
[14]
Self- consistent field, with exchange, for beryllium
Douglas Rayner Hartree and William Hartree. Self- consistent field, with exchange, for beryllium. Pro- ceedings of the Royal Society of London. Series A- Mathematical and Physical Sciences , 150(869):9–33, 1935
1935
-
[15]
The heisenberg representation of quantum computers, 1998
Daniel Gottesman. The heisenberg representation of quantum computers, 1998
1998
-
[16]
The term stabilizer bootstrapping, similar to our term, is used in [18] but only for quantum tomography
-
[17]
Beyond NISQ: The Megaquop Machine, Q2B 2024
John Preskill. Beyond NISQ: The Megaquop Machine, Q2B 2024
2024
-
[18]
Partial syn- drome measurement for hypergraph product codes
Noah Berthusen and Daniel Gottesman. Partial syn- drome measurement for hypergraph product codes. Quantum, 8:1345, 2024
2024
-
[19]
Stabilizer bootstrapping: A recipe for efficient agnos- tic tomography and magic estimation
Sitan Chen, Weiyuan Gong, Qi Ye, and Zhihan Zhang. Stabilizer bootstrapping: A recipe for efficient agnos- tic tomography and magic estimation. arXiv preprint arXiv:2408.06967, 2024. 7 Appendix Appendix A: Background
2024 arXiv
-
[20]
Quantum Machine Learning (QML) QML is a hybrid quantum-classical framework where variational quantum circuits are trained using classical op- timization methods to minimize an objective function, as illustrated in FIG. 8. During each training iteration, the feature vectors ( x...
-
[21]
Bayesian optimization Bayesian Optimization (BO) is an efficient framework for the global optimization of expensive black-box functions, particularly in scenarios where function evaluations are costly, time-consuming, or require significant computational resources. Unlike trad...
-
[22]
However, according to the Gottesman-Knill theorem, Clifford circuits can be efficiently simulated by classical computers in polynomial time
The stabilizer formalism Classical quantum computing tasks often require exponential resources. However, according to the Gottesman-Knill theorem, Clifford circuits can be efficiently simulated by classical computers in polynomial time. The Gottesman- Knill theorem states that...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.