Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Pitfalls when tackling the exponential concentration of parameterized quantum models

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Exponentially concentrated quantum measurements produce samples indistinguishable from fixed random noise, so no classical post-processing can recover the lost information.

desk verdict Sound core indistinguishability theorem, but the applied claims outrun the proofs—QNG in particular is currently a conjecture. read the letter →

arxiv 2507.22054 v2 pith:UCDCVXL5 submitted 2025-07-29 quant-ph

classification quant-ph MSC 81P6868Q12 PACS 03.67.Ac03.67.Lx
keywords exponentialconcentrationbarrenplateausparameterizedquantumcircuitsvariationalalgorithmsnaturalgradienthypothesistestingstatisticalindistinguishabilitymachinelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the right place to diagnose barren-plateau-style concentration is not the loss function's expectation value but the probabilities of individual measurement outcomes. Its central theorem states that when a measurement with polynomially many outcomes has exponentially concentrated outcome probabilities, polynomially many shots produce samples that are statistically indistinguishable from samples drawn from a fixed, variable-independent distribution. Because no classical post-processing can break this indistinguishability, any optimization strategy that only changes the classical part of a variational quantum algorithm — quantum natural gradient, sample-based CVaR losses, neural-network-assisted initialization, rescaled parameter shifts — inherits the failure. The paper therefore supplies a practical step-by-step diagnostic: identify the POVMs a procedure actually uses, check that they have polynomially many elements, and test whether their outcome probabilities concentrate. If they do, the procedure cannot learn from the quantum device with a polynomial measurement budget.

What carries the argument

The load-bearing object is outcome probability concentration (Definition 1): for every POVM element $M_k$, $\Pr_{\alpha\sim\mathcal D}(|p_k(\alpha)-\mu_k|\ge\delta)\le\beta/\delta^2$ with $\beta\in O(\exp(-n))$. The proof couples this to binary hypothesis testing: Lemma 1 gives the one-sample success probability $1/2+\|P-P'\|_1/4$, Lemma 2 bounds the $N$-sample product-distribution distance by the sum of single-sample distances, and Proposition 1 bounds the success probability by $1/2+N\|P_0-P'_0\|_1/4$. Choosing $\delta'=\beta^{1/4}$ and applying a union bound over the polynomially many POVM elements yields $\delta=|\mathcal M|\sqrt\beta$ and $\varepsilon=N|\mathcal M|\beta^{1/4}/4$, both exponentially vanishing. Corollary 1 then transfers indistinguishability to arbitrary post-processing maps by contradiction.

What would settle it

Run the binary hypothesis test behind Theorem 2 on a concrete concentrated model: choose a global-observable ansatz known to have exponentially concentrated outcomes, draw polynomially many shots at many random parameters, and test whether the empirical outcome histogram can be distinguished from the fixed distribution $P_{\mathrm{fixed}}$ with success probability exceeding $1/2+\exp(-n)$. If any polynomial-shot protocol, including one followed by quantum natural gradient or CVaR post-processing, reliably identifies the parameter value or produces a gradient estimate whose sign correlates with the true gradient at a rate that does not vanish exponentially in $n$, then Theorem 1 or Corollary 1 would fail.

Watch

Extended reading notes

Core claim

The central discovery is that exponential concentration at the level of measurement outcome probabilities implies statistical indistinguishability from a fixed, variable-independent distribution. Formally, for a POVM with $|\mathcal{M}|\in O(\operatorname{poly}(n))$, if each outcome probability $p_k(\alpha)=\operatorname{Tr}[\rho(\alpha)M_k]$ satisfies Definition 1 with $\beta\in O(\exp(-n))$, then after $N\in O(\operatorname{poly}(n))$ shots the samples $S_N(\alpha)$ are with probability at least $1-\delta$, $\delta\in O(\exp(-n))$, statistically indistinguishable from samples drawn from $P_{\mathrm{fixed}}=(\mu_1,\dots,\mu_{|\mathcal{M}|})$. Corollary 1 extends this to any classical post-processing map, because distinguishable outputs would yield a strategy to distinguish the underlying distributions. A direct corollary is that training a concentrated loss with vanilla gradient descent and polynomial shots is statistically indistinguishable from a random walk. The authors use this framework to argue that quantum natural gradient, sample-based CVaR optimization, classical neural-network-assisted initialization, and rescaled parameter-shift rules do not circumvent exponential concentration with finite measurement budgets, even though they may provide other training benefits.

Load-bearing premise

The negative conclusions about specific methods all rest on the premise that those methods' outcome probabilities actually concentrate exponentially under the parameter distribution they induce; if a method changes that distribution (as warm starting can), the conclusion stops applying, and for quantum natural gradient the paper itself notes that the proof is not supplied.

Editorial extensions

If this is right

  • Switching from vanilla gradients to quantum natural gradient, sample-based CVaR losses, classical-neural-network initialization, or rescaled parameter-shift rules does not remove the information bottleneck when the underlying outcome probabilities concentrate; such methods may still aid training in other ways.
  • Diagnosing barren plateaus solely through loss-variance scaling is insufficient, because an exponentially large prefactor can suppress the variance without changing the information content; outcome-probability concentration captures the finite-shot information content directly.
  • Training a concentrated landscape with vanilla gradient descent and polynomial shot budgets is provably a random walk: with probability exponentially close to one, each parameter update is statistically indistinguishable from a parameter-independent random variable.
  • The diagnostic extends beyond variational training to quantum kernel methods and quantum reservoir models, where the relevant variables are classical input data rather than trainable parameters.
  • The only ways out are to change the concentration properties themselves (for example, via identity initialization or warm starting) or to use measurements with exponentially many POVM elements, a regime the theorem explicitly does not cover.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The theorem implies that measure-first-estimate-later strategies such as classical shadows do not constitute a loophole: for a given target observable, their effective POVM is mathematically equivalent to directly measuring that observable, so the same indistinguishability bound applies regardless of the measurement protocol's randomization.
  • A testable extension is to report, for any proposed mitigation method, the induced distribution over measurement outcomes and to run the binary hypothesis test against the fixed distribution directly, rather than relying only on loss curves or variances.
  • The exponential-POVM caveat is the most promising open door in the paper: it suggests that entangled measurements which coarse-grain exponentially many outcomes into globally correlated patterns are the natural place to look for genuine avoidance of exponential concentration.
  • If the central claim is right, it sharpens the classical-simulability debate: under exponential concentration and polynomial shots, the quantum device contributes no learning signal at all, regardless of how powerful the classical post-processing is.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a shift in how exponential concentration is diagnosed in parameterized quantum models: instead of analyzing loss- or gradient-variance scaling, it analyzes concentration at the level of POVM outcome probabilities. A general procedure is introduced that covers loss evaluation, gradient-based training, natural gradient, CVaR-style sample processing, and neural-network-assisted initialization. The main theoretical result (Theorem 2, Appendix B) shows that if a POVM with polynomially many elements has exponentially concentrated outcome probabilities, then after polynomially many measurement shots the samples are statistically indistinguishable from samples from a fixed, variable-independent distribution, and no classical post-processing can restore distinguishability (Corollary 3). A corollary (Corollary 4) claims that vanilla gradient descent on a barren-plateau landscape is statistically indistinguishable from a random walk. The framework is then applied to argue that quantum natural gradient, sample-based CVaR optimization, classical neural-network-assisted initialization, and a rescaled parameter-shift rule do not overcome exponential concentration under finite measurement budgets.

Significance. The core hypothesis-testing formulation is valuable and clean: Theorem 2 provides a rigorous, parameter-free bridge between outcome-probability concentration and statistical indistinguishability, and Corollary 3 correctly captures the intuition that any classical post-processing of already-indistinguishable samples cannot recover information. The proof in Appendix B is elementary and sound, and the practical guidelines in Section IV are a useful diagnostic tool. The paper also honestly acknowledges the scope limitation to polynomial-size POVMs and includes a counterexample (Appendix D) showing that exponentially many POVM elements could in principle evade the theorem. If the applied claims were fully proven, the paper would be an important contribution to the barren-plateau literature. As it stands, the central framework is solid but the headline applications are only partially established.

major comments (3)
  1. [Section III and Appendix C] The claim that quantum natural gradient (QNG) does not overcome exponential concentration is not established. The paper states in Section III that 'the proof would be more complex' and provides no rigorous argument that the QNG update is information-less when the QGT is non-degenerate. The only numerical evidence, Figure 4a, is generated with the circuit in Eq. (C1), a single layer of X-rotations with a global Z observable; Appendix C explicitly computes g^+(θ) = 4I, making the QNG update rule (C2) exactly equivalent to gradient descent with an effective learning rate 4η. Thus the numerics do not test QNG's geometry at all. The informal discussion about QGT elements from deeper circuit layers does not address the possibility that the pseudo-inverse of the QGT could amplify or reweight gradient noise to produce a meaningful direction, nor does it account for the fact that QNG changes the parameter distribution over training while Definition 1 is distribution-dependent. The QNG conclusion should either be proven for a non-degenerate QGT or explicitly downgraded to a conjecture.
  2. [Section IV, neural-network-assisted initialization] The assertion that agnostic classical neural-network-assisted initialization 'does not address the root cause' and cannot overcome exponential concentration is not derived from the framework. The concentration notion in Definition 1 is defined with respect to a distribution D over the variables. When the trainable variables are the weights and biases of a classical neural network instead of the circuit parameters, the induced distribution over the circuit parameters is not analyzed. The numerical experiment in Figure 4c uses the same single-layer ansatz as the other panels, but no argument is given that the neural network's output distribution retains the exponential concentration properties. The paper itself concedes that warm starting can avoid concentration; the NN-induced distribution could equally change the concentration regime. As written, this conclusion is an assumption rather than a consequence of Theorem 2.
  3. [Appendix B, proof of Corollary 4] The random-walk proof reuses the initial concentration distribution D at every training step without accounting for the evolution of the parameter distribution. In Eqs. (B20)-(B25), the event A_ijk has probability at least 1 - 2√β only when the parameter point α_ijk is drawn from D, but after the first update the parameters θ^(k) are distributed according to a convolution of the initial distribution with the random-walk increments, which is generally not D. The union bound over training steps therefore does not follow from Theorem 2 unless the concentration property holds for the time-marginal distributions of the parameters, which is not established. This gap affects the generality of Corollary 2, which is presented as a standalone result.
minor comments (6)
  1. [Appendix B, Theorem 2 proof] The statement 'by denoting δ = |M|√β, we also have δ ∈ O(poly(n))' is a typo; δ = |M|√β is exponentially small, not polynomial.
  2. [Theorems 1 and 2, Corollaries 1 and 2] The informal statements use δ ∈ O(exp(-n)) and ε ∈ O(exp(-n)), while the formal proof gives δ = |M|√β and ε = N|M|β^{1/4}/4. For β = b^{-n}, these are exponentially decaying but with rates b^{-n/2} and b^{-n/4} up to polynomial factors, not necessarily b^{-n}. Please state the asymptotic rates consistently.
  3. [Section II and Corollary 4] The symbol Nℓ is used both for the number of physical quantities in the general procedure and, in Corollary 4, for the number of measurement sets 2NpNL. This reuse is confusing; please introduce distinct notation.
  4. [Appendix A, Lemma 1 proof] The sentence 'guess 1 if the two distributions are identical, this A is simply an empty set and the probability of making the right decision is simply 1/2' is incomplete and should be rewritten.
  5. [Figure 4] The caption lists n = 9, 11, 13, 15, 17 and three shot regimes, but the line styles and colors in panels a-d make it difficult to distinguish the different settings; please improve the figure legend and consider separate panels or clearer markers.
  6. [Section IV, Eq. (12)] The ratio ϵ_N(α) is defined as a function of α, but the condition 'we require ϵ_N(α) ≲ 1' is ambiguous; clarify whether this is required pointwise, in expectation, or for typical α.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: Theorem 1 follows from Definition 1 via standard hypothesis testing bounds, and the applied conclusions depend on external concentration premises rather than on the paper's own definitions.

full rationale

The central derivation chain is self-contained and not circular. Definition 1 states a Chebyshev-type tail bound on individual POVM outcome probabilities; Theorem 2 (formal version of Theorem 1) combines a union bound over POVM elements with the binary-hypothesis-testing bound of Proposition 1 to show that polynomial samples cannot distinguish P_alpha from P_fixed. This is a genuine probabilistic implication, not a restatement of the definition. Corollary 3 then applies an arbitrary post-processing map, and Corollary 4 repeats the argument with a union bound over training steps to obtain the random-walk conclusion; all steps are proved in Appendices A-B. The applications to specific methods (CVaR, neural-network initialization, scaled parameter-shift, and quantum natural gradient) rely on the external premise that the relevant circuits exhibit outcome-probability concentration, as cited from the barren-plateau literature (e.g., Ref. [41]); that is an input assumption, not a conclusion forced by the paper's own equations. There are self-citations (e.g., Refs. [10], [28], [34], [56], [79]), but they support background or auxiliary applications and are not load-bearing for Theorem 1. The quantum-natural-gradient discussion is explicitly informal ('Nonetheless, the proof would be more complex'), and Appendix C shows the numerical QNG experiment degenerates to g^+(theta)=4I, i.e., to plain gradient descent; this is a gap in evidence rather than a circular step. No fitted parameter is renamed as a prediction, and no uniqueness theorem is imported from the authors' prior work.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The central theorem is derived from the concentration definition (Definition 1) and standard probability inequalities, with no fitted parameters in the proof. The applied conclusions carry two extra burdens: (1) the relevant procedures are expressible with polynomially many POVM elements, and (2) the specific circuits analyzed have exponentially concentrating outcome probabilities, inherited from known barren-plateau results. Both are domain assumptions stated in the paper. Numerical hyperparameters (gamma, lambda) appear only in the demonstrations and are not load-bearing.

free parameters (2)
  • CVaR quantile hyperparameter gamma = 0.8
    Chosen for the numerical demo of sample-based CVaR optimization; the theoretical claim applies for any post-processing map, so this value is not load-bearing.
  • RPS scaling factor lambda = lambda = N d / (2 d^2 + N d - 2), d = 2^n
    Taken from the rescaled parameter-shift method [87] and used in the numerical demo; it is computed from the shot count and Hilbert-space dimension, not fitted by this paper.
assumptions (6)
  • standard math The success probability of binary hypothesis testing with N samples is at most 1/2 + N||P-P'||_1/4 (Lemma 1 and Proposition 1).
    Used in the proof of Theorem 2 to convert one-norm closeness into statistical indistinguishability.
  • standard math The union bound over POVM elements and over training steps is valid.
    Used to combine the concentration events for individual POVM elements in Theorem 2 and to bound the probability of a random-walk trajectory in Corollary 4.
  • domain assumption Standard parameterized quantum models can be represented with POVMs having |M| in O(poly(n)) elements ('Polynomial POVMs in disguise').
    The entire framework applies only to polynomial POVMs; exponential POVMs are outside its scope, as discussed in Section II and Section V.
  • domain assumption The analyzed circuits (e.g., global-Pauli observables) have exponentially concentrating outcome probabilities, inherited from known barren-plateau results such as Ref. [41].
    Used to apply the framework to QNG, CVaR, NN-init, and RPS; this property is not re-derived for these methods in the paper.
  • standard math The parameter-shift rule gives the gradient as a difference of two loss estimates (Refs. [7,93]).
    Used in Corollary 4 to show that gradient-descent updates are post-processing of concentrated measurement samples.
  • domain assumption Polynomial measurement budget, polynomially many quantities N_l, and polynomially many training steps.
    The indistinguishability result is asymptotic in n with poly(n) shots and poly(n) training steps.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Pitfalls when tackling the exponential concentration of parameterized quantum models." pith.science (2026). https://pith.science/paper/UCDCVXL5

@misc{pith2026250722054,
  author       = {Pith},
  title        = {Pith review of: Pitfalls when tackling the exponential concentration of parameterized quantum models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UCDCVXL5}},
  note         = {Machine review of arXiv:2507.22054}
}
read the original abstract

Identifying scalable circuit architectures remains a central challenge in variational quantum computing and quantum machine learning. Many approaches have been proposed to mitigate or avoid the barren plateau phenomenon or, more broadly, exponential concentration. However, due to the intricate interplay between quantum measurements and classical post-processing, we argue these techniques often fail to circumvent concentration effects in practice. Here, by analyzing concentration at the level of measurement outcome probabilities and leveraging tools from hypothesis testing, we develop a practical framework for diagnosing whether a parameterized quantum model is inhibited by exponential concentration. Applying this framework, we argue that several widely used methods (including quantum natural gradient, sample-based optimization, and certain neural-network-inspired initializations) do not overcome exponential concentration with finite measurement budgets, though they may still aid training in other ways.

Figures

Figures reproduced from arXiv: 2507.22054 by the authors.

Figure 1
Figure 1. Schematic representation of results. If mea￾surement outcomes (at the level of probabilities) are exponen￾tially concentrated then, with high probability, they contain no information about the trainable parameters and/or data inputs in the sense that they are indistinguishable from sam￾ples drawn from a variable-independent probability distribution (Theorem 1). It follows that further post-processing these mea￾surem… view at source ↗
Figure 2
Figure 2. Framework. Panel a) illustrates the description of the general procedure (as described in Section II). The procedure involves extracting information from variable-dependent quantum states through some quantum measurements. These measure￾ment outcomes are then processed to estimate some variable-dependent quantities, which can in turn be all processed together. Panel b) illustrates specific examples that fall within … view at source ↗
Figure 3
Figure 3. Training on the BP landscape with 15 qubits. In panel a), training trajectories with different measurement shots are present on a 2D cut of the landscape (chosen via PCA analysis [95]), showing random trajectories with polynomial measurement shots. In panels b) and c), the mean and variance of the updated parameters over different training trajectories are shown to align with the mean and variance of random walks. M… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Training curves. We plot the loss as a function of training steps for different shot budgets, various optimization methods and system sizes: Panel a) shows quantum natural gradient descent; Panel b) shows sample-based CVaR optimization; Panel c) shows classical neural …
Figure 5
Figure 5. Figure 5: Two different POVMs for estimating fidelity. Panel (a) illustrates the Loschmidt echo test, while panel (b) schematically depicts the SWAP test. able is mathematically equivalent to implementing the cor￾responding POVM of that observable. Consequently, our concentratio…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Trainability Beyond Linearity in Variational Quantum Objectives

    quant-ph 2026-04 unverdicted novelty 7.0 of 10

    The trainability boundary for variational quantum objectives is the affine regime; non-affine amplification-capable losses can mitigate barren plateaus when using coarse-grained statistics at polynomial widths.

Reference graph

Works this paper leans on

110 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [1]

    Given a procedureP, identify the quantities{ℓi(αi)} which require information to be extracted from quan- tum computers

  2. [2]

    For each quantity{ℓi(αi)}, identify the correspond- ingM (i) and check whether|M (i)|, the number of POVM elements inM (i), scales at most polynomi- ally with system sizen. Note that all barren plateau mitigation strategies we are aware of involve such polynomial-sized POVMs, even if some may initially appear to have exponen- tially many elements (see Sec...

  3. [3]

    Determine whether the outcome probabilities p(i) k (αi) = Tr h ρi(αi)M (i) k i exponentially concentrate with respect toαi (as per Definition 1). If this is the case, the procedurePsuffers from the concentration in the sense that the measurement out- comes, with probability exponentially close to 1, con- tain no information about the variablesαi. Based on...

  4. [4]

    Cerezo, A

    M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, Variational quantum algo- rithms, Nature Reviews Physics3, 625–644 (2021)

  5. [5]

    Bharti, A

    K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J. S. Kottmann, T. Menke,et al., Noisy intermediate- scale quantum algorithms, Reviews of Modern Physics94, 015004 (2022)

  6. [6]

    Abbas, R

    A. Abbas, R. King, H.-Y. Huang, W. J. Huggins, R. Movassagh, D. Gilboa, and J. R. McClean, On quan- tum backpropagation, information reuse, and cheating measurement collapse, arXiv preprint arXiv:2305.13362 (2023)

  7. [7]

    McArdle, S

    S. McArdle, S. Endo, A. Aspuru-Guzik, S. C. Benjamin, andX.Yuan,Quantumcomputationalchemistry,Reviews of Modern Physics92, 015003 (2020)

  8. [8]

    J. R. McClean, J. Romero, R. Babbush, and A. Aspuru- Guzik, The theory of variational hybrid quantum-classical algorithms, New Journal of Physics18, 023023 (2016)

Show all 110 references
  1. [9]

    Pérez-Salinas, A

    A. Pérez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, Data re-uploading for a universal quantum classifier, Quantum4, 226 (2020)

  2. [10]

    Mitarai, M

    K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, Quantum circuit learning, Physical Review A98, 032309 (2018)

  3. [11]

    Schuld, Supervised quantum machine learning mod- els are kernel methods, arXiv preprint arXiv:2101.11020 (2021)

    M. Schuld, Supervised quantum machine learning mod- els are kernel methods, arXiv preprint arXiv:2101.11020 (2021)

  4. [12]

    Havlíček, A

    V. Havlíček, A. D. Córcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, Supervised learning with quantum-enhanced feature spaces, Nature 567, 209 (2019)

  5. [13]

    Thanasilp, S

    S. Thanasilp, S. Wang, M. Cerezo, and Z. Holmes, Expo- nential concentration in quantum kernel methods, Nature Communications15, 5200 (2024)

  6. [14]

    Kübler, S

    J. Kübler, S. Buchholz, and B. Schölkopf, The inductive bias of quantum kernels, Advances in Neural Information Processing Systems34, 12661 (2021)

  7. [15]

    Huang, M

    H.-Y. Huang, M. Broughton, M. Mohseni, R. Babbush, S. Boixo, H. Neven, and J. R. McClean, Power of data in quantum machine learning, Nature Communications12, 1 (2021)

  8. [16]

    Gentinetta, A

    G. Gentinetta, A. Thomsen, D. Sutter, and S. Woerner, The complexity of quantum support vector machines, arXiv preprint arXiv:2203.00031 (2022)

  9. [17]

    B. Y. Gan, D. Leykam, and S. Thanasilp, A unified framework for trace-induced quan- tum kernels, arXiv preprint arXiv:2311.13552 https://doi.org/10.48550/arXiv.2311.13552 (2023)

  10. [18]

    Zimborás, B

    Z. Zimborás, B. Koczor, Z. Holmes, E.-M. Borrelli, A. Gilyén, H.-Y. Huang, Z. Cai, A. Acín, L. Aolita, L. Banchi,et al., Myths around quantum computation before full fault tolerance: What no-go theorems rule out and what they don’t, arXiv preprint arXiv:2501.05694 https://doi....

  11. [19]

    E. R. Anschuetz and B. T. Kiani, Quantum variational algorithms are swamped with traps, Nature Communica- tions13, 7760 (2022)

  12. [20]

    Larocca, N

    M. Larocca, N. Ju, D. García-Martín, P. J. Coles, and M. Cerezo, Theory of overparametrization in quantum neural networks, Nature Computational Science3, 542 (2023)

  13. [21]

    F. J. Schreiber, J. Eisert, and J. J. Meyer, Classical surro- gates for quantum learning models, Physical Review Let- ters131, 100803 (2023)

  14. [22]

    Sweke, E

    R. Sweke, E. Recio, S. Jerbi, E. Gil-Fuster, B. Fuller, J.Eisert,andJ.J.Meyer,Potentialandlimitationsofran- dom fourier features for dequantizing quantum machine learning, Quantum9, 1640 (2025)

  15. [23]

    Sahebi, A

    M. Sahebi, A. Barthe, Y. Suzuki, Z. Holmes, and M. Grossi, On dequantization of supervised quantum ma- chine learning via random fourier features, arXiv preprint arXiv:2505.15902 (2025)

  16. [24]

    M. S. Rudolph, T. Jones, Y. Teng, A. Angrisani, and Z. Holmes, Pauli propagation: A computational frame- 12 work for simulating quantum systems, arXiv preprint arXiv:2505.21606 (2025)

  17. [25]

    Angrisani, A

    A. Angrisani, A. Schmidhuber, M. S. Rudolph, M. Cerezo, Z. Holmes, and H.-Y. Huang, Classically estimating ob- servables of noiseless quantum circuits, arXiv preprint arXiv:2409.01706 (2024)

  18. [26]

    S. Shin, Y. S. Teo, and H. Jeong, Dequantizing quan- tummachinelearningmodelsusingtensornetworks,Phys. Rev. Res.6, 023218 (2024)

  19. [27]

    M. L. Goh, M. Larocca, L. Cincio, M. Cerezo, and F. Sauvage, Lie-algebraic classical simulations for quan- tum computing, arXiv preprint arXiv:2308.01432 (2023)

  20. [28]

    Kerenidis, J

    I. Kerenidis, J. Landman, and N. Mathur, Classical and quantum algorithms for orthogonal neural networks, arXiv preprint arXiv:2106.07198 (2021)

  21. [29]

    Schuster, C

    T. Schuster, C. Yin, X. Gao, and N. Y. Yao, A polynomial-time classical algorithm for noisy quantum circuits, arXiv preprint arXiv:2407.12768 https://doi.org/10.48550/arXiv.2407.12768 (2024)

  22. [30]

    Fontana, M

    E. Fontana, M. S. Rudolph, R. Duncan, I. Rungger, and C. Cîrstoiu, Classical simulations of noisy variational quantum circuits, npj Quantum Information11, 1 (2025)

  23. [31]

    Larocca, S

    M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Bia- monte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo, A review of barren plateaus in variational quantum computing, Nature Reviews Physics3, 625–644 (2025)

  24. [32]

    J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, andH.Neven,Barrenplateausinquantumneuralnetwork training landscapes, Nature Communications9, 1 (2018)

  25. [33]

    Fontana, D

    E. Fontana, D. Herman, S. Chakrabarti, N. Kumar, R. Yalovetzky, J. Heredge, S. H. Sureshbabu, and M. Pis- toia, Characterizing barren plateaus in quantum ansätze with the adjoint representation, Nature Communications 15, 7171 (2024)

  26. [34]

    Ragone, B

    M. Ragone, B. N. Bakalov, F. Sauvage, A. F. Kemper, C. Ortiz Marrero, M. Larocca, and M. Cerezo, A lie al- gebraic theory of barren plateaus for deep parameter- ized quantum circuits, Nature Communications15, 7172 (2024)

  27. [35]

    Cerezo, M

    M. Cerezo, M. Larocca, D. García-Martín, N. L. Diaz, P. Braccia, E. Fontana, M. S. Rudolph, P. Bermejo, A. Ijaz, S. Thanasilp,et al., Does provable absence of bar- renplateausimplyclassicalsimulability? Or, whyweneed to rethink variational quantum computing, arXiv preprint arX...

  28. [36]

    Suzuki and M

    Y. Suzuki and M. Li, Effect of alternating layered ansatzes on trainability of projected quantum kernel, arXiv preprint arXiv:2310.00361 (2023)

  29. [37]

    Xiong, G

    W. Xiong, G. Facelli, M. Sahebi, O. Agnel, T. Chotibut, S. Thanasilp, and Z. Holmes, On fundamental aspects of quantum extreme learning machines, arXiv preprint arXiv:2312.15124 (2023)

  30. [38]

    Xiong, Z

    W. Xiong, Z. Holmes, A. Angrisani, Y. Suzuki, T. Chotibut, and S. Thanasilp, Role of scrambling and noise in temporal information processing with quantum systems, arXiv preprint arXiv:2505.10080 (2025)

  31. [39]

    Shaydulin and S

    R. Shaydulin and S. M. Wild, Importance of kernel band- width in quantum machine learning, Physical Review A 106, 042407 (2022)

  32. [40]

    Arrasmith, Z

    A. Arrasmith, Z. Holmes, M. Cerezo, and P. J. Coles, Equivalence of quantum barren plateaus to cost concen- tration and narrow gorges, Quantum Science and Tech- nology7, 045015 (2022)

  33. [41]

    Arrasmith, M

    A. Arrasmith, M. Cerezo, P. Czarnik, L. Cincio, and P. J. Coles, Effect of barren plateaus on gradient-free optimiza- tion, Quantum5, 558 (2021)

  34. [42]

    Larocca, F

    M. Larocca, F. Sauvage, F. M. Sbahi, G. Verdon, P. J. Coles, and M. Cerezo, Group-invariant quantum machine learning, PRX Quantum3, 030341 (2022)

  35. [43]

    Schatzki, M

    L. Schatzki, M. Larocca, Q. T. Nguyen, F. Sauvage, and M. Cerezo, Theoretical guarantees for permutation- equivariant quantum neural networks, npj Quantum In- formation10, 12 (2024)

  36. [44]

    Cerezo, A

    M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shallow parametrized quantum circuits, Nature Communications 12, 1 (2021)

  37. [45]

    Larocca, P

    M. Larocca, P. Czarnik, K. Sharma, G. Muraleedharan, P. J. Coles, and M. Cerezo, Diagnosing Barren Plateaus with Tools from Quantum Optimal Control, Quantum6, 824 (2022)

  38. [46]

    Letcher, S

    A. Letcher, S. Woerner, and C. Zoufal, Tight and effi- cient gradient bounds for parameterized quantum circuits, Quantum8, 1484 (2024)

  39. [47]

    Basheer, Y

    A. Basheer, Y. Feng, C. Ferrie, and S. Li, Alternating layered variational quantum circuits can be classically op- timized efficiently using classical shadows, arXiv preprint arXiv:2208.11623 (2022)

  40. [48]

    Napp, Quantifying the barren plateau phenomenon for a model of unstructured variational ansätze, arXiv preprint arXiv:2203.06174 (2022)

    J. Napp, Quantifying the barren plateau phenomenon for a model of unstructured variational ansätze, arXiv preprint arXiv:2203.06174 (2022)

  41. [49]

    Zhang, S

    H.-K. Zhang, S. Liu, and S.-X. Zhang, Absence of barren plateaus in finite local-depth circuits with long-range en- tanglement, Physical Review Letters132, 150603 (2024)

  42. [50]

    Pesah, M

    A. Pesah, M. Cerezo, S. Wang, T. Volkoff, A. T. Sorn- borger, and P. J. Coles, Absence of barren plateaus in quantum convolutional neural networks, Physical Review X11, 041011 (2021)

  43. [51]

    Monbroussou, J

    L. Monbroussou, J. Landman, A. B. Grilo, R. Kukla, and E. Kashefi, Trainability and expressivity of hamming- weight preserving quantum circuits for machine learning, arXiv preprint arXiv:2309.15547 (2023)

  44. [52]

    S. Raj, I. Kerenidis, A. Shekhar, B. Wood, J. Dee, S. Chakrabarti, R. Chen, D. Herman, S. Hu, P. Minssen, et al., Quantum deep hedging, Quantum7, 1191 (2023)

  45. [53]

    N. L. Diaz, D. García-Martín, S. Kazi, M. Larocca, and M. Cerezo, Showcasing a barren plateau the- ory beyond the dynamical lie algebra, arXiv preprint arXiv:2310.11505 (2023)

  46. [54]

    A. A. Mele, A. Angrisani, S. Ghosh, S. Khatri, J. Eis- ert, D. S. França, and Y. Quek, Noise-induced shallow circuits and absence of barren plateaus, arXiv preprint arXiv:2403.13927 (2024). 13

  47. [55]

    Deshpande, M

    A. Deshpande, M. Hinsche, S. Najafi, K. Sharma, R. Sweke, and C. Zoufal, Dynamic parameterized quan- tum circuits: expressive and barren-plateau free, arXiv preprint arXiv:2411.05760 10.48550/arXiv.2411.05760 (2024)

  48. [56]

    Srimahajariyapong, S

    K. Srimahajariyapong, S. Thanasilp, and T. Chotibut, Connecting phases of matter to the flatness of the loss landscape in analog variational quantum algorithms, arXiv preprint arXiv:2506.13865 (2025)

  49. [57]

    Mhiri, R

    H. Mhiri, R. Puig, S. Lerch, M. S. Rudolph, T. Chotibut, S. Thanasilp, and Z. Holmes, A unify- ing account of warm start guarantees for patches of quantum landscapes, arXiv preprint arXiv:2502.07889 https://doi.org/10.48550/arXiv.2502.07889 (2025)

  50. [58]

    A. A. Mele, G. B. Mbeng, G. E. Santoro, M. Collura, and P. Torta, Avoiding barren plateaus via transferability of smooth solutions in a Hamiltonian variational ansatz, Physical Review A106, L060401 (2022)

  51. [59]

    R. Puig, M. Drudis, S. Thanasilp, and Z. Holmes, Varia- tional quantum simulation: A case study for understand- ing warm starts, PRX Quantum6, 010317 (2025)

  52. [60]

    S. Y. Chang, S. Thanasilp, B. L. Saux, S. Val- lecorsa, and M. Grossi, Latent style-based quantum gan for high-quality image generation, arXiv preprint arXiv:2406.02668 (2024)

  53. [61]

    Y. Wang, B. Qi, C. Ferrie, and D. Dong, Trainabil- ity enhancement of parameterized quantum circuits via reduced-domain parameter initialization, Physical Review Applied22, 054005 (2024)

  54. [62]

    Park and N

    C.-Y. Park and N. Killoran, Hamiltonian variational ansatz without barren plateaus, Quantum8, 1239 (2024)

  55. [63]

    C.-Y. Park, M. Kang, and J. Huh, Hardware-efficient ansatz without barren plateaus in any depth, arXiv preprint arXiv:2403.04844 (2024)

  56. [64]

    Zhang, L

    K. Zhang, L. Liu, M.-H. Hsieh, and D. Tao, Escaping from the barren plateau via Gaussian initializations in deep variational quantum circuits, inAdvances in Neu- ral Information Processing Systems(2022)

  57. [65]

    Tangpanitanon, S

    J. Tangpanitanon, S. Thanasilp, N. Dangniam, M.-A. Lemonde, and D. G. Angelakis, Expressibility and train- ability of parametrized analog quantum systems for ma- chine learning applications, Physical Review Research2, 043364 (2020)

  58. [66]

    Shi and Y

    X. Shi and Y. Shang, Avoiding barren plateaus via Gaussian mixture model, arXiv preprint arXiv:2402.13501 (2024)

  59. [67]

    C. Cao, Y. Zhou, S. Tannu, N. Shannon, and R. Joynt, Exploiting many-body localization for scalable variational quantum simulation, arXiv preprint arXiv:2404.17560 (2024)

  60. [68]

    Grant, L

    E. Grant, L. Wossnig, M. Ostaszewski, and M. Benedetti, Aninitializationstrategyforaddressingbarrenplateausin parametrized quantum circuits, Quantum3, 214 (2019)

  61. [69]

    L.FriedrichandJ.Maziero,Avoidingbarrenplateauswith classical deep neural networks, Physical Review A106, 042433 (2022)

  62. [70]

    Miao, C.-Y

    J. Miao, C.-Y. Hsieh, and S.-X. Zhang, Neural-network- encoded variational quantum algorithms, Physical Review Applied21, 014053 (2024)

  63. [71]

    A. Rad, A. Seif, and N. M. Linke, Surviving the bar- ren plateau in variational quantum circuits with Bayesian learning initialization, arXiv preprint arXiv:2203.02464 (2022)

  64. [72]

    Faílde, J

    D. Faílde, J. D. Viqueira, M. M. Juane, and A. Gómez, Using differential evolution to avoid local minima in variational quantum algorithms, Scientific Reports13, 10.1038/s41598-023-43404-3 (2023)

  65. [73]

    Kashif and S

    M. Kashif and S. Al-kuwari, Resqnets: A residual ap- proach for mitigating barren plateaus in quantum neural networks (2023), arXiv:2305.03527 [quant-ph]

  66. [74]

    P. K. Barkoutsos, G. Nannicini, A. Robert, I. Tavernelli, and S. Woerner, Improving variational quantum optimiza- tion using cvar, Quantum4, 256 (2020)

  67. [75]

    Stokes, J

    J. Stokes, J. Izaac, N. Killoran, and G. Carleo, Quantum natural gradient, Quantum4, 269 (2020)

  68. [76]

    C. O. Marrero, M. Kieferová, and N. Wiebe, Entanglement-induced barren plateaus, PRX Quan- tum2, 040316 (2021)

  69. [77]

    Sharma, M

    K. Sharma, M. Cerezo, L. Cincio, and P. J. Coles, Train- ability of dissipative perceptron-based quantum neural networks, Physical Review Letters128, 180505 (2022)

  70. [78]

    T. L. Patti, K. Najafi, X. Gao, and S. F. Yelin, Entangle- ment devised barren plateau mitigation, Physical Review Research3, 033090 (2021)

  71. [79]

    S. Wang, E. Fontana, M. Cerezo, K. Sharma, A. Sone, L. Cincio, and P. J. Coles, Noise-induced barren plateaus in variational quantum algorithms, Nature Communica- tions12, 1 (2021)

  72. [80]

    Holmes, K

    Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, Con- necting ansatz expressibility to gradient magnitudes and barren plateaus, PRX Quantum3, 010313 (2022)

  73. [81]

    Khatri, R

    S. Khatri, R. LaRose, A. Poremba, L. Cincio, A. T. Sorn- borger, and P. J. Coles, Quantum-assisted quantum com- piling, Quantum3, 140 (2019)

  74. [82]

    M. S. Rudolph, S. Lerch, S. Thanasilp, O. Kiss, O. Shaya, S. Vallecorsa, M. Grossi, and Z. Holmes, Trainability bar- riers and opportunities in quantum generative modeling, npj Quantum Information10, 116 (2024)

  75. [83]

    Kieferova, O

    M. Kieferova, O. M. Carlos, and N. Wiebe, Quantum gen- erative training using rényi divergences, arXiv preprint arXiv:2106.09567 (2021)

  76. [84]

    Thanaslip, S

    S. Thanaslip, S. Wang, N. A. Nghiem, P. J. Coles, and M. Cerezo, Subtleties in the trainability of quantum ma- chine learning models, Quantum Machine Intelligence5, 21 (2023)

  77. [85]

    Holmes, A

    Z. Holmes, A. Arrasmith, B. Yan, P. J. Coles, A. Albrecht, and A. T. Sornborger, Barren plateaus preclude learning scramblers, Physical Review Letters126, 190501 (2021)

  78. [86]

    E. C. Martín, K. Plekhanov, and M. Lubasch, Barren plateaus in quantum tensor network optimization, Quan- tum7, 974 (2023)

  79. [87]

    E. R. Anschuetz, A unified theory of quantum neural network loss landscapes, arXiv preprint arXiv:2408.11901 14 (2024)

  80. [88]

    Crognaletti, M

    G. Crognaletti, M. Grossi, and A. Bassi, Estimates of loss functionconcentrationinnoisyparametrizedquantumcir- cuits, arXiv preprint arXiv:2410.01893 (2024)

  81. [89]

    R. Mao, G. Tian, and X. Sun, Towards determining the presence of barren plateaus in some chemically inspired variational quantum algorithms, Communications Physics 7, 342 (2024)

  82. [90]

    Y. S. Teo, Optimized numerical gradient and hessian esti- mation for variational quantum algorithms, Physical Re- view A107, 10.1103/physreva.107.042421 (2023)

  83. [91]

    Fujii and K

    K. Fujii and K. Nakajima, Harnessing disordered- ensemble quantum dynamics for machine learning, Phys- ical Review Applied8, 024030 (2017)

  84. [92]

    Nakajima, K

    K. Nakajima, K. Fujii, M. Negoro, K. Mitarai, and M. Kitagawa, Boosting computational power through spa- tial multiplexing in quantum reservoir computing, Physi- cal Review Applied11, 034021 (2019)

  85. [93]

    Mujal, R

    P. Mujal, R. Martí nez-Peña, J. Nokkala, J. García-Beni, G. L. Giorgi, M. C. Soriano, and R. Zambrini, Opportuni- ties in quantum reservoir computing and extreme learning machines, Advanced Quantum Technologies4, 2100027 (2021)

  86. [94]

    Ghosh, A

    S. Ghosh, A. Opala, M. Matuszewski, T. Paterek, and T. C. Liew, Reconstructing quantum states with quantum reservoir networks, IEEE Transactions on Neural Net- works and Learning Systems32, 3148 (2020)

  87. [95]

    F. Hu, G. Angelatos, S. A. Khan, M. Vives, E. Türeci, L. Bello, G. E. Rowlands, G. J. Ribeill, and H. E. Tureci, Tackling sampling noise in physical systems for machine learning applications: Fundamental limits and eigentasks, Physical Review X13, 041020 (2023)

  88. [96]

    Schuld, V

    M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Kil- loran, Evaluating analytic gradients on quantum hard- ware, Physical Review A99, 032331 (2019)

  89. [97]

    Singh, S

    H. Singh, S. Majumder, and S. Mishra, Benchmarking of different optimizers in the variational quantum algorithms for applications in quantum chemistry, The Journal of Chemical Physics159(2023)

  90. [98]

    M. S. Rudolph, S. Sim, A. Raza, M. Stechly, J. R. Mc- Clean, E. R. Anschuetz, L. Serrano, and A. Perdomo- Ortiz, Orqviz: visualizing high-dimensional landscapes in variational quantum algorithms, arXiv preprint arXiv:2111.04695 (2021)

  91. [99]

    Wierichs, C

    D. Wierichs, C. Gogolin, and M. Kastoryano, Avoiding local minima in variational quantum eigensolvers with the natural gradient optimizer, Physical Review Research2, 043246 (2020)

  92. [100]

    T. Haug, K. Bharti, and M. Kim, Capacity and quantum geometry of parametrized quantum circuits, PRX Quan- tum2, 10.1103/prxquantum.2.040309 (2021)

  93. [101]

    Huang, R

    H.-Y. Huang, R. Kueng, and J. Preskill, Predicting many properties of a quantum system from very few measure- ments, Nature Physics16, 1050 (2020)

  94. [102]

    Elben, S

    A. Elben, S. T. Flammia, H.-Y. Huang, R. Kueng, J. Preskill, B. Vermersch, and P. Zoller, The ran- domized measurement toolbox, Nature Review Physics 10.1038/s42254-022-00535-2 (2022)

  95. [103]

    Huang, M

    H.-Y. Huang, M. Broughton, J. Cotler, S. Chen, J. Li, M. Mohseni, H. Neven, R. Babbush, R. Kueng, J. Preskill, and J. R. McClean, Quantum advantage in learning from experiments, Science376, 1182 (2022)

  96. [104]

    Cerezo and P

    M. Cerezo and P. J. Coles, Higher order derivatives of quantum neural networks with barren plateaus, Quantum Science and Technology6, 035006 (2021). 15 Appendix Table of Contents A Hypothesis Testing 15 1 One sample . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....

  97. [105]

    right decision betweenH0andH 1

    One sample Lemma 1.Consider two probability distributionsPandP ′ over some finite setI. Suppose we are given a single sample Sdrawn from eitherPorP ′ with equal probability. We have the following two hypotheses: •Null HypothesisH 0:Sis drawn fromP. •Alternative hypothesisH 1:S...

  98. [106]

    NO i=2 Pi # +P ′ 1 ⊗

    Many samples We proceed to the scenario withNmany samples where all samples are drawn from eitherPorP′. The goal remains similar in the sense that given this set of samples we have to make a decision which of the two distributions the samples are drawn from. To see how Lemma 1...

  99. [107]

    right decision betweenH 0 andH 1

    Statistical indistinguishability We define statistical indistinguishability based on the success probability of a binary hypothesis testing task. Further- more, we introduce the notion of statistical indistinguishability at two levels. Namely, Definition 3 concerns the level o...

  100. [108]

    right decision betweenH0andH 1

    From Eq. (B7), we have Pr α (Ek)⩾1− p β , β∈ O 1 bn .(B8) Now, the probability that allEk occur can be bounded with the union bound as Pr α   |M|\ k=1 Ek   = 1−Pr α   |M|[ k=1 ¯Ek   (B9) ⩾1− |M|X k=1 Pr α ¯Ek (B10) ⩾1− |M| p β ,(B11) where ¯Ek is conjugate event ofE k,...

  101. [109]

    right decision betweenH0andH 1

    Counter example We provide a counterexample of a distribution with exponentially many elements that exhibits exponential concentra- tion according to Definition 1, yet remains distinguishable from its fixed distribution. Consider two probability distributionsPα andP fixed over...

  102. [110]

    right decision betweenH0andH 1

    Indistinguishable example We also present another example where the distribution remains indistinguishable, despite having exponentially many elements. To see this, consider the same setup, but now choosef(α) =1 M 2. In this case, the hypothesis testing problem is between the ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.