REVIEW 3 major objections 6 minor 1 cited by
Pitfalls when tackling the exponential concentration of parameterized quantum models
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Exponentially concentrated quantum measurements produce samples indistinguishable from fixed random noise, so no classical post-processing can recover the lost information.
desk verdict Sound core indistinguishability theorem, but the applied claims outrun the proofs—QNG in particular is currently a conjecture. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is outcome probability concentration (Definition 1): for every POVM element $M_k$, $\Pr_{\alpha\sim\mathcal D}(|p_k(\alpha)-\mu_k|\ge\delta)\le\beta/\delta^2$ with $\beta\in O(\exp(-n))$. The proof couples this to binary hypothesis testing: Lemma 1 gives the one-sample success probability $1/2+\|P-P'\|_1/4$, Lemma 2 bounds the $N$-sample product-distribution distance by the sum of single-sample distances, and Proposition 1 bounds the success probability by $1/2+N\|P_0-P'_0\|_1/4$. Choosing $\delta'=\beta^{1/4}$ and applying a union bound over the polynomially many POVM elements yields $\delta=|\mathcal M|\sqrt\beta$ and $\varepsilon=N|\mathcal M|\beta^{1/4}/4$, both exponentially vanishing. Corollary 1 then transfers indistinguishability to arbitrary post-processing maps by contradiction.
What would settle it
Run the binary hypothesis test behind Theorem 2 on a concrete concentrated model: choose a global-observable ansatz known to have exponentially concentrated outcomes, draw polynomially many shots at many random parameters, and test whether the empirical outcome histogram can be distinguished from the fixed distribution $P_{\mathrm{fixed}}$ with success probability exceeding $1/2+\exp(-n)$. If any polynomial-shot protocol, including one followed by quantum natural gradient or CVaR post-processing, reliably identifies the parameter value or produces a gradient estimate whose sign correlates with the true gradient at a rate that does not vanish exponentially in $n$, then Theorem 1 or Corollary 1 would fail.
Extended reading notes
Core claim
The central discovery is that exponential concentration at the level of measurement outcome probabilities implies statistical indistinguishability from a fixed, variable-independent distribution. Formally, for a POVM with $|\mathcal{M}|\in O(\operatorname{poly}(n))$, if each outcome probability $p_k(\alpha)=\operatorname{Tr}[\rho(\alpha)M_k]$ satisfies Definition 1 with $\beta\in O(\exp(-n))$, then after $N\in O(\operatorname{poly}(n))$ shots the samples $S_N(\alpha)$ are with probability at least $1-\delta$, $\delta\in O(\exp(-n))$, statistically indistinguishable from samples drawn from $P_{\mathrm{fixed}}=(\mu_1,\dots,\mu_{|\mathcal{M}|})$. Corollary 1 extends this to any classical post-processing map, because distinguishable outputs would yield a strategy to distinguish the underlying distributions. A direct corollary is that training a concentrated loss with vanilla gradient descent and polynomial shots is statistically indistinguishable from a random walk. The authors use this framework to argue that quantum natural gradient, sample-based CVaR optimization, classical neural-network-assisted initialization, and rescaled parameter-shift rules do not circumvent exponential concentration with finite measurement budgets, even though they may provide other training benefits.
Load-bearing premise
The negative conclusions about specific methods all rest on the premise that those methods' outcome probabilities actually concentrate exponentially under the parameter distribution they induce; if a method changes that distribution (as warm starting can), the conclusion stops applying, and for quantum natural gradient the paper itself notes that the proof is not supplied.
Editorial extensions
If this is right
- Switching from vanilla gradients to quantum natural gradient, sample-based CVaR losses, classical-neural-network initialization, or rescaled parameter-shift rules does not remove the information bottleneck when the underlying outcome probabilities concentrate; such methods may still aid training in other ways.
- Diagnosing barren plateaus solely through loss-variance scaling is insufficient, because an exponentially large prefactor can suppress the variance without changing the information content; outcome-probability concentration captures the finite-shot information content directly.
- Training a concentrated landscape with vanilla gradient descent and polynomial shot budgets is provably a random walk: with probability exponentially close to one, each parameter update is statistically indistinguishable from a parameter-independent random variable.
- The diagnostic extends beyond variational training to quantum kernel methods and quantum reservoir models, where the relevant variables are classical input data rather than trainable parameters.
- The only ways out are to change the concentration properties themselves (for example, via identity initialization or warm starting) or to use measurements with exponentially many POVM elements, a regime the theorem explicitly does not cover.
Reading between the lines
- The theorem implies that measure-first-estimate-later strategies such as classical shadows do not constitute a loophole: for a given target observable, their effective POVM is mathematically equivalent to directly measuring that observable, so the same indistinguishability bound applies regardless of the measurement protocol's randomization.
- A testable extension is to report, for any proposed mitigation method, the induced distribution over measurement outcomes and to run the binary hypothesis test against the fixed distribution directly, rather than relying only on loss curves or variances.
- The exponential-POVM caveat is the most promising open door in the paper: it suggests that entangled measurements which coarse-grain exponentially many outcomes into globally correlated patterns are the natural place to look for genuine avoidance of exponential concentration.
- If the central claim is right, it sharpens the classical-simulability debate: under exponential concentration and polynomial shots, the quantum device contributes no learning signal at all, regardless of how powerful the classical post-processing is.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a shift in how exponential concentration is diagnosed in parameterized quantum models: instead of analyzing loss- or gradient-variance scaling, it analyzes concentration at the level of POVM outcome probabilities. A general procedure is introduced that covers loss evaluation, gradient-based training, natural gradient, CVaR-style sample processing, and neural-network-assisted initialization. The main theoretical result (Theorem 2, Appendix B) shows that if a POVM with polynomially many elements has exponentially concentrated outcome probabilities, then after polynomially many measurement shots the samples are statistically indistinguishable from samples from a fixed, variable-independent distribution, and no classical post-processing can restore distinguishability (Corollary 3). A corollary (Corollary 4) claims that vanilla gradient descent on a barren-plateau landscape is statistically indistinguishable from a random walk. The framework is then applied to argue that quantum natural gradient, sample-based CVaR optimization, classical neural-network-assisted initialization, and a rescaled parameter-shift rule do not overcome exponential concentration under finite measurement budgets.
Significance. The core hypothesis-testing formulation is valuable and clean: Theorem 2 provides a rigorous, parameter-free bridge between outcome-probability concentration and statistical indistinguishability, and Corollary 3 correctly captures the intuition that any classical post-processing of already-indistinguishable samples cannot recover information. The proof in Appendix B is elementary and sound, and the practical guidelines in Section IV are a useful diagnostic tool. The paper also honestly acknowledges the scope limitation to polynomial-size POVMs and includes a counterexample (Appendix D) showing that exponentially many POVM elements could in principle evade the theorem. If the applied claims were fully proven, the paper would be an important contribution to the barren-plateau literature. As it stands, the central framework is solid but the headline applications are only partially established.
major comments (3)
- [Section III and Appendix C] The claim that quantum natural gradient (QNG) does not overcome exponential concentration is not established. The paper states in Section III that 'the proof would be more complex' and provides no rigorous argument that the QNG update is information-less when the QGT is non-degenerate. The only numerical evidence, Figure 4a, is generated with the circuit in Eq. (C1), a single layer of X-rotations with a global Z observable; Appendix C explicitly computes g^+(θ) = 4I, making the QNG update rule (C2) exactly equivalent to gradient descent with an effective learning rate 4η. Thus the numerics do not test QNG's geometry at all. The informal discussion about QGT elements from deeper circuit layers does not address the possibility that the pseudo-inverse of the QGT could amplify or reweight gradient noise to produce a meaningful direction, nor does it account for the fact that QNG changes the parameter distribution over training while Definition 1 is distribution-dependent. The QNG conclusion should either be proven for a non-degenerate QGT or explicitly downgraded to a conjecture.
- [Section IV, neural-network-assisted initialization] The assertion that agnostic classical neural-network-assisted initialization 'does not address the root cause' and cannot overcome exponential concentration is not derived from the framework. The concentration notion in Definition 1 is defined with respect to a distribution D over the variables. When the trainable variables are the weights and biases of a classical neural network instead of the circuit parameters, the induced distribution over the circuit parameters is not analyzed. The numerical experiment in Figure 4c uses the same single-layer ansatz as the other panels, but no argument is given that the neural network's output distribution retains the exponential concentration properties. The paper itself concedes that warm starting can avoid concentration; the NN-induced distribution could equally change the concentration regime. As written, this conclusion is an assumption rather than a consequence of Theorem 2.
- [Appendix B, proof of Corollary 4] The random-walk proof reuses the initial concentration distribution D at every training step without accounting for the evolution of the parameter distribution. In Eqs. (B20)-(B25), the event A_ijk has probability at least 1 - 2√β only when the parameter point α_ijk is drawn from D, but after the first update the parameters θ^(k) are distributed according to a convolution of the initial distribution with the random-walk increments, which is generally not D. The union bound over training steps therefore does not follow from Theorem 2 unless the concentration property holds for the time-marginal distributions of the parameters, which is not established. This gap affects the generality of Corollary 2, which is presented as a standalone result.
minor comments (6)
- [Appendix B, Theorem 2 proof] The statement 'by denoting δ = |M|√β, we also have δ ∈ O(poly(n))' is a typo; δ = |M|√β is exponentially small, not polynomial.
- [Theorems 1 and 2, Corollaries 1 and 2] The informal statements use δ ∈ O(exp(-n)) and ε ∈ O(exp(-n)), while the formal proof gives δ = |M|√β and ε = N|M|β^{1/4}/4. For β = b^{-n}, these are exponentially decaying but with rates b^{-n/2} and b^{-n/4} up to polynomial factors, not necessarily b^{-n}. Please state the asymptotic rates consistently.
- [Section II and Corollary 4] The symbol Nℓ is used both for the number of physical quantities in the general procedure and, in Corollary 4, for the number of measurement sets 2NpNL. This reuse is confusing; please introduce distinct notation.
- [Appendix A, Lemma 1 proof] The sentence 'guess 1 if the two distributions are identical, this A is simply an empty set and the probability of making the right decision is simply 1/2' is incomplete and should be rewritten.
- [Figure 4] The caption lists n = 9, 11, 13, 15, 17 and three shot regimes, but the line styles and colors in panels a-d make it difficult to distinguish the different settings; please improve the figure legend and consider separate panels or clearer markers.
- [Section IV, Eq. (12)] The ratio ϵ_N(α) is defined as a function of α, but the condition 'we require ϵ_N(α) ≲ 1' is ambiguous; clarify whether this is required pointwise, in expectation, or for typical α.
Circularity Check
No circular derivation: Theorem 1 follows from Definition 1 via standard hypothesis testing bounds, and the applied conclusions depend on external concentration premises rather than on the paper's own definitions.
full rationale
The central derivation chain is self-contained and not circular. Definition 1 states a Chebyshev-type tail bound on individual POVM outcome probabilities; Theorem 2 (formal version of Theorem 1) combines a union bound over POVM elements with the binary-hypothesis-testing bound of Proposition 1 to show that polynomial samples cannot distinguish P_alpha from P_fixed. This is a genuine probabilistic implication, not a restatement of the definition. Corollary 3 then applies an arbitrary post-processing map, and Corollary 4 repeats the argument with a union bound over training steps to obtain the random-walk conclusion; all steps are proved in Appendices A-B. The applications to specific methods (CVaR, neural-network initialization, scaled parameter-shift, and quantum natural gradient) rely on the external premise that the relevant circuits exhibit outcome-probability concentration, as cited from the barren-plateau literature (e.g., Ref. [41]); that is an input assumption, not a conclusion forced by the paper's own equations. There are self-citations (e.g., Refs. [10], [28], [34], [56], [79]), but they support background or auxiliary applications and are not load-bearing for Theorem 1. The quantum-natural-gradient discussion is explicitly informal ('Nonetheless, the proof would be more complex'), and Appendix C shows the numerical QNG experiment degenerates to g^+(theta)=4I, i.e., to plain gradient descent; this is a gap in evidence rather than a circular step. No fitted parameter is renamed as a prediction, and no uniqueness theorem is imported from the authors' prior work.
Assumptions & free parameters
free parameters (2)
- CVaR quantile hyperparameter gamma =
0.8
- RPS scaling factor lambda =
lambda = N d / (2 d^2 + N d - 2), d = 2^n
assumptions (6)
- standard math The success probability of binary hypothesis testing with N samples is at most 1/2 + N||P-P'||_1/4 (Lemma 1 and Proposition 1).
- standard math The union bound over POVM elements and over training steps is valid.
- domain assumption Standard parameterized quantum models can be represented with POVMs having |M| in O(poly(n)) elements ('Polynomial POVMs in disguise').
- domain assumption The analyzed circuits (e.g., global-Pauli observables) have exponentially concentrating outcome probabilities, inherited from known barren-plateau results such as Ref. [41].
- standard math The parameter-shift rule gives the gradient as a difference of two loss estimates (Refs. [7,93]).
- domain assumption Polynomial measurement budget, polynomially many quantities N_l, and polynomially many training steps.
Cite this review
Pith. "Pith review of Pitfalls when tackling the exponential concentration of parameterized quantum models." pith.science (2026). https://pith.science/paper/UCDCVXL5
@misc{pith2026250722054,
author = {Pith},
title = {Pith review of: Pitfalls when tackling the exponential concentration of parameterized quantum models},
year = {2026},
howpublished = {\url{https://pith.science/paper/UCDCVXL5}},
note = {Machine review of arXiv:2507.22054}
}
read the original abstract
Identifying scalable circuit architectures remains a central challenge in variational quantum computing and quantum machine learning. Many approaches have been proposed to mitigate or avoid the barren plateau phenomenon or, more broadly, exponential concentration. However, due to the intricate interplay between quantum measurements and classical post-processing, we argue these techniques often fail to circumvent concentration effects in practice. Here, by analyzing concentration at the level of measurement outcome probabilities and leveraging tools from hypothesis testing, we develop a practical framework for diagnosing whether a parameterized quantum model is inhibited by exponential concentration. Applying this framework, we argue that several widely used methods (including quantum natural gradient, sample-based optimization, and certain neural-network-inspired initializations) do not overcome exponential concentration with finite measurement budgets, though they may still aid training in other ways.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Trainability Beyond Linearity in Variational Quantum Objectives
The trainability boundary for variational quantum objectives is the affine regime; non-affine amplification-capable losses can mitigate barren plateaus when using coarse-grained statistics at polynomial widths.
Reference graph
Works this paper leans on
-
[1]
Given a procedureP, identify the quantities{ℓi(αi)} which require information to be extracted from quan- tum computers
-
[2]
For each quantity{ℓi(αi)}, identify the correspond- ingM (i) and check whether|M (i)|, the number of POVM elements inM (i), scales at most polynomi- ally with system sizen. Note that all barren plateau mitigation strategies we are aware of involve such polynomial-sized POVMs, even if some may initially appear to have exponen- tially many elements (see Sec...
-
[3]
Determine whether the outcome probabilities p(i) k (αi) = Tr h ρi(αi)M (i) k i exponentially concentrate with respect toαi (as per Definition 1). If this is the case, the procedurePsuffers from the concentration in the sense that the measurement out- comes, with probability exponentially close to 1, con- tain no information about the variablesαi. Based on...
-
[4]
Cerezo, A
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, Variational quantum algo- rithms, Nature Reviews Physics3, 625–644 (2021)
2021
-
[5]
Bharti, A
K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J. S. Kottmann, T. Menke,et al., Noisy intermediate- scale quantum algorithms, Reviews of Modern Physics94, 015004 (2022)
2022
- [6]
-
[7]
McArdle, S
S. McArdle, S. Endo, A. Aspuru-Guzik, S. C. Benjamin, andX.Yuan,Quantumcomputationalchemistry,Reviews of Modern Physics92, 015003 (2020)
2020
-
[8]
J. R. McClean, J. Romero, R. Babbush, and A. Aspuru- Guzik, The theory of variational hybrid quantum-classical algorithms, New Journal of Physics18, 023023 (2016)
2016
Show all 110 references
-
[9]
Pérez-Salinas, A
A. Pérez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, Data re-uploading for a universal quantum classifier, Quantum4, 226 (2020)
2020
-
[10]
Mitarai, M
K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, Quantum circuit learning, Physical Review A98, 032309 (2018)
2018
-
[11]
Schuld, Supervised quantum machine learning mod- els are kernel methods, arXiv preprint arXiv:2101.11020 (2021)
M. Schuld, Supervised quantum machine learning mod- els are kernel methods, arXiv preprint arXiv:2101.11020 (2021)
2021 arXiv
-
[12]
Havlíček, A
V. Havlíček, A. D. Córcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, Supervised learning with quantum-enhanced feature spaces, Nature 567, 209 (2019)
2019
-
[13]
Thanasilp, S
S. Thanasilp, S. Wang, M. Cerezo, and Z. Holmes, Expo- nential concentration in quantum kernel methods, Nature Communications15, 5200 (2024)
2024
-
[14]
Kübler, S
J. Kübler, S. Buchholz, and B. Schölkopf, The inductive bias of quantum kernels, Advances in Neural Information Processing Systems34, 12661 (2021)
2021
-
[15]
Huang, M
H.-Y. Huang, M. Broughton, M. Mohseni, R. Babbush, S. Boixo, H. Neven, and J. R. McClean, Power of data in quantum machine learning, Nature Communications12, 1 (2021)
2021
-
[16]
Gentinetta, A
G. Gentinetta, A. Thomsen, D. Sutter, and S. Woerner, The complexity of quantum support vector machines, arXiv preprint arXiv:2203.00031 (2022)
2022 arXiv
- [17]
-
[18]
Zimborás, B
Z. Zimborás, B. Koczor, Z. Holmes, E.-M. Borrelli, A. Gilyén, H.-Y. Huang, Z. Cai, A. Acín, L. Aolita, L. Banchi,et al., Myths around quantum computation before full fault tolerance: What no-go theorems rule out and what they don’t, arXiv preprint arXiv:2501.05694 https://doi....
-
[19]
E. R. Anschuetz and B. T. Kiani, Quantum variational algorithms are swamped with traps, Nature Communica- tions13, 7760 (2022)
2022
-
[20]
Larocca, N
M. Larocca, N. Ju, D. García-Martín, P. J. Coles, and M. Cerezo, Theory of overparametrization in quantum neural networks, Nature Computational Science3, 542 (2023)
2023
-
[21]
F. J. Schreiber, J. Eisert, and J. J. Meyer, Classical surro- gates for quantum learning models, Physical Review Let- ters131, 100803 (2023)
2023
-
[22]
Sweke, E
R. Sweke, E. Recio, S. Jerbi, E. Gil-Fuster, B. Fuller, J.Eisert,andJ.J.Meyer,Potentialandlimitationsofran- dom fourier features for dequantizing quantum machine learning, Quantum9, 1640 (2025)
2025
-
[23]
Sahebi, A
M. Sahebi, A. Barthe, Y. Suzuki, Z. Holmes, and M. Grossi, On dequantization of supervised quantum ma- chine learning via random fourier features, arXiv preprint arXiv:2505.15902 (2025)
2025
-
[24]
M. S. Rudolph, T. Jones, Y. Teng, A. Angrisani, and Z. Holmes, Pauli propagation: A computational frame- 12 work for simulating quantum systems, arXiv preprint arXiv:2505.21606 (2025)
2025 arXiv
-
[25]
Angrisani, A
A. Angrisani, A. Schmidhuber, M. S. Rudolph, M. Cerezo, Z. Holmes, and H.-Y. Huang, Classically estimating ob- servables of noiseless quantum circuits, arXiv preprint arXiv:2409.01706 (2024)
2024
-
[26]
S. Shin, Y. S. Teo, and H. Jeong, Dequantizing quan- tummachinelearningmodelsusingtensornetworks,Phys. Rev. Res.6, 023218 (2024)
2024
-
[27]
M. L. Goh, M. Larocca, L. Cincio, M. Cerezo, and F. Sauvage, Lie-algebraic classical simulations for quan- tum computing, arXiv preprint arXiv:2308.01432 (2023)
2023
-
[28]
Kerenidis, J
I. Kerenidis, J. Landman, and N. Mathur, Classical and quantum algorithms for orthogonal neural networks, arXiv preprint arXiv:2106.07198 (2021)
2021 arXiv
- [29]
-
[30]
Fontana, M
E. Fontana, M. S. Rudolph, R. Duncan, I. Rungger, and C. Cîrstoiu, Classical simulations of noisy variational quantum circuits, npj Quantum Information11, 1 (2025)
2025
-
[31]
Larocca, S
M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Bia- monte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo, A review of barren plateaus in variational quantum computing, Nature Reviews Physics3, 625–644 (2025)
2025
-
[32]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, andH.Neven,Barrenplateausinquantumneuralnetwork training landscapes, Nature Communications9, 1 (2018)
2018
-
[33]
Fontana, D
E. Fontana, D. Herman, S. Chakrabarti, N. Kumar, R. Yalovetzky, J. Heredge, S. H. Sureshbabu, and M. Pis- toia, Characterizing barren plateaus in quantum ansätze with the adjoint representation, Nature Communications 15, 7171 (2024)
2024
-
[34]
Ragone, B
M. Ragone, B. N. Bakalov, F. Sauvage, A. F. Kemper, C. Ortiz Marrero, M. Larocca, and M. Cerezo, A lie al- gebraic theory of barren plateaus for deep parameter- ized quantum circuits, Nature Communications15, 7172 (2024)
2024
-
[35]
Cerezo, M
M. Cerezo, M. Larocca, D. García-Martín, N. L. Diaz, P. Braccia, E. Fontana, M. S. Rudolph, P. Bermejo, A. Ijaz, S. Thanasilp,et al., Does provable absence of bar- renplateausimplyclassicalsimulability? Or, whyweneed to rethink variational quantum computing, arXiv preprint arX...
2023 arXiv
-
[36]
Suzuki and M
Y. Suzuki and M. Li, Effect of alternating layered ansatzes on trainability of projected quantum kernel, arXiv preprint arXiv:2310.00361 (2023)
2023 arXiv
-
[37]
Xiong, G
W. Xiong, G. Facelli, M. Sahebi, O. Agnel, T. Chotibut, S. Thanasilp, and Z. Holmes, On fundamental aspects of quantum extreme learning machines, arXiv preprint arXiv:2312.15124 (2023)
2023 arXiv
-
[38]
Xiong, Z
W. Xiong, Z. Holmes, A. Angrisani, Y. Suzuki, T. Chotibut, and S. Thanasilp, Role of scrambling and noise in temporal information processing with quantum systems, arXiv preprint arXiv:2505.10080 (2025)
2025 arXiv
-
[39]
Shaydulin and S
R. Shaydulin and S. M. Wild, Importance of kernel band- width in quantum machine learning, Physical Review A 106, 042407 (2022)
2022
-
[40]
Arrasmith, Z
A. Arrasmith, Z. Holmes, M. Cerezo, and P. J. Coles, Equivalence of quantum barren plateaus to cost concen- tration and narrow gorges, Quantum Science and Tech- nology7, 045015 (2022)
2022
-
[41]
Arrasmith, M
A. Arrasmith, M. Cerezo, P. Czarnik, L. Cincio, and P. J. Coles, Effect of barren plateaus on gradient-free optimiza- tion, Quantum5, 558 (2021)
2021
-
[42]
Larocca, F
M. Larocca, F. Sauvage, F. M. Sbahi, G. Verdon, P. J. Coles, and M. Cerezo, Group-invariant quantum machine learning, PRX Quantum3, 030341 (2022)
2022
-
[43]
Schatzki, M
L. Schatzki, M. Larocca, Q. T. Nguyen, F. Sauvage, and M. Cerezo, Theoretical guarantees for permutation- equivariant quantum neural networks, npj Quantum In- formation10, 12 (2024)
2024
-
[44]
Cerezo, A
M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shallow parametrized quantum circuits, Nature Communications 12, 1 (2021)
2021
-
[45]
Larocca, P
M. Larocca, P. Czarnik, K. Sharma, G. Muraleedharan, P. J. Coles, and M. Cerezo, Diagnosing Barren Plateaus with Tools from Quantum Optimal Control, Quantum6, 824 (2022)
2022
-
[46]
Letcher, S
A. Letcher, S. Woerner, and C. Zoufal, Tight and effi- cient gradient bounds for parameterized quantum circuits, Quantum8, 1484 (2024)
2024
-
[47]
Basheer, Y
A. Basheer, Y. Feng, C. Ferrie, and S. Li, Alternating layered variational quantum circuits can be classically op- timized efficiently using classical shadows, arXiv preprint arXiv:2208.11623 (2022)
2022 arXiv
-
[48]
Napp, Quantifying the barren plateau phenomenon for a model of unstructured variational ansätze, arXiv preprint arXiv:2203.06174 (2022)
J. Napp, Quantifying the barren plateau phenomenon for a model of unstructured variational ansätze, arXiv preprint arXiv:2203.06174 (2022)
2022 arXiv
-
[49]
Zhang, S
H.-K. Zhang, S. Liu, and S.-X. Zhang, Absence of barren plateaus in finite local-depth circuits with long-range en- tanglement, Physical Review Letters132, 150603 (2024)
2024
-
[50]
Pesah, M
A. Pesah, M. Cerezo, S. Wang, T. Volkoff, A. T. Sorn- borger, and P. J. Coles, Absence of barren plateaus in quantum convolutional neural networks, Physical Review X11, 041011 (2021)
2021
-
[51]
Monbroussou, J
L. Monbroussou, J. Landman, A. B. Grilo, R. Kukla, and E. Kashefi, Trainability and expressivity of hamming- weight preserving quantum circuits for machine learning, arXiv preprint arXiv:2309.15547 (2023)
2023 arXiv
-
[52]
S. Raj, I. Kerenidis, A. Shekhar, B. Wood, J. Dee, S. Chakrabarti, R. Chen, D. Herman, S. Hu, P. Minssen, et al., Quantum deep hedging, Quantum7, 1191 (2023)
2023
-
[53]
N. L. Diaz, D. García-Martín, S. Kazi, M. Larocca, and M. Cerezo, Showcasing a barren plateau the- ory beyond the dynamical lie algebra, arXiv preprint arXiv:2310.11505 (2023)
2023 arXiv
-
[54]
A. A. Mele, A. Angrisani, S. Ghosh, S. Khatri, J. Eis- ert, D. S. França, and Y. Quek, Noise-induced shallow circuits and absence of barren plateaus, arXiv preprint arXiv:2403.13927 (2024). 13
2024 arXiv
- [55]
-
[56]
Srimahajariyapong, S
K. Srimahajariyapong, S. Thanasilp, and T. Chotibut, Connecting phases of matter to the flatness of the loss landscape in analog variational quantum algorithms, arXiv preprint arXiv:2506.13865 (2025)
2025
- [57]
-
[58]
A. A. Mele, G. B. Mbeng, G. E. Santoro, M. Collura, and P. Torta, Avoiding barren plateaus via transferability of smooth solutions in a Hamiltonian variational ansatz, Physical Review A106, L060401 (2022)
2022
-
[59]
R. Puig, M. Drudis, S. Thanasilp, and Z. Holmes, Varia- tional quantum simulation: A case study for understand- ing warm starts, PRX Quantum6, 010317 (2025)
2025
-
[60]
S. Y. Chang, S. Thanasilp, B. L. Saux, S. Val- lecorsa, and M. Grossi, Latent style-based quantum gan for high-quality image generation, arXiv preprint arXiv:2406.02668 (2024)
2024 arXiv
-
[61]
Y. Wang, B. Qi, C. Ferrie, and D. Dong, Trainabil- ity enhancement of parameterized quantum circuits via reduced-domain parameter initialization, Physical Review Applied22, 054005 (2024)
2024
-
[62]
Park and N
C.-Y. Park and N. Killoran, Hamiltonian variational ansatz without barren plateaus, Quantum8, 1239 (2024)
2024
-
[63]
C.-Y. Park, M. Kang, and J. Huh, Hardware-efficient ansatz without barren plateaus in any depth, arXiv preprint arXiv:2403.04844 (2024)
2024 arXiv
-
[64]
Zhang, L
K. Zhang, L. Liu, M.-H. Hsieh, and D. Tao, Escaping from the barren plateau via Gaussian initializations in deep variational quantum circuits, inAdvances in Neu- ral Information Processing Systems(2022)
2022
-
[65]
Tangpanitanon, S
J. Tangpanitanon, S. Thanasilp, N. Dangniam, M.-A. Lemonde, and D. G. Angelakis, Expressibility and train- ability of parametrized analog quantum systems for ma- chine learning applications, Physical Review Research2, 043364 (2020)
2020
-
[66]
Shi and Y
X. Shi and Y. Shang, Avoiding barren plateaus via Gaussian mixture model, arXiv preprint arXiv:2402.13501 (2024)
2024 arXiv
-
[67]
C. Cao, Y. Zhou, S. Tannu, N. Shannon, and R. Joynt, Exploiting many-body localization for scalable variational quantum simulation, arXiv preprint arXiv:2404.17560 (2024)
2024
-
[68]
Grant, L
E. Grant, L. Wossnig, M. Ostaszewski, and M. Benedetti, Aninitializationstrategyforaddressingbarrenplateausin parametrized quantum circuits, Quantum3, 214 (2019)
2019
-
[69]
L.FriedrichandJ.Maziero,Avoidingbarrenplateauswith classical deep neural networks, Physical Review A106, 042433 (2022)
2022
-
[70]
Miao, C.-Y
J. Miao, C.-Y. Hsieh, and S.-X. Zhang, Neural-network- encoded variational quantum algorithms, Physical Review Applied21, 014053 (2024)
2024
-
[71]
A. Rad, A. Seif, and N. M. Linke, Surviving the bar- ren plateau in variational quantum circuits with Bayesian learning initialization, arXiv preprint arXiv:2203.02464 (2022)
2022 arXiv
-
[72]
Faílde, J
D. Faílde, J. D. Viqueira, M. M. Juane, and A. Gómez, Using differential evolution to avoid local minima in variational quantum algorithms, Scientific Reports13, 10.1038/s41598-023-43404-3 (2023)
2023 doi
-
[73]
Kashif and S
M. Kashif and S. Al-kuwari, Resqnets: A residual ap- proach for mitigating barren plateaus in quantum neural networks (2023), arXiv:2305.03527 [quant-ph]
2023 arXiv
-
[74]
P. K. Barkoutsos, G. Nannicini, A. Robert, I. Tavernelli, and S. Woerner, Improving variational quantum optimiza- tion using cvar, Quantum4, 256 (2020)
2020
-
[75]
Stokes, J
J. Stokes, J. Izaac, N. Killoran, and G. Carleo, Quantum natural gradient, Quantum4, 269 (2020)
2020
-
[76]
C. O. Marrero, M. Kieferová, and N. Wiebe, Entanglement-induced barren plateaus, PRX Quan- tum2, 040316 (2021)
2021
-
[77]
Sharma, M
K. Sharma, M. Cerezo, L. Cincio, and P. J. Coles, Train- ability of dissipative perceptron-based quantum neural networks, Physical Review Letters128, 180505 (2022)
2022
-
[78]
T. L. Patti, K. Najafi, X. Gao, and S. F. Yelin, Entangle- ment devised barren plateau mitigation, Physical Review Research3, 033090 (2021)
2021
-
[79]
S. Wang, E. Fontana, M. Cerezo, K. Sharma, A. Sone, L. Cincio, and P. J. Coles, Noise-induced barren plateaus in variational quantum algorithms, Nature Communica- tions12, 1 (2021)
2021
-
[80]
Holmes, K
Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, Con- necting ansatz expressibility to gradient magnitudes and barren plateaus, PRX Quantum3, 010313 (2022)
2022
-
[81]
Khatri, R
S. Khatri, R. LaRose, A. Poremba, L. Cincio, A. T. Sorn- borger, and P. J. Coles, Quantum-assisted quantum com- piling, Quantum3, 140 (2019)
2019
-
[82]
M. S. Rudolph, S. Lerch, S. Thanasilp, O. Kiss, O. Shaya, S. Vallecorsa, M. Grossi, and Z. Holmes, Trainability bar- riers and opportunities in quantum generative modeling, npj Quantum Information10, 116 (2024)
2024
-
[83]
Kieferova, O
M. Kieferova, O. M. Carlos, and N. Wiebe, Quantum gen- erative training using rényi divergences, arXiv preprint arXiv:2106.09567 (2021)
2021 arXiv
-
[84]
Thanaslip, S
S. Thanaslip, S. Wang, N. A. Nghiem, P. J. Coles, and M. Cerezo, Subtleties in the trainability of quantum ma- chine learning models, Quantum Machine Intelligence5, 21 (2023)
2023
-
[85]
Holmes, A
Z. Holmes, A. Arrasmith, B. Yan, P. J. Coles, A. Albrecht, and A. T. Sornborger, Barren plateaus preclude learning scramblers, Physical Review Letters126, 190501 (2021)
2021
-
[86]
E. C. Martín, K. Plekhanov, and M. Lubasch, Barren plateaus in quantum tensor network optimization, Quan- tum7, 974 (2023)
2023
-
[87]
E. R. Anschuetz, A unified theory of quantum neural network loss landscapes, arXiv preprint arXiv:2408.11901 14 (2024)
2024 arXiv
-
[88]
Crognaletti, M
G. Crognaletti, M. Grossi, and A. Bassi, Estimates of loss functionconcentrationinnoisyparametrizedquantumcir- cuits, arXiv preprint arXiv:2410.01893 (2024)
2024 arXiv
-
[89]
R. Mao, G. Tian, and X. Sun, Towards determining the presence of barren plateaus in some chemically inspired variational quantum algorithms, Communications Physics 7, 342 (2024)
2024
-
[90]
Y. S. Teo, Optimized numerical gradient and hessian esti- mation for variational quantum algorithms, Physical Re- view A107, 10.1103/physreva.107.042421 (2023)
2023 doi
-
[91]
Fujii and K
K. Fujii and K. Nakajima, Harnessing disordered- ensemble quantum dynamics for machine learning, Phys- ical Review Applied8, 024030 (2017)
2017
-
[92]
Nakajima, K
K. Nakajima, K. Fujii, M. Negoro, K. Mitarai, and M. Kitagawa, Boosting computational power through spa- tial multiplexing in quantum reservoir computing, Physi- cal Review Applied11, 034021 (2019)
2019
-
[93]
Mujal, R
P. Mujal, R. Martí nez-Peña, J. Nokkala, J. García-Beni, G. L. Giorgi, M. C. Soriano, and R. Zambrini, Opportuni- ties in quantum reservoir computing and extreme learning machines, Advanced Quantum Technologies4, 2100027 (2021)
2021
-
[94]
Ghosh, A
S. Ghosh, A. Opala, M. Matuszewski, T. Paterek, and T. C. Liew, Reconstructing quantum states with quantum reservoir networks, IEEE Transactions on Neural Net- works and Learning Systems32, 3148 (2020)
2020
-
[95]
F. Hu, G. Angelatos, S. A. Khan, M. Vives, E. Türeci, L. Bello, G. E. Rowlands, G. J. Ribeill, and H. E. Tureci, Tackling sampling noise in physical systems for machine learning applications: Fundamental limits and eigentasks, Physical Review X13, 041020 (2023)
2023
-
[96]
Schuld, V
M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Kil- loran, Evaluating analytic gradients on quantum hard- ware, Physical Review A99, 032331 (2019)
2019
-
[97]
Singh, S
H. Singh, S. Majumder, and S. Mishra, Benchmarking of different optimizers in the variational quantum algorithms for applications in quantum chemistry, The Journal of Chemical Physics159(2023)
2023
-
[98]
M. S. Rudolph, S. Sim, A. Raza, M. Stechly, J. R. Mc- Clean, E. R. Anschuetz, L. Serrano, and A. Perdomo- Ortiz, Orqviz: visualizing high-dimensional landscapes in variational quantum algorithms, arXiv preprint arXiv:2111.04695 (2021)
2021 arXiv
-
[99]
Wierichs, C
D. Wierichs, C. Gogolin, and M. Kastoryano, Avoiding local minima in variational quantum eigensolvers with the natural gradient optimizer, Physical Review Research2, 043246 (2020)
2020
-
[100]
T. Haug, K. Bharti, and M. Kim, Capacity and quantum geometry of parametrized quantum circuits, PRX Quan- tum2, 10.1103/prxquantum.2.040309 (2021)
2021 doi
-
[101]
Huang, R
H.-Y. Huang, R. Kueng, and J. Preskill, Predicting many properties of a quantum system from very few measure- ments, Nature Physics16, 1050 (2020)
2020
-
[102]
Elben, S
A. Elben, S. T. Flammia, H.-Y. Huang, R. Kueng, J. Preskill, B. Vermersch, and P. Zoller, The ran- domized measurement toolbox, Nature Review Physics 10.1038/s42254-022-00535-2 (2022)
2022 doi
-
[103]
Huang, M
H.-Y. Huang, M. Broughton, J. Cotler, S. Chen, J. Li, M. Mohseni, H. Neven, R. Babbush, R. Kueng, J. Preskill, and J. R. McClean, Quantum advantage in learning from experiments, Science376, 1182 (2022)
2022
-
[104]
Cerezo and P
M. Cerezo and P. J. Coles, Higher order derivatives of quantum neural networks with barren plateaus, Quantum Science and Technology6, 035006 (2021). 15 Appendix Table of Contents A Hypothesis Testing 15 1 One sample . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
2021
-
[105]
right decision betweenH0andH 1
One sample Lemma 1.Consider two probability distributionsPandP ′ over some finite setI. Suppose we are given a single sample Sdrawn from eitherPorP ′ with equal probability. We have the following two hypotheses: •Null HypothesisH 0:Sis drawn fromP. •Alternative hypothesisH 1:S...
-
[106]
NO i=2 Pi # +P ′ 1 ⊗
Many samples We proceed to the scenario withNmany samples where all samples are drawn from eitherPorP′. The goal remains similar in the sense that given this set of samples we have to make a decision which of the two distributions the samples are drawn from. To see how Lemma 1...
-
[107]
right decision betweenH 0 andH 1
Statistical indistinguishability We define statistical indistinguishability based on the success probability of a binary hypothesis testing task. Further- more, we introduce the notion of statistical indistinguishability at two levels. Namely, Definition 3 concerns the level o...
-
[108]
right decision betweenH0andH 1
From Eq. (B7), we have Pr α (Ek)⩾1− p β , β∈ O 1 bn .(B8) Now, the probability that allEk occur can be bounded with the union bound as Pr α |M|\ k=1 Ek = 1−Pr α |M|[ k=1 ¯Ek (B9) ⩾1− |M|X k=1 Pr α ¯Ek (B10) ⩾1− |M| p β ,(B11) where ¯Ek is conjugate event ofE k,...
-
[109]
right decision betweenH0andH 1
Counter example We provide a counterexample of a distribution with exponentially many elements that exhibits exponential concentra- tion according to Definition 1, yet remains distinguishable from its fixed distribution. Consider two probability distributionsPα andP fixed over...
-
[110]
right decision betweenH0andH 1
Indistinguishable example We also present another example where the distribution remains indistinguishable, despite having exponentially many elements. To see this, consider the same setup, but now choosef(α) =1 M 2. In this case, the hypothesis testing problem is between the ...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.