{"id":"57bb31ad-9ed0-441d-ac10-b1d8418b1cfa","arxiv_id":"2608.09810","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Under barren plateau gradient decay, SPSA-optimized VQEs need exponentially more iterations and measurements to reach a fixed relative gradient-energy accuracy.","lead":"Variational quantum eigensolvers use tunable quantum circuits, but 'barren plateaus' flatten their energy landscapes as qubits grow. This paper proves that the SPSA optimizer then needs exponentially more iterations and measurement shots to reach a fixed relative accuracy.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exponential-resource claim rests on an imported BP scaling (EΘ=O(N2^{-n}) with zero mean) that is unverified for the RealAmplitudes ansatz and on treating a sufficient bound as a requirement.","rationale":"The reader's weakest assumption is the right one: the paper's own theorems are valid sufficient bounds, but the exponential consequence is carried by an imported scaling. I agree with the CONDITIONAL verdict because the core SPSA analysis (Theorems 1-2) is mathematically coherent and the flaw is in the interpretation and domain of application, not in the algebra. I would add that Corollary 1's validity condition (12) also limits the fixed-SNR statement: for N growing with n, the admissible SNR target shrinks polynomially, so the 'fixed SNR' reading of M=Ω(2^{3n/2}) needs qualification. This reinforces CONDITIONAL rather than REJECT. A Monte Carlo check of the actual gradient moments for the demonstrated ansatz is the cleanest way to decide whether the exponential resource conclusion applies to the simulations shown.","tokens_in":20938,"tokens_out":18149,"duration_ms":169135,"concrete_test":"Classically simulate the RealAmplitudes ansatz of Figure 1 (n=4,6,8,10; N fixed and N=n) with parameters drawn from the same initialization distribution, and Monte Carlo estimate EΘ[∂ℓ f] and EΘ[(∂ℓ f)^2] for each parameter ℓ. Check whether EΘ[(∂ℓ f)^2] decays as Θ(2^{-n}) and whether the mean is negligible relative to the root-second-moment. If either fails, the exponential iteration and measurement-budget conclusion for this ansatz is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central exponential-resource claim is an application, not a theorem: Theorem 2 gives only a sufficient convergence bound, and the exponential growth in (17)-(18) is obtained by substituting the external BP scalings EΘ[(∂ℓf)^2]=O(2^{-n}) and EΘ[||∇f||²]=O(N2^{-n}) with the extra assumption that the mean gradient vanishes (Section 2, after Corollary 1). This is the load-bearing step. The paper does not derive or verify this scaling for the RealAmplitudes ansatz used in the simulations, nor does the SPSA trajectory stay in the parameter distribution over which EΘ is taken. If the mean gradient is not negligible or the second moment decays more slowly, the iteration and budget exponents in (17)-(18) need not be exponential. Even granting the scaling, the wording 'required' overstates what a sufficient-condition proof can establish.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies SPSA optimization for variational quantum eigensolvers under finite-shot measurement noise. It proves a Hessian Lipschitz bound for VQE cost landscapes (Proposition 1), a measurement-noise model with bias/variance bounds for the SPSA gradient estimator (Lemma 1 and Theorem 1), an SNR-based measurement-budget formula (Corollary 1), and a convergence guarantee with explicit schedule μ_t=μ0 T^{-1/2}, c_t=c0 T^{-1/8}, M_t=M0 T^{1/4} (Theorem 2). Using the barren-plateau scaling EΘ[(∂ℓ f)^2]=O(2^{-n}) and EΘ[||∇f||^2]=O(N 2^{-n}), the paper claims exponential growth in the required iteration count and total measurement budget.","tokens_in":21125,"tokens_out":10568,"duration_ms":81746,"significance":"The paper's non-asymptotic characterizations are a useful contribution: the bias/variance decomposition is explicit, the convergence rate O(T^{-1/2}) with a total measurement budget scaling is clearly derived, and the numerical experiments illustrate a contrast between SPSA and a constrained SPSA-IHT variant. If the sufficient bounds are treated as sufficient conditions, the results provide a rigorous framework for quantifying how landscape flatness affects SPSA resource counts. The central exponential-resource claim, however, is conditional on imported BP scalings and on interpreting sufficient bounds as requirements; the paper does not provide lower bounds or verify the scaling for its own ansatz.","major_comments":[{"comment":"The paper repeatedly converts a sufficient condition into a 'required' cost. Theorem 2 proves only that if T=Ω(κ^2/(ε2^2(EΘ[||∇f||^2])^2)), then min_t E[||∇f(θ_t)||^2] ≤ O(ε2 EΘ[||∇f||^2]); it does not show that smaller T fails to reach the target. Therefore the Abstract's statement that BP causes an exponential increase in 'the number of iterations required' is not a logical consequence of the theorem. The claims in the Conclusion and the paragraph after Theorem 2 should be reworded to 'sufficient iteration count' and 'sufficient measurement budget,' or supplemented with a lower bound.","section":"Section 2, Eq. (17)-(18), and Abstract"},{"comment":"The exponential-resource conclusion rests on the imported scalings EΘ[(∂ℓ f)^2]=O(2^{-n}) and EΘ[||∇f||^2]=O(N 2^{-n}), which require the mean gradient to vanish under the parameter distribution. These conditions are not derived for the RealAmplitudes ansatz used in the simulations, nor are they numerically verified for that ansatz. As the headline result is the exponential growth in Eq. (17)-(18), the paper should either prove or cite a result that the specific ansatz and parameter distribution satisfy these scalings, or state the conclusion as conditional on that scaling.","section":"Section 2, after Corollary 1; Section 3"},{"comment":"Corollary 1's measurement-budget conclusion M=Ω(ε1||w||2^2/(EΘ[(∂ℓ f)^2])^{3/2}) is only valid when ε1 satisfies condition (12). For large N, the right-hand side of (12) tends to 0 polynomially in N (e.g., O(N^{-6}) for ||w||1=O(1)), so a fixed target signal-to-noise ratio ε1 becomes inadmissible as N grows. The subsequent claim that BP forces M=Ω(2^{3n/2}) for a fixed SNR therefore does not follow for large N unless ε1 is allowed to shrink with N. The validity regime of Eq. (13) must be stated explicitly.","section":"Section 2, Corollary 1, Eq. (12)-(13)"},{"comment":"The convergence target in Theorem 2 is a fraction ε2 of EΘ[||∇f||^2], which is itself exponentially small under BP. Because Eq. (16) only upper-bounds the minimum gradient energy over the trajectory, the paper does not establish that the SPSA trajectory initially lies above this target; if the initial gradient energy is already of order EΘ[||∇f||^2], the sufficient T from Eq. (17) is not indicative of the actual number of iterations needed. An exponential iteration-complexity claim requires a corresponding lower bound or a trajectory-dependent target.","section":"Section 2, Theorem 2 and Eq. (17)"}],"minor_comments":[{"comment":"Equation (17) contains an ill-formed expression 'T=Ω(κ^2, (ε2)^2(... )^2)'; this should be T=Ω(κ^2/((ε2)^2(EΘ[||∇f||^2])^2)).","section":"Section 2, Eq. (17)"},{"comment":"The identity Σ_{t=0}^{T-1} c_t^4 μ_t = c0 μ0 is incorrect; the right-hand side should be c0^4 μ0. Consequently, the constant κ in Theorem 2 and Eq. (87) should contain c0^4 rather than c0. The O(T^{-1/2}) rate is unaffected.","section":"Appendix E, proof of Theorem 2"},{"comment":"The schedules in (14) depend on the total horizon T, so T must be known in advance to implement the algorithm. This should be stated explicitly.","section":"Section 2, Eq. (14)"},{"comment":"The simulations use a fixed shot budget M=1000 per Pauli term per evaluation, whereas Theorem 2 analyzes M_t=M0 T^{1/4} increasing with the horizon; the relationship between the simulated setting and the theorem should be clarified.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a careful set of sufficient bounds and the proofs in the appendices appear internally consistent. The main weakness is interpretive: the headline exponential-resource claim goes beyond what the sufficient-condition framework can establish. With appropriate rewording and clarification of the validity regimes, the manuscript could be publishable. The citation practice is standard and the topic fits the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zhen, quick take on 2608.09810. This is a solid theory paper. The genuinely new content is in the non-asymptotic SPSA estimator bounds (Theorem 1) and the SNR-based measurement budget (Corollary 1). The bias/variance decomposition is rigorous and the algebra in Appendices A–E checks out; the Hessian Lipschitz constant for VQE and the measurement-noise lemma are useful on their own. The 3/2 exponent in the measurement budget is a real insight: because the perturbation radius must shrink as the gradient energy decays, shot noise gets amplified, and that's not just a restatement of standard BP lore.\n\nWhere the paper gets wobbly is the interpretation. Theorem 2 and the subsequent budget bound are sufficient conditions, not necessary ones. The paper repeatedly says 'required' when all it has shown is that this particular schedule with these constants guarantees the target. That's a real overstatement. Corollary 1 also has a validity condition on ε1 that depends on N and γ; for large N the admissible SNR target shrinks, so the M = Ω(2^{3n/2}) statement is not uniform in the way the abstract implies. The exponential iteration/budget conclusion itself follows by substituting the external BP scaling EΘ[‖∇f‖^2] = O(N2^{-n}) with a zero-mean assumption—imported from the BP literature, not derived or verified for the RealAmplitudes ansatz used in the simulations. The numerical section is illustrative only; no error bars, no comparison to the predicted exponents. I don't see circularity or fitted parameters—the scaling is genuinely substituted from outside—but the trajectory of SPSA need not stay in the parameter distribution over which EΘ is taken.\n\nIf the authors revise to say 'a sufficient iteration bound under this schedule is...' and spell out the validity regime, the paper is a good contribution. As it stands, the headline is the weakest part.\n\nWho should read it: people working on VQE trainability and SPSA in particular. It deserves peer review; a good referee will force the language fix and maybe a more careful numerical section. I'd cite it in follow-up work on SPSA, yes.","headline":"A clean sufficient-condition analysis of SPSA under barren plateaus, with honest bounds, but the exponential-resource headline overstates what a sufficient bound can prove.","tokens_in":21655,"tokens_out":1995,"would_cite":true,"duration_ms":18272,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","68Q12"],"pacs":["03.67.-a"],"model":"deepseek-v4-flash","headline":"The paper proves that SPSA, under barren-plateau scaling, requires exponentially many iterations and measurements to reach fixed relative gradient energy.","keywords":["variational quantum eigensolver","SPSA","barren plateau","finite-shot measurement","gradient estimation","signal-to-noise ratio","measurement budget","convergence guarantee"],"falsifier":"Measure, for the RealAmplitudes ansatz at $n=10,20,30$, the empirical gradient energy $E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]$ under the distribution of SPSA iterates rather than under the initialization ensemble; if it decays more slowly than $2^{-n}$, or if the mean gradient is not negligible compared with the second moment, the $\\Omega(2^{3n/2})$ and $\\Omega(2^{5n/2})$ resource conclusions of Corollary 1 and Theorem 2 are contradicted.","tokens_in":20713,"feed_emoji":"⚛️","tokens_out":10014,"duration_ms":78799,"temperature":0.7,"pith_summary":"The paper aims to close the gap between the geometry of barren plateaus and the actual cost of running a variational quantum eigensolver. Its central contention is that exponentially vanishing gradients translate, through finite-shot measurement noise, into exponentially large iteration counts and measurement budgets for SPSA, the standard two-evaluation stochastic optimizer used on quantum hardware. The author derives non-asymptotic bias and variance bounds for the SPSA gradient estimator, introduces a signal-to-noise ratio for gradient reliability, and proves a convergence rate of $O(T^{-1/2})$ for the minimum expected gradient energy. Substituting the conventional barren-plateau scaling $E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]=O(N2^{-n})$ converts those bounds into exponential growth in the number of qubits, making high-accuracy VQE optimization impractical at large $n$.","feed_headline":"SPSA under barren plateaus needs exponentially many shots","feed_subtitle":"A finite-shot analysis shows iteration and measurement budgets blow up as gradient energy decays","key_machinery":"The load-bearing object is the SPSA gradient estimator $\\hat g_\\ell(\\theta)=(\\hat f(\\theta+c\\Delta)-\\hat f(\\theta-c\\Delta))/(2c\\Delta_\\ell)$, computed from two noisy objective evaluations with a Rademacher perturbation vector $\\Delta$. The argument decomposes this estimator into the true partial derivative, cross terms from other gradient components, a finite-difference remainder controlled by the Hessian Lipschitz constant, and measurement noise. A signal-to-noise ratio, defined as $E_\\Theta[(\\partial_\\ell f)^2]/E_\\Theta[\\mathrm{Var}(\\hat g_\\ell-\\partial_\\ell f)]$, converts the variance bound into a measurement-budget requirement. The paper balances the bias growth in $c$ against the $1/(2c^2M)$ measurement variance by setting $c^2=\\Theta(\\sqrt{E_\\Theta[(\\partial_\\ell f)^2]})$, which is what raises the shot-noise exponent from $1$ to $3/2$. The convergence proof then uses the biased-descent lemma with the explicitly scheduled step size, perturbation radius, and per-iteration shot count, yielding the $O(T^{-1/2})$ rate and the $5/2$ exponent in the total budget.","core_discovery":"On the paper's own terms, the discovery is a quantitative pipeline from landscape flatness to optimizer cost. For the ideal VQE objective $f(\\theta)=\\langle\\phi_0|U^\\dagger(\\theta)HU(\\theta)|\\phi_0\\rangle$ with $H=\\sum_{\\alpha=1}^L w_\\alpha P_\\alpha$, the Hessian is Lipschitz with constant $8\\|w\\|_1 N^{3/2}$, and the finite-shot measurement noise is conditionally unbiased with variance at most $\\|w\\|_2^2/M$. These facts bound the bias of each SPSA gradient component by $4c^2 N^3\\|w\\|_1$ and its variance by the sum of the SPSA intrinsic perturbation variance, the measurement variance $1/(2c^2M)$, and higher-order Hessian terms. Choosing $c^2=\\gamma\\sqrt{E_\\Theta[(\\partial_\\ell f)^2]}$ and requiring an SNR of $\\epsilon_1$ forces a per-Pauli measurement budget $M=\\Omega(\\epsilon_1\\|w\\|_2^2/(E_\\Theta[(\\partial_\\ell f)^2])^{3/2})$, which under the barren-plateau condition $E_\\Theta[(\\partial_\\ell f)^2]=O(2^{-n})$ is $\\Omega(2^{3n/2})$. For the full trajectory, with $\\mu_t=\\mu_0 T^{-1/2}$, $c_t=c_0 T^{-1/8}$, and $M_t=M_0 T^{1/4}$, the paper proves $\\min_{0\\le t\\le T-1}\\mathbb{E}[\\|\\nabla f(\\theta_t)\\|_2^2]\\le \\kappa T^{-1/2}$, and achieving the relative accuracy $\\epsilon_2 E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]$ requires $T=\\Omega(\\kappa^2/(\\epsilon_2^2(E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2])^2))$ iterations and $N_{\\rm SPSA}=\\Omega(LM_0\\kappa^{5/2}/(\\epsilon_2^{5/2}(E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2])^{5/2}))$ total measurements; with $E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]=O(N2^{-n})$, both grow exponentially in $n$.","pith_inferences":["An implication beyond the paper: the same SNR-based argument should apply to any two-evaluation stochastic gradient estimator, so the $3/2$ and $5/2$ exponents are likely structural for derivative-free optimizers on flat quantum landscapes, not special to SPSA.","If the gradient second moment decays polynomially rather than exponentially, the paper's formulas predict polynomial iteration and measurement costs; this gives a concrete resource-based ranking for barren-plateau mitigation strategies by the gradient energy they preserve.","A testable extension would track the parameter distribution actually visited by SPSA iterates; if that distribution differs from the symmetric initialization ensemble used in barren-plateau theory, the predicted $\\Omega(2^{3n/2})$ and $\\Omega(2^{5n/2})$ budgets may over- or under-state the true cost."],"forward_implications":["For a fixed Hamiltonian and a fixed relative accuracy $\\epsilon_2$, the iteration count $T$ scales as the square of the inverse intrinsic gradient energy, so any exponential decay of the gradient energy translates directly into exponentially many SPSA steps.","The per-gradient measurement cost for a single reliable step scales as the $3/2$ power of the inverse gradient second moment, worse than the $1/M$ shot-noise scaling because the perturbation radius must shrink as the landscape flattens.","Under the standard scaling $E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]=O(N2^{-n})$, the total measurement budget grows like $\\Omega(2^{5n/2})$ up to polynomial factors, so no fixed polynomial shot budget can keep SPSA trainable in the barren-plateau regime.","The same theorem says that any initialization or ansatz modification that increases the intrinsic gradient energy automatically improves both iteration and measurement complexity, because both are polynomial in the inverse of that energy."],"supporting_citations":[{"why":"Introduces barren plateaus in quantum neural network training landscapes, motivating the exponentially vanishing gradient premise.","marker":"[7]"},{"why":"Supplies the standard barren-plateau variance condition $E[(\\partial_\\ell f)^2]=O(2^{-n})$ used throughout the scaling analysis.","marker":"[10]"},{"why":"Provides the reduced-domain initialization theorem used for the SPSA-IHT comparison and the mitigated-gradient-energy contrast in the simulations.","marker":"[24]"},{"why":"Baseline study of optimizer performance in variational quantum algorithms under limited measurements, motivating the finite-shot SPSA analysis.","marker":"[35]"},{"why":"Introduces the hardware-efficient VQE ansatz used in the numerical simulations and provides the practical SPSA adoption context.","marker":"[3]"},{"why":"Parameter-shift rule baseline whose cost grows with parameter count, against which SPSA's two-evaluation property is framed.","marker":"[45]"}],"fun_headline_variants":["Barren plateaus force exponential SPSA measurement cost","SPSA needs exponentially more shots under barren plateaus","Barren plateaus blow up SPSA iteration and shot budgets","Finite-shot SPSA succumbs to exponential barren plateau blowup","Exponential measurement budget for SPSA under barren plateaus"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The exponential conclusion rests on the imported barren-plateau scaling $E_\\Theta[(\\partial_\\ell f)^2]=O(2^{-n})$ and $E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]=O(N2^{-n})$ for the parameter distribution that SPSA actually encounters, together with the zero-mean-gradient condition that makes variance equal the second moment; if those fail along the trajectory, the exponential iteration and measurement bounds do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Barren plateaus force exponential SPSA measurement cost","SPSA needs exponentially more shots under barren plateaus","Barren plateaus blow up SPSA iteration and shot budgets","Finite-shot SPSA succumbs to exponential barren plateau blowup","Exponential measurement budget for SPSA under barren plateaus"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000327,"raw_usage":{"total_tokens":1974,"prompt_tokens":1236,"completion_tokens":738,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":852,"completion_tokens_details":{"reasoning_tokens":651}},"tokens_in":852,"tokens_out":738,"duration_ms":6011,"temperature":1.0,"reasoning_tokens":651,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:21:47.630015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, for the RealAmplitudes ansatz at $n=10,20,30$, the empirical gradient energy $E_\\Theta[\\|\\nabla f(\\theta)\\|_2^2]$ under the distribution of SPSA iterates rather than under the initialization ensemble; if it decays more slowly than $2^{-n}$, or if the mean gradient is not negligible compared with the second moment, the $\\Omega(2^{3n/2})$ and $\\Omega(2^{5n/2})$ resource conclusions of Corollary 1 and Theorem 2 are contradicted.","supporting_citations":[{"cited_title":"Barren plateaus in variational quantum computing","cited_arxiv_id":null,"evidence_quote":"Supplies the standard barren-plateau variance condition $E[(\\partial_\\ell f)^2]=O(2^{-n})$ used throughout the scaling analysis."},{"cited_title":"Trainability enhancement of parameterized quantum circuits via reduced-domain parameter initialization.Physical Review Applied, 22(5):054005, 2024","cited_arxiv_id":null,"evidence_quote":"Provides the reduced-domain initialization theorem used for the SPSA-IHT comparison and the mitigated-gradient-energy contrast in the simulations."},{"cited_title":"Performance comparison of optimization methods on variational quantum algorithms.Physical Review A, 107(3):032407, 2023","cited_arxiv_id":null,"evidence_quote":"Baseline study of optimizer performance in variational quantum algorithms under limited measurements, motivating the finite-shot SPSA analysis."}],"review_version":1}