Pith. sign in

REVIEW 3 major objections 6 minor 68 references

Shot noise forces gradient attacks on quantum classifiers to cost a measurement budget that scales as d^{5/2} or worse.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 07:05 UTC pith:LMXHNP63

load-bearing objection Solid scaling law for the shot cost of gradient attacks on QML; the main limit is the hardness-regime scope the authors already state, not a hole in the derivation. the 3 major comments →

arxiv 2607.11095 v1 pith:LMXHNP63 submitted 2026-07-13 quant-ph cs.CRcs.LG

When cheap gradients fail: the measurement cost of attacking quantum classifiers

classification quant-ph cs.CRcs.LG
keywords quantum machine learningadversarial robustnessparameter-shift ruleshot noisevariational quantum circuitsgradient cost ratiofinite-shot estimation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that finite quantum measurement statistics act as a built-in defense against gradient-based test-time attacks on variational quantum classifiers. Because every gradient component must be reconstructed from repeated circuit runs under any unbiased estimator, a white-box attacker who cannot classically simulate the model must spend a total shot budget that grows at least quadratically with input dimension—and as d^{5/2} under natural perturbation scaling. Simulations up to 784 dimensions recover the predicted shot-noise geometry once gradient norms are folded out, and a matched classical network keeps a constant, dimension-independent gradient cost. The relative cost of attacking therefore diverges as models scale, but only when the forward map is classically hard and the attacker is denied free backpropagation through simulation. A reader cares because this turns ordinary measurement noise into a structural security asymmetry against the cheap-gradient principle of classical autodiff.

Core claim

Under attacker-favoring assumptions, single-step gradient attacks on quantum classifiers require a total measurement budget R = Θ(ε d²), becoming Θ(d^{5/2}) when the perturbation budget scales as √d. For the tested deep circuits without barren-plateau mitigation the realized budget grows as about d^{3} because gradient norms decay with dimension; folding the measured norm back in recovers the parameter-free d^{3/2} geometry of shot noise. Against a classical baseline whose gradient cost ratio is O(1), the quantum ratio grows polynomially, so the attacker’s relative cost diverges with scale—provided the forward map is classically hard to simulate.

What carries the argument

The shot-budget scaling law for unbiased gradient extraction: each of d input components is recovered by an unbiased estimator (parameter-shift or equivalent) with per-component variance Θ(1/s), so constant attack efficacy forces s = Θ(ε d) shots per dimension and total R = d s = Θ(ε d²). Measurement grouping cannot remove this cost in expressive circuits with near-maximal dynamical Lie algebras.

Load-bearing premise

The defense only binds when the attacker cannot classically simulate the circuit and backpropagate for free; every simulation and the hardware demo in the paper sit on the simulable side of that line.

What would settle it

Force a white-box attacker to extract gradients solely by measurement (no simulation) on a plateau-mitigated quantum classifier across several large input dimensions, and check whether the total shots needed to hold attack success fixed still scale at least as d^{5/2}; clear sub-quadratic scaling under unbiased estimation would refute the floor.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Robustness evaluations of quantum classifiers should report and normalize by total shot budgets, not only perturbation size.
  • The classical cheap-gradient principle fails on quantum hardware: the gradient-to-inference cost ratio diverges polynomially with input dimension.
  • Even a single high-accuracy FGSM sample at d=784 already costs on the order of hours of sequential device time; dataset-wide attacks become prohibitive.
  • The barrier is operative precisely where QML is expected to be useful—when the forward map cannot be simulated classically.
  • Defenses that amplify or exploit shot sensitivity (budget-aware training, controlled noise) complement architectural robustness guarantees.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Plateau-mitigated architectures that keep gradient norms stable would sit at the milder d^{5/2} floor rather than the measured ~d^{3}, so trainability improvements may also weaken this defensive effect.
  • Biased zero-order estimators such as SPSA might lower the exponent in low-curvature regimes; whether they can do so for the expressive, classically hard models where the defense matters remains open.
  • If multi-copy joint measurements or co-measurement schemes ever beat the expressivity–measurement-efficiency trade-off, the irreducible Θ(d) per-gradient floor could be challenged.
  • The same measurement-cost asymmetry may appear in any quantum sensing or metrology setting where an adversary must reconstruct a high-dimensional score from finite shots.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that finite-shot measurement noise is an intrinsic, dimension-dependent cost on white-box gradient extraction for variational quantum classifiers, and therefore a built-in defense against gradient-based test-time attacks when the forward map is classically hard to simulate. Under attacker-favoring assumptions (local linearity, i.i.d. shot noise, small noise), Proposition 3 gives a single-step budget R=Θ(ε d²), hence Θ(d^{5/2}) under ε∝√d and ∥g*∥=Θ(1); a correlated-noise relaxation (Corollary 2) and a success-probability route (Proposition 4) recover the same quadratic total-shot structure. Iterative C&W is cast as SGLD with β∝s, yielding only a sufficient poly(d) certificate. Simulations to d=784 recover the parameter-free d^{3/2} geometry after folding out measured gradient norms, while bare budgets grow ~d^{3.00} for the tested plateau-prone circuits; a matched classical CNN has ρ≈5 independent of d, so the quantum/classical cost ratio diverges polynomially. A d=12 hardware run on ibm_boston tracks the simulator at matched budgets.

Significance. If the scaling law holds in the classically hard regime the authors target, the result is a genuine structural contribution: it recasts the classical cheap-gradient principle (Baur–Strassen) as failing on quantum hardware for unbiased gradient extraction, and it quantifies a physical cost orthogonal to dynamical robustness guarantees. Strengths include an attacker-favoring derivation that still yields a high floor; an empirical ledger that closes without free parameters once ∥g*∥ and the landscape factor ζ(d) are measured (recovering ~d^{1.46} vs d^{3/2}); a flat classical baseline; architecture and estimator robustness checks; and a hardware comparison that isolates shot noise from device bias. The explicit scoping—that experiments establish the scaling law, not a deployed defense at simulable sizes—is good practice and should be retained.

major comments (3)
  1. Section 4.1.3 and Table 4: absolute values of ρ_quantum (and the claim that quantum overhead exceeds classical for d≳30) depend on the free normalizations R_fwd=100 and ΔL⋆=1 unit of attack loss (Appendix E). The exponent is independent of these choices and is the load-bearing claim, but the absolute crossover and the large ratios at high d are not. Please either (i) report a short sensitivity sweep over R_fwd and ΔL⋆, or (ii) present Table 4 primarily as an exponent comparison and mark absolute ρ_quantum as illustrative under a stated convention.
  2. Section 3.2 / Appendix F: the SGLD certificate produces a very high sufficient degree (schematically ~d^{16} with large log factors). The manuscript correctly calls this sufficient rather than necessary and uses β∝s mainly as a modeling frame for the empirical sweet spot, but the main-text contribution list still pairs it with the single-step lower-envelope law. Please tighten the main-text wording so that iterative attacks are clearly a qualitative/sufficient analysis, not a matching lower bound, and so that the single-step Θ(d^{5/2}) floor remains the primary quantitative claim.
  3. Section 4.1.1 and the σ² bookkeeping: per-shot variance is measured flat (σ²≈3) only through d=121 and assumed flat to d=784 when building s⋆ grids and interpreting high-d coefficients. Because H(d)∝σ²/∥g*∥, a mild growth of σ² with d would change the bare exponent attribution. Please either measure σ² at the larger dimensions used in the power-law fits or state the flat-σ² assumption as an explicit limitation on the d>121 points and re-fit with a sensitivity band.
minor comments (6)
  1. Abstract and §1: the phrase “built-in defense” is accurate only under the classical-hardness scoping stated later. Consider one early sentence in the abstract that the defensive reading applies when the forward map is classically hard, matching the careful “Regime of applicability” paragraph.
  2. Notation table (Table 2): ζ(d) and κ_vMF appear in the theory before they are fully defined in the experimental coefficient rule; a one-line forward pointer would help.
  3. Figure 2 waterfall axes: “height = absolute shortfall” is clear in the caption but the color/marker density at high s makes the 1/s decay hard to read; a supplementary 1/s collapse plot (already partly in Fig. 3) would help.
  4. Hardware §4.2 uses fixed ε=1.0 at d=12 rather than ε∝√d; the text already treats this as a single-dimension robustness check, but a short explicit note that it does not test the ε-scaling of Proposition 3 would avoid over-reading “reproduces the effect.”
  5. Appendix L correctly leaves SPSA/biased estimators open; a single sentence in the main-text limitations list pointing to that appendix would make the unbiased scope of the exponent more visible to readers who skip appendices.
  6. Typos / polish: “ass→∞” spacing in a few places; “plateau-pronecircuits” missing space in the abstract block; ensure arXiv ID and journal metadata are consistent in the final submission.

Circularity Check

0 steps flagged

No significant circularity: single-step R=Θ(ε d²) follows from vMF concentration under stated Assumptions 1–3; empirical excess over the d^{5/2} floor is accounted for by independently measured ∥g*∥ decay, recovering the parameter-free geometry.

full rationale

The central derivation (Proposition 3) obtains the loss surplus ΔL(s)≈ε(d−1)σ²/(2s∥g*∥) from the von Mises–Fisher concentration of the normalized PSR estimator under local linearity, i.i.d. Gaussian shot noise and the small-noise regime; inverting for fixed gap yields s=Θ(ε d/∥g*∥) and R=Θ(ε d²). The ∥g*∥=Θ(1) specialization that produces the d^{5/2} floor under ε∝√d is an explicit modeling premise (realized by plateau-mitigated architectures), not a quantity fitted to force the claim. Empirically the bare coefficient grows as d^{2.00} because the tested deep circuits exhibit measured ∥g*∥∝d^{−0.73}; folding that independently measured norm back into the coefficient recovers the parameter-free geometric factor d^{1.46}≈d^{3/2}, which is a consistency check rather than a circular prediction. The classical ρ≈5 baseline is the external Baur–Strassen theorem; the expressivity–measurement-efficiency trade-off that prevents amortization is cited from Chinzei et al. (external); the SGLD mapping for iterative attacks rests on Raginsky et al. No load-bearing step reduces by construction to a fitted input or a self-citation chain. The hardness-regime scoping is an explicit limitation, not a circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 9 axioms · 2 invented entities

The central scaling rests on standard shot-noise / PSR unbiasedness plus three attacker-favoring modeling assumptions, the ε∝√d convention, single-copy measurement, and the literature expressivity–measurement-efficiency trade-off that blocks amortization. Empirical exponents additionally depend on measured model properties (∥g*∥ decay, ζ) rather than free fit knobs that define the claim. No new physical entity is postulated; the 'defense' is an interpretation of measurement cost under classical hardness.

free parameters (4)
  • κ in ε=κ√d (simulation: κ=1/32)
    Sets the absolute perturbation scale; scaling exponents are designed to be robust to κ, but absolute shot counts and practical-cost estimates depend on it.
  • Target loss surplus Δ / ΔL⋆ for ρ_quantum extraction
    Choice of fixed attack-efficacy tolerance shifts the vertical scale of ρ_quantum while leaving the reported exponent unchanged (Appendix E).
  • Forward-inference budget R_fwd=100
    Normalization for the gradient cost ratio; again shifts absolute ρ but not the d-exponent.
  • Per-shot variance σ² (≈3 measured to d=121, assumed flat above)
    Enters H(d) and critical shot counts; flatness above d=121 is an extrapolation used in grid construction and scaling interpretation.
axioms (9)
  • domain assumption Assumption 1: loss is locally linear inside the ε-ball
    Used in Proposition 3 and success-probability geometry; paper notes curvature mildly favors the attacker (ζ<1) rather than invalidating the baseline.
  • domain assumption Assumption 2: i.i.d. Gaussian gradient components with Var=σ²/s
    Drives vMF concentration; relaxed in Corollary 2 to zero-mean error with tr Σ=Θ(d).
  • domain assumption Assumption 3: small-noise regime s≫σ²/∥g*∥²
    Justifies large-κ vMF approximation; experiments place grids from the critical count upward.
  • domain assumption Unbiased per-component gradient estimation with variance Θ(σ²/s); co-measurement collapses to O(1) in near-maximal dynamical Lie algebra circuits
    Fixes the Θ(d) circuit count and thus the exponent; cited from expressivity–measurement-efficiency trade-off [29]. Biased estimators excluded from the main bound (Appendix L).
  • domain assumption Attacker processes one copy at a time; no cross-circuit entanglement / joint multi-copy measurement
    Stated in Section 1; rules out collective measurement strategies that might change the resource accounting.
  • domain assumption ε∝√d from ℓ2 norm concentration of natural noise / per-coordinate imperceptibility
    Appendix H; converts R=Θ(ε d²) into the d^{5/2} headline floor.
  • domain assumption Defense is operative only when the forward map is classically hard to simulate
    Section 1 scoping; without it a white-box attacker backpropagates through simulation at O(1) cost (Remark 1).
  • standard math PSR / two-eigenvalue generators and standard shot-noise variance bounds (quantum Cramér–Rao / Born-rule sampling)
    Background from Mitarai/Schuld parameter-shift literature and measurement theory; used throughout Section 2.
  • standard math Baur–Strassen / cheap-gradient principle: classical ρ≤5 independent of d
    External classical baseline for the diverging cost-ratio claim (Remark 1, Section 4.1.3).
invented entities (2)
  • Landscape factor ζ(d) independent evidence
    purpose: Ratio of realized clean-attack loss rise to first-order prediction; absorbs local curvature so the geometric shot-noise factor can be isolated.
    Defined and measured in §4.1.1 / Appendix I; not a new physical object, but a paper-specific diagnostic. Independent evidence is the transverse line-scan measurement itself.
  • Quantum gradient cost ratio ρ_quantum(d)=R_bwd/R_fwd independent evidence
    purpose: Dimensionless comparison of attack gradient cost to one forward inference, enabling classical vs quantum overhead contrast.
    Definitional construct analogous to classical cost ratio; value is measured/derived, not postulated as a new force or particle.

pith-pipeline@v1.1.0-grok45 · 46481 in / 4612 out tokens · 48924 ms · 2026-07-14T07:05:34.189029+00:00 · methodology

0 comments
read the original abstract

Adversarial perturbations threaten machine learning classifiers, including variational quantum classifiers. We show that finite quantum measurement statistics (shot noise) act as a built-in defense against gradient-based test-time attacks whose cost scales unfavorably for the attacker. Because every gradient component must be inferred from repeated circuit executions under any unbiased gradient-estimation rule, white-box extraction consumes a dimension-dependent measurement budget that measurement grouping cannot remove in expressive circuits. Under stated assumptions, single-step attacks need at least quadratically many shots in the input dimension $d$, growing as $d^{5/2}$ under norm-concentration scaling, with a sufficient-budget analysis for iterative attacks via stochastic gradient Langevin dynamics. Simulations up to 784 input dimensions validate the law: the realized total budget is the $d^{5/2}$ geometric floor for plateau-mitigated models and grows as $d^{3.00}$ for the tested deep circuits, whose gradient norms decay with dimension absent barren-plateau mitigation; folding the measured gradient norm back in recovers the parameter-free $d^{3/2}$ shot-noise geometry. Against a matched classical baseline whose attack overhead is dimension-independent (the cheap-gradient principle of automatic differentiation), the quantum gradient cost ratio grows empirically as $d^{3.00}$, so the attacker's relative cost diverges as the model scales. Experiments on a 156-qubit IBM processor (ibm_boston, 4-qubit circuits, $d=12$) reproduce the effect: at matched budgets the device attack tracks the ideal within a few percent, with the high-shot gradient faithful to the exact one. The defense operates precisely when the forward map is classically hard to simulate: only then is a white-box attacker denied the simulate-and-backpropagate shortcut and must pay the measurement cost we quantify.

Figures

Figures reproduced from arXiv: 2607.11095 by Bacui Li, Chandra Thapa, Tansu Alpcan, Udaya Parampalli.

Figure 1
Figure 1. Figure 1: Geometry of a single-step shot-noise attack. The sample x lies at perpendicular distance lx from a locally linear decision boundary, with the true gradient g ∗ (blue) pointing toward it. Shot noise deflects the estimated gradient gˆ (red dashed) by angle α; the perturbation of strength ϵ (orange) follows gˆ. The attack succeeds (i.e., x+δ crosses the boundary) only when α ≤ τ = arccos(lx/ϵ). The dotted lin… view at source ↗
Figure 2
Figure 2. Figure 2: FGSM shot-scaling on MNIST and Fashion-MNIST. Waterfall plots versus shot budget and model size, on log-spaced grids at fixed multiples h = s/s⋆ of the critical count, reaching h = 12 at every q. Axes: x = shots per input dimension s (log scale, including both PSR shifts); depth = qubit/class count q ∈ {4, . . . , 10}; height = absolute shortfall of the s-shot attack from the perfect-gradient attack; a mar… view at source ↗
Figure 3
Figure 3. Figure 3: Per-dimension shot-scaling coefficients. Log–log plot of the through￾origin coefficients Bq (fit window h ≥ 1) versus input dimension d. MNIST: Bq ≈ 4.7 × 10−3 d 2.00 (R2 = 0.999); Fashion-MNIST: Bq ≈ 4.5 × 10−3 d 2.07 (R2 = 0.999). Dotted: the aggregate reference ∝ d 2.23 obtained by inserting the cohort-mean gradient￾norm decay into the coefficient rule (§4.1.1). Both datasets exhibit superlinear growth … view at source ↗
Figure 4
Figure 4. Figure 4: Gradient cost ratio ρ(d) on log–log axes. Classical CNN (blue): flat at ρ ≈ 5, confirming O(1) scaling (Baur–Strassen). Quantum (red/orange): polynomial growth ∝ d 3.00 (MNIST), d 3.07 (Fashion-MNIST). Dashed gray: Θ(d 2.5 ) geometric￾floor reference, anchored at the MNIST d=121 point. The quantum–classical overhead ratio grows as Θ(d 3.00), with absolute value set by the forward-inference normalization ( … view at source ↗
Figure 5
Figure 5. Figure 5: Shot-noise robustness: simulation vs. experiment on ibm_boston (d = 12, N = 100, ϵ = 1.0). (a) Adversarial accuracy and (b) loss gap to the exact￾optimal attack, ∆L(s) = L(δ ∗ ) − L(δs), versus attacker per-parameter shot budget s (log axis) under ℓ2 FGSM. The simulation reaches the exact-gradient floor (10%, dashed) as s → ∞; the experiment plateaus above it, its device-biased gradient bounded by the meas… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 3 canonical work pages

  1. [1]

    Discussion and conclusion The picture that emerges from theory and experiment is that finite quantum measurement imposes a shot-cost floor on gradient extraction, and that this floor reshapes adversarial attack efficacy as the input dimension grows. For single-step attacks we derive a simple law tying optimization error to the per-dimension shot budget,∆L...

  2. [2]

    Verifying these for the model yieldsd-dependent constants M=O(cd+ 1),b=O(c 2d2),B=O(c), and gradient-noise varianceσ 2 g =O(c 2σ2 Σd/s), withm= 1andκ 0 dimension-independent, where the bounded measurement-covariance condition∥Σ v∥ ≤σ 2 Σ keeps the observable normO(1)under unitary conjugation. Combining the bound with the change of variable (F.3) translate...

  3. [3]

    Goodfellow I J, Shlens J and Szegedy C 20143rd International Conference on Learning Representations, ICLR 2015 - Conference Track ProceedingsPublisher: InternationalConference on Learning Representations, ICLR URLhttps://arxiv.org/abs/1412.6572v3

  4. [4]

    URLhttps://arxiv.org/abs/1608.04644v2

    Carlini N and Wagner D 2017Proceedings - IEEE Symposium on Security and Privacy39–57 ISSN 10816011 iSBN: 9781509055326 Publisher: Institute of Electrical and Electronics Engineers Inc. URLhttps://arxiv.org/abs/1608.04644v2

  5. [5]

    Szegedy C, Zaremba W, Sutskever I, Bruna J, Erhan D, Goodfellow I and Fergus R 2013 Intriguing properties of neural networksICLR

  6. [6]

    1080/23746149.2023.2165452

    Melnikov A, Kordzanganeh M, Alodjants A and Lee R K 2023Advances in Physics: X82165452 ISSN 23746149 publisher: Taylor & Francis URLhttps://www.tandfonline.com/doi/abs/10. 1080/23746149.2023.2165452

  7. [7]

    Liu N and Wittek P 2020Physical Review A101062331 ISSN 24699934 publisher: American Physical Society URLhttps://journals.aps.org/pra/abstract/10.1103/PhysRevA.101. 062331

  8. [8]

    1103/PhysRevA.103.042427

    Liao H, Convy I, Huggins W J and Whaley K B 2021Physical Review A103042427 ISSN 24699934 publisher: American Physical Society URLhttps://journals.aps.org/pra/abstract/10. 1103/PhysRevA.103.042427

  9. [9]

    Mahloujifar S, Diochnos D I and Mahmoody M 201833rd AAAI Conference on Artificial Intelligence, AAAI 2019, 31st Innovative Applications of Artificial Intelligence Conference, IAAI 2019 and the 9th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 20194536–4543 ISSN 2159-5399 iSBN: 9781577358091 Publisher: AAAI Press URL https://arxiv...

  10. [10]

    West M T, Erfani S M, Leckie C, Sevior M, Hollenberg L C L and Usman M 2023Physical Review Research5023186 ISSN 26431564 publisher: American Physical Society (APS) URL https://journals.aps.org/prresearch/abstract/10.1103/PhysRevResearch.5.023186

  11. [11]

    Du Y, Hsieh M H, Liu T, Tao D and Liu N 2021Physical Review Research3023153 publisher: American Physical Society URLhttps://link.aps.org/doi/10.1103/PhysRevResearch.3. 023153

  12. [12]

    org/abstract/document/10545406

    Kundu S, Choudhury N, Das S, Raha A and Basu K 2024 QNAD: Quantum Noise Injection for Adversarial Defense in Deep Neural Networks2024 IEEE International Symposium on Hardware Oriented Security and Trust (HOST)pp 1–11 iSSN: 2765-8406 URLhttps://ieeexplore.ieee. org/abstract/document/10545406

  13. [13]

    Ahmed T, Kashif M, Marchisio A and Shafique M 2025Scientific Reports1533654 ISSN 2045- 2322 publisher: Nature Publishing Group URLhttps://www.nature.com/articles/s41598- 025-17769-6

  14. [14]

    Zhang H F, Chen Z Y, Wang P, Guo L L, Wang T L, Yang X Y, Zhao R Z, Zhao Z A, Zhang S, Du L, Tao H R, Jia Z L, Kong W C, Liu H Y, Vasilakos A V, Yang Y, Wu Y C, Guan J, Duan P and Guo G P 2026Science China Physics, Mechanics & Astronomy69arXiv:2505.16714 [quant-ph] URLhttp://arxiv.org/abs/2505.16714

  15. [15]

    Dowling N, West M T, Southwell A, Nakhl A C, Sevior M, Usman M and Modi K 2026npj Quantum Information1216 URLhttps://doi.org/10.1038/s41534-025-01129-3

  16. [16]

    Mitarai K, Negoro M, Kitagawa M and Fujii K 2018Physical Review A98032309 ISSN 2469-9926, 2469-9934 URLhttps://link.aps.org/doi/10.1103/PhysRevA.98.032309

  17. [17]

    Schuld M, Bergholm V, Gogolin C, Izaac J and Killoran N 2019Physical Review A99032331 ISSN 2469-9926, 2469-9934 URLhttps://link.aps.org/doi/10.1103/PhysRevA.99.032331

  18. [18]

    Helstrom C W 1969Journal of Statistical Physics1231–252 ISSN 1572-9613 URLhttps: //doi.org/10.1007/BF01007479

  19. [19]

    Braunstein S L and Caves C M 1994Physical Review Letters723439–3443 publisher: American Physical Society URLhttps://link.aps.org/doi/10.1103/PhysRevLett.72.3439

  20. [20]

    Scriva G, Astrakhantsev N, Pilati S and Mazzola G 2024Physical Review A109032408 arXiv:2308.00044 Measurement cost of gradient attacks on QML55

  21. [21]

    Kübler J M, Arrasmith A, Cincio L and Coles P J 2020Quantum4263 publisher: Verein zur Förderung des Open Access Publizierens in den Quantenwissenschaften URLhttps://quantum- journal.org/papers/q-2020-05-11-263/

  22. [22]

    Bittel L, Watty J and Kliesch M 2022 Fast gradient estimation for variational quantum algorithms arXiv:2210.06484 [quant-ph] URLhttp://arxiv.org/abs/2210.06484

  23. [23]

    Ito K and Fujii K 2023 SantaQlaus: A resource-efficient method to leverage quantum shot- noise for optimization of variational quantum algorithms arXiv:2312.15791 [quant-ph] URL http://arxiv.org/abs/2312.15791

  24. [24]

    Kreplin D A and Roth M 2024Quantum81385 publisher: Verein zur Förderung des Open Access Publizierens in den Quantenwissenschaften URLhttps://quantum-journal.org/papers/q- 2024-06-25-1385/

  25. [25]

    org/document/9605341/

    Ma Z, Gokhale P, Zheng T X, Zhou S, Yu X, Jiang L, Maurer P and Chong F T 2021 Adaptive Circuit Learning for Quantum Metrology2021 IEEE International Conference on Quantum Computing and Engineering (QCE)pp 419–430 URLhttps://ieeexplore.ieee. org/document/9605341/

  26. [26]

    Athalye A, Carlini N and Wagner D 2018 Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial ExamplesProceedings of the 35th International Conference on Machine Learning(PMLR) pp 274–283 iSSN: 2640-3498 URLhttps : / / proceedings.mlr.press/v80/athalye18a.html

  27. [27]

    Cohen J M, Rosenfeld E and Kolter J Z 2019 Certified adversarial robustness via randomized smoothingProceedings of the 36th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Researchvol 97) (PMLR) pp 1310–1320

  28. [28]

    Wong E and Kolter J Z 2018 Provable defenses against adversarial examples via the convex outer adversarial polytopeProceedings of the 35th International Conference on Machine Learning (ICML)(Proceedings of Machine Learning Researchvol 80) (PMLR) pp 5286–5295 arXiv:1711.00851

  29. [29]

    Baur W and Strassen V 1983Theoretical Computer Science22317–330

  30. [30]

    Griewank A and Walther A 2008Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation2nd ed (SIAM)

  31. [31]

    Chinzei K, Yamano S, Tran Q H, Endo Y and Oshima H 2025npj Quantum Information1179 arXiv:2406.18316

  32. [32]

    Wierichs D, Izaac J, Wang C and Lin C Y Y 2022Quantum6677 URLhttps://quantum- journal.org/papers/q-2022-03-30-677/

  33. [33]

    Banchi L, Branford D and Waghela C 2026Quantum Science and TechnologyAccepted manuscript; arXiv:2510.05289 [quant-ph] URLhttps://doi.org/10.1088/2058-9565/ae73ad

  34. [34]

    Periyasamy M, Plinge A, Mutschler C, Scherer D D and Mauerer W 2024 Guided-SPSA: Simultaneous Perturbation Stochastic Approximation Assisted by the Parameter Shift Rule 2024 IEEE International Conference on Quantum Computing and Engineering (QCE)vol 01 pp 1504–1515 URLhttps://ieeexplore.ieee.org/document/10821406/

  35. [35]

    Hoffmann T and Brown D 2022 Gradient Estimation with Constant Scaling for Hybrid Quantum Machine Learning arXiv:2211.13981 [quant-ph] URLhttp://arxiv.org/abs/2211.13981

  36. [36]

    Heidari M, Naved M A, Honjani Z, Xie W, Grama A J and Szpankowski W 2024 Quantum Shadow Gradient Descent for Variational Quantum Algorithms arXiv:2310.06935 [quant-ph] URLhttp://arxiv.org/abs/2310.06935

  37. [37]

    Lockwood O 2022 An Empirical Review of Optimization Techniques for Quantum Variational Circuits arXiv:2202.01389 [quant-ph] URLhttp://arxiv.org/abs/2202.01389

  38. [38]

    Biggio B, Corona I, Maiorca D, Nelson B, Šrndić N, Laskov P, Giacinto G and Roli F 2013 Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)8190 LNAI387–402ISSN03029743iSBN:9783642409936 Publisher: Springer, Berlin, Heidelberg URLhttps://link.springer.com/chapter/10.1007/ 978-...

  39. [39]

    Raginsky M, Rakhlin A and Telgarsky M 2017 Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysisProceedings of the 2017 Conference on Learning Theory(PMLR) pp 1674–1703 iSSN: 2640-3498 URLhttps://proceedings.mlr.press/v65/ raginsky17a.html

  40. [40]

    Xu P, Chen J, Zou D and Gu Q 2018 Global Convergence of Langevin Dynamics Based Algorithms for Nonconvex OptimizationAdvances in Neural Information Processing Systems 31 (NeurIPS 2018)(arXiv) arXiv:1707.06618 [stat] URLhttp://arxiv.org/abs/1707.06618

  41. [41]

    Nesterov Y and Spokoiny V 2017Foundations of Computational Mathematics17527–566

  42. [42]

    Jamieson K G, Nowak R D and Recht B 2012 Query complexity of derivative-free optimization Advances in Neural Information Processing Systems (NeurIPS)vol 25 pp 2672–2680 arXiv:1209.2434

  43. [43]

    Duchi J C, Jordan M I, Wainwright M J and Wibisono A 2015IEEE Transactions on Information Theory612788–2806

  44. [44]

    Pérez-Salinas A, Cervera-Lierta A, Gil-Fuster E and Latorre J I 2020Quantum4226 (Preprint 1907.02085) URLhttps://quantum-journal.org/papers/q-2020-02-06-226/

  45. [45]

    Havlíček V, Córcoles A D, Temme K, Harrow A W, Kandala A, Chow J M and Gambetta J M 2019Nature 2019 567:7747567209–212 ISSN 1476-4687 publisher: Nature Publishing Group URLhttps://www.nature.com/articles/s41586-019-0980-2

  46. [46]

    Schuld M, Bocharov A, Svore K M and Wiebe N 2020Physical Review A101032308 (Preprint 1804.00633) URLhttps://link.aps.org/doi/10.1103/PhysRevA.101.032308

  47. [47]

    Madry A, Makelov A, Schmidt L, Tsipras D and Vladu A 20186th International Conference on Learning Representations, ICLR 2018 - Conference Track ProceedingsPublisher: International Conference on Learning Representations, ICLR URLhttps://arxiv.org/abs/1706.06083v4

  48. [48]

    1103/PhysRevResearch.2.033212

    Lu S, Duan L M and Deng D L 2020Physical Review Research2033212 ISSN 26431564 publisher: American Physical Society URLhttps://journals.aps.org/prresearch/abstract/10. 1103/PhysRevResearch.2.033212

  49. [49]

    Baydin A G, Pearlmutter B A, Radul A A and Siskind J M 2018Journal of Machine Learning Research181–43 arXiv:1502.05767

  50. [50]

    Feynman R P 1982International Journal of Theoretical Physics21467–488 ISSN 1572-9575 URL https://doi.org/10.1007/BF02650179

  51. [51]

    Bernstein E and Vazirani U 1997SIAM Journal on Computing261411–1473 ISSN 0097-5397 publisher: Society for Industrial and Applied Mathematics URLhttps://epubs.siam.org/ doi/10.1137/S0097539796300921

  52. [52]

    cambridge.org/core/books/asymptotic-statistics/A3C7DAD3F7E66A1FA60E9C8FE132EE1D

    van der Vaart A W 1998Asymptotic Statistics(Cambridge University Press) URLhttps://www. cambridge.org/core/books/asymptotic-statistics/A3C7DAD3F7E66A1FA60E9C8FE132EE1D

  53. [53]

    Yen T C, Ganeshram A and Izmaylov A F 2023npj Quantum Information9URLhttps: //www.nature.com/articles/s41534-023-00683-y

  54. [54]

    Crawford O, van Straaten B, Wang D, Parks T, Campbell E and Brierley S 2021Quantum5385 URLhttps://quantum-journal.org/papers/q-2021-01-20-385/

  55. [55]

    Mari A, Bromley T R and Killoran N 2021Phys. Rev. A103012405 URLhttps://link.aps. org/doi/10.1103/PhysRevA.103.012405

  56. [56]

    Kaminishi E, Mori T, Sugawara Met al.2026Scientific Reports169390 URLhttps://doi.org/ 10.1038/s41598-026-40123-3

  57. [57]

    Dauphin Y N, Pascanu R, Gulcehre C, Cho K, Ganguli S and Bengio Y 2014 Identifying and attacking the saddle point problem in high-dimensional non-convex optimizationAdvances in Neural Information Processing Systems (NeurIPS)vol 27 pp 2933–2941 arXiv:1406.2572

  58. [58]

    Li H, Xu Z, Taylor G, Studer C and Goldstein T 2018 Visualizing the loss landscape of neural nets Advances in Neural Information Processing Systems (NeurIPS)vol 31 arXiv:1712.09913

  59. [59]

    Pesah A, Cerezo M, Wang S, Volkoff T, Sornborger A T and Coles P J 2021Physical Review X 11041011

  60. [60]

    Cerezo M, Sone A, Volkoff T, Cincio L and Coles P J 2021Nature Communications12 Measurement cost of gradient attacks on QML57

  61. [61]

    Larocca M, Thanasilp S, Wang S, Sharma K, Biamonte J, Coles P J, Cincio L, McClean J R, Holmes Z and Cerezo M 2025Nature Reviews Physics7174–189

  62. [62]

    Kverne C, Akewar M, DiBrita N S, Huo Y, Patel T and Bhimani J 2026 Variational quantum algorithms are lipschitz smooth Preprint, under review at ICLR 2026 openReview: beg6QFuff4

  63. [63]

    wiley.com/doi/abs/10.1002/9780470316979.ch7

    Mardia K V and Jupp P E 1999 Tests on von Mises DistributionsDirectional Statistics(John Wiley & Sons, Ltd) pp 119–142 ISBN 978-0-470-31697-9 section: 7 URLhttps://onlinelibrary. wiley.com/doi/abs/10.1002/9780470316979.ch7

  64. [64]

    Laurent B and Massart P 2000Annals of Statistics281302–1338

  65. [65]

    Vershynin R 2018High-Dimensional Probability: An Introduction with Applications in Data ScienceCambridge Series in Statistical and Probabilistic Mathematics (Cambridge University Press)

  66. [66]

    Alabdulkareem A and Honorio J 2021 Information-theoretic lower bounds for zero-order stochastic gradient estimationIEEE International Symposium on Information Theory (ISIT)pp 2316–2321 arXiv:2003.13881

  67. [67]

    Spall J C 1992IEEE Transactions on Automatic Control37332–341 ISSN 0018-9286

  68. [68]

    Majumder R, Khan S M, Ahmed F, Khan Z, Ngeni F, Comert G, Mwakalonge J, Michalaka D and Chowdhury M 2021arXiv preprint arXiv:2108.01125(Preprint2108.01125)