Pith. sign in

REVIEW 3 major objections 5 minor 12 references

Evaluation of derivatives using approximate generalized parameter shift rule

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper introduces the approximate generalized parameter shift rule (aGPSR), showing that replacing the full set of spectral-gap equations with a small set of pseudo-gap equations yields derivatives accurate to $O(\alpha^{2K})$ while…

desk verdict aGPSR is a genuinely new construction with clean noiseless numerics, but the practical NISQ savings claim rests on an unanalysed bias-variance tradeoff. read the letter →

arxiv 2505.18090 v2 pith:MXUKPXPS submitted 2025-05-23 quant-ph

classification quant-ph MSC 81P68
keywords approximategeneralizedparametershiftrulequantumgradientsvariationaleigensolverneutralatomcomputingspectralgapsshotnoisemachinelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the cost of exact generalized parameter shift derivative estimation can be cut from solving $S$ equations, where $S$ grows as $(2^N)(2^N-1)/2$, to solving $K \ll S$ equations without materially losing accuracy. It does this by allowing arbitrary pseudo-gaps in a truncated sine linear system and showing that the resulting derivative coincides with the exact one up to a bias of order $\alpha^{2K}$, where $\alpha$ controls the shift size. If true, this makes analytic gradients practical for arbitrary-device Hamiltonians such as neutral-atom arrays, where ordinary parameter shift rules fail and exact GPSR requires thousands of expectation values. Demonstrated on variational quantum eigensolver tasks for 3 to 6 qubits, aGPSR reaches the same target energy as exact GPSR while using 7 to 504 times fewer expectation calls. The paper therefore positions aGPSR as an accuracy-controlled, shot-budget-friendly derivative rule for variational quantum algorithms.

What carries the argument

The central mechanism is the truncated sine system $F_i = 4\sum_{k=1}^K \sin(\delta_i\gamma_k/2) R'_k$, solved with Cramer's rule so that each $R'_k$ is a ratio of determinants. The paper defines functions $\eta_{ks}$ and $\xi_s = \sum_k \gamma_k \eta_{ks}$, where $\xi_s$ is the effective replacement for the true spectral gap $\Delta_s$. The key identity is the expansion $\xi_s = \Delta_s + O(\alpha^{2K})$ of Eq. (16), obtained by expanding determinant ratios in powers of the shift scale $\alpha$. The error terms factor as products of the form $(\gamma_k^2 - \Delta_s^2)$, which is why choosing pseudo-gaps close to the real spectral gaps suppresses the bias and why including more equations raises the error order to $2K$.

What would settle it

Take a six-qubit neutral-atom generator with strong interactions ($J/\Omega > 1$), fix $K=4$ pseudo-gaps, and compare aGPSR derivatives against exact GPSR at a grid of points $x$ using a finite shot budget such as 1000 shots per expectation value. If the mean relative error exceeds the noiseless $O(\alpha^{8})$ prediction by more than the shot-noise variance allows, or if the variance-minimizing shifts from the paper's Section 3 fail to hold the VQE energy at the target with those shots, the bias-variance trade-off behind the claimed savings does not hold on real hardware.

Watch

Extended reading notes

Core claim

The paper's central claim is that the exact generalized parameter shift rule can be replaced, for arbitrary Hermitian generators $\hat G$ with many spectral gaps, by a truncated version with $K \ll S$ equations. Working from the exact GPSR system $F_i = 4\sum_s \sin(\delta_i \Delta_s/2) R_s$, the paper substitutes $K$ arbitrary pseudo-gaps $\{\gamma_k\}$ and solves the smaller system for coefficients $R'_k$. The resulting approximate derivative is $\frac{df}{dx}\big|_K = \sum_s (\Delta_s + O(\alpha^{2K})) R_s$, where $\alpha$ is the common scale of the shifts $\delta_s$; this is Eq. (16) and is the load-bearing identity. It means the approximate derivative inherits the exact GPSR coefficients $R_s$ up to a controllable bias of order $\alpha^{2K}$, independently of the chosen pseudo-gaps to leading order. The paper then demonstrates on a neutral-atom Hamiltonian that $K=4$ to $K=20$ equations reproduce exact derivative curves, and in VQE runs on 3 to 6 qubits $K=4$ reaches the same target energy with 7, 30, 124, and 504 times fewer expectation calls for the analog ansatz.

Load-bearing premise

The method's practical accuracy hinges on choosing parameter shifts small enough that the inherited error from the truncated equations is negligible while the shifted measurement differences still remain accurate with a limited number of shots; the paper's variance analysis starts after that bias is controlled, and its VQE savings are demonstrated on noiseless analytical expectation values.

Editorial extensions

If this is right

  • For a given generator spectrum, a fixed small $K$ can hold relative derivative error near a target value while system size grows, whereas exact GPSR needs exponentially many shifted points.
  • VQE training on analog and digital ansatze can reach the same optimized energy with orders-of-magnitude fewer expectation calls, moving the practical bottleneck away from gradient evaluation.
  • The same truncated-system rule applies to any parametrized unitary $\exp(-ix\hat G/2)$ with an arbitrary Hermitian generator, so quantum machine learning and other gradient-based quantum algorithms inherit the savings.
  • Because $K$ is small, the search for variance-minimizing shift values $\delta_s$ becomes tractable, giving a practical route to noise-robust gradient estimation under limited shot budgets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the $O(\alpha^{2K})$ expansion suggests an adaptive scheme the paper mentions only as future work: start with $K=1$ or $2$ early in optimization and increase $K$ when gradients become small, which would make the reported savings grow beyond the fixed-$K$ factors.
  • Editorial extension: since error terms factor as $(\gamma_k^2 - \Delta_s^2)$, one could target specific dominant spectral gaps if approximate gap information is available from classical pre-simulation, using the error function $Q_K$ to decide how many gaps must be resolved for a target derivative accuracy.
  • Editorial caution: the reported 7 to 504 times savings are demonstrated on noiseless analytical expectation values; on real hardware with finite shots, the variance of aGPSR must be compared with exact GPSR variance at equal total shot budget, which the paper leaves as the key open test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes an approximate generalized parameter shift rule (aGPSR) for estimating derivatives of expectation values generated by arbitrary generators with many spectral gaps. Instead of solving the full S-by-S GPSR linear system, aGPSR solves a K-by-K system using K 'pseudo-gaps' and approximates the derivative as a weighted sum of K coefficients. The central claim is Eq. (16): with K pseudo-gap equations, the approximate derivative equals the exact GPSR derivative plus an O(alpha^{2K}) bias, where alpha scales the shift sizes. The authors give explicit bias expansions for K=1,2,3, propose pseudo-gap selection by uniform sampling over [0, Delta_max], and illustrate the method on neutral-atom Hamiltonians. In VQE experiments for 3 to 6 qubits, they report reducing the number of expectation calls by factors from 7 to 504 while reaching the same target energy.

Significance. If the central claim holds, aGPSR is potentially valuable: exact GPSR requires 2S shifted expectation values with S growing exponentially in the number of qubits, while aGPSR uses a small, user-selected K. The explicit error expansions for K=1,2,3 and the numerical agreement with exact derivatives for the tested neutral-atom cases are concrete strengths, and the open-source Qadence implementation makes the method testable. However, the practical, hardware-oriented claim of large computational savings is currently supported only by noiseless simulations; the finite-shot bias-variance tradeoff that would determine actual NISQ performance is not analyzed. Because the method's usefulness hinges on this tradeoff, the paper needs additional work before the savings claim can be accepted.

major comments (3)
  1. The variance analysis is performed only for the exact S-by-S GPSR system, and the text states that the same procedure can be applied to aGPSR 'as usual'. For K<S, however, the bias condition of Eq. (16) requires small shifts delta_s = alpha delta'_s with alpha small. In that regime M_{ij} = 4 sin(delta_i gamma_j / 2) is approximately 2 alpha delta'_i gamma_j, a rank-one matrix, so the inverse M^{-1} has entries growing as alpha^{-(2K-1)} (for K=2, as alpha^{-3}). Since each measured F_s has shot variance 2 sigma_0^2 / N_shots that is independent of alpha, the variance of each R'_k, Var(R'_k) = (2 sigma_0^2 / N_shots) sum_i (M^{-1})_{ki}^2, diverges as alpha^{-2(2K-1)}. The divergent parts cancel in the mean, giving the O(alpha^{2K}) bias, but they do not cancel in the variance. The paper never minimizes bias plus variance jointly for K<S, and the VQE experiments in §4.2 use noiseless analytical expectation values. Consequently, the reported 7–504x reduction in expectation calls does not account for the additional shots required to hold total error fixed, and the central hardware-oriented savings claim is unverified.
  2. The proof that the approximation error is O(alpha^{2K}) for general K is not a valid induction. The text shows explicit expansions only for K=1,2,3, then argues by assuming the relation for K=S-1 and noting exactness for K=S, concluding that the relation holds for all K. This does not establish a base case and induction step from K to K+1, nor does exactness at K=S imply the O(alpha^{2S-2}) error at K=S-1. Since Eq. (16) is the mathematical foundation for using arbitrarily small K, the general claim needs either a rigorous proof or an explicit restriction to the small-K cases with controlled error bounds. This is a load-bearing point for the method's purported scalability.
  3. The manuscript does not provide a complete, a priori procedure for choosing K and the pseudo-gap set {gamma_k}. The uniform sampling proposal in §2.3 gives a direction, but in the numerical studies K and the step a are selected per test case (e.g., K=4 with a=4 in the weak-interaction case, K=8 or K=20 in the strong-interaction cases), and Fig. 5b shows that the relative error can be non-monotonic in K (the K=16 curve is singled out as an exception). Without a criterion for selecting K and the pseudo-gaps for a new Hamiltonian, the method's practical reliability and reproducibility are not fully established.
minor comments (5)
  1. There is an incomplete sentence: 'Having outlined general formulations for aGPSR, we follow up by analysing the relationship between error terms in Eq. 16 and' ends abruptly before the next subsection.
  2. The heading '2.2 Approximate GPSR' appears twice, and the heading '4.1 Derivative calculations' also appears twice; these duplicated headings should be removed.
  3. The relative error r defined in Eq. (34) divides by f'_exact(x_i), which can be zero or very small for oscillatory derivative curves; the metric should be regularized or replaced by an absolute error measure to avoid unstable values.
  4. The caption of Fig. 5 says 'Comparison of derivatives df/dx calculated for different K values with zero initial state |0>' but the figure actually shows scaling of K and S and relative errors; this caption is a copy-paste error and should be corrected.
  5. Reference [13] is described as a patent application 'submitted' in the references but as 'filed' in the ethics declaration; the wording should be made consistent.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: aGPSR is a self-contained approximate-derivative construction whose accuracy claim is an analytic expansion result, not a fitted quantity presented as a prediction.

full rationale

The paper's central derivation is self-contained and non-circular. It starts from the exact GPSR formula of Ref. [6], treats it as ground truth, and constructs aGPSR by truncating the linear system to K equations with pseudo-gaps. Equations (5)-(16) are derived by linear algebra (Cramer's rule) and Taylor expansion in the small shift parameter alpha. The key claim that xi_s = Delta_s + O(alpha^{2K}) is an analytic error bound derived from the determinant structure, not an assumption equivalent to the target derivative. No parameter is fitted to the derivative values or to the VQE target energy: K is a user-chosen truncation size, gamma_k and delta_k are method parameters, and the VQE energy is used only as a benchmark for comparing methods. Self-citations appear (Ref. [6] by co-author Elfving; Refs. [10], [13] by Pasqal authors), but none is load-bearing in a circular sense: Ref. [6] provides the external exact GPSR starting point, and Refs. [10], [13] concern implementation and patent disclosure, not the derivation's validity. The skeptical concern about finite-shot bias-variance trade-off for K < S is a genuine correctness or validation gap, but it is not circularity: an unverified practical premise is different from a claim that reduces by construction to its own inputs. The paper even acknowledges that the variance minimization section analyzes the exact S-by-S case and states the aGPSR extension follows 'as usual', which is a completeness limitation rather than a circular step. Therefore the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The method rests on the exact GPSR formula as external input, on a small set of hand-chosen parameters (K, gamma, delta), and on the assumption that small-shift expansions remain valid for the finite shifts used in practice. No new physical entity is introduced; the pseudo-gap set is a computational device rather than a postulated observable.

free parameters (3)
  • Number of pseudo-gap equations K = K=1 (digital VQE), K=4 (analog VQE), K=2,4,8,20 (derivative benchmarks)
    Chosen by hand per task to balance accuracy and cost; no data-driven or theoretical stopping rule is provided.
  • Pseudo-gap set {gamma_k} = e.g., uniform step a=4 in Figure 3; otherwise not fully specified
    The paper suggests uniform sampling in [0, Delta_max] but does not give a complete recipe for reproducing all experiments.
  • Shift values {delta_s} = Not specified for VQE; equidistant in [pi/4, pi/2] in Figure 1; optimized in Section 3
    Shift values control the bias-variance trade-off and are tuned per experiment, but exact values for the main results are not reported.
assumptions (4)
  • domain assumption The exact GPSR derivative formula (Eq. 2) from Ref [6] is correct for arbitrary Hermitian generators.
    The paper adopts the generalized parameter shift rule as ground truth without re-deriving it; if this formula is wrong, aGPSR inherits the error.
  • domain assumption Shifted expectation values f(x +/- delta_s) are independent and identically distributed normal variables with constant variance sigma_0^2.
    Used in Section 3 to derive variance formulas; it ignores x-dependence, correlations between shifted shots, and higher-order shot statistics.
  • ad hoc to paper Pseudo-gaps can be chosen arbitrarily and the leading O(alpha^{2K}) error remains representative for finite shifts.
    Section 2.2 claims 'we can in principle use any values {gamma_s}', but this relies on neglecting higher-order terms; for the finite shifts used in experiments it is an assumption, not a proven bound.
  • domain assumption The spectral gap distribution can be approximated by uniform sampling on [0, Delta_max].
    Section 2.3 proposes this heuristic without a proof of optimality or robustness, and Figure 2 illustrates it only schematically.
invented entities (1)
  • Pseudo-gaps {gamma_k}
    purpose: Replace the true spectral gaps to reduce the GPSR linear system from S equations to K equations
    Pseudo-gaps are algorithmic hyperparameters with no physical meaning and no falsifiable prediction outside the method itself.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluation of derivatives using approximate generalized parameter shift rule." pith.science (2026). https://pith.science/paper/MXUKPXPS

@misc{pith2026250518090,
  author       = {Pith},
  title        = {Pith review of: Evaluation of derivatives using approximate generalized parameter shift rule},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MXUKPXPS}},
  note         = {Machine review of arXiv:2505.18090}
}
read the original abstract

Parameter shift rules are instrumental for derivatives estimation in a wide range of quantum algorithms, especially in the context of Quantum Machine Learning. Application of single-gap parameter shift rule is often not possible in algorithms running on noisy intermediate-scale quantum (NISQ) hardware due to noise effects and interaction between device qubits. In such cases, generalized parameter shift rules must be applied yet are computationally expensive for larger systems. In this paper we present the approximate generalized parameter rule (aGPSR) that can handle arbitrary device Hamiltonians and provides an accurate derivative estimation while significantly reducing the computational requirements. When applying aGPSR for a variational quantum eigensolver test case ranging from 3 to 6 qubits, the number of expectation calls is reduced by a factor ranging from 7 to 504 while reaching the exact same target energy, demonstrating its huge computational savings capabilities.

Figures

Figures reproduced from arXiv: 2505.18090 by the authors.

Figure 1
Figure 1. Error functions QK(∆) for different values of K. preferably from high-probabilty regions. Therefore, the selected K values of γk would be optimal for the most accurate estimate of the derivative. However, such sampling is too costly to implement for larger systems, as the eigenspectrum scales exponentially. On the other hand, we could still use efficient algorithms [7] to estimate the largest and smallest eigenvalue… view at source ↗
Figure 2
Figure 2. Illustration of uniform sampling procedure of pseudo-gaps [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Comparison of derivatives df dx calculated for different K values with zero initial state |ψ0⟩. From the results of aGPSR calculations for a random initial state presented in [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Comparison of derivatives df dx calculated for different K values with random initial state |ψ0⟩. Let us investigate the scaling of the number K denoting the minimum required number of equations in aGPSR framework to achieve some target mean relative error r of the der…
Figure 5
Figure 5. Figure 5: Comparison of derivatives df dx calculated for different K values with zero initial state |ψ0⟩. 4.2 Variational quantum eigensolver In this section, we demonstrate the efficiency of aGPSR when applied within a variational quantum eigensolver (VQE) task [9]. A VQE task …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references · 11 canonical work pages

  1. [1]

    Solving nonlinear differential equations with differentiable quantum circuits

    Oleksandr Kyriienko, Annie E Paine, and Vincent E Elfving. Solving nonlinear differential equations with differentiable quantum circuits. Physical Review A , 103(5):052416, 2021

  2. [2]

    Quantum extremal learning

    Savvas Varsamopoulos, Evan Philip, Vincent E Elfving, Herman WT van Vlijmen, Sairam Menon, Ann Vos, Natalia Dyubankova, Bert Torfs, and Anthony Rowe. Quantum extremal learning. Quantum Machine Intelligence , 6(2):42, 2024

  3. [3]

    Mitarai, M

    K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii. Quantum circuit learning. Phys. Rev. A , 98:032309, 2018

  4. [4]

    Evaluating analytic gradients on quantum hardware

    Maria Schuld, Ville Bergholm, Christian Gogolin, Josh Izaac, and Nathan Killoran. Evaluating analytic gradients on quantum hardware. Phys. Rev. A , 99:032331, 2019

  5. [5]

    Noise-robust optimization of quantum machine learning models for polymer properties using a simulator and validated on the ionq quantum computer

    Yuki Ishiyama, Ryutaro Nagai, Shunsuke Mieda, Yuki Takei, Yuichiro Minato, and Yutaka Natsume. Noise-robust optimization of quantum machine learning models for polymer properties using a simulator and validated on the ionq quantum computer. Scientific Reports , 12(1):19003, 2022

  6. [6]

    Generalized quantum circuit differentiation rules

    Oleksandr Kyriienko and Vincent E Elfving. Generalized quantum circuit differentiation rules. Physical Review A , 104(5):052417, 2021

  7. [7]

    Lehoucq, Danny C

    Richard B. Lehoucq, Danny C. Sorensen, and Chao Yang. ARPACK Users' Guide: Solution of Large-Scale Eigenvalue Problems with Implicitly Restarted Arnoldi Methods . SIAM, 1998

  8. [8]

    Quantum computing with neutral atoms

    Lo \" i c Henriet, Lucas Beguin, Adrien Signoles, Thierry Lahaye, Antoine Browaeys, Georges-Olivier Reymond, and Christophe Jurczak. Quantum computing with neutral atoms. Quantum , 4:327, September 2020

Show all 12 references
  1. [9]

    Love, Alán Aspuru-Guzik, and Jeremy L

    Alberto Peruzzo, Jarrod McClean, Peter Shadbolt, Man-Hong Yung, Xiao-Qi Zhou, Peter J. Love, Alán Aspuru-Guzik, and Jeremy L. O’Brien. A variational eigenvalue solver on a photonic quantum processor. Nature Communications , 5(1), July 2014

  2. [10]

    Moutinho, Roland Guichard, Vytautas Abramavicius, Aleksander Wennersteen, Gert-Jan Both, Anton Quelle, Caroline de Groot, Gergana V

    Dominik Seitz, Niklas Heim, João P. Moutinho, Roland Guichard, Vytautas Abramavicius, Aleksander Wennersteen, Gert-Jan Both, Anton Quelle, Caroline de Groot, Gergana V. Velikova, Vincent E. Elfving, and Mario Dagrada. Qadence: a differentiable interface for digital and analog ...

  3. [11]

    Schreiber, Jens Eisert, and Johannes Jakob Meyer

    Franz J. Schreiber, Jens Eisert, and Johannes Jakob Meyer. Classical surrogates for quantum learning models. Phys. Rev. Lett. , 131:100803, Sep 2023

  4. [12]

    Resource frugal optimizer for quantum machine learning

    Charles Moussa, Max Hunter Gordon, Michal Baczyk, M Cerezo, Lukasz Cincio, and Patrick J Coles. Resource frugal optimizer for quantum machine learning. Quantum Science and Technology , 8(4):045019, August 2023

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.