REVIEW 3 major objections 5 minor 12 references
Evaluation of derivatives using approximate generalized parameter shift rule
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper introduces the approximate generalized parameter shift rule (aGPSR), showing that replacing the full set of spectral-gap equations with a small set of pseudo-gap equations yields derivatives accurate to $O(\alpha^{2K})$ while…
desk verdict aGPSR is a genuinely new construction with clean noiseless numerics, but the practical NISQ savings claim rests on an unanalysed bias-variance tradeoff. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the truncated sine system $F_i = 4\sum_{k=1}^K \sin(\delta_i\gamma_k/2) R'_k$, solved with Cramer's rule so that each $R'_k$ is a ratio of determinants. The paper defines functions $\eta_{ks}$ and $\xi_s = \sum_k \gamma_k \eta_{ks}$, where $\xi_s$ is the effective replacement for the true spectral gap $\Delta_s$. The key identity is the expansion $\xi_s = \Delta_s + O(\alpha^{2K})$ of Eq. (16), obtained by expanding determinant ratios in powers of the shift scale $\alpha$. The error terms factor as products of the form $(\gamma_k^2 - \Delta_s^2)$, which is why choosing pseudo-gaps close to the real spectral gaps suppresses the bias and why including more equations raises the error order to $2K$.
What would settle it
Take a six-qubit neutral-atom generator with strong interactions ($J/\Omega > 1$), fix $K=4$ pseudo-gaps, and compare aGPSR derivatives against exact GPSR at a grid of points $x$ using a finite shot budget such as 1000 shots per expectation value. If the mean relative error exceeds the noiseless $O(\alpha^{8})$ prediction by more than the shot-noise variance allows, or if the variance-minimizing shifts from the paper's Section 3 fail to hold the VQE energy at the target with those shots, the bias-variance trade-off behind the claimed savings does not hold on real hardware.
Extended reading notes
Core claim
The paper's central claim is that the exact generalized parameter shift rule can be replaced, for arbitrary Hermitian generators $\hat G$ with many spectral gaps, by a truncated version with $K \ll S$ equations. Working from the exact GPSR system $F_i = 4\sum_s \sin(\delta_i \Delta_s/2) R_s$, the paper substitutes $K$ arbitrary pseudo-gaps $\{\gamma_k\}$ and solves the smaller system for coefficients $R'_k$. The resulting approximate derivative is $\frac{df}{dx}\big|_K = \sum_s (\Delta_s + O(\alpha^{2K})) R_s$, where $\alpha$ is the common scale of the shifts $\delta_s$; this is Eq. (16) and is the load-bearing identity. It means the approximate derivative inherits the exact GPSR coefficients $R_s$ up to a controllable bias of order $\alpha^{2K}$, independently of the chosen pseudo-gaps to leading order. The paper then demonstrates on a neutral-atom Hamiltonian that $K=4$ to $K=20$ equations reproduce exact derivative curves, and in VQE runs on 3 to 6 qubits $K=4$ reaches the same target energy with 7, 30, 124, and 504 times fewer expectation calls for the analog ansatz.
Load-bearing premise
The method's practical accuracy hinges on choosing parameter shifts small enough that the inherited error from the truncated equations is negligible while the shifted measurement differences still remain accurate with a limited number of shots; the paper's variance analysis starts after that bias is controlled, and its VQE savings are demonstrated on noiseless analytical expectation values.
Editorial extensions
If this is right
- For a given generator spectrum, a fixed small $K$ can hold relative derivative error near a target value while system size grows, whereas exact GPSR needs exponentially many shifted points.
- VQE training on analog and digital ansatze can reach the same optimized energy with orders-of-magnitude fewer expectation calls, moving the practical bottleneck away from gradient evaluation.
- The same truncated-system rule applies to any parametrized unitary $\exp(-ix\hat G/2)$ with an arbitrary Hermitian generator, so quantum machine learning and other gradient-based quantum algorithms inherit the savings.
- Because $K$ is small, the search for variance-minimizing shift values $\delta_s$ becomes tractable, giving a practical route to noise-robust gradient estimation under limited shot budgets.
Reading between the lines
- Editorial extension: the $O(\alpha^{2K})$ expansion suggests an adaptive scheme the paper mentions only as future work: start with $K=1$ or $2$ early in optimization and increase $K$ when gradients become small, which would make the reported savings grow beyond the fixed-$K$ factors.
- Editorial extension: since error terms factor as $(\gamma_k^2 - \Delta_s^2)$, one could target specific dominant spectral gaps if approximate gap information is available from classical pre-simulation, using the error function $Q_K$ to decide how many gaps must be resolved for a target derivative accuracy.
- Editorial caution: the reported 7 to 504 times savings are demonstrated on noiseless analytical expectation values; on real hardware with finite shots, the variance of aGPSR must be compared with exact GPSR variance at equal total shot budget, which the paper leaves as the key open test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an approximate generalized parameter shift rule (aGPSR) for estimating derivatives of expectation values generated by arbitrary generators with many spectral gaps. Instead of solving the full S-by-S GPSR linear system, aGPSR solves a K-by-K system using K 'pseudo-gaps' and approximates the derivative as a weighted sum of K coefficients. The central claim is Eq. (16): with K pseudo-gap equations, the approximate derivative equals the exact GPSR derivative plus an O(alpha^{2K}) bias, where alpha scales the shift sizes. The authors give explicit bias expansions for K=1,2,3, propose pseudo-gap selection by uniform sampling over [0, Delta_max], and illustrate the method on neutral-atom Hamiltonians. In VQE experiments for 3 to 6 qubits, they report reducing the number of expectation calls by factors from 7 to 504 while reaching the same target energy.
Significance. If the central claim holds, aGPSR is potentially valuable: exact GPSR requires 2S shifted expectation values with S growing exponentially in the number of qubits, while aGPSR uses a small, user-selected K. The explicit error expansions for K=1,2,3 and the numerical agreement with exact derivatives for the tested neutral-atom cases are concrete strengths, and the open-source Qadence implementation makes the method testable. However, the practical, hardware-oriented claim of large computational savings is currently supported only by noiseless simulations; the finite-shot bias-variance tradeoff that would determine actual NISQ performance is not analyzed. Because the method's usefulness hinges on this tradeoff, the paper needs additional work before the savings claim can be accepted.
major comments (3)
- The variance analysis is performed only for the exact S-by-S GPSR system, and the text states that the same procedure can be applied to aGPSR 'as usual'. For K<S, however, the bias condition of Eq. (16) requires small shifts delta_s = alpha delta'_s with alpha small. In that regime M_{ij} = 4 sin(delta_i gamma_j / 2) is approximately 2 alpha delta'_i gamma_j, a rank-one matrix, so the inverse M^{-1} has entries growing as alpha^{-(2K-1)} (for K=2, as alpha^{-3}). Since each measured F_s has shot variance 2 sigma_0^2 / N_shots that is independent of alpha, the variance of each R'_k, Var(R'_k) = (2 sigma_0^2 / N_shots) sum_i (M^{-1})_{ki}^2, diverges as alpha^{-2(2K-1)}. The divergent parts cancel in the mean, giving the O(alpha^{2K}) bias, but they do not cancel in the variance. The paper never minimizes bias plus variance jointly for K<S, and the VQE experiments in §4.2 use noiseless analytical expectation values. Consequently, the reported 7–504x reduction in expectation calls does not account for the additional shots required to hold total error fixed, and the central hardware-oriented savings claim is unverified.
- The proof that the approximation error is O(alpha^{2K}) for general K is not a valid induction. The text shows explicit expansions only for K=1,2,3, then argues by assuming the relation for K=S-1 and noting exactness for K=S, concluding that the relation holds for all K. This does not establish a base case and induction step from K to K+1, nor does exactness at K=S imply the O(alpha^{2S-2}) error at K=S-1. Since Eq. (16) is the mathematical foundation for using arbitrarily small K, the general claim needs either a rigorous proof or an explicit restriction to the small-K cases with controlled error bounds. This is a load-bearing point for the method's purported scalability.
- The manuscript does not provide a complete, a priori procedure for choosing K and the pseudo-gap set {gamma_k}. The uniform sampling proposal in §2.3 gives a direction, but in the numerical studies K and the step a are selected per test case (e.g., K=4 with a=4 in the weak-interaction case, K=8 or K=20 in the strong-interaction cases), and Fig. 5b shows that the relative error can be non-monotonic in K (the K=16 curve is singled out as an exception). Without a criterion for selecting K and the pseudo-gaps for a new Hamiltonian, the method's practical reliability and reproducibility are not fully established.
minor comments (5)
- There is an incomplete sentence: 'Having outlined general formulations for aGPSR, we follow up by analysing the relationship between error terms in Eq. 16 and' ends abruptly before the next subsection.
- The heading '2.2 Approximate GPSR' appears twice, and the heading '4.1 Derivative calculations' also appears twice; these duplicated headings should be removed.
- The relative error r defined in Eq. (34) divides by f'_exact(x_i), which can be zero or very small for oscillatory derivative curves; the metric should be regularized or replaced by an absolute error measure to avoid unstable values.
- The caption of Fig. 5 says 'Comparison of derivatives df/dx calculated for different K values with zero initial state |0>' but the figure actually shows scaling of K and S and relative errors; this caption is a copy-paste error and should be corrected.
- Reference [13] is described as a patent application 'submitted' in the references but as 'filed' in the ethics declaration; the wording should be made consistent.
Circularity Check
No significant circularity: aGPSR is a self-contained approximate-derivative construction whose accuracy claim is an analytic expansion result, not a fitted quantity presented as a prediction.
full rationale
The paper's central derivation is self-contained and non-circular. It starts from the exact GPSR formula of Ref. [6], treats it as ground truth, and constructs aGPSR by truncating the linear system to K equations with pseudo-gaps. Equations (5)-(16) are derived by linear algebra (Cramer's rule) and Taylor expansion in the small shift parameter alpha. The key claim that xi_s = Delta_s + O(alpha^{2K}) is an analytic error bound derived from the determinant structure, not an assumption equivalent to the target derivative. No parameter is fitted to the derivative values or to the VQE target energy: K is a user-chosen truncation size, gamma_k and delta_k are method parameters, and the VQE energy is used only as a benchmark for comparing methods. Self-citations appear (Ref. [6] by co-author Elfving; Refs. [10], [13] by Pasqal authors), but none is load-bearing in a circular sense: Ref. [6] provides the external exact GPSR starting point, and Refs. [10], [13] concern implementation and patent disclosure, not the derivation's validity. The skeptical concern about finite-shot bias-variance trade-off for K < S is a genuine correctness or validation gap, but it is not circularity: an unverified practical premise is different from a claim that reduces by construction to its own inputs. The paper even acknowledges that the variance minimization section analyzes the exact S-by-S case and states the aGPSR extension follows 'as usual', which is a completeness limitation rather than a circular step. Therefore the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- Number of pseudo-gap equations K =
K=1 (digital VQE), K=4 (analog VQE), K=2,4,8,20 (derivative benchmarks)
- Pseudo-gap set {gamma_k} =
e.g., uniform step a=4 in Figure 3; otherwise not fully specified
- Shift values {delta_s} =
Not specified for VQE; equidistant in [pi/4, pi/2] in Figure 1; optimized in Section 3
assumptions (4)
- domain assumption The exact GPSR derivative formula (Eq. 2) from Ref [6] is correct for arbitrary Hermitian generators.
- domain assumption Shifted expectation values f(x +/- delta_s) are independent and identically distributed normal variables with constant variance sigma_0^2.
- ad hoc to paper Pseudo-gaps can be chosen arbitrarily and the leading O(alpha^{2K}) error remains representative for finite shifts.
- domain assumption The spectral gap distribution can be approximated by uniform sampling on [0, Delta_max].
invented entities (1)
-
Pseudo-gaps {gamma_k}
Cite this review
Pith. "Pith review of Evaluation of derivatives using approximate generalized parameter shift rule." pith.science (2026). https://pith.science/paper/MXUKPXPS
@misc{pith2026250518090,
author = {Pith},
title = {Pith review of: Evaluation of derivatives using approximate generalized parameter shift rule},
year = {2026},
howpublished = {\url{https://pith.science/paper/MXUKPXPS}},
note = {Machine review of arXiv:2505.18090}
}
read the original abstract
Parameter shift rules are instrumental for derivatives estimation in a wide range of quantum algorithms, especially in the context of Quantum Machine Learning. Application of single-gap parameter shift rule is often not possible in algorithms running on noisy intermediate-scale quantum (NISQ) hardware due to noise effects and interaction between device qubits. In such cases, generalized parameter shift rules must be applied yet are computationally expensive for larger systems. In this paper we present the approximate generalized parameter rule (aGPSR) that can handle arbitrary device Hamiltonians and provides an accurate derivative estimation while significantly reducing the computational requirements. When applying aGPSR for a variational quantum eigensolver test case ranging from 3 to 6 qubits, the number of expectation calls is reduced by a factor ranging from 7 to 504 while reaching the exact same target energy, demonstrating its huge computational savings capabilities.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Solving nonlinear differential equations with differentiable quantum circuits
Oleksandr Kyriienko, Annie E Paine, and Vincent E Elfving. Solving nonlinear differential equations with differentiable quantum circuits. Physical Review A , 103(5):052416, 2021
2021
-
[2]
Savvas Varsamopoulos, Evan Philip, Vincent E Elfving, Herman WT van Vlijmen, Sairam Menon, Ann Vos, Natalia Dyubankova, Bert Torfs, and Anthony Rowe. Quantum extremal learning. Quantum Machine Intelligence , 6(2):42, 2024
work page 2024
-
[3]
K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii. Quantum circuit learning. Phys. Rev. A , 98:032309, 2018
work page 2018
-
[4]
Evaluating analytic gradients on quantum hardware
Maria Schuld, Ville Bergholm, Christian Gogolin, Josh Izaac, and Nathan Killoran. Evaluating analytic gradients on quantum hardware. Phys. Rev. A , 99:032331, 2019
work page 2019
-
[5]
Yuki Ishiyama, Ryutaro Nagai, Shunsuke Mieda, Yuki Takei, Yuichiro Minato, and Yutaka Natsume. Noise-robust optimization of quantum machine learning models for polymer properties using a simulator and validated on the ionq quantum computer. Scientific Reports , 12(1):19003, 2022
work page 2022
-
[6]
Generalized quantum circuit differentiation rules
Oleksandr Kyriienko and Vincent E Elfving. Generalized quantum circuit differentiation rules. Physical Review A , 104(5):052417, 2021
work page 2021
-
[7]
Richard B. Lehoucq, Danny C. Sorensen, and Chao Yang. ARPACK Users' Guide: Solution of Large-Scale Eigenvalue Problems with Implicitly Restarted Arnoldi Methods . SIAM, 1998
work page 1998
-
[8]
Quantum computing with neutral atoms
Lo \" i c Henriet, Lucas Beguin, Adrien Signoles, Thierry Lahaye, Antoine Browaeys, Georges-Olivier Reymond, and Christophe Jurczak. Quantum computing with neutral atoms. Quantum , 4:327, September 2020
work page 2020
Show all 12 references
-
[9]
Love, Alán Aspuru-Guzik, and Jeremy L
Alberto Peruzzo, Jarrod McClean, Peter Shadbolt, Man-Hong Yung, Xiao-Qi Zhou, Peter J. Love, Alán Aspuru-Guzik, and Jeremy L. O’Brien. A variational eigenvalue solver on a photonic quantum processor. Nature Communications , 5(1), July 2014
2014
-
[10]
Moutinho, Roland Guichard, Vytautas Abramavicius, Aleksander Wennersteen, Gert-Jan Both, Anton Quelle, Caroline de Groot, Gergana V
Dominik Seitz, Niklas Heim, João P. Moutinho, Roland Guichard, Vytautas Abramavicius, Aleksander Wennersteen, Gert-Jan Both, Anton Quelle, Caroline de Groot, Gergana V. Velikova, Vincent E. Elfving, and Mario Dagrada. Qadence: a differentiable interface for digital and analog ...
2025
-
[11]
Schreiber, Jens Eisert, and Johannes Jakob Meyer
Franz J. Schreiber, Jens Eisert, and Johannes Jakob Meyer. Classical surrogates for quantum learning models. Phys. Rev. Lett. , 131:100803, Sep 2023
2023
-
[12]
Resource frugal optimizer for quantum machine learning
Charles Moussa, Max Hunter Gordon, Michal Baczyk, M Cerezo, Lukasz Cincio, and Patrick J Coles. Resource frugal optimizer for quantum machine learning. Quantum Science and Technology , 8(4):045019, August 2023
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.