REVIEW 6 minor 25 references
Imaginarity as a necessary resource for trainability in QAOA
T0 review · 0 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read For diagonal-cost QAOA with a real mixer, the final mixer-angle gradient is an exact signed sum of cost differences times imaginary coherences, so zero imaginarity forces zero gradient.
desk verdict Sound, honestly-scoped resource-theoretic bound on QAOA's final mixer gradient; the derivation is elementary, but the connection is real and the paper deserves refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the mixer-edge imaginarity $I_{\ell_1}^{(S_B)}(\rho) = \sum_{(z,z')\in S_B} |\mathrm{Im}\,\rho_{zz'}|$, the $\ell_1$-norm of imaginarity restricted to pairs of bitstrings that the mixer Hamiltonian couples. The argument works by differentiating $\rho_\beta = e^{-i\beta H_B}\rho_0 e^{i\beta H_B}$; the identity $\partial_\beta \rho_\beta = -i[H_B,\rho_\beta]$ turns the gradient into a trace of $H_C$ against a commutator, and since $H_C$ is diagonal while a real Hermitian mixer gives a real antisymmetric commutator, only imaginary parts of the state survive. For the transverse-field mixer $H_B=\sum_i X_i$, the mixer edges are exactly pairs at Hamming distance one, so the selected resource is Hamming-one imaginarity and the coarse bound becomes $|\partial_\beta F|\le \Delta_1(H_C)\, I_{\ell_1}^{(1)}(\rho_\beta)$, with $\Delta_1$ the largest one-bit-flip cost change.
What would settle it
An exact numerical search over random diagonal costs, real mixers, and random pre-final states can test the identity directly: compute $\partial_\beta F$ from finite differences and compare it with the right-hand side of Eq. (10). Any instance with zero mixer-edge imaginarity and a nonzero derivative, or any violation of the bound Eq. (11), would refute the paper's central claim; the paper's simulations report none among its sampled Max-Cut instances.
Extended reading notes
Core claim
The central discovery is the exact identity $\partial_\beta F = -\sum_{(z,z')\in S_B} (C_z-C_{z'})(H_B)_{zz'} \, \mathrm{Im}(\rho_\beta)_{zz'}$, where $S_B$ is the set of ordered computational-basis pairs directly coupled by the mixer. This is not an inequality; it is an equality derived from the commutator $[H_C,H_B]$ and the Heisenberg evolution of the state under the final mixer. Taking absolute values gives the bound $|\partial_\beta F| \le \sum_{S_B} |C_z-C_{z'}||(H_B)_{zz'}||\mathrm{Im}(\rho_\beta)_{zz'}|$, so all mixer-edge imaginary coherences vanishing forces the gradient to vanish. The paper stresses that the converse fails because signed contributions can cancel. For terminal noise, the same identity applies with the cost Hamiltonian replaced by the adjoint-propagated observable $H_C^{(\lambda)} = \mathcal{E}_\lambda^\dagger(H_C)$, provided this observable remains diagonal; the paper verifies this for phase-flip, depolarizing, and amplitude-damping channels.
Load-bearing premise
The core identity assumes the objective Hamiltonian is diagonal in the computational basis, the mixer's entries in that basis are real, and the state just before the final mixer is independent of the final angle; the noisy extension further assumes the channel acts once after the whole circuit and that the adjoint-propagated cost stays diagonal, conditions that hold for the three channels tested but are not proven for mid-circuit noise or off-diagonal costs.
Editorial extensions
If this is right
- If the state entering the final mixer has zero imaginarity on all mixer edges, the final mixer angle cannot be trained by gradient methods; the optimizer must adjust other angles or the state must be changed.
- For transverse-field QAOA on Max-Cut, the gradient magnitude is bounded by the maximum graph degree times Hamming-one imaginarity, giving a computable per-instance trainability diagnostic.
- The bound is tight enough to be useful in practice: in the paper's exact simulations, every noisy instance satisfies $|\partial_\beta F_\lambda| \le B^{(\lambda)}_{\mathrm{adj}}(\rho_\beta)$ pointwise, with equality approached when cancellations are small.
- Terminal phase-flip noise at the dephasing point can erase the final state's imaginarity while leaving the gradient unaffected, so final-state imaginarity alone is not a reliable noise monitor.
- Ordinary Hamming-one coherence gives a larger, less specific bound than imaginarity, because the derivative depends only on imaginary off-diagonal components.
Reading between the lines
- Editorial inference: the exact identity suggests a measurement strategy—estimating the signed imaginary coherences on mixer edges would predict not just whether the gradient can be nonzero, but its sign and approximate magnitude, which a standard absolute-value bound cannot do.
- Editorial inference: the same commutator argument should apply to intermediate mixer angles if the cost observable is propagated through later layers; the paper notes this as open, and the likely obstruction is that the propagated observable may acquire off-diagonal terms, so the bound would need correction terms.
- Editorial inference: if typical random instances have exponentially small mixer-edge imaginarity, the final-angle gradient is exponentially suppressed, giving a concrete route from this identity to barren-plateau scaling that the paper's finite-size numerics do not yet probe.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper considers QAOA with a diagonal cost Hamiltonian H_C and a mixer H_B that is real in the computational basis. It proves an exact identity (Eq. (10)) for the derivative of the cost expectation with respect to the final mixer angle: the derivative equals a signed sum over mixer edges of cost differences, mixer matrix elements, and imaginary coherences of the state immediately before the final mixer. This yields the bound Eq. (11) and a corollary bounding the gradient by the maximum one-bit cost difference times the Hamming-one imaginarity. For Max-Cut, the maximum degree of the problem graph enters. The paper extends the bound to terminal noise by writing the noisy expectation with the adjoint-propagated cost observable; Proposition 2 gives a weighted bound that holds whenever that observable remains diagonal, and the authors verify the diagonality explicitly for phase-flip, depolarizing, and amplitude-damping channels. Numerical simulations of depth-one QAOA on small regular-graph Max-Cut instances confirm the bound pointwise. The paper is careful to state that imaginarity is necessary but not sufficient, that the result concerns only the final mixer angle, and that the noisy extension assumes terminal noise.
Significance. The result is a clean, parameter-free necessary condition for the final mixer-angle gradient in QAOA. It identifies mixer-edge imaginarity as the specific resource selected by the mixer, which is a useful diagnostic for when a parameter update is impossible. The use of the adjoint channel to handle terminal noise is elegant and correctly distinguishes the pre-channel state from the final noisy state; the phase-flip lambda=1/2 example makes this distinction concrete. The proof is self-contained, the bounds are explicit and checkable, and the numerical comparison is a genuine verification rather than a fit. The limitations are stated honestly: only one gradient signal is addressed, the noise model is terminal, and no asymptotic scaling claim is made. These limitations are appropriate for the paper's scope.
minor comments (6)
- [Title and Abstract] The title and the abstract's phrase 'trainability in QAOA' are broader than the proven statement, which concerns only the gradient with respect to the final mixer angle under a noiseless or terminal-noise model; consider adding a qualifier such as 'final-angle gradient' so that the scope matches the body.
- [Abstract and Section 3] The abstract states that imaginarity is necessary for a nonzero gradient, but in the terminal-noise setting the relevant quantity is the imaginarity of the pre-channel state with respect to the adjoint-propagated observable, not the final-state imaginarity; the phase-flip lambda=1/2 discussion in Section 3 makes this clear, so a one-sentence clarification in the abstract would prevent a misleading reading.
- [Section 3 (Eq. (24))] The sum defining I^(1)_ell1(rho) is written over d_H(z,z')=1 without stating the ordering convention; since Eq. (9) defines S_B as ordered pairs and Eq. (14) uses that convention, please state explicitly that the Hamming-one sum is over ordered pairs (equivalently twice the unordered sum) to avoid a factor-of-two ambiguity.
- [Appendix (Proof of Proposition 1)] The step 'Pairing the ordered terms ...' is terse; a sentence explaining that the real antisymmetric commutator and the Hermiticity of rho_beta select the imaginary part would make the proof easier to follow.
- [Figure 1 caption] The phrase 'all lambda overlap' is informal; consider replacing it with 'the points for all six noise strengths coincide' for clarity.
- [Section 5 (Corollary 3 proof)] The use of |E_G| for the number of graph edges alongside the edge-set notation E_G is slightly confusing; introducing a symbol such as m for the edge count would improve readability.
Circularity Check
No circularity: the gradient identity is a direct exact derivation, the bounds are triangle inequalities, and the numerics are verification rather than fit.
full rationale
The paper's central claim is an exact analytical identity, not a fitted or self-referential construction. Proposition 1 follows directly from the product rule, cyclicity of the trace, the diagonality of H_C, and the reality of H_B: Eq. (10) is obtained by substituting ∂βρβ = -i[H_B, ρβ] into ∂βF = Tr(H_C ∂βρβ) and pairing ordered terms, and Eq. (11) is the triangle inequality. No parameter is fitted, and the numerical simulations evaluate gradients analytically from the adjoint-channel expression, then compare them pointwise with the bound; this is a verification of a proven inequality, not an inference of the claim from data. Proposition 2 is likewise derived from the definition of the adjoint channel and is explicitly conditional on the adjoint-propagated observable remaining diagonal, a condition the paper states and checks for the three terminal channels used. The paper also explicitly acknowledges limitations, including that phase-flip noise at λ = 1/2 removes final-state imaginarity while the gradient can remain nonzero, and that the bound does not automatically extend to mid-circuit noise or off-diagonal cost Hamiltonians. The only notable self-citation is reference [25], a review co-authored by K. Blekos, but it is used only as background on QAOA benchmarks and is not load-bearing for any derivation. There is no reduction of the result to its own inputs, no equation equivalent to an assumption by construction, and no fitted input renamed as a prediction. The paper is self-contained against external benchmarks and its proofs are explicit.
Assumptions & free parameters
assumptions (4)
- domain assumption Resource theory of imaginarity: real density matrices are free, and the ℓ1 norm of imaginarity is a valid measure (refs [15-17,22]).
- domain assumption The QAOA ansatz prepares the state as in Eq. (4) from the uniform superposition.
- domain assumption The mixer Hamiltonian H_B is Hermitian and real in the computational basis.
- standard math Cyclicity of the trace and the Hilbert-Schmidt adjoint channel definition (Eq. (17)).
Cite this review
Pith. "Pith review of Imaginarity as a necessary resource for trainability in QAOA." pith.science (2026). https://pith.science/paper/OXEQYPIQ
@misc{pith2026260805093,
author = {Pith},
title = {Pith review of: Imaginarity as a necessary resource for trainability in QAOA},
year = {2026},
howpublished = {\url{https://pith.science/paper/OXEQYPIQ}},
note = {Machine review of arXiv:2608.05093}
}
read the original abstract
The quantum approximate optimization algorithm (QAOA) tackles combinatorial problems by tuning a quantum circuit in a classical loop, often guided by gradients. We show that the gradient used to tune the circuit's final parameter is bounded by imaginarity, which weights phase relationships between candidate solutions by how strongly the circuit connects them and how differently the problem scores them. Imaginarity is necessary but not sufficient for a nonzero gradient. We extend the bound to three common noise models and compare it numerically with the gradient in Max-Cut simulations.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Preskill, Quantum computing in the nisq era and be- yond, Quantum2, 79 (2018)
J. Preskill, Quantum computing in the nisq era and be- yond, Quantum2, 79 (2018)
2018
-
[3]
Abbas, A
A. Abbas, A. Ambainis, B. Augustino, A. B¨ artschi, H. Buhrman, C. Coffrin, G. Cortiana, V. Dunjko, D. J. Egger, B. G. Elmegreen, N. Franco, F. Fratini, B. Fuller, J. Gacon, C. Gonciulea, S. Gribling, S. Gupta, S. Had- field, R. Heese, G. Kircher, T. Kleinert, T. Koch, G. Ko- rpas, S. Lenk, J. Marecek, V. Markov, G. Mazzola, S. Mensa, N. Mohseni, G. Nanni...
2024
-
[4]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Bab- bush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nature Communications9, 4812 (2018)
2018
-
[5]
Cerezo, A
M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shal- low parametrized quantum circuits, Nature Communica- tions12, 1791 (2021)
2021
-
[6]
Arrasmith, M
A. Arrasmith, M. Cerezo, P. Czarnik, L. Cincio, and P. J. Coles, Effect of barren plateaus on gradient-free optimization, Quantum5, 558 (2021)
2021
-
[7]
Larocca, S
M. Larocca, S. Thanasilp, S. Wang, K. Sharma, J. Bia- monte, P. J. Coles, L. Cincio, J. R. McClean, Z. Holmes, and M. Cerezo, Barren plateaus in variational quantum computing, Nature Reviews Physics7, 174 (2025)
2025
-
[8]
M. Hillery, Coherence as a resource in decision problems: The Deutsch–Jozsa algorithm and a variation, Phys. Rev. A93, 012111 (2016). 9
work page 2016
Show all 25 references
-
[9]
Naseri, T
M. Naseri, T. V. Kondra, S. Goswami, M. Fellous-Asiani, and A. Streltsov, Entanglement and coherence in the Bernstein–Vazirani algorithm, Phys. Rev. A106, 062429 (2022)
2022
-
[10]
Ahnefeld, T
F. Ahnefeld, T. Theurer, D. Egloff, J. M. Matera, and M. B. Plenio, Coherence as a resource for Shor’s algo- rithm, Phys. Rev. Lett.129, 120501 (2022)
2022
-
[11]
Y.-C. Liu, J. Shang, and X. Zhang, Coherence depletion in quantum algorithms, Entropy21, 260 (2019)
2019
-
[12]
C. Feng, L. Chen, and L.-J. Zhao, Coherence and en- tanglement in Grover and Harrow–Hassidim–Lloyd algo- rithm, Physica A: Statistical Mechanics and its Applica- tions626, 129048 (2023)
2023
-
[13]
A. N. C´ aliz, J. Riu, J. Bosch, P. Torrente, J. Miralles, and A. Riera, A coherent approach to quantum-classical optimization, Communications Physics8, 197 (2025)
2025
-
[14]
Garc ´ ıa Sarmina, J
B. Garc ´ ıa Sarmina, J. Saavedra Benavides, G.-H. Sun, and S.-H. Dong, Probing entanglement and parameter sensitivity in QAOA via quantum fisher information, Quantum Review Letters2, 1 (2026)
2026
-
[15]
Hickey and G
A. Hickey and G. Gour, Quantifying the imaginarity of quantum mechanics, Journal of Physics A: Mathematical and Theoretical51, 414009 (2018)
2018
-
[16]
K.-D. Wu, T. V. Kondra, S. Rana, C. M. Scandolo, G.- Y. Xiang, C.-F. Li, G.-C. Guo, and A. Streltsov, Oper- ational resource theory of imaginarity, Phys. Rev. Lett. 126, 090401 (2021)
2021
-
[17]
K.-D. Wu, T. V. Kondra, S. Rana, C. M. Scandolo, G.-Y. Xiang, C.-F. Li, G.-C. Guo, and A. Streltsov, Resource theory of imaginarity: Quantification and state conver- sion, Phys. Rev. A103, 032401 (2021)
2021
-
[18]
T. Haug, K. Bharti, and D. E. Koh, Pseudorandom uni- taries are neither real nor sparse nor noise-robust, Quan- tum9, 1759 (2025)
2025
-
[19]
Miyazaki and K
J. Miyazaki and K. Matsumoto, Imaginarity-free quan- tum multiparameter estimation, Quantum6, 665 (2022)
2022
-
[20]
L. Ye, Z. Wu, and N. Zhou, Coherence and imaginar- ity as resources in quantum circuit complexity, Advanced Quantum Technologies9, e01007 (2026)
2026
-
[21]
Xuan, Z.-X
D.-P. Xuan, Z.-X. Shen, W. Zhou, S.-M. Fei, and Z.-X. Wang, Quantum-imaginarity-based quantum speed limit, Phys. Rev. A112, 052202 (2025)
2025
-
[22]
Q. Chen, T. Gao, and F. Yan, Measures of imaginarity and quantum state order, Science China Physics, Me- chanics & Astronomy66, 280312 (2023)
2023
-
[23]
Chang, R
W.-L. Chang, R. Wong, W.-Y. Chung, Y.-H. Chen, J.- C. Chen, and A. V. Vasilakos, Quantum speedup for the maximum cut problem (2023), arXiv:2305.16644 [quant- ph]
2023 arXiv
-
[24]
Zhou, S.-T
L. Zhou, S.-T. Wang, S. Choi, H. Pichler, and M. D. Lukin, Quantum approximate optimization algorithm: Performance, mechanism, and implementation on near- term devices, Physical Review X10, 021067 (2020)
2020
-
[25]
Blekos, D
K. Blekos, D. Brand, A. Ceschini, C.-H. Chou, R.-H. Li, K. Pandya, and A. Summer, A review on quan- tum approximate optimization algorithm and its vari- ants, Physics Reports1068, 1–66 (2024)
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.