REVIEW 3 major objections 3 minor 74 references
A fermionic, particle-number-preserving circuit family claims to be trainable at every particle number k — gradient variance Θ(k²/n⁵) — while staying classically hard to simulate, with a parallel parameter-shift rule cutting training cost b
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:17 UTC pith:XUP4FTUH
load-bearing objection The brick-wall trainability proof has a load-bearing gradient-formula error, but the parallel PSR and the honest hardness ladder make this worth refereeing. the 3 major comments →
Scalable Quantum Machine Learning: Trainability, Expressivity and Efficiency
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's core discovery is that interleaving Rz phase gates with RBS gates changes the dynamical Lie algebra from so(n) to u(n) without leaving the passive-FLO / particle-number-preserving framework, so the architecture directly parametrizes U(n) via a rectangular mesh of nearest-neighbor rotations (making Haar initialization possible) and yields a two-body readout gradient variance Θ(k²/n⁵), polynomial in n at every particle number k — while the non-Gaussian magic-state input keeps the circuit's output classically hard at k=Θ(n) and at operating points k=60 and k=90. A second discovery is the multi-layer parallel parameter-shift rule, which exploits the commuting structure of disjoint ga
What carries the argument
The load-bearing objects are the unitary brick-wall circuit — n columns of nearest-neighbor RBS(θ)·(Rz(ϕ)⊗I) gates plus a final Rz layer, which directly parametrizes U(n) in the single-particle sector via the Givens/QR-type decomposition — and its all-to-all counterpart, the unitary butterfly. Trainability is carried by a second-moment (Haar-twirl) computation on the antisymmetric subspace Λ²C^n, which decomposes Λ²⊗Λ² into three irreducibles of dimension Θ(n⁴); combined with a locality lemma bounding the commutator of any single gate generator with the delocalized 2-RDM by Θ(k²/n), this yields the Θ(k²/n⁵) variance. Gradient efficiency is carried by the multi-layer parallel parameter-shift
Load-bearing premise
The proof of the unconditional Θ(k²/n⁵) gradient variance assumes that the derivative with respect to a gate buried in the brick-wall can be reduced to a fixed-generator commutator against the full Haar-random unitary W, but in a product of Givens rotations the derivative actually involves a generator conjugated by the partial circuit before the gate, so the Weingarten step in Eq. (28) does not follow as written.
What would settle it
For n=8, 16, and 32 with k=2, compute the exact gradient variance of the two-body correlator under Haar-initialized brick-wall parameters by full state-vector simulation; if Var[∂L/∂θ] deviates from c·k²/n⁵ by more than a constant factor, the unconditional trainability claim fails. Independently, at n=8, k=2, test the 2k-shift parallel-PSR estimator against exact gradients for random circuit parameters; the paper itself notes that the k-shift estimator is biased by an O(1) relative error at this size, so a machine-precision match for the 2k-shift estimator is a direct check of the unbiasedness
If this is right
- If the central claim is correct, no exponential barren plateau occurs for two-body readouts at any particle number, including k=n/2 where average-case Fermion-Sampling hardness applies.
- At the concrete operating points k=60 (sampling) and k=90 (two-body expectation values), best-known classical simulation exceeds 10²⁴ operations at every n, while a gradient step costs only k(8n+4) or 8k log n circuit evaluations.
- The parallel parameter-shift rule is exact in expectation for any particle-number-preserving observable, with the shift coefficients depending only on k and computed once in O(k²) classical time.
- The architecture is claimed to escape the polynomial-DLA-simulability tension: a polynomial dynamical Lie algebra u(n) does not imply classical simulability when the readout is sample-based rather than a g-compatible expectation value.
- Training cost scales as O(kn) for the brick-wall and O(k log n) for the butterfly, both far below the superpolynomial Hilbert-space dimension, so quantum advantage is not precluded by trainability.
Where Pith is reading between the lines
- If Conjecture 17 (an approximate 2-design on Λ²C^n) is resolved, the butterfly's sharp Θ(k²/n⁵) rate would become unconditional; the appendix's parity-decoupling recursion suggests the spectral-gap method may extend to that sector.
- A testable extension: for constant k, the claimed Θ(1/n⁵) gradient-variance scaling could be verified on small systems via exact state-vector simulation, and the parallel PSR estimator should reproduce exact gradients to machine precision at n≈8–32 with k=2–3 (the paper itself reports a biased k-shift alternative at that size).
- The Dicke-state encoding noted in the paper offers a single-state route to both sampling and expectation-value hardness with form-rank exactly k, potentially unifying the paired and triplet magic-block constructions at the price of deeper state preparation.
- The factor-3n/(8k) reduction in circuit evaluations may compose with stochastic gradient methods to yield further per-step savings, since the parallel PSR's random sign vectors already act as a variance-controlled random direction probe.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a k-particle, particle-number-preserving quantum neural network family (the 'unitary brick-wall' and the 'unitary butterfly') built from RBS and Rz gates, with a magic-state data encoding. It claims three simultaneous properties: (i) trainability, via DLA u(n) and, for the brick-wall, Haar initialization, with gradient variance Θ(k²/n⁵) for two-body correlators (unconditional for the brick-wall, conditional for the butterfly); (ii) expressivity, via a ladder of classical-hardness statements for sampling and two-body expectation values controlled by the particle number k and calibrated at B=30 magic blocks (k=60 and 90); and (iii) efficiency, via a multi-layer parallel parameter-shift rule computing all O(n²) gradients from O(kn) (brick-wall) or O(k log n) (butterfly) circuit evaluations. The paper also includes a generalization bound and a discussion of the Cerezo et al. simulability framework.
Significance. The multi-layer parallel parameter-shift rule (Theorem 19) and the explicit hardness calibration are genuinely useful contributions and are largely independent of the trainability proof. The paper is also unusually honest: Conjecture 17, Conjecture 33, the butterfly reachability caveat, and the form-rank-3 fragility are all stated explicitly. However, the central trainability claim is not established by the arguments given. The brick-wall gradient formula used in Theorem 10 is algebraically incorrect; the key localization lemma contains an unproved and in fact false equidistribution assertion; and the butterfly's unconditional lower bound applies a 2-design theorem at the wrong moment level. Because the paper's headline contribution is precisely the simultaneous trainability-plus-hardness guarantee, these gaps are load-bearing.
major comments (3)
- [§IV.B, Theorem 10, Eq. (28)] Eq. (28) is not the derivative of L for a gate in the interior of the Givens mesh. Writing W = A e^{-iθG} B, the exact derivative is i tr_{Λ²}(G^(2)[A^(2)† M^(2) A^(2), B^(2) ρ₂(x) B^(2)†]) (up to sign), with the partial circuits A and B appearing. It is not i tr_{Λ²}([G^(2), Λ²(W)† M^(2) Λ²(W)] ρ₂(x)). The sentence preceding Eq. (28) ('we do not split W into independent factors') addresses the distribution of A and B, not their algebraic presence in the derivative. Equations (29)–(40) therefore compute the Haar average of a different quantity. Since Theorem 10 and the abstract's 'unconditional Θ(k²/n⁵)' claim rest on this computation, the central trainability result is unproven.
- [§IV.B, Lemma 8, Eq. (23)] The 'row equidistribution' step is false as stated. For H the Walsh–Hadamard matrix, the column of Λ²(H) indexed by (c,d) is supported only on row pairs {a,b} satisfying ⟨a⊕b, c⊕d⟩=1; the complementary rows have exactly zero entries. Thus the diagonal of Λ²(H)ρ_in² Λ²(H)† is not uniform over the Λ² basis. A concrete n=4 example with ρ_in = |{0,1}⟩⟨{0,1}| gives diagonal entries 1/4 on four of six rows and 0 on the other two, contradicting the claimed '1/d ∥ρ∥²(1+o(1)) uniformly in P'. Consequently the proof of Eq. (23), and hence Step 4 of Theorem 10, lacks support; the Θ(k²/n) commutator norm may be true, but it is not established by the argument given.
- [§IV.C, Theorem 15 vs. Theorem 13/Conjecture 17] Theorem 15 claims an unconditional polynomial lower bound from Theorem 13's approximate-2-design property. But the gradient variance requires the two-copy second moment of Λ²(W), i.e. E[(Λ²(W)†)⊗2 M⊗2 (Λ²(W))⊗2] on Λ²C^n ⊗ Λ²C^n. Theorem 13 controls only the one-copy adjoint action E[Λ²(W) Y Λ²(W)†] on End(Λ²C^n), whose fixed space is one-dimensional, not the three-dimensional decomposition in Eq. (30). The needed statement is exactly Conjecture 17. Thus Theorem 15's unconditional lower bound is not derived from the provided spectral-gap result, and the butterfly's 'unconditional' no-barren-plateau claim is also unsupported as written.
minor comments (3)
- [Abstract and §IV summary] The abstract says the brick-wall 'directly parametrizes U(n) via Givens rotations, enabling Haar initialization' and that this gives Θ(k²/n⁵) unconditionally. Given the state of the proof, the word 'unconditional' is only accurate for the architectural parametrization, not for the variance bound. The summary in §IV.D repeats the same conflation.
- [§V, Theorem 19] The parallel parameter-shift rule is presented clearly, and the Fourier-degree reduction to 2k is convincing. However, the practical variance of the random-sign estimator is not analyzed; the statement that the coefficients carry only a poly(k) overhead would benefit from a bound on the estimator variance, since the gradient-variance guarantee alone does not control the number of shots needed per gradient step.
- [General] There are several notational overloads: 'ρ₂,g2' is introduced without a precise definition of the traceless projection subscript, and 'Λ²(W)' is used both for the exterior square of the single-particle unitary and for the induced action on two-particle operators. These should be disambiguated in a revision.
Circularity Check
No circularity found: the trainability, gradient-cost, and hardness claims are derived from explicit circuit structure and independent external results, not from their own conclusions.
full rationale
I walked the claimed derivation chain. The brick-wall/butterfly universality and DLA u(n) claims are proved in Propositions 4 and 5 by explicit Givens/QR and Lie-closure arguments (with [34,35] cited only for the standard commutator fact), so Haar initialization is a consequence, not an assumed input. Theorem 10's Θ(k²/n⁵) is computed from the Haar second-moment identity on Λ²Cⁿ, the 2-RDM norm, and generator locality; none of these constants is fitted to the target variance, and the result is not the same as the initialization assumption. The butterfly unconditional lower bound is proved via the spectral-gap calculation in Appendix A, and the sharp rate is explicitly labeled conditional on Conjecture 17 rather than being asserted unconditionally. The parallel PSR estimator is proved from the Fourier-degree bound and an explicit quadrature condition (Theorem 19), so the k(8n+4) cost is derived, not assumed. Hardness claims rest on independent results (Fermion Sampling [3], FLO-extent algorithms [1], mixed-Pfaffian tractability [2], hyper-Pfaffian VNP-completeness [54]); the k=60/90 operating points are calibrations chosen after those results, not predictions fitted to them. The paper's own citations to prior RBS/QML work [23,39] supply background architecture but are not the load-bearing step for the new results; moreover the universality/DLA facts are re-derived in the text. There is a mathematical concern in Eq. (28) about conjugating the gate generator by partial circuits instead of the full W, but that is a correctness/derivation gap, not a case of an output being equal to an input by construction, so it does not count as circularity under the stated rules.
Axiom & Free-Parameter Ledger
free parameters (5)
- particle number k =
60 (sampling), 90 (expectation values)
- magic block type =
|Ψ4⟩ paired vs |Ψ6⟩ triplet
- block count B =
30
- data encoding scale f =
φ_i=π x_i/max|x_i| (example)
- shift nodes r_q and weights a_q =
Chebyshev-Gauss nodes
axioms (8)
- standard math Weingarten formula and Schur’s lemma on U(n) apply to the full circuit unitary W
- domain assumption Haar initialization maps to Givens angles of the brick-wall exactly and efficiently
- domain assumption Fermion Sampling hardness assumptions of Oszmaniec et al. [3] (polynomial-hierarchy noncollapse, avg-case #P-hardness)
- domain assumption FLO-extent algorithms [1], mixed-Pfaffian estimator [2], hyper-Pfaffian VNP-completeness [54]
- ad hoc to paper Conjecture 17: butterfly is an O(1/n)-approximate 2-design on Λ²C^n
- ad hoc to paper Conjecture 33: sparse-input anticoncentration and worst-to-average-case reduction
- ad hoc to paper Lemma 8 row equidistribution of the spread 2-RDM under Walsh-Hadamard
- domain assumption Generic x non-degeneracy
read the original abstract
Designing scalable parameterized quantum circuits for machine learning faces three fundamental obstacles: barren plateaus that prevent gradient-based training, the absence of provable guarantees that the learned function class is classically hard, and prohibitive circuit evaluations per gradient step. We propose the unitary brick-wall: a $k$-particle fermionic architecture for nearest-neighbor hardware, combining Reconfigurable Beam Splitter gates with interleaved single-qubit phase gates and a non-Gaussian magic-state encoding, where the particle number $k$ is a tunable dial trading classical simulation hardness against training cost. Trainable. The brick-wall has dynamical Lie algebra $\mathfrak{u}(n)$ and directly parametrizes $U(n)$ via Givens rotations, enabling Haar initialization. Two-body correlator readouts achieve gradient variance $\Theta(k^2/n^5)$. Expressive. Classical hardness is controlled by the particle number $k$: best-known classical algorithms for sampling and for two-body expectation values run in time $2^{\Theta(k)}\mathrm{poly}(n)$, worst-case #P-hardness holds at $k = n^{\epsilon}$, and average-case hardness applies at $k = \Theta(n)$. Efficient. A multi-layer parallel parameter-shift rule computes all $O(n^2)$ gradients from $k(8n+4)$ circuit evaluations per gradient step, a factor $3n/(8k)$ reduction over the $3n^2$ evaluations required by the standard parameter-shift rule. The unitary butterfly variant targets all-to-all hardware, with depth $2\log n$ and $n\log n$ parameters. It achieves similar hardness guarantees at $8k\log n$ evaluations per gradient step, the same factor $3n/(8k)$ reduction. Its trainability is established at two levels: the absence of exponential barren plateaus is unconditional, whereas the sharp $\Theta(k^2/n^5)$ rate holds under a two-particle approximate-2-design conjecture.
Reference graph
Works this paper leans on
-
[1]
MMD), which inherits polynomial gradient variance from the local-observable analysis applied to Mercer features of the kernel
Trainability of the readouts For sample-based readouts, trainability is controlled by the sample complexity of the distributional loss (e.g. MMD), which inherits polynomial gradient variance from the local-observable analysis applied to Mercer features of the kernel. For expectation-value readouts, the situation depends critically on body number, and the ...
-
[2]
fermion sampling with less magic
andb∈R C are classical parame- ters, trained alongside the quantum parameters (θ,ϕ). Apply softmax to obtain class probabilities, and train against labels with cross-entropy. This is a hybrid ar- chitecture: the quantum circuit computes the two-body correlator vector, and a classical linear head maps it to predictions. In practice one can restrict to a su...
2049
-
[3]
Reardon-Smith, M
O. Reardon-Smith, M. Oszmaniec, and K. Korzekwa, Im- proved simulation of quantum circuits dominated by free fermionic operations, Quantum 8, 1549 (2024)
2024
-
[4]
C. Oh, M. Oszmaniec, O. Reardon-Smith, and Z. Zimbor´ as, Classical simulation of free-fermionic dy- namics and quantum chemistry with magic input, arxiv:2604.26813 (2026)
Pith/arXiv arXiv 2026
-
[5]
Oszmaniec, N
M. Oszmaniec, N. Dangniam, M. E. S. Morales, and Z. Zimbor´ as, Fermion sampling: a robust quantum com- putational advantage scheme using fermionic linear op- tics and magic input states, PRX Quantum3, 020328 (2022)
2022
-
[6]
A. W. Harrow, A. Hassidim, and S. Lloyd, Quantum al- gorithm for linear systems of equations, Physical Review Letters103, 150502 (2009)
2009
-
[7]
Kerenidis and A
I. Kerenidis and A. Prakash, Quantum recommendation systems, inInnovations in Theoretical Computer Science (ITCS)(2017) pp. 49:1–49:21
2017
-
[8]
A. Gily´ en, Y. Su, G. H. Low, and N. Wiebe, Quantum singular value transformation and beyond: exponential improvements for quantum matrix arithmetics, inPro- ceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing (STOC)(2019) pp. 193–204, arXiv:1806.01838
Pith/arXiv arXiv 2019
-
[9]
S. Lloyd, M. Mohseni, and P. Rebentrost, Quantum algo- rithms for supervised and unsupervised machine learning, arXiv:1307.0411 (2013)
Pith/arXiv arXiv 2013
-
[10]
Lloyd, M
S. Lloyd, M. Mohseni, and P. Rebentrost, Quantum prin- cipal component analysis, Nature Physics10, 631 (2014)
2014
-
[11]
Kerenidis, J
I. Kerenidis, J. Landman, A. Luongo, and A. Prakash, q- means: A quantum algorithm for unsupervised machine learning, inAdvances in Neural Information Processing Systems (NeurIPS), Vol. 32 (2019)
2019
-
[12]
E. Tang, A quantum-inspired classical algorithm for rec- ommendation systems, inProceedings of the 51st An- nual ACM SIGACT Symposium on Theory of Computing (STOC)(2019) pp. 217–228, arXiv:1807.04271
Pith/arXiv arXiv 2019
-
[13]
E. Tang, Quantum principal component analysis only achieves an exponential speedup because of its state preparation assumptions, Physical Review Letters127, 060503 (2021), arXiv:1811.00414
Pith/arXiv arXiv 2021
-
[14]
N.-H. Chia, A. Gily´ en, T. Li, H.-H. Lin, E. Tang, and C. Wang, Sampling-based sublinear low-rank ma- trix arithmetic framework for dequantizing quantum machine learning, Journal of the ACM69, 1 (2022), arXiv:1910.06151
Pith/arXiv arXiv 2022
-
[15]
H. Zhao, A. Zlokapa, H. Neven, R. Babbush, J. Preskill, J. R. McClean, and H.-Y. Huang, Exponential quan- tum advantage in processing massive classical data, arXiv preprint arXiv:2604.07639 (2026)
Pith/arXiv arXiv 2026
-
[16]
Mitarai, M
K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, Quantum circuit learning, Physical Review A98, 032309 (2018)
2018
-
[17]
Schuld and N
M. Schuld and N. Killoran, Quantum machine learning 33 in feature Hilbert spaces, Physical Review Letters122, 040504 (2019)
2019
-
[18]
Benedetti, E
M. Benedetti, E. Lloyd, S. Sack, and M. Fiorentini, Pa- rameterized quantum circuits as machine learning mod- els, Quantum Science and Technology4, 043001 (2019)
2019
-
[19]
P´ erez-Salinas, A
A. P´ erez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, Data re-uploading for a universal quantum classifier, Quantum4, 226 (2020)
2020
-
[20]
Biamonte, P
J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, Quantum machine learning, Na- ture549, 195 (2017)
2017
-
[21]
Cerezo, A
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio,et al., Variational quantum algorithms, Nature Reviews Physics3, 625 (2021)
2021
-
[22]
E. Farhi and H. Neven, Classification with quantum neu- ral networks on near term processors, arXiv:1802.06002 (2018)
Pith/arXiv arXiv 2018
-
[23]
Havl ´ ıˇ cek, A
V. Havl ´ ıˇ cek, A. D. C´ orcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, Super- vised learning with quantum-enhanced feature spaces, Nature567, 209 (2019)
2019
-
[24]
Abbas, D
A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, The power of quantum neural networks, Nature Computational Science1, 403 (2021)
2021
-
[25]
Landman, N
J. Landman, N. Mathur, Y. Y. Li, M. Strahm, S. Kazdaghli, A. Prakash, and I. Kerenidis, Quantum methods for neural networks and application to medical image classification, Quantum6, 881 (2022)
2022
-
[26]
Dunjko, J
V. Dunjko, J. M. Taylor, and H. J. Briegel, Quantum- enhanced machine learning, Physical Review Letters 117, 130501 (2016)
2016
-
[27]
Jerbi, C
S. Jerbi, C. Gyurik, S. Marshall, H. Briegel, and V. Dun- jko, Parametrized quantum policies for reinforcement learning, inAdvances in Neural Information Processing Systems (NeurIPS), Vol. 34 (2021) pp. 28362–28375
2021
-
[28]
E. A. Cherrat, S. Raj, I. Kerenidis, A. Shekhar, B. Wood, J. Dee, S. Chakrabarti, R. Chen, D. Herman, S. Hu, et al., Quantum deep hedging, Quantum7, 1191 (2023)
2023
-
[29]
Schuld, I
M. Schuld, I. Sinayskiy, and F. Petruccione, The quest for a quantum neural network, Quantum Information Pro- cessing13, 2567 (2014)
2014
-
[30]
J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Bab- bush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nature Communications9, 4812 (2018)
2018
-
[31]
I. Cong, S. Choi, and M. D. Lukin, Quantum convolu- tional neural networks, Nature Physics15, 1273 (2019)
2019
-
[32]
Pesah, M
A. Pesah, M. Cerezo, S. Wang, T. Volkoff, A. T. Sorn- borger, and P. J. Coles, Absence of barren plateaus in quantum convolutional neural networks, Physical Review X11, 041011 (2021)
2021
-
[33]
Grant, M
E. Grant, M. Benedetti, S. Cao, A. Hallam, J. Lock- hart, V. Stojevic, A. G. Green, and S. Severini, Hierar- chical quantum classifiers, npj Quantum Information4, 65 (2018)
2018
-
[34]
J. J. Meyer, M. Mularski, E. Gil-Fuster, A. A. Mele, F. Arzani, A. Wilms, and J. Eisert, Exploiting symmetry in variational quantum machine learning, PRX Quantum 4, 010328 (2023)
2023
-
[35]
Larocca, F
M. Larocca, F. Sauvage, F. M. Sbahi, G. Verdon, P. J. Coles, and M. Cerezo, Group-invariant quantum machine learning, PRX Quantum3, 030341 (2022)
2022
-
[36]
Fontana, D
E. Fontana, D. Herman, S. Chakrabarti, N. Kumar, R. Yalovetzky, J. Heredge, S. H. Sureshbabu, and M. Pis- toia, Characterizing barren plateaus in quantum ans¨ atze with the adjoint representation, Nature Communications 15, 7171 (2024)
2024
-
[37]
Larocca, P
M. Larocca, P. Czarnik, K. Sharma, G. Muraleedharan, P. J. Coles, and M. Cerezo, Diagnosing barren plateaus with tools from quantum optimal control, Quantum6, 824 (2022)
2022
-
[38]
Cerezo, M
M. Cerezo, M. Larocca, D. Garc ´ ıa-Mart ´ ın, N. L. Diaz, P. Braccia, E. Fontana, M. S. Rudolph, P. Bermejo, A. Ijaz, S. Thanasilp,et al., Does provable absence of bar- ren plateaus imply classical simulability?, Nature Com- munications16, 7907 (2025)
2025
-
[39]
Bermejo, P
P. Bermejo, P. Braccia, M. S. Rudolph, Z. Holmes, L. Cincio, and M. Cerezo, Quantum convolutional neu- ral networks are (effectively) classically simulable, PRX Quantum 7, 020304 (2026)
2026
-
[40]
E. R. Anschuetz, A. Bauer, B. T. Kiani, and S. Lloyd, Efficient classical algorithms for simulating symmetric quantum systems, Quantum7, 1189 (2023)
2023
-
[41]
I. Kerenidis, J. Landman, and N. Mathur, Quantum and classical algorithms for orthogonal neural networks, arXiv:2106.07198 (2022)
Pith/arXiv arXiv 2022
-
[42]
E. A. Cherrat, I. Kerenidis, N. Mathur, J. Landman, M. Strahm, and Y. Y. Li, Quantum vision transformers, Quantum8, 1265 (2024)
2024
-
[43]
Thakkar, S
S. Thakkar, S. Kazdaghli, N. Mathur, I. Kerenidis, A. J. Ferreira-Martins, and S. Brito, Improved financial fore- casting via quantum machine learning, Quantum Ma- chine Intelligence6, 27 (2024)
2024
-
[44]
Kazdaghli, I
S. Kazdaghli, I. Kerenidis, J. Kieckbusch, and P. Teare, Improved clinical data imputation via classical and quan- tum determinantal point processes, eLife12, (RP89947) (2023)
2023
-
[45]
N. Mathur, P. K. Barkoutsos, M. Yamada, M. Roetteler, and I. Kerenidis, Scalable on-hardware training of quan- tum neural networks and application to clinical data im- putation, arXiv:2606.03517 (2026)
Pith/arXiv arXiv 2026
-
[46]
Jozsa and A
R. Jozsa and A. Miyake, Matchgates and classical simu- lation of quantum circuits, Proceedings of the Royal So- ciety A464, 3089 (2008)
2008
-
[47]
Monbroussou, E
L. Monbroussou, E. Z. Mamon, J. Landman, A. B. Grilo, R. Kukla, and E. Kashefi, Trainability and expressivity of Hamming-weight preserving quantum circuits for ma- chine learning, Quantum9, 1745 (2025)
2025
-
[48]
Schuld, V
M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Kil- loran, Evaluating analytic gradients on quantum hard- ware, Physical Review A99, 032331 (2019)
2019
-
[49]
Gacon, C
J. Gacon, C. Zoufal, G. Carleo, and S. Woerner, Si- multaneous perturbation stochastic approximation of the quantum fisher information, inQuantum, Vol. 5 (2021) p. 567
2021
-
[50]
Wierichs, J
D. Wierichs, J. Izaac, C. Wang, and C.-Y. Lin, General parameter-shift rules for quantum gradients, Quantum6, 677 (2022)
2022
-
[51]
B. Coyle, S. Raj, V. Umathe, E. A. Cher- rat, and E. Kashefi, Adaptive directional gradients for parameterised quantum circuits, arXiv preprint arXiv:2606.09734 (2026)
Pith/arXiv arXiv 2026
-
[52]
Coyle, S
B. Coyle, S. Raj, N. Mathur, E. A. Cherrat, N. Jain, S. Kazdaghli, and I. Kerenidis, Training-efficient density quantum machine learning, npj Quantum Information 11, 172 (2025)
2025
-
[53]
B¨ artschi and S
A. B¨ artschi and S. Eidenbenz, Short-depth circuits for 34 Dicke state preparation, IEEE International Conference on Quantum Computing and Engineering (QCE) , 87 (2022)
2022
-
[54]
R. M. S. Farias, T. O. Maciel, G. Camilo, R. Lin, S. Ramos-Calderer, and L. Aolita, Quantum encoder for fixed-Hamming-weight subspaces, Physical Review Ap- plied23, 044014 (2025)
2025
-
[55]
Schuld, R
M. Schuld, R. Sweke, and J. J. Meyer, Effect of data encoding on the expressive power of variational quantum- machine-learning models, Physical Review A103, 032430 (2021)
2021
-
[56]
C. Ikenmeyer and M. Walter, Hyperpfaffians and geomet- ric complexity theory, arxiv:1912.09389 (2019)
Pith/arXiv arXiv 1912
-
[57]
Hebenstreit, R
M. Hebenstreit, R. Jozsa, B. Kraus, S. Strelchuk, and M. Yoganathan, All pure fermionic non-Gaussian states are magic states for matchgate computations, Physical Review Letters123, 080503 (2019)
2019
-
[58]
L. G. Valiant, Quantum computers that can be simu- lated classically in polynomial time, inProceedings of the 33rd Annual ACM Symposium on Theory of Computing (STOC)(2001) pp. 114–123
2001
-
[59]
B. M. Terhal and D. P. DiVincenzo, Classical simulation of noninteracting-fermion quantum circuits, Physical Re- view A65, 032325 (2002)
2002
-
[60]
Knill,Fermionic linear optics and matchgates, Tech
E. Knill,Fermionic linear optics and matchgates, Tech. Rep. LAUR-01-4472 (Los Alamos National Laboratory,
-
[61]
Cerezo, A
M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, Cost function dependent barren plateaus in shal- low parametrized quantum circuits, Nature Communica- tions12, 1791 (2021)
2021
-
[62]
J. W. Cooley and J. W. Tukey, An algorithm for the ma- chine calculation of complex Fourier series, Mathematics of Computation19, 297 (1965)
1965
-
[63]
N. Jain, J. Landman, N. Mathur, and I. Kerenidis, Quantum Fourier networks for solving parametric PDEs, Quantum Science and Technology9, 035026 (2024)
2024
-
[64]
Coyle, D
B. Coyle, D. Mills, V. Danos, and E. Kashefi, The Born supremacy: quantum advantage and training of an Ising Born machine, npj Quantum Information6, 60 (2020)
2020
-
[65]
S. R. White, Density matrix formulation for quantum renormalization groups, Physical Review Letters69, 2863 (1992)
1992
-
[66]
Nesterov, Efficiency of coordinate descent methods on huge-scale optimization problems, SIAM Journal on Optimization22, 341 (2012)
Y. Nesterov, Efficiency of coordinate descent methods on huge-scale optimization problems, SIAM Journal on Optimization22, 341 (2012)
2012
-
[67]
M. C. Caro, H.-Y. Huang, M. Cerezo, K. Sharma, A. Sornborger, L. Cincio, and P. J. Coles, Generalization in quantum machine learning from few training data, Na- ture Communications13, 4919 (2022). Appendix A: Spectral Gap of the Butterfly’s Second-Moment Operator This appendix proves Theorem 13 of Section IV: the second-moment operator of the unitary butt...
2022
-
[68]
Lloyd and C
S. Lloyd and C. Weedbrook, Quantum generative ad- versarial learning, Physical Review Letters121, 040502 (2018)
2018
-
[69]
Schuld and F
M. Schuld and F. Petruccione,Supervised Learning with Quantum Computers, 2nd ed. (Springer, 2021)
2021
-
[71]
Setup LetW n =U K · · ·U1 be the unitary butterfly onn= 2K modes, where layerℓapplies independent Haar-U(2) gates onn/2 disjoint pairs at stride 2ℓ−1. This is theproof modelfor the spectral gap: each two-qubit gate is treated as a Haar-randomU(2) element restricted to the single- excitation subspace, which is a sufficient condition for the fixed-point str...
-
[72]
Proof.Tracing over the second factor setsd=band sums overb: [Tr2 Φ2[Y]] a,c = X b,e,f,g,h E[WaeWbf W ∗ cgW ∗ bh]Y (e,f),(g,h)
Step A: Partial trace identity Lemma 38.For any exact1-designνand anyY∈ End(Cn ⊗C n): Tr2(Φν 2[Y]) = Tr 1(Φν 2[Y]) = (tr(Y)/n)I n. Proof.Tracing over the second factor setsd=band sums overb: [Tr2 Φ2[Y]] a,c = X b,e,f,g,h E[WaeWbf W ∗ cgW ∗ bh]Y (e,f),(g,h) . Unitarity gives P b Wbf W ∗ bh =δ f hdeterministically, so the expectation factors: P b E[· · ·] =...
-
[73]
Step B: Wick pairings vanish on traceless inputs Lemma 40.For any exact1-designνand traceless A, B∈End(C n): the factored-1-design contribution to Φ2[A⊗B]is zero. Proof.The 1-design second-moment formula E[WaeW ∗ cg] =δ acδeg/n, applied independently to each factor, yields the single contraction (δac/n)(δbd/n) tr(A) tr(B) = 0 since tr(A) = tr(B) = 0
-
[74]
Lemma 41(Closed basis).Φ Wn 2 [Pij]∈span{P kl :k < l}for alli < j
Step C: T ransfer matrix in the antisymmetric projector basis For modesi̸=jdefineP ij =|ψ ij⟩⟨ψij|with|ψ ij⟩= |ij⟩ − |ji⟩. Lemma 41(Closed basis).Φ Wn 2 [Pij]∈span{P kl :k < l}for alli < j. This follows from the two layer formulas below by in- duction over layers: each layer maps span{P kl :k < l} into itself (Facts 42 and 43), so the compositionW n does ...
-
[2001]
arXiv:quant-ph/0108033
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.