REVIEW 3 major objections 3 minor 29 references
Expressive Power and Limitations of Multi-photon Quantum Neural Networks
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Multi-photon QNNs gain from extra photons only up to a mode-count threshold.
desk verdict A useful formal analysis of MPQNN expressivity that overstates its main conclusion: the claimed photon-number threshold at n=m−2 is not supported and, for m=3,n=2, false. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the assignment map $\pi_g(y_1,\ldots,y_m)(x)=g(y_1(x),\ldots,y_m(x))$, which expresses every MPQNN output as a degree-constrained polynomial in $m$ nonnegative, unit-sum trigonometric polynomials. The proof chain is: Lemma 4 characterizes the first column of the effective linear-optical unitary as an arbitrary normalized vector of degree-$L$ trigonometric polynomials; Fejér–Riesz converts each nonnegative $y_j$ into $|u_{j,1}|^2$; a concrete $y$ and $g$ are built from Chebyshev antiderivatives so that the derivative map has rank $2dL+1$, and an open mapping argument lifts this local surjectivity to the inclusion $P_{dL}\subseteq H_g$ for $d=\min\{n,m-2\}$; Jackson's inequality converts the inclusion into the stated approximation rates. This machinery is what produces the threshold: the constructed inclusion saturates at $d=m-2$ for fixed observables, while the trainable-observable family obtains a second, independent inclusion from convex-geometry arguments that keeps growing with $n$.
What would settle it
A direct check of the proof's key inclusion: take $m=4$, $n=2$, $L=1$ and the polynomial $g$ constructed in Lemma 6, then verify numerically whether every real trigonometric polynomial of degree $2$ can be written as $\alpha\pi_g(y)$ for some $y\in\Delta$ and $\alpha\in\mathbb{R}$; a single degree-$2$ polynomial that cannot be represented would invalidate $P_{dL}\subseteq H_g$. At the model level, train a fixed-observable MPQNN (without rescaling) for $m=4$ on a fixed smooth target at $n=2$ and $n=5$; systematic improvement at $n=5$ would contradict the claimed saturation of expressivity past $m-2$ for the physical model.
Extended reading notes
Core claim
The central discovery is a representation theorem plus two rate bounds. Theorem 1 states that an MPQNN with $n$ identical photons, $m$ modes, and $L$ layers can output exactly the functions $h(x)=g(y_1(x),\ldots,y_m(x))$, where $g\in\mathbb{R}[x_1,\ldots,x_m]$ has total degree at most $n$ and each $y_j$ is a real trigonometric polynomial of degree at most $L$ with $y_j\ge 0$ and $\sum_j y_j=1$. For a fixed observable, the paper proves there exists a single polynomial $g$ such that every $K$-times differentiable $2\pi$-periodic $f$ satisfies $\inf_{h\in H_g}\|f-h\|_\infty \le C_K\|f^{(K)}\|_\infty/(dL)^K$ with $d=\min\{n,m-2\}$, where $H_g$ is the rescaled family $\{\alpha\pi_g(y):y\in\Delta,\alpha\in\mathbb{R}\}$. For a trainable observable, the corresponding family $H$ (arbitrary $g$ of degree at most $n$) satisfies $\inf_{h\in H}\|f-h\|_\infty \le C_K\|f^{(K)}\|_\infty/d^K$ with $d=\min\{nL,\max\{(m-2)L,n\lfloor(m-1)/2\rfloor\}\}$. The paper reads these bounds as: in the fixed-observable case photon number is a resource only up to the mode-dependent threshold $m-2$; in the trainable-observable case it is an unlimited resource.
Load-bearing premise
The fixed-observable bound holds for the rescaled hypothesis space $H_g=\{\alpha\pi_g(y):y\in\Delta,\alpha\in\mathbb{R}\}$, not for the physical output of an MPQNN with a fixed observable, and the paper does not show the physical model approximates the unscaled target at the same rate.
Editorial extensions
If this is right
- In the fixed-observable case, setting $n=m-2$ already saturates the proved spectral reach; values $n>m-2$ do not lower the bound even though the underlying Fock space is much larger.
- In the trainable-observable case, the bound tends to zero as $n\to\infty$, so photon number is a genuine hyperparameter for expressivity without changing the interferometer structure.
- The trainable-observable model pays for this with $\binom{n+m-1}{m-1}$ extra classical weights, an exponentially growing cost in $n$.
- Setting $n=1$ reduces the MPQNN hypothesis space to the familiar degree-$L$ trigonometric-polynomial family of data-re-uploading QNNs.
- The numerical simulations, using the paper's dynamic-programming simulator, show test loss decreasing as photon number and layer number increase, matching the predicted qualitative scaling.
Reading between the lines
- The fixed-observable saturation is most plausibly a parameter-counting effect: the trainable part of the interferometer has $\Theta(mL)$ parameters independent of $n$, so beyond $n\approx m-2$ the extra spectral dimensions have no trainable degrees of freedom to steer them; a direct probe would be to count how many independent Fourier coefficients can actually be tuned at $n>m-2$.
- Because Theorem 2 is an existence statement for one specially constructed $g$, it does not by itself tell an experimenter which fixed observable to pick; a testable extension would be to check whether randomly initialized fixed observables also show the same saturation, or whether only the optimally chosen one does.
- The representation theorem suggests a classical proxy for studying MPQNN expressivity: optimize over polynomials $g$ of degree $n$ applied to the boundary of the convex set of nonnegative degree-$L$ trigonometric polynomials; if the proxy reproduces the $(dL)^{-K}$ rates, it would let practitioners estimate achievable errors for large $n$ without simulating bosonic amplitudes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the expressivity of multi-photon quantum neural networks (MPQNNs), in which n identical photons pass through L layers of alternating phase-encoding and trainable linear-optical unitaries, followed by measurement of a diagonal observable in the Fock basis. Theorem 1 characterizes the set of possible outputs as h = π_g(y_1,...,y_m), where g is a real polynomial of total degree at most n and (y_1,...,y_m) ∈ Δ is the family of nonnegative trigonometric polynomials of degree at most L summing to 1. For a fixed observable, the paper claims (Theorem 2) an approximation bound inf_{h∈H_g} ||f−h||_∞ ≤ C_K ||f^{(K)}||_∞/(dL)^K with d = min{n, m−2}, and interprets this as a photon-number threshold m−2 above which additional photons do not enhance expressivity. For a trainable observable, Theorem 3 claims a bound that continues to improve as n grows. The proofs combine Jackson's inequality with explicit constructions of trigonometric-polynomial subspaces and an open-mapping argument. Numerical simulations of trainable-observable MPQNNs are reported.
Significance. The questions addressed are timely, and the hypothesis-space characterization in Theorem 1 is a potentially useful contribution. The paper is self-contained and its overall strategy—Jackson's inequality plus explicit trigonometric-polynomial realization—is appropriate for the problem. The dynamic-programming simulation algorithm in Appendix D is also a practical contribution. However, the central fixed-observable threshold claim is not established and, as stated, is false. The sign error in the derivative computation of Lemma 6 invalidates the proof of Theorem 2 as written; the inference from an upper bound to a saturation threshold is logically unsupported; and the paper's own Sard count plus an explicit m=3, n=2 construction contradict the claimed threshold m−2. In addition, the fixed-observable hypothesis space H_g includes a free rescaling coefficient α that is not present in the physical output defined in Eq. (7), so the fixed-observable approximation guarantee is for an augmented model. The trainable-observable bound appears more plausible and may be salvageable, but the advertised limitation result cannot be accepted in its current form.
major comments (3)
- [Appendix B, Eqs. (B18)–(B21)] The derivative computation in Lemma 6 contains a sign error. With g_j = ∂g/∂x_j, the construction gives g_j(y) = (1/ε) q'_j(cos(Lx)) for j ≤ d and g_{d+1}(y) = −(1/ε) ∑_{i=1}^d q'_i(cos(Lx)). Therefore the two sums displayed in Eq. (B20) cancel exactly, and the computation yields Dπ_g(y)(T) = 0, not P_{dL}. Consequently Lemma 6 does not prove P_{dL} ⊆ H_g, and the proof of Theorem 2 collapses at this point. A corrected construction (for example, arranging the y_j so that the individual derivatives do not cancel, or using a different balancing argument) is needed before the bound can be accepted.
- [Section IV and abstract (threshold claim)] The claimed photon-number threshold m−2 is not a consequence of Theorem 2. Theorem 2 is an existence statement: it shows that one particular polynomial g achieves an error bound with d = min{n, m−2}. It does not rule out the possibility that another fixed observable g' realizes P_N with N > (m−2)L. The paper's own Sard-count bound in Eq. (B30) gives N ≤ min{nL, (m−1)L + ⌊(m−1)/2⌋}, which for m=3, n=2, L=1 permits N=2, not just N=1. Moreover, an explicit counterexample refutes the threshold: for m=3, n=2, L=1, take y_1 = 1/3 + ε cos x, y_2 = 1/3 + ε sin x, y_3 = 1/3 − ε(cos x + sin x), and g = A[(x_1−1/3)^2 + (x_2−1/3)^2 − ε^2]. Then y ∈ Δ, g(y)=0, and Dπ_g(y)(T) = 2Aε(cos x P_1 + sin x P_1) = P_2. By the same open-mapping argument used in Lemma 6, P_2 ⊆ H_g. Since n=1 realises only P_1, increasing n from 1 to 2 changes the expressivity even though m−2 = 1. The saturation claim must therefore be revised or removed.
- [Section III, Eq. (13)] The fixed-observable hypothesis space H_g includes a free rescaling coefficient α multiplying π_g(y), but the physical MPQNN output in Eq. (7) contains no such coefficient. All fixed-observable approximation guarantees in Theorem 2 are therefore for the augmented family H_g, not for the fixed-observable MPQNN as defined. This is a load-bearing distinction: without α, the output range is bounded by the range of g on the simplex, so a fixed-observable MPQNN cannot approximate arbitrary continuous periodic functions to arbitrary accuracy. The paper should either prove the stated error decay for appropriately rescaled target functions within the physical model, or explicitly and prominently state that the theorem concerns the affine-rescaled model.
minor comments (3)
- [Section III] There is a typo: 'resacle coefficient' should be 'rescale coefficient'.
- [Appendix B, after Eq. (B30)] The sentence describing Eq. (B30) as 'asymptotically the same as min{n,m−2}L' is inaccurate: for fixed m and n→∞, the bound behaves like (m−1)L + ⌊(m−1)/2⌋, which differs from (m−2)L by about L.
- [Section VI] The numerical experiments train both the linear-optical parameters and the measured observable, so Fig. 2 does not directly validate the fixed-observable threshold claim of Section IV. The text should clarify which theoretical claim the simulations are intended to support.
Circularity Check
No significant circularity: the expressivity bounds are derived from explicit constructions and Jackson's inequality; the advertised saturation threshold is an over-inference from an upper bound, not a circular step.
full rationale
The paper's derivation chain is self-contained. Theorem 1 is proven from the permanent expansion of Fock amplitudes and the Fejer-Riesz factorization, with Lemma 4 giving an explicit synthesis of the first column of the linear-optical unitary. Theorem 2's proof constructs a concrete polynomial g (antiderivatives of Chebyshev polynomials) and a concrete interior point y in Delta, verifies D pi_g(y)(T) = P_{dL}, and applies Sussmann's open mapping theorem to obtain P_{dL} subset of H_g; the error bound then follows from Jackson's inequality. None of the constants or degree cutoffs are fitted to the target f or to numerical data; the d = min{n, m-2} arises from the degree of the constructed g, and the bound is an upper bound. The rescale coefficient alpha in Eq. (13) is an explicitly declared modeling convention cited to external works [9,10,25], not a hidden fit. Self-citations [19,20] give background and context and are not load-bearing; no uniqueness theorem by the authors is invoked to forbid alternatives. The main caveat is Section IV's interpretation that the min implies a hard saturation threshold at n = m-2; Theorem 2 only establishes existence of one g with that error bound and does not rule out better fixed observables using photons beyond m-2. This is a correctness or over-claim issue, not circularity: the claimed threshold is not forced by definition of the model's output, and the proof does not reduce the saturation statement to the bound. Hence no circular step qualifies under the hard-rule standard.
Assumptions & free parameters
assumptions (7)
- standard math Fejer-Riesz theorem: every nonnegative trigonometric polynomial is the modulus squared of a trigonometric polynomial
- standard math Jackson's inequality for periodic functions
- standard math Sussmann's open mapping theorem for smooth maps
- standard math Sard's theorem and dimension counting for smooth maps
- standard math Any compact subset of R^{2d} lies in some (2d+1)-vertex simplex
- domain assumption Universal linear optical networks can implement any unitary (Reck/Clements decomposition)
- ad hoc to paper Admissibility of adding a rescale coefficient α to the fixed-observable hypothesis space
Cite this review
Pith. "Pith review of Expressive Power and Limitations of Multi-photon Quantum Neural Networks." pith.science (2026). https://pith.science/paper/YVIQVUCG
@misc{pith2026260801365,
author = {Pith},
title = {Pith review of: Expressive Power and Limitations of Multi-photon Quantum Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/YVIQVUCG}},
note = {Machine review of arXiv:2608.01365}
}
read the original abstract
Quantum neural networks (QNNs) have shown promise in leveraging quantum computation for machine learning tasks. Utilizing multiple identical photons as input, multi-photon quantum neural networks (MPQNNs) have the potential to enhance the expressivity through increasing the photon number. However, how precisely the expressivity of an MPQNN is affected by an increase in photon number, and whether it can be infinitely enhanced by increasing the photon number, remains unexplored. In this work, we quantitatively estimate the expressivity of this model by deriving upper bounds on approximation error in two cases. In the case of a fixed observable, there exists a threshold that scales linearly with the mode number. Below the threshold, the expressivity of an MPQNN can be enhanced polynomially by increasing the photon number. Above the threshold, however, increasing the photon number does not affect the expressivity. In the case of a trainable observable, the expressivity can always be enhanced polynomially by increasing the photon number. These findings are then validated by numerical simulations. Our work elucidates the performance enhancement of multi-photon quantum feature in QNNs, as well as its limitations, offering guidance for leveraging multi-photon advantages in quantum machine learning.
Figures
Reference graph
Works this paper leans on
-
[1]
Cerezo, G
M. Cerezo, G. Verdon, H.-Y. Huang, L. Cincio, and P. J. Coles, Challenges and opportunities in quantum machine learning, Nature Computational Science2, 567 (2022)
2022
-
[2]
Abbas, D
A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, The power of quantum neural networks, Nature Computational Science1, 403 (2021)
2021
-
[3]
K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, 12 Algorithm 1:calculation of amplitudes Input:n, m,s,t, U Output:⟨t| ˆU|s⟩ F unctionAMP(s, t): ifamp[s, t]is definedthen returnamp[s, t]; end if Pm i=1 ti = 0then amp[s, t]←1; returnamp[s, t]; else find the smallest indexithat makess i >0; calculates− i...
work page 2022
-
[4]
Cerezo, A
M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, and P. J. Coles, Variational quantum algo- rithms, Nature Reviews Physics3, 625 (2021)
2021
-
[5]
P´ erez-Salinas, A
A. P´ erez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, Data re-uploading for a universal quantum classifier, Quantum4, 226 (2020)
2020
-
[6]
A. P´ erez-Salinas, D. L´ opez-N´ u˜ nez, A. Garc´ ıa-S´ aez, P. Forn-D´ ıaz, and J. I. Latorre, One qubit as a universal approximant, Phys. Rev. A104, 012405 (2021)
work page 2021
-
[7]
A. P´ erez-Salinas, M. Yaghubi Rad, A. Barthe, and V. Dunjko, Universal approximation of continuous func- tions with minimal quantum circuits, Phys. Rev. Res.7, 043282 (2025)
work page 2025
- [8]
Show all 29 references
-
[9]
Z. Yu, H. Yao, M. Li, and X. Wang, Power and limi- tations of single-qubit native quantum neural networks, inAdvances in Neural Information Processing Systems, Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc.,
-
[10]
Z. Yu, Q. Chen, Y. Jiao, Y. Li, X. Lu, X. Wang, and J. Z. Yang, Non-asymptotic approximation error bounds of parameterized quantum circuits, inAdvances in Neu- ral Information Processing Systems, Vol. 37, edited by A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Pa- quet, J. ...
-
[11]
Aaronson and A
S. Aaronson and A. Arkhipov, The computational com- plexity of linear optics, inProceedings of the Forty- Third Annual ACM Symposium on Theory of Comput- ing, STOC ’11 (Association for Computing Machinery, New York, NY, USA, 2011) p. 333–342
2011
-
[12]
A. W. Harrow and A. Montanaro, Quantum computa- tional supremacy, Nature549, 203 (2017)
2017
-
[13]
Zhong, H
H.-S. Zhong, H. Wang, Y.-H. Deng, M.-C. Chen, L.-C. Peng, Y.-H. Luo, J. Qin, D. Wu, X. Ding, Y. Hu, P. Hu, X.-Y. Yang, W.-J. Zhang, H. Li, Y. Li, X. Jiang, L. Gan, G. Yang, L. You, Z. Wang, L. Li, N.-L. Liu, C.-Y. Lu, and J.-W. Pan, Quantum computational advantage using photon...
2020 doi
-
[14]
Zhong, Y.-H
H.-S. Zhong, Y.-H. Deng, J. Qin, H. Wang, M.-C. Chen, L.-C. Peng, Y.-H. Luo, D. Wu, S.-Q. Gong, H. Su, Y. Hu, P. Hu, X.-Y. Yang, W.-J. Zhang, H. Li, Y. Li, X. Jiang, L. Gan, G. Yang, L. You, Z. Wang, L. Li, N.-L. Liu, J. J. Renema, C.-Y. Lu, and J.-W. Pan, Phase-programmable g...
2021
-
[15]
L. S. Madsen, F. Laudenbach, M. F. Askarani, F. Rortais, T. Vincent, J. F. F. Bulmer, F. M. Miatto, L. Neuhaus, L. G. Helt, M. J. Collins, A. E. Lita, T. Gerrits, S. W. Nam, V. D. Vaidya, M. Menotti, I. Dhand, Z. Vernon, N. Quesada, and J. Lavoie, Quantum computational ad- van...
2022
-
[16]
H.-L. Liu, H. Su, Y.-H. Deng, S.-Q. Gong, Y.-C. Gu, H.-Y. Tang, M.-H. Jia, Q. Wei, Y.-K. Song, D.-Z. Wang, M.-Y. Zheng, F.-X. Chen, L.-B. Li, S.-Y. Ren, X.-Z. Zhu, M.-H. Wang, Y.-J. Chen, Y.-F. Liu, L.-S. Song, P.-Y. Yang, J.-S. Chen, H. An, L. Zhang, L. Gan, G.-w. Yang, J.-M....
2026
-
[17]
B. Y. Gan, D. Leykam, and D. G. Angelakis, Fock state- enhanced expressivity of quantum machine learning mod- els, EPJ Quantum Technology9, 16 (2022)
2022
-
[18]
M. F. X. Mauser, S. Four, L. M. Predl, R. Albiero, F. Ceccarelli, R. Osellame, P. Petersen, B. Daki´ c, I. Agresti, and P. Walther, Experimental data re- uploading with provable enhanced learning capabilities (2025), arXiv:2507.05120 [quant-ph]
2025 arXiv
-
[19]
Y. Wang, Z. Yin, T. Haug, C. Pentangelo, S. Piacentini, A. Crespi, F. Ceccarelli, R. Osellame, and P. Walther, Multiple photons enhance quantum machine learning, npj Quantum Information 10.1038/s41534-026-01302-2 (2026)
2026 doi
-
[20]
Y. Wang, S. Xue, Y. Wang, Y. Liu, J. Ding, W. Shi, D. Wang, Y. Liu, X. Fu, G. Huang, A. Huang, M. Deng, and J. Wu, Quantum generative adversarial learning in photonics, Opt. Lett.48, 5197 (2023)
2023
-
[21]
M. Reck, A. Zeilinger, H. J. Bernstein, and P. Bertani, Experimental realization of any discrete unitary opera- tor, Phys. Rev. Lett.73, 58 (1994)
1994
-
[22]
W. R. Clements, P. C. Humphreys, B. J. Metcalf, W. S. Kolthammer, and I. A. Walmsley, Optimal design for uni- versal multiport interferometers, Optica3, 1460 (2016)
2016
-
[23]
de Guise, O
H. de Guise, O. Di Matteo, and L. L. S´ anchez-Soto, Sim- ple factorization of unitary transformations, Phys. Rev. A97, 022328 (2018)
2018
-
[24]
Neufeld, P
A. Neufeld, P. Schmocker, and V. K. Tran, Approxima- tion rates of quantum neural networks for periodic func- tions via jackson’s inequality (2025), arXiv:2511.16149 [quant-ph]. 13
2025
-
[25]
J. Tang, J. Zhang, and X. Sun, Saqnn: Spectral adap- tive quantum neural network as a universal approximator (2026), arXiv:2602.09718 [quant-ph]
2026
-
[26]
G. G. Lorentz,Approximation of Functions(Holt, Rine- hart and Winston, New York, 1966)
1966
-
[27]
Riesz and B
F. Riesz and B. Sz.-Nagy,Functional Analysis, 2nd ed. (Dover Publications, New York, NY, 1990)
1990
-
[28]
H. J. Sussmann, High-order open mapping theorems, in Directions in Mathematical Systems Theory and Opti- mization, edited by A. Rantzer and C. I. Byrnes (Springer Berlin Heidelberg, Berlin, Heidelberg, 2003) pp. 293–316
2003
-
[29]
J. M. Lee,Introduction to Smooth Manifolds, 2nd ed., Graduate Texts in Mathematics, Vol. 218 (Springer, New York, NY, 2012)
2012
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.