{"id":"981d64e3-27bd-47ce-a185-3945b7b00b3a","arxiv_id":"2608.01365","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For multi-photon quantum neural networks, approximation error decays polynomially with photon number only until n reaches m-2 when the observable is fixed; with a trainable observable, more photons always help.","lead":"This paper derives upper bounds on how well multi-photon quantum neural networks can approximate functions, and shows that adding photons helps only up to a threshold when the measurement observable is fixed. The result gives photonic QML designers a rule for when more photons buy more expressivity.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"For m=3, n=2 an explicit degree-2 fixed observable gives P_2 ⊆ H_g, so the claimed threshold at n=m−2 is false and the saturation claim must be revised.","rationale":"I read the paper as trying to establish quantitative upper bounds on MPQNN approximation error and then interpreting these bounds as a hard threshold in photon number for fixed observables. The load-bearing condition for that interpretation is that increasing n cannot improve expressivity beyond n=m−2. This condition is false. The proof of Theorem 2 only supplies an upper bound with d=min{n,m−2}; there is no matching lower bound, and the Sard dimension count in Appendix B actually permits polynomial coverage up to degree about (m−1)L. For m=3, n=2, L=1, an explicit quadratic observable gives P_2⊆H_g, while n=1 gives only P_1, so expressivity strictly improves above the claimed threshold. I also checked the apparent sign error in Lemma 6's displayed calculation (Eq. B20): it is not an error, because A·P_L + (−A)·P_L = A·P_L as a Minkowski sum with independent P_L factors. Similarly, the rescaling coefficient α in H_g is not fatal, because α can be absorbed into the fixed observable weights. The reader's rationale correctly flagged the degree-coverage mismatch as an overreach, but the reader's formal weakest_assumption emphasized the rescaling coefficient; hence my agreement is partial. Because the central limitation claim is demonstrably wrong, the appropriate verdict is REJECT, although the upper-bound results may survive a major revision that removes or corrects the saturation claim.","tokens_in":16647,"tokens_out":24174,"duration_ms":216723,"concrete_test":"Verify the m=3, L=1 construction: set y_1=1/3+ε cos x, y_2=1/3+ε sin x, y_3=1/3−ε(cos x+sin x), and g=A[(x_1−1/3)^2+(x_2−1/3)^2−ε^2]. Compute the tangent map Dπ_g(y)(v)=2Aε(cos x·v_1+sin x·v_2) for v∈T with v_1+v_2+v_3=0, and check that as v_1,v_2 range over P_1 the image is P_2. If this computation is correct, then P_2⊆H_g for n=2, m=3, which disproves the claimed saturation threshold at n=m−2=1.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central limitation claim (Section IV and abstract) is that for a fixed observable, photon number n enhances expressivity only up to n=m−2, after which adding photons has no effect. The proof establishes only an upper bound with d=min{n,m−2}; it does not prove that no better choice of g exists. In fact, the paper's own Sard count (Appendix B, Eq. B30) permits P_N with N≈(m−1)L, already larger than the claimed dL=(m−2)L. The saturation claim is not merely unsupported; for m=3 it is false. Take L=1, small ε>0, y_1=1/3+ε cos x, y_2=1/3+ε sin x, y_3=1/3−ε(cos x+sin x), and g(x)=A[(x_1−1/3)^2+(x_2−1/3)^2−ε^2]. Then y∈Δ, g(y)=0, and Dπ_g(y)(T)=2Aε(cos x P_1+sin x P_1)=P_2. By the same open-mapping argument as Lemma 6, P_2⊆H_g. Thus with n=2,m=3 a fixed-observable MPQNN realizes all of P_2, whereas any degree-1 observable (n=1) only realizes P_1. Since n=2 exceeds the paper's threshold m−2=1, this directly contradicts the claim that increasing n beyond m−2 does not affect expressivity. The correct threshold is at least m−1, and the saturation statement cannot stand as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the expressivity of multi-photon quantum neural networks (MPQNNs), in which n identical photons pass through L layers of alternating phase-encoding and trainable linear-optical unitaries, followed by measurement of a diagonal observable in the Fock basis. Theorem 1 characterizes the set of possible outputs as h = π_g(y_1,...,y_m), where g is a real polynomial of total degree at most n and (y_1,...,y_m) ∈ Δ is the family of nonnegative trigonometric polynomials of degree at most L summing to 1. For a fixed observable, the paper claims (Theorem 2) an approximation bound inf_{h∈H_g} ||f−h||_∞ ≤ C_K ||f^{(K)}||_∞/(dL)^K with d = min{n, m−2}, and interprets this as a photon-number threshold m−2 above which additional photons do not enhance expressivity. For a trainable observable, Theorem 3 claims a bound that continues to improve as n grows. The proofs combine Jackson's inequality with explicit constructions of trigonometric-polynomial subspaces and an open-mapping argument. Numerical simulations of trainable-observable MPQNNs are reported.","tokens_in":16938,"tokens_out":10638,"duration_ms":95406,"significance":"The questions addressed are timely, and the hypothesis-space characterization in Theorem 1 is a potentially useful contribution. The paper is self-contained and its overall strategy—Jackson's inequality plus explicit trigonometric-polynomial realization—is appropriate for the problem. The dynamic-programming simulation algorithm in Appendix D is also a practical contribution. However, the central fixed-observable threshold claim is not established and, as stated, is false. The sign error in the derivative computation of Lemma 6 invalidates the proof of Theorem 2 as written; the inference from an upper bound to a saturation threshold is logically unsupported; and the paper's own Sard count plus an explicit m=3, n=2 construction contradict the claimed threshold m−2. In addition, the fixed-observable hypothesis space H_g includes a free rescaling coefficient α that is not present in the physical output defined in Eq. (7), so the fixed-observable approximation guarantee is for an augmented model. The trainable-observable bound appears more plausible and may be salvageable, but the advertised limitation result cannot be accepted in its current form.","major_comments":[{"comment":"The derivative computation in Lemma 6 contains a sign error. With g_j = ∂g/∂x_j, the construction gives g_j(y) = (1/ε) q'_j(cos(Lx)) for j ≤ d and g_{d+1}(y) = −(1/ε) ∑_{i=1}^d q'_i(cos(Lx)). Therefore the two sums displayed in Eq. (B20) cancel exactly, and the computation yields Dπ_g(y)(T) = 0, not P_{dL}. Consequently Lemma 6 does not prove P_{dL} ⊆ H_g, and the proof of Theorem 2 collapses at this point. A corrected construction (for example, arranging the y_j so that the individual derivatives do not cancel, or using a different balancing argument) is needed before the bound can be accepted.","section":"Appendix B, Eqs. (B18)–(B21)"},{"comment":"The claimed photon-number threshold m−2 is not a consequence of Theorem 2. Theorem 2 is an existence statement: it shows that one particular polynomial g achieves an error bound with d = min{n, m−2}. It does not rule out the possibility that another fixed observable g' realizes P_N with N > (m−2)L. The paper's own Sard-count bound in Eq. (B30) gives N ≤ min{nL, (m−1)L + ⌊(m−1)/2⌋}, which for m=3, n=2, L=1 permits N=2, not just N=1. Moreover, an explicit counterexample refutes the threshold: for m=3, n=2, L=1, take y_1 = 1/3 + ε cos x, y_2 = 1/3 + ε sin x, y_3 = 1/3 − ε(cos x + sin x), and g = A[(x_1−1/3)^2 + (x_2−1/3)^2 − ε^2]. Then y ∈ Δ, g(y)=0, and Dπ_g(y)(T) = 2Aε(cos x P_1 + sin x P_1) = P_2. By the same open-mapping argument used in Lemma 6, P_2 ⊆ H_g. Since n=1 realises only P_1, increasing n from 1 to 2 changes the expressivity even though m−2 = 1. The saturation claim must therefore be revised or removed.","section":"Section IV and abstract (threshold claim)"},{"comment":"The fixed-observable hypothesis space H_g includes a free rescaling coefficient α multiplying π_g(y), but the physical MPQNN output in Eq. (7) contains no such coefficient. All fixed-observable approximation guarantees in Theorem 2 are therefore for the augmented family H_g, not for the fixed-observable MPQNN as defined. This is a load-bearing distinction: without α, the output range is bounded by the range of g on the simplex, so a fixed-observable MPQNN cannot approximate arbitrary continuous periodic functions to arbitrary accuracy. The paper should either prove the stated error decay for appropriately rescaled target functions within the physical model, or explicitly and prominently state that the theorem concerns the affine-rescaled model.","section":"Section III, Eq. (13)"}],"minor_comments":[{"comment":"There is a typo: 'resacle coefficient' should be 'rescale coefficient'.","section":"Section III"},{"comment":"The sentence describing Eq. (B30) as 'asymptotically the same as min{n,m−2}L' is inaccurate: for fixed m and n→∞, the bound behaves like (m−1)L + ⌊(m−1)/2⌋, which differs from (m−2)L by about L.","section":"Appendix B, after Eq. (B30)"},{"comment":"The numerical experiments train both the linear-optical parameters and the measured observable, so Fig. 2 does not directly validate the fixed-observable threshold claim of Section IV. The text should clarify which theoretical claim the simulations are intended to support.","section":"Section VI"}],"recommendation":"reject","confidential_remarks":"The paper contains a useful characterization theorem and a practical simulation algorithm, but its central advertised limitation is false as stated. The explicit m=3, n=2 counterexample is decisive, and the sign error in Lemma 6 means the main proof is invalid as written. I do not think a local revision can preserve the claimed threshold m−2; the fixed-observable section needs substantive reworking, possibly toward a threshold of m−1, and the central claim in the abstract must change. I recommend rejection, while noting that a substantially revised manuscript focused on the trainable-observable bound and a corrected fixed-observable statement could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, this is not a throwaway. The hypothesis-space characterization in Theorem 1 is genuinely new and useful: it gives a clean description of all MPQNN outputs as polynomial functions of the mode probabilities, generalizing the m=2 single-photon result. The trainable-observable bound in Theorem 3 is also a real contribution, as is the dynamic-programming simulation algorithm. If all you take from this paper is the positive approximation bounds, they are worth having.\n\nThe problem is the headline. The paper claims that with a fixed observable, increasing the photon number n beyond m−2 does not affect expressivity. The proof establishes only that for a specially constructed g, the space P_{min(n,m-2)L} sits inside H_g. That is an upper-bound construction, not an impossibility result. The paper's own Sard argument (Eq. B30) allows degree coverage up to roughly (m−1)L, already larger than dL = (m−2)L. And the limitation claim is not just unsupported; it is false as stated. For m=3, n=2, L=1, take y_1 = 1/3+ε cos x, y_2 = 1/3+ε sin x, y_3 = 1/3−ε(cos x+sin x), and g = A[(x_1−1/3)^2+(x_2−1/3)^2−ε^2]. Then g(y)=0 and Dπ_g(y)(T) = 2Aε(cos x P_1 + sin x P_1) = P_2, so the open-mapping argument gives P_2 ⊆ H_g. With n=1 you only get P_1. So the second photon does change expressivity, contradicting the threshold at m−2=1. The correct threshold appears to be m−1, consistent with the Sard bound.\n\nThere is also a concrete error in the proof of Lemma 6: the displayed equations (B18)–(B21) cancel to zero because the negative term for the (d+1)-th variable is dropped. The construction is repairable—write the image as a sum of q'_j(cos Lx)(v_j−v_{d+1})—but the printed derivation is wrong.\n\nOne minor caveat on Theorem 2: the guarantee is for H_g = {απ_g(y)}, with a free rescale coefficient α not present in the physical output (7). This is a standard normalization trick, and the paper cites precedent, but it should be stated more carefully.\n\nThe numerical section tests only the trainable-observable case, so it does not validate the fixed-observable saturation. Bottom line: the paper has valuable formal content, but the central limitation claim needs to be either proved (I doubt it) or rephrased as a property of the constructed observable rather than of the model. I would send it to peer review—the hypothesis-space theorem alone justifies that—but with a clear instruction to the authors to fix the overreach.","headline":"A useful formal analysis of MPQNN expressivity that overstates its main conclusion: the claimed photon-number threshold at n=m−2 is not supported and, for m=3,n=2, false.","tokens_in":17506,"tokens_out":7592,"would_cite":true,"duration_ms":60502,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","41A10","41A25"],"pacs":[],"model":"deepseek-v4-flash","headline":"Multi-photon QNNs gain from extra photons only up to a mode-count threshold.","keywords":["multi-photon quantum neural networks","expressivity","approximation error bounds","trigonometric polynomial approximation","linear optical networks","Fock state encoding","photon-number threshold","Jackson's inequality"],"falsifier":"A direct check of the proof's key inclusion: take $m=4$, $n=2$, $L=1$ and the polynomial $g$ constructed in Lemma 6, then verify numerically whether every real trigonometric polynomial of degree $2$ can be written as $\\alpha\\pi_g(y)$ for some $y\\in\\Delta$ and $\\alpha\\in\\mathbb{R}$; a single degree-$2$ polynomial that cannot be represented would invalidate $P_{dL}\\subseteq H_g$. At the model level, train a fixed-observable MPQNN (without rescaling) for $m=4$ on a fixed smooth target at $n=2$ and $n=5$; systematic improvement at $n=5$ would contradict the claimed saturation of expressivity past $m-2$ for the physical model.","tokens_in":16409,"feed_emoji":"⚛️","tokens_out":9225,"duration_ms":78758,"temperature":0.7,"pith_summary":"This paper asks whether adding more photons to a data-re-uploading quantum neural network built from a linear optical network reliably increases what the network can learn to approximate. Its answer is quantitative and conditional on how the measurement is handled: for a fixed measurement observable, the proved approximation-rate bound improves with photon number $n$ only until $n$ reaches $m-2$, where $m$ is the number of modes, and then stops improving; for a trainable observable, the bound keeps improving as $n$ grows. The argument rests on a complete description of the model's outputs as $g(y_1,\\ldots,y_m)$, a polynomial of degree at most $n$ applied to a collection of nonnegative trigonometric polynomials of degree at most $L$ that sum to one. If the bounds are tight in spirit, they give a practical design rule: below the threshold, photons buy polynomial spectral power essentially for free; above it, the circuit has no free trainable parameters left to use them.","feed_headline":"Extra photons stop helping a fixed-readout QNN past a mode cutoff","feed_subtitle":"With a trainable readout, extra photons always help; with a fixed readout, the gain saturates at $m-2$ photons.","key_machinery":"The load-bearing object is the assignment map $\\pi_g(y_1,\\ldots,y_m)(x)=g(y_1(x),\\ldots,y_m(x))$, which expresses every MPQNN output as a degree-constrained polynomial in $m$ nonnegative, unit-sum trigonometric polynomials. The proof chain is: Lemma 4 characterizes the first column of the effective linear-optical unitary as an arbitrary normalized vector of degree-$L$ trigonometric polynomials; Fejér–Riesz converts each nonnegative $y_j$ into $|u_{j,1}|^2$; a concrete $y$ and $g$ are built from Chebyshev antiderivatives so that the derivative map has rank $2dL+1$, and an open mapping argument lifts this local surjectivity to the inclusion $P_{dL}\\subseteq H_g$ for $d=\\min\\{n,m-2\\}$; Jackson's inequality converts the inclusion into the stated approximation rates. This machinery is what produces the threshold: the constructed inclusion saturates at $d=m-2$ for fixed observables, while the trainable-observable family obtains a second, independent inclusion from convex-geometry arguments that keeps growing with $n$.","core_discovery":"The central discovery is a representation theorem plus two rate bounds. Theorem 1 states that an MPQNN with $n$ identical photons, $m$ modes, and $L$ layers can output exactly the functions $h(x)=g(y_1(x),\\ldots,y_m(x))$, where $g\\in\\mathbb{R}[x_1,\\ldots,x_m]$ has total degree at most $n$ and each $y_j$ is a real trigonometric polynomial of degree at most $L$ with $y_j\\ge 0$ and $\\sum_j y_j=1$. For a fixed observable, the paper proves there exists a single polynomial $g$ such that every $K$-times differentiable $2\\pi$-periodic $f$ satisfies $\\inf_{h\\in H_g}\\|f-h\\|_\\infty \\le C_K\\|f^{(K)}\\|_\\infty/(dL)^K$ with $d=\\min\\{n,m-2\\}$, where $H_g$ is the rescaled family $\\{\\alpha\\pi_g(y):y\\in\\Delta,\\alpha\\in\\mathbb{R}\\}$. For a trainable observable, the corresponding family $H$ (arbitrary $g$ of degree at most $n$) satisfies $\\inf_{h\\in H}\\|f-h\\|_\\infty \\le C_K\\|f^{(K)}\\|_\\infty/d^K$ with $d=\\min\\{nL,\\max\\{(m-2)L,n\\lfloor(m-1)/2\\rfloor\\}\\}$. The paper reads these bounds as: in the fixed-observable case photon number is a resource only up to the mode-dependent threshold $m-2$; in the trainable-observable case it is an unlimited resource.","pith_inferences":["The fixed-observable saturation is most plausibly a parameter-counting effect: the trainable part of the interferometer has $\\Theta(mL)$ parameters independent of $n$, so beyond $n\\approx m-2$ the extra spectral dimensions have no trainable degrees of freedom to steer them; a direct probe would be to count how many independent Fourier coefficients can actually be tuned at $n>m-2$.","Because Theorem 2 is an existence statement for one specially constructed $g$, it does not by itself tell an experimenter which fixed observable to pick; a testable extension would be to check whether randomly initialized fixed observables also show the same saturation, or whether only the optimally chosen one does.","The representation theorem suggests a classical proxy for studying MPQNN expressivity: optimize over polynomials $g$ of degree $n$ applied to the boundary of the convex set of nonnegative degree-$L$ trigonometric polynomials; if the proxy reproduces the $(dL)^{-K}$ rates, it would let practitioners estimate achievable errors for large $n$ without simulating bosonic amplitudes."],"forward_implications":["In the fixed-observable case, setting $n=m-2$ already saturates the proved spectral reach; values $n>m-2$ do not lower the bound even though the underlying Fock space is much larger.","In the trainable-observable case, the bound tends to zero as $n\\to\\infty$, so photon number is a genuine hyperparameter for expressivity without changing the interferometer structure.","The trainable-observable model pays for this with $\\binom{n+m-1}{m-1}$ extra classical weights, an exponentially growing cost in $n$.","Setting $n=1$ reduces the MPQNN hypothesis space to the familiar degree-$L$ trigonometric-polynomial family of data-re-uploading QNNs.","The numerical simulations, using the paper's dynamic-programming simulator, show test loss decreasing as photon number and layer number increase, matching the predicted qualitative scaling."],"supporting_citations":[{"why":"Supplies the identification of data-re-uploading QNN outputs with real trigonometric polynomials of bounded degree, which Theorem 1 generalizes to the multi-photon setting.","marker":"[8]"},{"why":"Provides the single-photon precursor result whose Lemma 3 is generalized by the paper's Lemma 4, along with the rescaling convention used for fixed-observable hypothesis spaces.","marker":"[9]"},{"why":"Supplies Jackson's inequality, which converts the polynomial-inclusion results into concrete approximation-error bounds.","marker":"[26]"},{"why":"Supplies the Fejér–Riesz theorem used to realize each nonnegative trigonometric polynomial $y_j$ as the modulus square of a degree-$L$ trigonometric polynomial.","marker":"[27]"},{"why":"Supplies the open mapping theorem used to prove that the local surjectivity at the constructed point lifts to the global inclusion $P_{dL}\\subseteq H_g$.","marker":"[28]"},{"why":"Supplies Sard's theorem used in the dimension argument giving a necessary condition for $P_N\\subseteq H_g$ and supporting the fixed-observable threshold.","marker":"[29]"},{"why":"Provides the earlier numerical evidence that Fock-state inputs increase QNN expressivity, the phenomenon this paper makes quantitative.","marker":"[17]"},{"why":"The recent finding that multiple photons reduce data consumption in learning, which the paper's threshold result explicitly complements.","marker":"[19]"}],"fun_headline_variants":["Fixed readout: photon advantage stops at m−2 photons","Trainable readout: every photon helps; fixed caps at m−2","MPQNNs: photons always help if readout is trainable","Photon boost in MPQNNs: capped for fixed readout, not trainable","For fixed observables, photon count stops mattering at m−2"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The fixed-observable bound holds for the rescaled hypothesis space $H_g=\\{\\alpha\\pi_g(y):y\\in\\Delta,\\alpha\\in\\mathbb{R}\\}$, not for the physical output of an MPQNN with a fixed observable, and the paper does not show the physical model approximates the unscaled target at the same rate.","fun_headline_variants_meta":{"raw":{"variants":["Fixed readout: photon advantage stops at m−2 photons","Trainable readout: every photon helps; fixed caps at m−2","MPQNNs: photons always help if readout is trainable","Photon boost in MPQNNs: capped for fixed readout, not trainable","For fixed observables, photon count stops mattering at m−2"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001311,"raw_usage":{"total_tokens":5430,"prompt_tokens":1122,"completion_tokens":4308,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":738,"completion_tokens_details":{"reasoning_tokens":4209}},"tokens_in":738,"tokens_out":4308,"duration_ms":28204,"temperature":1.0,"reasoning_tokens":4209,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:08:33.213812+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check of the proof's key inclusion: take $m=4$, $n=2$, $L=1$ and the polynomial $g$ constructed in Lemma 6, then verify numerically whether every real trigonometric polynomial of degree $2$ can be written as $\\alpha\\pi_g(y)$ for some $y\\in\\Delta$ and $\\alpha\\in\\mathbb{R}$; a single degree-$2$ polynomial that cannot be represented would invalidate $P_{dL}\\subseteq H_g$. At the model level, train a fixed-observable MPQNN (without rescaling) for $m=4$ on a fixed smooth target at $n=2$ and $n=5$; systematic improvement at $n=5$ would contradict the claimed saturation of expressivity past $m-2$ for the physical model.","supporting_citations":[{"cited_title":"Schuld, R","cited_arxiv_id":null,"evidence_quote":"Supplies the identification of data-re-uploading QNN outputs with real trigonometric polynomials of bounded degree, which Theorem 1 generalizes to the multi-photon setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the single-photon precursor result whose Lemma 3 is generalized by the paper's Lemma 4, along with the rescaling convention used for fixed-observable hypothesis spaces."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Jackson's inequality, which converts the polynomial-inclusion results into concrete approximation-error bounds."},{"cited_title":"Riesz and B","cited_arxiv_id":null,"evidence_quote":"Supplies the Fejér–Riesz theorem used to realize each nonnegative trigonometric polynomial $y_j$ as the modulus square of a degree-$L$ trigonometric polynomial."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the open mapping theorem used to prove that the local surjectivity at the constructed point lifts to the global inclusion $P_{dL}\\subseteq H_g$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Sard's theorem used in the dimension argument giving a necessary condition for $P_N\\subseteq H_g$ and supporting the fixed-observable threshold."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the earlier numerical evidence that Fock-state inputs increase QNN expressivity, the phenomenon this paper makes quantitative."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The recent finding that multiple photons reduce data consumption in learning, which the paper's threshold result explicitly complements."}],"review_version":2}