REVIEW 2 major objections 4 minor 1 cited by
Universality of physical neural networks with multivariate nonlinearity
T0 review · 2 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A physical neural network whose input is encoded in tunable parameters is a universal approximator exactly when its multivariate nonlinearity has no identically vanishing mixed derivative.
desk verdict The universality theorem is solid and novel; the 'provably universal' free-space architecture has a load-bearing gap at the block-diagonal decoupling step. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multivariate nonlinear encoding function σ together with the non-degeneracy condition (4): no mixed partial derivative of σ vanishes identically. The theorem's proof uses Riesz representation and Hahn-Banach: if the span is not dense, a nonzero measure annihilates it, and differentiating under the measure forces all moments to vanish unless some mixed derivative of σ is identically zero. For the optical implementation, the load-bearing identity is the Neumann-series inverse M(x)=r_m(I-r_m T(x)ST(x))^{-1}, which describes repeated reflections between the mirror and the scattering structure; it makes the encoding genuinely multivariate, and the proof shows that almost
What would settle it
Run an approximation experiment with an encoder that has no cross-input derivatives, e.g., σ(x)=cos x1+cos x2, and target the function f(x)=x1x2 on the unit square. The theorem predicts the approximation error stays bounded away from zero no matter how many copies r, because every mixed derivative of every term is identically zero; if the error can be driven to zero, the necessary direction of the criterion is wrong.
Extended reading notes
Core claim
The central claim is Theorem T1: for σ∈C∞, the class of networks f(x)=Σ_j c_j σ(a_j∘x+b_j) is dense in C(Ω;R^m) if and only if no multi-index α has ∂^ασ/∂x^α≡0. Sufficiency follows by showing that any measure annihilating the span must have all polynomial moments zero, hence be zero; necessity follows by constructing a nonzero functional that vanishes on every such network when some derivative of σ is identically zero. In the free-space optics setting, the encoding function is σ(x;S)=e^T r_m(I-r_m T(x)ST(x))^{-1}e, produced by multiple passes between a mirror and a spatial light modulator. The proof shows that the mixed derivative ∂^{2rd}/(∂x_1^{2r}⋯∂x_d^{2r}) σ(x;S) at x=0 is a nonvanishing
Load-bearing premise
The proof of universality for the free-space architecture assumes the input copies on the spatial light modulator are far enough apart not to interact, so the scattering matrix is block-diagonal and the output splits into independent terms; if those copies couple, the output is no longer a sum of independent responses and the theorem does not apply—Remark R8 only conjectures a perturbative extension without proof.
Editorial extensions
If this is right
- Any physical system that can be mapped to the sum-of-copies form (S1) with a non-degenerate σ is a universal function approximator as the number of input copies grows, so nonlinearity must couple input components to arbitrary order—not merely be non-polynomial.
- In the proposed free-space mPNN, almost every symmetric unitary scattering matrix S satisfies the non-degeneracy condition, so the design is universal without fine-tuning S.
- Numerical experiments on MNIST and Fashion-MNIST show accuracy scaling with the number of input copies: trained S reaches 98.42% and 90.19%, respectively, while random S reaches 97.64% and 89.35%.
- Temporal multiplexing with a nonzero reference wave preserves universality under intensity detection, converting the discrete sum into a time integral and offering effective system sizes up to r∼10^7 on integrated photonic hardware at 100 GHz switching over 100 ms.
- The theorem extends to several different encoding functions σ_j (Theorem T2) and to complex-valued σ and coefficients, broadening the class of physical systems covered.
Reading between the lines
- My inference: the criterion gives a practical diagnostic—measure or compute the mixed partial derivatives of a candidate device's encoding map; if any mixed derivative is identically zero over the operating region, the device is provably not universal no matter how large, so engineering effort is better spent breaking that degeneracy.
- My inference: the genericity result for random symmetric unitary S suggests that untrained or randomly structured physical media could serve as universal encoders, potentially simplifying hardware by removing the need to train the scattering matrix.
- My inference: the temporal multiplexing construction implies the same scaling laws seen with spatial copies should appear when a single physical channel is time-multiplexed; comparing accuracy versus effective r for both routes would provide a direct experimental test of the paper's scaling prediction.
- My inference: the block-diagonal assumption used in the proof is likely stronger than necessary; a numerical study that gradually turns on inter-copy couplings would empirically map how much interaction the sum-form approximation tolerates before universality degrades.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives a universality criterion for physical neural networks (PNNs) whose input–output relation can be written as f(x) = Σ_j c_j σ(a_j ∘ x + b_j), where σ is a multivariate nonlinear encoding function. The main theorem (T1, SI S1.2) states that such networks are universal in C(Ω; R^m) if and only if no partial derivative ∂^ασ/∂x^α vanishes identically; the proof uses Riesz representation and Hahn–Banach duality. The authors then propose a free-space optical implementation based on a multiple-pass cavity with a spatial light modulator and a scattering matrix S, and prove (Theorems T3–T5, SI S3) that for generic symmetric unitary S the resulting encoding function is non-degenerate, thereby concluding that the proposed architecture is provably universal. Numerical experiments on MNIST and Fashion-MNIST show accuracy that improves with the number of input copies. A temporal multiplexing extension is also proposed, with a formal universality statement for intensity detection with a reference wave.
Significance. If the claims are fully established, the paper makes a valuable contribution: it gives a sharp, checkable criterion for universality of a class of PNNs that is distinct from standard ANN universality, and it demonstrates how that criterion can guide optical hardware design. The proof of Theorem T1 is rigorous and self-contained, and the generic non-degeneracy argument for scattering matrices is careful and nontrivial. The numerical experiments support the scaling prediction of the theorem in the setting they simulate. However, the step from the generic non-degeneracy of σ to universality of the actual device in Fig. 2a contains a load-bearing assumption that is not proven; this limits the current version's central hardware claim.
major comments (2)
- [SI S3.4, Eq. (S72), Remark R8] The proof that the free-space architecture in Fig. 2a is universal assumes that the r input copies are spatially separated so that the scattering matrix S decouples into independent sub-blocks, yielding Eq. (S72). This is an assumption, not a consequence of the device as drawn. In a single-scatterer geometry common to all copies, S has generally non-zero off-diagonal blocks coupling different copies; the Neumann series in Eq. (6) then contains cross terms such as T_j S_{jk} T_k S_{kl} T_l, so the output is not of the independent-sum form (S1) required by Theorem T1/T2. Remark R8 explicitly defers the needed argument to a perturbative expectation without proof. Therefore the headline claim that the proposed device is "provably universal" is currently established only for an idealized segmented/isolated-channel design, not for the single scattering structure shown in Fig. 2a. This is load-
- [Main text, 'Numerical experiments'; SI S4, Eq. (S86)] The numerical experiments adopt the same block-diagonal ansatz for S from the outset, with the text stating "we assume that S is a block-matrix, corresponding physically to input copies having some spatial separation." Thus the experiments do not test the common-scatterer configuration that the universality claim concerns; they validate only the restricted block-diagonal model. If the authors retain the claim that the Fig. 2a device is universal, they need either a rigorous proof that the cross-coupling terms are negligible or controlled, or a separate experiment/simulation with a dense (non-block-diagonal) S. As written, the simulation results cannot resolve the gap identified above.
minor comments (4)
- [SI S5, Proposition P3] The proof states that ψδ approximates ψ, but the expansion in Eq. (S98) actually gives ψδ = (ψ + ψ*)/2 + O(δ), i.e., it approximates Re ψ. Since the target f is real-valued and ψ approximates f, this is sufficient, but the text should explicitly say that the construction approximates the real part of ψ and that Theorem T1's complex-coefficient extension is being used.
- [SI S1.3, Theorem T2 proof] The induction proof for pairwise different σ_j is compressed, particularly the notation g_j(x), the scaling of c_j and a_j with δ, and the handling of the O(δ) terms. A more formal statement of the induction invariant would improve readability and verifiability.
- [Main text, Eq. (5)-(6)] The equality r_m = t_m is described as the case treated in the proof, but Remark R4 notes that for a lossless mirror t_m ≠ r_m in general. The authors state the general case is a straightforward extension; this is fine, but the main text could briefly signal that the Neumann-series simplification is a special (though physically realizable) choice.
- [General] The phrase "dense" in the main text for the set of matrices S is weaker than what is proved in SI S3.2, where the set is shown to have full Lebesgue measure. The authors could state the stronger property to avoid under-selling the result.
Circularity Check
No circularity: the universality theorem is proved from Riesz/Hahn-Banach and polynomial density; the optical-system proof derives generic non-degeneracy for the stated scattering model, and the block-decoupling assumption is explicit, not a disguised input.
full rationale
Theorem T1 and its proof are self-contained: universality is reduced to the density of the span of functions {σ(a∘x+b)}, and the Riesz representation theorem plus Hahn-Banach yield a measure-annihilation criterion. Non-degeneracy is shown to be equivalent to the vanishing of all polynomial moments, with the density of polynomials closing the argument. The necessity direction constructs an explicit continuous functional that annihilates all approximants when a derivative of σ vanishes identically. The free-space implementation is derived from the stated optical model via the Neumann series for the cavity, and Theorems T3/T4 prove generic non-degeneracy of the resulting encoding function for almost all unitary symmetric S. The multi-copy recombination proof explicitly assumes (SI S3.4, Eq. S72) that input copies are spatially separated so S decouples into blocks; this is an acknowledged modeling assumption (Remarks R8, R11) and not a circular reliance on the theorem, since the universality proof is carried out for that decoupled model. Numerical experiments simulate the same block model but do not fit constants to the theorem; the observed scaling is an independent empirical check. There are no load-bearing self-citations, no imported uniqueness theorem, and no fitted parameter renamed as a prediction.
Assumptions & free parameters
free parameters (1)
- Mirror reflectivity r_m =
1/2
assumptions (5)
- standard math Riesz representation theorem and Hahn-Banach theorem
- standard math Multivariate polynomials are dense in C(Ω) for compact Ω
- domain assumption The physical system's transfer matrix S is unitary and symmetric (lossless, reciprocal, monochromatic, paraxial optics)
- domain assumption Input copies are spatially separated so that S is block-diagonal, decoupling the system into independent sub-blocks
- domain assumption σ ∈ C^∞(R^d) and the Neumann series for the encoding block converges
Cite this review
Pith. "Pith review of Universality of physical neural networks with multivariate nonlinearity." pith.science (2026). https://pith.science/paper/NWPE3LJN
@misc{pith2026250905420,
author = {Pith},
title = {Pith review of: Universality of physical neural networks with multivariate nonlinearity},
year = {2026},
howpublished = {\url{https://pith.science/paper/NWPE3LJN}},
note = {Machine review of arXiv:2509.05420}
}
read the original abstract
The enormous energy demand of artificial intelligence is driving the development of alternative hardware for deep learning. Physical neural networks try to exploit physical systems to perform machine learning more efficiently. In particular, optical systems can calculate with light using negligible energy. While their computational capabilities were long limited by the linearity of optical materials, nonlinear computations have recently been demonstrated through modified input encoding. Despite this breakthrough, our inability to determine if physical neural networks can learn arbitrary relationships between data -- a key requirement for deep learning known as universality -- hinders further progress. Here we present a fundamental theorem that establishes a universality condition for physical neural networks. It provides a powerful mathematical criterion that imposes device constraints, detailing how inputs should be encoded in the tunable parameters of the physical system. Based on this result, we propose a scalable architecture using free-space optics that is provably universal and achieves high accuracy on image classification tasks. Further, by combining the theorem with temporal multiplexing, we present a route to potentially huge effective system sizes in highly practical but poorly scalable on-chip photonic devices. Our theorem and scaling methods apply beyond optical systems and inform the design of a wide class of universal, energy-efficient physical neural networks, justifying further efforts in their development.
Figures
Forward citations
Cited by 1 Pith paper
-
Tutorial: A practical guide to the alignment of defocused spatial light modulators for fast diffractive neural networks
A semi-automatic procedure for aligning defocused SLMs to enable spatial multiplexing and pixel-level conjugation in diffractive neural networks for faster parallel processing.
Reference graph
Works this paper leans on
-
[1]
author author A. de Vries ,\ title title The growing energy footprint of artificial intelligence ,\ @noop journal journal Joule \ volume 7 ,\ pages 2191 ( year 2023 ) NoStop
work page 2023
-
[2]
author author D. Marković , author A. Mizrahi , author D. Querlioz ,\ and\ author J. Grollier ,\ title title Physics for neuromorphic computing ,\ @noop journal journal Nat. Rev. Phys. \ volume 2 ,\ pages 499 ( year 2020 ) NoStop
work page 2020
-
[3]
author author G. Wetzstein , author A. Ozcan , author S. Gigan , author S. Fan , author D. Englund , author M. Solja c i \'c , author C. Denz , author D. A. \ Miller ,\ and\ author D. Psaltis ,\ title title Inference in artificial intelligence with deep optics and photonics ,\ @noop journal journal Nature \ volume 588 ,\ pages 39 ( year 2020 ) NoStop
work page 2020
-
[4]
author author B. J. \ Shastri , author A. N. \ Tait , author T. Ferreira de Lima , author W. H. P. \ Pernice , author H. Bhaskaran , author C. D. \ Wright ,\ and\ author P. R. \ Prucnal ,\ title title Photonics for artificial intelligence and neuromorphic computing ,\ @noop journal journal Nat. Photon. \ volume 15 ,\ pages 102 ( year 2021 ) NoStop
work page 2021
-
[5]
author author A. Momeni , author B. Rahmani , author B. Scellier , author L. G. \ Wright , author P. L. \ McMahon , author C. C. \ Wanjura , author Y. Li , author A. Skalli , author N. G. \ Berloff , author T. Onodera , author I. Oguz , author F. Morichetti , author P. del Hougne , author M. L. \ Gallo , author A. Sebastian , author A. Mirhoseini , author...
arXiv 2024
-
[6]
author author D. Psaltis \ and\ author N. Farhat ,\ title title Optical information processing based on an associative-memory model of neural nets with thresholding and feedback ,\ @noop journal journal Opt. Lett. \ volume 10 ,\ pages 98 ( year 1985 ) NoStop
work page 1985
-
[7]
author author X. Lin , author Y. Rivenson , author N. T. \ Yardimci , author M. Veli , author Y. Luo , author M. Jarrahi ,\ and\ author A. Ozcan ,\ title title All-optical machine learning using diffractive deep neural networks ,\ @noop journal journal Science \ volume 361 ,\ pages 1004 ( year 2018 ) NoStop
work page 2018
-
[8]
author author L. G. \ Wright , author T. Onodera , author M. M. \ Stein , author T. Wang , author D. T. \ Schachter , author Z. Hu ,\ and\ author P. L. \ McMahon ,\ title title Deep physical neural networks trained with backpropagation ,\ @noop journal journal Nature \ volume 601 ,\ pages 549 ( year 2022 ) NoStop
work page 2022
Show all 32 references
-
[9]
author author P. L. \ McMahon ,\ title title The physics of optical computing ,\ @noop journal journal Nat. Rev. Phys. \ volume 5 ,\ pages 717 ( year 2023 ) NoStop
2023
-
[10]
Eliezer , author U
author author Y. Eliezer , author U. R \"u hrmair , author N. Wisiol , author S. Bittner ,\ and\ author H. Cao ,\ title title Tunable nonlinear optical mapping in a multiple-scattering cavity ,\ @noop journal journal Proc. Nat. Acad. Sci. \ volume 120 ,\ pages e2305027120 ( ye...
2023
-
[11]
author author C. C. \ Wanjura \ and\ author F. Marquardt ,\ title title Fully nonlinear neuromorphic computing with linear wave scattering ,\ @noop journal journal Nat. Phys. \ volume 20 ,\ pages 1434 ( year 2024 ) NoStop
2024
-
[12]
Xia , author K
author author F. Xia , author K. Kim , author Y. Eliezer , author S. Han , author L. Shaughnessy , author S. Gigan ,\ and\ author H. Cao ,\ title title Nonlinear optical encoding enabled by recurrent linear scattering ,\ @noop journal journal Nat. Photon. \ volume 18 ,\ pages ...
2024
-
[13]
Yildirim , author N
author author M. Yildirim , author N. U. \ Dinc , author I. Oguz , author D. Psaltis ,\ and\ author C. Moser ,\ title title Nonlinear processing with linear optics ,\ @noop journal journal Nat. Photon. \ volume 18 ,\ pages 1076 ( year 2024 ) NoStop
2024
-
[14]
Hornik , author M
author author K. Hornik , author M. Stinchcombe ,\ and\ author H. White ,\ title title Multilayer feedforward networks are universal approximators ,\ @noop journal journal Neural Netw. \ volume 2 ,\ pages 359 ( year 1989 ) NoStop
1989
-
[15]
author author P. L. \ McMahon ,\ title title Nonlinear computation with linear systems ,\ @noop journal journal Nat. Phys. \ volume 20 ,\ pages 1365 ( year 2024 ) NoStop
2024
-
[16]
Athale \ and\ author D
author author R. Athale \ and\ author D. Psaltis ,\ title title Optical computing: past and future ,\ @noop journal journal Opt. Photonics News \ volume 27 ,\ pages 32 ( year 2016 ) NoStop
2016
-
[17]
Silva , author F
author author A. Silva , author F. Monticone , author G. Castaldi , author V. Galdi , author A. Al \`u ,\ and\ author N. Engheta ,\ title title Performing mathematical operations with metamaterials ,\ @noop journal journal Science \ volume 343 ,\ pages 160 ( year 2014 ) NoStop
2014
-
[18]
Cordaro , author H
author author A. Cordaro , author H. Kwon , author D. Sounas , author A. F. \ Koenderink , author A. Al \`u ,\ and\ author A. Polman ,\ title title High-index dielectric metasurfaces performing mathematical operations ,\ @noop journal journal Nano Lett. \ volume 19 ,\ pages 84...
2019
-
[19]
Zangeneh-Nejad , author D
author author F. Zangeneh-Nejad , author D. L. \ Sounas , author A. Al \`u ,\ and\ author R. Fleury ,\ title title Analogue computing with metamaterials ,\ @noop journal journal Nat. Rev. Mater. \ volume 6 ,\ pages 207 ( year 2021 ) NoStop
2021
-
[20]
Mohammadi Estakhri , author B
author author N. Mohammadi Estakhri , author B. Edwards ,\ and\ author N. Engheta ,\ title title Inverse-designed metastructures that solve equations ,\ @noop journal journal Science \ volume 363 ,\ pages 1333 ( year 2019 ) NoStop
2019
-
[21]
Cordaro , author B
author author A. Cordaro , author B. Edwards , author V. Nikkhah , author A. Al \`u , author N. Engheta ,\ and\ author A. Polman ,\ title title Solving integral equations in free space with inverse-designed ultrathin optical metagratings ,\ @noop journal journal Nat. Nanotechn...
2023
-
[22]
Yang , author L
author author L. Yang , author L. Zhang ,\ and\ author R. Ji ,\ title title On-chip optical matrix-vector multiplier ,\ in\ @noop booktitle Optics and Photonics for Information Processing VII ,\ Vol.\ volume 8855 \ ( organization SPIE ,\ year 2013 )\ p.\ pages 88550F NoStop
2013
-
[23]
Nikkhah , author A
author author V. Nikkhah , author A. Pirmoradi , author F. Ashtiani , author B. Edwards , author F. Aflatouni ,\ and\ author N. Engheta ,\ title title Inverse-designed low-index-contrast structures on a silicon photonics platform for vector--matrix multiplication ,\ @noop jour...
2024
-
[24]
author author S. L. \ Brunton \ and\ author J. N. \ Kutz ,\ @noop title Data-Driven Science and Engineering: Machine Learning, Dynamical Systems, and Control \ ( publisher Cambridge University Press ,\ year 2022 ) NoStop
2022
-
[25]
Shen , author N
author author Y. Shen , author N. C. \ Harris , author S. Skirlo , author M. Prabhu , author T. Baehr-Jones , author M. Hochberg , author X. Sun , author S. Zhao , author H. Larochelle , author D. Englund ,\ and\ author M. Solja c i \'c ,\ title title Deep learning with cohere...
2017
-
[26]
Pai , author Z
author author S. Pai , author Z. Sun , author T. W. \ Hughes , author T. Park , author B. Bartlett , author I. A. D. \ Williamson , author M. Minkov , author M. Milanizadeh , author N. Abebe , author F. Morichetti , author A. Melloni , author S. Fan , author O. Solgaard ,\ and...
2023
-
[27]
Hua , author E
author author S. Hua , author E. Divita , author S. Yu , author B. Peng , author C. Roques-Carmes , author Z. Su , author Z. Chen , author Y. Bai , author J. Zou , author Y. Zhu , author Y. Xu , author C.-k. \ Lu , author Y. Di , author H. Chen , author L. Jiang , author L. Wa...
2025
-
[28]
Yu , author Z
author author H. Yu , author Z. Huang , author S. Lamon , author B. Wang , author H. Ding , author J. Lin , author Q. Wang , author H. Luan , author M. Gu ,\ and\ author Q. Zhang ,\ title title All-optical image transportation through a multimode fibre using a miniaturized dif...
2025
-
[29]
Bandyopadhyay , author A
author author S. Bandyopadhyay , author A. Sludds , author S. Krastanov , author R. Hamerly , author N. Harris , author D. Bunandar , author M. Streshinsky , author M. Hochberg ,\ and\ author D. Englund ,\ title title Single-chip photonic deep neural network with forward-only ...
2024
-
[30]
LeCun , author L
author author Y. LeCun , author L. Bottou , author Y. Bengio ,\ and\ author P. Haffner ,\ title title Gradient-based learning applied to document recognition ,\ @noop journal journal Proc. IEEE \ volume 86 ,\ pages 2278 ( year 1998 ) NoStop
1998
-
[31]
Xiao , author K
author author H. Xiao , author K. Rasul ,\ and\ author R. Vollgraf ,\ title title Fashion- MNIST : a novel image dataset for benchmarking machine learning algorithms ,\ @noop journal journal arXiv:1708.07747 \ ( year 2017 ) NoStop
2017 arXiv
-
[32]
Walter Rudin, Real and Complex Analysis, 3rd ed., McGraw--Hill, 1987
1987
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.