Pith. sign in

REVIEW 3 major objections 5 minor 77 references

Embedding two quantum emitters in an inverse-designed silicon structure yields strong optical nonlinearity at sub-nanowatt intensity, enabling all-optical neural networks to solve nonlinear classification and reinforcement learning tasks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 12:44 UTC pith:RKKOXNTK

load-bearing objection Credible device physics and a full-wave-verified classification demo, but the RL and LLM claims rest on digital simulations and favorable extrapolations that need much more support. the 3 major comments →

arxiv 2601.01690 v2 pith:RKKOXNTK submitted 2026-01-04 physics.optics physics.app-phphysics.comp-ph

Quantum Nonlinearity for Optical Neural Computing

classification physics.optics physics.app-phphysics.comp-ph
keywords quantum emittersoptical neural networksnonlinear activationinverse designsaturable absorptionexpressive powerall-optical computingphotonic LLM
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proposes an all-optical neural network architecture whose nonlinear activation is a pair of quantum emitters (two-level systems) embedded in an adjoint-optimized silicon nanophotonic structure. Because emitters saturate after absorbing only a few photons, the device produces a strong nonlinear response at intensities below 1 nW/µm², seven orders of magnitude lower than conventional materials require. Using physics-aware training in which linear layers are realized as energy-preserving transmission matrices and the activation curve is taken from nonlinear FDFD simulation, the authors demonstrate nonlinear spiral classification and reinforcement learning tasks (Atari Pong and HalfCheetah) that linear optical networks cannot perform. They also define an expressive-power metric based on the depth-wise growth of data-manifold curvature, showing that the quantum activation reaches r≈1.2 per layer at nW intensity, and they estimate that all-optical LLMs built this way could run on a few watts with sublinear power scaling.

Core claim

The central claim is that the nonlinearity bottleneck of optical neural networks is not fundamental: a saturable quantum emitter, enhanced by a nanophotonic structure designed with adjoint optimization, provides a strongly nonlinear input-output response at intensities below 1 nW/µm². Two emitters are arranged in a two-port geometry so that their saturating change in scattering produces a transmission change |Δt| = 2, a contrast impossible with a single emitter. The paper verifies the activation by full-wave nonlinear FDFD simulation and uses its simulated transmission curve as the activation function of a digitally trained network whose weights are constrained to be energy-preserving comple

What carries the argument

The central object is the 'quantum activation unit': two quantum emitters (modeled as two-level systems with Rabi-frequency-dependent dipole moments, parameters taken from silicon-vacancy color centers) embedded in a 5×1 µm² inverse-designed silicon structure. Its nonlinearity comes from saturable absorption: at low input the emitters scatter resonantly; at higher input they saturate and become nearly transparent, and interference between the two emitters yields a transmission change of magnitude 2. The supporting theoretical framework is a metric of expressive power that measures the growth factor r with which a nonlinear activation folds the curvature of an input trajectory per network lay

Load-bearing premise

The central claim hinges on the assumption that each trained complex energy-preserving weight matrix can be realized with high fidelity by the modular adjoint-optimized passive blocks, so that the network performance demonstrated in digital simulation (and, for classification, full-wave FDFD) survives physical implementation.

What would settle it

Measure the intensity-dependent transmission of a realized two-emitter device: if the output does not show a saturable transition with a transmission change approaching |Δt|=2 at input intensities around 1 nW/µm² (or an r-factor at least matching the digital baseline), the seven-order claim is falsified. Alternatively, in simulation, replace the two emitters with linear scatterers in the full-wave classification network; if accuracy does not collapse to the linear baseline, the nonlinearity is not doing the claimed work.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • All-optical neural networks equipped with this activation can solve nonlinear classification (spiral, MNIST, FashionMNIST) and reinforcement learning (Pong, HalfCheetah) tasks that linear optical networks cannot.
  • The required intensity for strong nonlinearity is below 1 nW/µm², seven orders of magnitude lower than graphene saturable absorbers and eleven orders lower than silicon Kerr nonlinearity for matching a digital expressive-power baseline.
  • System-level estimates show that all-optical LLMs could run on under 2.6 W, with optical power scaling as P ∝ N_param^0.66, sublinear in model size.
  • The expressive-power framework (curvature growth factor r per layer) provides a quantitative way to compare any physical nonlinearity against digital baselines.
  • The architecture is modular: linear blocks are inverse-designed to realize trained energy-preserving matrices, and quantum activation units are inserted between them, verified by full-wave FDFD simulation for classification.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct experimental test would be to fabricate the 5×1 µm² device with two deterministically placed SiV centers and measure its transmission curve; if the |Δt|=2 contrast does not appear at nW-level intensities with realistic dephasing, the seven-order advantage would shrink.
  • Because the RL demonstrations are trained digitally with the abstracted activation function rather than verified with full-wave simulation, a natural next step is to run the modular mapping through FDFD for a small RL policy to check that the learned control policy survives hardware transfer.
  • The sublinear power scaling suggests that if this nonlinearity is realized, the energy advantage of optical computing over electronics grows with model size; this could renew interest in optics for large-scale inference even if memory and fan-out constraints remain.
  • The expressive-power metric could be reused as a standard benchmark for any proposed optical activation, giving the community a single number (r at a given intensity) to compare.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes embedding quantum emitters (modeled as saturable two-level systems) in inverse-designed silicon nanophotonic structures to create an ultra-low-intensity nonlinear activation for all-optical neural networks. Using nonlinear FDFD simulations, the authors extract an input–output transmission curve, define an expressive-power metric based on the curvature growth factor r(I), and train optical neural networks via physics-aware backpropagation for nonlinear classification, Atari Pong, and HalfCheetah. Full-wave FDFD verification is reported for the classification network, while the reinforcement-learning agents are trained and evaluated digitally using the abstracted activation curve. The authors further use their expressive-power framework to compare the required intensity with conventional Kerr and graphene saturable-absorber nonlinearities, claiming a seven-orders-of-magnitude advantage, and extrapolate via Eq. (1) to sub-watt all-optical large-language-model inference.

Significance. The device concept addresses a real bottleneck in optical neural networks, and the FDFD modeling of the two-emitter activation unit, including the full-wave classification benchmark, is a concrete and useful contribution. The expressive-power framework is an interesting attempt to compare nonlinear platforms on a common basis, and the paper is candid about practical challenges such as cryogenic operation, inhomogeneous broadening, and bandwidth in the Discussion. However, the reinforcement-learning and large-scale language-model claims are not supported by the same level of verification: the RL policies are only simulated digitally, and the power estimates rely on the unvalidated assumption that modular mapping preserves network performance and that every activation is driven at I_min. If these gaps are addressed, or the claims are appropriately softened, this would be a valuable contribution to the field.

major comments (3)
  1. [§II.C, Fig. 3] The reinforcement-learning results are presented as demonstrating all-optical agents, but the main text verifies them only in the digital domain. Section II.C trains PPO/SAC policies in PyTorch using the abstracted activation curve; no full-wave nonlinear FDFD or experimental verification is reported for the Pong or HalfCheetah networks, despite the Introduction's statement that full-wave simulations verify 'as well as a reinforcement learning tasks.' The load-bearing assumption is that the modular mapping in §II.B—adjoint-optimized passive blocks implementing target energy-preserving matrices plus the single-unit activation—preserves network behavior at scale. This is not automatically true: inverse-designed unitary blocks have finite fidelity, and inter-module reflections, mode mismatch, or crosstalk can effectively change the activation. Please provide end-to-end FDFD verification for
  2. [§II.D, Eq. (1)] The large-language-model power estimate is not sufficiently justified. Equation (1) assumes each optical neuron occupies A = 0.1 µm² and is driven at I_min, and it counts only 3L_seq d_model input ports per layer. This neglects (i) insertion loss in the inverse-designed linear blocks, so driving every nonlinear unit at I_min requires higher input power; (ii) the need for repeated nonlinearities in attention (softmax) and feed-forward blocks, which are not captured by the curvature-growth model; and (iii) the fact that I_min was derived from a single simulated activation curve, not from a system-level power budget. Please state explicitly that this is an idealized lower bound, define all symbols, and provide a sensitivity analysis (e.g., with 1–3 dB insertion loss or with I_min varied over its uncertainty). Without this, the sub-watt and sublinear-scaling claims are premature.
  3. [§II.A / §II.D] The 'seven orders of magnitude' improvement is computed within the authors' own expressive-power framework, using r_digital = 1.045–1.095 as the baseline and r(I) from a single simulated transmission curve. This is an internal benchmark, not an externally validated one. The growth factor may depend on the input distribution and on how the activation is embedded in a deep network; a single-curve metric does not guarantee that the corresponding task performance transfers to hardware. Please validate the r(I) framework by computing growth factors for multiple input ensembles, or by correlating r(I) with the classification/RL accuracy after the modular mapping. As written, the claim 'reduces the required intensity by seven orders of magnitude' is supported only by a self-consistent calculation, not by a falsifiable cross-platform measurement.
minor comments (5)
  1. [§II.B, Fig. 2] The statement that the two emitters induce a maximal change in transmission coefficient |Δt| = 2 should specify the normalization and units (field amplitude vs. intensity), and how the 'impossible with a single emitter' proof is affected by losses in the inverse-designed structure.
  2. [§II.D, Eq. (1)] The notation 'L' is used both for network depth in the expressive-power discussion and for the number of transformer layers in Eq. (1); please disambiguate. Also clarify whether L_seq is the maximum context length or the actual inference sequence length, and whether Eq. (1) includes any optical power for the output linear readout.
  3. [§II.C, Fig. 3(d)] The human-level performance line (reward = 9.3) is taken from [2] for DQN on Atari; the comparison would be clearer if it is stated that this is an approximate reference and not a direct benchmark against the same optical agent.
  4. [§II.B and §II.D] The text reports r ≈ 1.2 at I = 1 nW/µm² in §II.B but later uses 'merely ~0.5 nW/µm²' for the same quantum activation in §II.D. These two numbers should be reconciled, or the dependence of I_min on the chosen r threshold should be stated explicitly.
  5. [Supplementary notes] The main text relies heavily on Supplementary Notes S1–S12 (e.g., design details, expressive-power derivation, LLM architectures). Please ensure that all referenced supplementary material is included in the posted version so that reviewers and readers can verify these details; the current arXiv version appears to omit them.

Circularity Check

0 steps flagged

No significant circularity: activation is physically characterized, comparison is a forward calculation, and the classification claim is full-wave verified.

full rationale

The paper's central derivation chain is not circular. The quantum activation unit is characterized by nonlinear FDFD simulations using a stated two-level-system dipole model (dz ∝ Ω/Γ0 / (1+2(Ω/Γ0)^2)) with parameters from SiV- color centers; this is first-principles physics, not a fit to the paper's conclusions. The expressive-power metric r is adopted from the cited external theory of Poole et al. and Raghu et al., and the intensity thresholds for silicon and graphene are computed from the same forward metric using known material parameters. The claim that quantum activation reaches the digital baseline at ~0.5 nW/µm² is therefore a quantitative result of the framework, not a restatement of an input. The classification network is additionally verified end-to-end with full-wave FDFD, providing independent support for the modular design flow. The only notable gap is that the reinforcement-learning sections (Pong, HalfCheetah) are trained and tested digitally with the abstracted activation curve, with no full-wave or experimental verification of the modular photonic mapping for those larger networks; this is a missing support for a strong claim, not a circular reduction. Similarly, the LLM power estimate (Eq. 1) is a straightforward extrapolation from I_min and assumed neuron area. No step in the paper is equivalent to its inputs by construction, and no load-bearing conclusion rests on a self-citation chain. The self-citation to the group's earlier FDTD modeling paper is a technical reference to an external-standard model, not an imported uniqueness or ansatz. Overall, no significant circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 0 invented entities

No new physical entities are introduced; the quantum emitters are existing solid-state defects (SiV- centers). The free parameters are the digital baseline, the per-neuron area, and the fitted scaling exponent, each of which directly influences the headline energy claims.

free parameters (3)
  • r_digital baseline (1.045–1.095) = 1.045–1.095
    Chosen as the target expressive-power growth for digital MLP activations; I_min for all platforms is defined relative to it. Different choices would move the comparison.
  • Effective neuron area A=0.1 µm² = 0.1 µm²
    Assumed for all-optical LLM power estimate (Eq. 1); taken from Ref. [6] for a different architecture and not justified for this device.
  • Scaling exponent 0.66 = 0.66
    Fitted to 9 LLM data points in Fig. 4(c); describes the data rather than a derived law.
axioms (6)
  • domain assumption Two-level saturable emitter model (dz ∝ Ω/Γ0 / (1+2(Ω/Γ0)^2))
    Used in Section II.B to simulate the quantum activation; assumes lifetime-limited linewidth, no dephasing, and steady-state behavior.
  • standard math Curvature-based expressive power framework from Refs. 39,40
    The paper adopts the total-curvature growth factor r as a measure of expressive power and assumes it governs depth-wise network expressivity.
  • domain assumption Nonlinear FDFD simulation fidelity
    Used for all device and network verification; assumes 2D effective models are valid for on-chip structures (Section II.B).
  • ad hoc to paper Passive adjoint-optimized blocks can realize target energy-preserving matrices
    Section II.B: 'we then run a separate adjoint-based optimization to realize a compact block that implements this target matrix with high fidelity.' This is assumed without scaling analysis.
  • ad hoc to paper Modular physics-aware training preserves performance when mapped to hardware
    The paper trains with an abstracted activation and then maps weights; the RL results are only verified in the abstracted model, not the full-wave model.
  • ad hoc to paper LLM power model Eq. (1) is valid
    Assumes each optical neuron is an independent channel with area A driven at I_min, ignoring inter-channel coupling and system losses.

pith-pipeline@v1.3.0-alltime-deepseek · 11727 in / 15200 out tokens · 154338 ms · 2026-08-03T12:44:44.525474+00:00 · methodology

0 comments
read the original abstract

The rapid scaling of deep neural networks comes at the cost of unsustainable power consumption. While optical neural networks offer an alternative, their capabilities remain constrained by the lack of efficient optical nonlinearities. To address this, we propose an optical neural computing architecture by embedding quantum emitters in inverse-designed nanophotonic structures. Due to their saturability, quantum emitters exhibit exceptionally strong nonlinearity compared with conventional materials. Using physics-aware training, we numerically demonstrate that the proposed architecture can solve complex tasks, including nonlinear classification and reinforcement learning, within all-optical neural networks. To enable fair comparison across different platforms, we introduce a framework that quantitatively links nonlinearity to a network's expressive power. Analysis shows that our quantum activation operates at $\text{nW}/\mu\text{m}^2$ intensity, which is seven orders of magnitude below the nonlinearity threshold of conventional optical materials. Looking ahead to large language models, we estimate the nonlinearity-limited optical power, which scales sublinearly with model size. Our results indicate that quantum nanophotonics may provide a route toward sustainable AI inference.

Figures

Figures reproduced from arXiv: 2601.01690 by Guoming Huang, Jungmin Kim, Ming Zhou, Qingyi Zhou, Yutian Tao, Zewei Shao, Zongfu Yu.

Figure 1
Figure 1. Figure 1: , a closed trajectory of data points serves as the input. The total curvature K of this trajectory provides a robust measure of curve complexity, and is monitored as the curve propagates through successive layers. Intuitively, the total curvature measures the degree of “folding” applied to the data manifold, a capability that is fundamentally [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

77 extracted references · 11 linked inside Pith

  1. [1]

    structural nonlinearity

    and game-playing [2] to protein design [3] and language processing [4]. Such progress has been driven by the continuous scaling of the model size. However, this scaling trend imposes an energy cost toward unsustainable levels [5]. A growing effort has been directed towards finding alternative computing paradigms. In particular, following the pioneering wo...

  2. [2]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in neural information processing systems25(2012)

  3. [3]

    V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski,et al., Human-level control through deep reinforcement learning, nature518, 529 (2015)

  4. [4]

    Jumper, R

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ ´ ıdek, A. Potapenko,et al., Highly accurate protein structure prediction with alphafold, nature596, 583 (2021)

  5. [5]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, Attention is all you need, Advances in neural information processing systems30(2017)

  6. [6]

    Patterson, J

    D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, Carbon emissions and large neural network training, arXiv preprint arXiv:2104.10350 (2021)

  7. [7]

    Y. Shen, N. C. Harris, S. Skirlo, M. Prabhu, T. Baehr-Jones, M. Hochberg, X. Sun, S. Zhao, H. Larochelle, D. Englund, et al., Deep learning with coherent nanophotonic circuits, Nature photonics11, 441 (2017)

  8. [8]

    X. Lin, Y. Rivenson, N. T. Yardimci, M. Veli, Y. Luo, M. Jarrahi, and A. Ozcan, All-optical machine learning using diffractive deep neural networks, Science361, 1004 (2018)

  9. [9]

    Shekhar, W

    S. Shekhar, W. Bogaerts, L. Chrostowski, J. E. Bowers, M. Hochberg, R. Soref, and B. J. Shastri, Roadmapping the next generation of silicon photonics, Nature Communications15, 751 (2024)

  10. [10]

    Wetzstein, A

    G. Wetzstein, A. Ozcan, S. Gigan, S. Fan, D. Englund, M. Soljaˇ ci´ c, C. Denz, D. A. Miller, and D. Psaltis, Inference in artificial intelligence with deep optics and photonics, Nature588, 39 (2020)

  11. [11]

    Wang, S.-Y

    T. Wang, S.-Y. Ma, L. G. Wright, T. Onodera, B. C. Richard, and P. L. McMahon, An optical neural network using less than 1 photon per multiplication, Nature Communications13, 123 (2022)

  12. [12]

    W. Shi, Z. Huang, T. Fu, and H. Chen, Review of nonlinear activation functions in optical neural networks, Advanced Photonics7, 064004 (2025)

  13. [13]

    R. W. Boyd, A. L. Gaeta, and E. Giese, Nonlinear optics, inSpringer Handbook of Atomic, Molecular, and Optical Physics (Springer, 2008) pp. 1097–1110

  14. [14]

    Feldmann, N

    J. Feldmann, N. Youngblood, C. D. Wright, H. Bhaskaran, and W. H. Pernice, All-optical spiking neurosynaptic networks with self-learning capabilities, Nature569, 208 (2019)

  15. [15]

    B. Wu, H. Li, W. Tong, J. Dong, and X. Zhang, Low-threshold all-optical nonlinear activation function based on a ge/si hybrid structure in a microring resonator, Optical Materials Express12, 970 (2022)

  16. [16]

    T. Wu, Y. Li, L. Ge, and L. Feng, Field-programmable photonic nonlinearity, Nature Photonics , 1 (2025)

  17. [17]

    A. Jha, C. Huang, and P. R. Prucnal, Reconfigurable all-optical nonlinear activation functions for neuromorphic photonics, Optics letters45, 4819 (2020)

  18. [18]

    Yanagimoto, B

    R. Yanagimoto, B. A. Ash, M. M. Sohoni, M. M. Stein, Y. Zhao, F. Presutti, M. Jankowski, L. G. Wright, T. Onodera, and P. L. McMahon, Programmable on-chip nonlinear photonics, Nature , 1 (2025)

  19. [19]

    I. A. Williamson, T. W. Hughes, M. Minkov, B. Bartlett, S. Pai, and S. Fan, Reprogrammable electro-optic nonlinear activation functions for optical neural networks, IEEE Journal of Selected Topics in Quantum Electronics26, 1 (2019)

  20. [20]

    M. M. Pour Fard, I. A. Williamson, M. Edwards, K. Liu, S. Pai, B. Bartlett, M. Minkov, T. W. Hughes, S. Fan, and T.-A. Nguyen, Experimental realization of arbitrary activation functions for optical neural networks, Optics Express28, 12138 (2020)

  21. [21]

    Yildirim, N

    M. Yildirim, N. U. Dinc, I. Oguz, D. Psaltis, and C. Moser, Nonlinear processing with linear optics, Nature Photonics18, 1076 (2024)

  22. [22]

    F. Xia, K. Kim, Y. Eliezer, S. Han, L. Shaughnessy, S. Gigan, and H. Cao, Nonlinear optical encoding enabled by recurrent linear scattering, Nature Photonics18, 1067 (2024)

  23. [23]

    C. C. Wanjura and F. Marquardt, Fully nonlinear neuromorphic computing with linear wave scattering, Nature Physics 20, 1434 (2024)

  24. [24]

    Lodahl, S

    P. Lodahl, S. Mahmoodian, and S. Stobbe, Interfacing single photons and single quantum dots with photonic nanostruc- tures, Reviews of Modern Physics87, 347 (2015)

  25. [25]

    C. Zhu, T. Wang, P. L. McMahon, and D. Soh, Quantum optical neural networks using atom-cavity interactions to provide all-optical nonlinearity, arXiv preprint arXiv:2511.06167 (2025). 9

  26. [26]

    Canora, X

    R. Canora, X. Xu, Z. Niu, H. Alaeian, and S. Du, Engineering nonlinear activation functions for all-optical neural networks via quantum interference, arXiv preprint arXiv:2504.04009 (2025)

  27. [27]

    Y. Zuo, B. Li, Y. Zhao, Y. Jiang, Y.-C. Chen, P. Chen, G.-B. Jo, J. Liu, and S. Du, All-optical neural network with nonlinear activation functions, Optica6, 1132 (2019)

  28. [28]

    A. Ryou, J. Whitehead, M. Zhelyeznyakov, P. Anderson, C. Keskin, M. Bajcsy, and A. Majumdar, Free-space optical neural network based on thermal atomic nonlinearity, Photonics Research9, B128 (2021)

  29. [29]

    Huang, W

    Z. Huang, W. Shi, S. Wu, Y. Wang, S. Yang, and H. Chen, Pre-sensor computing with compact multilayer optical neural network, Science Advances10, eado8516 (2024)

  30. [30]

    K. Ohno, F. J. Heremans, L. C. Bassett, B. A. Myers, D. M. Toyli, A. C. B. Jayich, C. J. Palmstrøm, and D. D. Awschalom, Engineering shallow spins in diamond with nitrogen delta-doping, Applied Physics Letters101, 082413 (2012)

  31. [31]

    Y.-C. Chen, B. Griffiths, L. Weng, S. S. Nicley, S. N. Ishmael, Y. Lekhai, S. Johnson, C. J. Stephen, B. L. Green, G. W. Morley,et al., Laser writing of individual nitrogen-vacancy defects in diamond with near-unity yield, Optica6, 662 (2019)

  32. [32]

    A. M. Day, J. R. Dietz, M. Sutula, M. Yeh, and E. L. Hu, Laser writing of spin defects in nanophotonic cavities, Nature Materials22, 696 (2023)

  33. [33]

    C. M. Lalau-Keraly, S. Bhargava, O. D. Miller, and E. Yablonovitch, Adjoint shape optimization applied to electromagnetic design, Optics express21, 21693 (2013)

  34. [34]

    Kulce, D

    O. Kulce, D. Mengu, Y. Rivenson, and A. Ozcan, All-optical information-processing capacity of diffractive surfaces, Light: Science & Applications10, 25 (2021)

  35. [35]

    D. A. Miller, Why optics needs thickness, Science379, 41 (2023)

  36. [36]

    Li and F

    Y. Li and F. Monticone, The spatial complexity of optical computing: toward space-efficient design, Nature Communica- tions16, 8588 (2025)

  37. [37]

    Onodera, M

    T. Onodera, M. M. Stein, B. A. Ash, M. M. Sohoni, M. Bosch, R. Yanagimoto, M. Jankowski, T. P. McKenna, T. Wang, G. Shvets,et al., Arbitrary control over multimode wave propagation for machine learning, Nature Physics , 1 (2025)

  38. [38]

    S. Yu, X. Piao, and N. Park, Nonlinear unitary circuits for photonic neural networks, ACS Photonics (2025)

  39. [39]

    Mont´ ufar, R

    G. Mont´ ufar, R. Pascanu, K. Cho, and Y. Bengio, On the number of linear regions of deep neural networks, Advances in neural information processing systems27(2014)

  40. [40]

    Poole, S

    B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli, Exponential expressivity in deep neural networks through transient chaos, Advances in neural information processing systems29(2016)

  41. [41]

    Raghu, B

    M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, and J. Sohl-Dickstein, On the expressive power of deep neural networks, in international conference on machine learning(PMLR, 2017) pp. 2847–2854

  42. [42]

    Y. Shi, J. Ren, G. Chen, W. Liu, C. Jin, X. Guo, Y. Yu, and X. Zhang, Nonlinear germanium-silicon photodiode for activation and monitoring in photonic neuromorphic networks, Nature Communications13, 6048 (2022)

  43. [43]

    Q. Bao, H. Zhang, Y. Wang, Z. Ni, Y. Yan, Z. X. Shen, K. P. Loh, and D. Y. Tang, Atomic-layer graphene as a saturable absorber for ultrafast pulsed lasers, Advanced Functional Materials19, 3077 (2009)

  44. [44]

    J. Shi, P. Yu, F. Liu, P. He, R. Wang, L. Qin, J. Zhou, X. Li, J. Zhou, X. Sui,et al., 3r mos2 with broken inversion symmetry: a promising ultrathin nonlinear optical device, Advanced Materials29, 1701486 (2017)

  45. [45]

    B. Liu, K. Liang, Q. Zhou, A. R. Khan, Z. Lu, T. Yildirim, X. Sun, S. Rahman, Y. Liu, Z. Yu,et al., Giant second harmonic generation in two-dimensional tellurene with synthesis and thickness engineering, Applied physics reviews12 (2025)

  46. [46]

    Hanamura, Rapid radiative decay and enhanced optical nonlinearity of excitons in a quantum well, Physical Review B 38, 1228 (1988)

    E. Hanamura, Rapid radiative decay and enhanced optical nonlinearity of excitons in a quantum well, Physical Review B 38, 1228 (1988)

  47. [47]

    Javadi, I

    A. Javadi, I. S¨ ollner, M. Arcari, S. L. Hansen, L. Midolo, S. Mahmoodian, G. Kirˇ sansk˙ e, T. Pregnolato, E. Lee, J. Song, et al., Single-photon non-linear optics with a quantum dot in a waveguide, Nature communications6, 8655 (2015)

  48. [48]

    J. Volz, M. Scheucher, C. Junge, and A. Rauschenbeutel, Nonlinearπphase shift for single fibre-guided photons interacting with a single resonator-enhanced atom, Nature Photonics8, 965 (2014)

  49. [49]

    Shomroni, S

    I. Shomroni, S. Rosenblum, Y. Lovsky, O. Bechler, G. Guendelman, and B. Dayan, All-optical routing of single photons by a one-atom switch controlled by a single photon, Science345, 903 (2014)

  50. [50]

    Hacker, S

    B. Hacker, S. Welte, G. Rempe, and S. Ritter, A photon–photon quantum gate based on a single atom in an optical resonator, Nature536, 193 (2016)

  51. [51]

    D. M. Lukin, C. Dory, M. A. Guidry, K. Y. Yang, S. D. Mishra, R. Trivedi, M. Radulaski, S. Sun, D. Vercruysse, G. H. Ahn,et al., 4h-silicon-carbide-on-insulator for integrated quantum and nonlinear photonics, Nature Photonics14, 330 (2020)

  52. [52]

    Q. Zhou, S. Gangaraj, M. Zhou, and Z. Yu, Simulating quantum emitters in arbitrary photonic environments using fdtd: beyond the semi-classical regime, arXiv preprint arXiv:2410.16118 (2024)

  53. [53]

    Wang and S

    H. Wang and S. Fan, Lorentz–drude dipoles in the radiative limit and their modeling in finite-difference time-domain methods, Annalen der Physik , e00156 (2025)

  54. [54]

    Nikkhah, A

    V. Nikkhah, A. Pirmoradi, F. Ashtiani, B. Edwards, F. Aflatouni, and N. Engheta, Inverse-designed low-index-contrast structures on a silicon photonics platform for vector–matrix multiplication, Nature Photonics18, 501 (2024)

  55. [55]

    Silver, A

    D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Pan- neershelvam, M. Lanctot,et al., Mastering the game of go with deep neural networks and tree search, nature529, 484 (2016)

  56. [56]

    Hwangbo, J

    J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter, Learning agile and dynamic motor skills for legged robots, Science Robotics4, eaau5872 (2019)

  57. [57]

    Towers, A

    M. Towers, A. Kwiatkowski, J. Terry, J. U. Balis, G. De Cola, T. Deleu, M. Goul˜ ao, A. Kallinteris, M. Krimmel, A. KG, 10 et al., Gymnasium: A standard interface for reinforcement learning environments, arXiv preprint arXiv:2407.17032 (2024)

  58. [58]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347 (2017)

  59. [59]

    Haarnoja, A

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, inInternational conference on machine learning(Pmlr, 2018) pp. 1861–1870

  60. [60]

    M. Dinu, F. Quochi, and H. Garcia, Third-order nonlinearities in silicon at telecom wavelengths, Applied physics letters 82, 2954 (2003)

  61. [61]

    Dulkeith, Y

    E. Dulkeith, Y. A. Vlasov, X. Chen, N. C. Panoiu, and R. M. Osgood Jr, Self-phase-modulation in submicron silicon-on- insulator photonic wires, Optics express14, 5524 (2006)

  62. [62]

    K. Y. Lau, X. Liu, and J. Qiu, A comparison for saturable absorbers: Carbon nanotube versus graphene, Advanced Photonics Research3, 2200023 (2022)

  63. [63]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever,et al., Language models are unsupervised multitask learners, OpenAI blog1, 9 (2019)

  64. [64]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in neural information processing systems33, 1877 (2020)

  65. [65]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi` ere, N. Goyal, E. Hambro, F. Azhar, et al., Llama: Open and efficient foundation language models, arXiv preprint arXiv:2302.13971 (2023)

  66. [66]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al., Llama 2: Open foundation and fine-tuned chat models, arXiv preprint arXiv:2307.09288 (2023)

  67. [67]

    Dubey, A

    A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan,et al., The llama 3 herd of models, arXiv e-prints , arXiv (2024)

  68. [68]

    X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu,et al., Deepseek llm: Scaling open-source language models with longtermism, arXiv preprint arXiv:2401.02954 (2024)

  69. [69]

    A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Dengr, C. Ruan, D. Dai, D. Guo,et al., Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, arXiv preprint arXiv:2405.04434 (2024)

  70. [70]

    A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan,et al., Deepseek-v3 technical report, arXiv preprint arXiv:2412.19437 (2024)

  71. [71]

    Anderson, S.-Y

    M. Anderson, S.-Y. Ma, T. Wang, L. Wright, and P. McMahon, Optical transformers, Transactions on Machine Learning Research (2023)

  72. [72]

    Englund, D

    D. Englund, D. Fattal, E. Waks, G. Solomon, B. Zhang, T. Nakaoka, Y. Arakawa, Y. Yamamoto, and J. Vuˇ ckovi´ c, Controlling the spontaneous emission rate of single quantum dots in a two-dimensional photonic crystal, Physical review letters95, 013904 (2005)

  73. [73]

    Laucht, J

    A. Laucht, J. Villas-Bˆ oas, S. Stobbe, N. Hauke, F. Hofbauer, G. B¨ ohm, P. Lodahl, M.-C. Amann, M. Kaniber, and J. Finley, Mutual coupling of two semiconductor quantum dots via an optical nanocavity, Physical Review B—Condensed Matter and Materials Physics82, 075305 (2010)

  74. [74]

    Y.-C. Chen, P. S. Salter, S. Knauer, L. Weng, A. C. Frangeskou, C. J. Stephen, S. N. Ishmael, P. R. Dolan, S. Johnson, B. L. Green,et al., Laser writing of coherent colour centres in diamond, Nature Photonics11, 77 (2017)

  75. [75]

    Laferri` ere, E

    P. Laferri` ere, E. Yeung, I. Miron, D. B. Northeast, S. Haffouz, J. Lapointe, M. Korkusinski, P. J. Poole, R. L. Williams, and D. Dalacu, Unity yield of deterministically positioned quantum dot single photon sources, Scientific Reports12, 6376 (2022)

  76. [76]

    Sipahigil, K

    A. Sipahigil, K. D. Jahnke, L. J. Rogers, T. Teraji, J. Isoya, A. S. Zibrov, F. Jelezko, and M. D. Lukin, Indistinguishable photons from separated silicon-vacancy centers in diamond, Physical Review Letters113, 113602 (2014)

  77. [77]

    L. J. Rogers, K. D. Jahnke, T. Teraji, L. Marseglia, C. M¨ uller, B. Naydenov, H. Schauffert, C. Kranz, J. Isoya, L. P. McGuinness, and F. Jelezko, Multiple intrinsically identical single-photon emitters in the solid state, Nature Communica- tions5, 4739 (2014)