REVIEW 4 major objections 8 minor 1 cited by
Perfecting Imperfect Physical Neural Networks with Transferable Robustness using Sharpness-Aware Training
T0 review · 4 major / 8 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read By training physical neural networks to sit in flat regions of the loss landscape, Sharpness-Aware Training keeps them accurate despite imperfect models, fabrication error, and environmental drift.
desk verdict A serious and useful application of SAM to physical neural networks, with real hardware evidence, but the universality and transfer claims outrun the current data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the geometry of the loss landscape over control parameters. SAT minimizes $L_1 = L(y, y_{\mathrm{target}};\Theta) + \alpha \|\partial L/\partial\Theta\|^2$, penalizing the gradient norm to force parameters into uniformly low-loss neighborhoods. To avoid Hessian computation, it uses the sharpness-aware minimization trick: compute $\Delta\Theta = \rho\,\partial L/\partial\Theta / \|\partial L/\partial\Theta\|_2$, take a forward pass at $\Theta+\Delta\Theta$ (the point of maximum loss in the neighborhood), and update from the gradient at that point. For parameters like alignment angles with no explicit model, the same two-step procedure runs with finite-difference gradient estimates. The two-step differentiation is what carries the argument: it converts training from 'match the model' to 'find a flat minimum,' and the flatness of the minimum is what the paper claims transfers to real hardware.
What would settle it
Train SAT offline on a deliberately coarse model and deploy it on several chips whose per-device tuning curves lie outside the modeled spread; if any chip's accuracy degrades as steeply as under standard backpropagation, the landscape-transfer premise fails. A direct check is to measure the curvature of the physical loss at the deployed parameters: if it is large on hardware while small in the model, the geometry has not transferred.
Extended reading notes
Core claim
SAT changes what is being optimized: instead of finding parameters that match a model, it finds parameters whose neighborhood has uniformly low loss. The paper connects sharp minima to hardware fragility and flat minima to robustness, arguing that the loss landscape geometry, not model fidelity, is what transfers from the digital training environment to the physical chip. On an MRR weight bank, SAT-trained parameters reach 97.0% MNIST handwritten-digit accuracy on hardware where standard backpropagation falls to 80.0%, and survive a 2 degree C temperature shift at 91.0% versus 52.0%. In a free-space diffractive network, SAT holds 98.0% accuracy at 1 degree rotation misalignment where standard training drops to 43.0%. In simulated MZI (Mach-Zehnder interferometer) meshes with fabrication error, offline SAT reaches 94.1% accuracy compared with 92.3% for dual-adaptive online training, with the largest Hessian eigenvalue (a standard curvature measure) falling from 246.02 to 0.63; adding SAT to the online framework reaches 96.1% and enables transfer to other error-perturbed devices at 95.2% accuracy. The paper's central assertion is that flatness is the transferable quantity, so SAT works whether or not the physical model is explicitly known.
Load-bearing premise
Flat minima found in the approximate digital model remain flat and low-loss on the real physical system, despite fabrication variation, thermal crosstalk, and all the side effects the training model ignores.
Editorial extensions
If this is right
- Offline training becomes viable without an exact digital model, since flat minima tolerate the model–system gap.
- Trained PNN parameters can be deployed across nominally identical devices, removing the device-specificity of online training.
- Deployed PNNs can operate through thermal drift, alignment shifts, and other perturbations without retraining.
- SAT can be stacked on existing online gradient-estimation frameworks to improve both accuracy and robustness.
- The extra cost of two-step differentiation remains small relative to measurement-heavy training, and a zero-extra-cost sharpness surrogate exists.
Reading between the lines
- Inference: since SAT minimizes worst-case loss in a neighborhood, it should also dampen correlated global errors such as uniform temperature offsets; a testable extension is to stress trained models with simultaneous thermal, alignment, and fabrication errors rather than one error at a time.
- Inference: the paper's three demonstrations are all photonic; the universality claim would be strengthened or bounded by applying the same two-step flatness objective to non-optical analog substrates, such as analog electronic or spintronic networks, where parameter-to-weight maps have different smoothness.
- Inference: the component-level term $\partial W/\partial\Theta$ suggests SAT should also reduce sensitivity to control-electronics drift, such as current-source or bias-voltage drift, which the experiments probe only through temperature changes.
- Inference: treating flatness as a transferable quantity implies a trade-off: at some level of modeling error, even a flat minimum's basin will not contain the true hardware loss; a natural follow-up is to characterize the maximum model-error magnitude SAT can tolerate.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Sharpness-Aware Training (SAT) for physical neural networks (PNNs), adapting Sharpness-Aware Minimization to train parameters in flat regions of the loss landscape. The authors claim that SAT enables offline training with imperfect digital models to outperform online training, provides robustness to post-deployment perturbations, and allows transfer of trained parameters across devices. They demonstrate the method on a microring-resonator (MRR) weight-bank experiment, a diffractive optical NN experiment, and an MZI-mesh simulation, comparing against standard backpropagation, physical-aware training (PAT), dual-adaptive training (DAT), and optical pruning.
Significance. If the claims are substantiated, SAT would be a valuable and practical tool for training analog hardware without requiring exact models, with potential benefits for deployment robustness and model transfer. The paper's strengths include the two hardware demonstrations, the use of a deliberately simplified model in the MRR experiment, and the comparison with several existing training methods. The work also transparently reports training times and Hessian eigenvalue measurements. However, the evidence is currently insufficient to support the universality and transferability claims, and the central algorithm contains a mis-specified equation that must be corrected.
major comments (4)
- [Section 2.1, Eq. (4)] Equation (4) as written cannot be correct: the derivative of (Theta + Delta Theta) with respect to Theta is the identity matrix, not the gradient of the sharpness-aware loss. The SAM update should use the gradient of the loss evaluated at Theta + Delta Theta, i.e., grad L(Theta + Delta Theta), and the role of alpha1 relative to the alpha in Eq. (2) is undefined. Because this equation defines the proposed training scheme, the algorithm description needs to be corrected and clarified before the experiments can be reproduced.
- [Sections 2.2, 2.3, 2.4] The core claim that flat minima found in the approximate digital model remain flat and low-loss on the physical hardware is not established. The MRR experiment uses a deliberately simplified model (identical MRRs, no thermal crosstalk) and tests only temperature drift; the diffractive experiment (Section 2.3) tests robustness only to the same rotation/shift/scale perturbations used in the adversarial update; the MZI simulation (Section 2.4) uses i.i.d. Gaussian phase/splitting errors and targets sampled from the same distribution. No experiment or analysis addresses correlated, non-Gaussian, or unanticipated error sources, so the universality and device-transfer claims are under-supported. This is the load-bearing premise of the offline training claim in Section 2.2, and it should be explicitly tested or theoretically qualified.
- [Figures 2h, 2j, 3e] The reported inference accuracies appear to be single measurements without error bars, repeats, or statistical analysis. Given the small hardware scale (four MRRs, a single chip, one free-space setup), the quantitative comparisons (e.g., 97.0% vs 80.0% at 22C, 91.0% vs 52.0% without TEC, 98.0% vs 43.0% at 1 degree rotation) should be supported by repeated trials or measurement uncertainty to justify the claim that SAT outperforms standard BP and online training.
- [Section 2.4, Figure 4d] The transfer experiment is in-distribution: target devices are generated from the same zero-mean Gaussian error model (sigma_bs = sigma_ps = 0.15) used in training. Real fabrication errors are typically correlated across devices and may have systematic offsets or non-Gaussian statistics. The paper needs a sensitivity analysis over error statistics (e.g., varying the noise distribution, adding correlations or biases) to support the claim that SAT enables reliable transfer to other devices.
minor comments (8)
- [Section 2.1, Eq. (2)] The statement that computing the second term in Eq. (2) requires computing second-order derivatives, the Hessian matrix, is imprecise: the term ||grad L||^2 is first-order; its gradient involves the Hessian. Please clarify the connection between this penalty and loss sharpness, and how the SAM approximation avoids the Hessian.
- [Eq. (2)] The notation L1 for the total loss and L for the task loss is not defined and is confusing; please define the relationship between the two.
- [Section 2.1] The claim that, for the first time, sharp minima are associated with poor robustness of physical systems is difficult to verify and should be softened or supported with a citation.
- [Figure 1d] The decomposition of grad L into partial W over partial Theta and partial L over partial W is a chain-rule identity rather than a derivation of two distinct stability objectives; the text should clarify what is new here.
- [Abstract and Discussion] The term 'universally applicable' is too strong given that all demonstrations are optical NNs on MNIST classification with small-scale hardware; consider limiting the claim to the tested platforms.
- [Section 2.4] The MZI demonstration is carried out on a digital-optical hybrid network where 'online' training with PAT/DAT is simulated rather than performed on hardware; the text sometimes blurs this distinction and should be explicit that Figure 4 reports simulations.
- [Throughout] The manuscript contains several typos and informal notations, including 'singal-end' (Section 2.2), 'develope' (Methods 4.3), 'F A' for EDFA (Methods 4.2), and the informal citation 'Ziyang et.al.' in Section 2.4; these should be corrected.
- [Figure 2j and Section 2.2] The comparison with optical pruning is said to come from simulations, but Figure 2j presents the pruning result alongside experimental numbers for BP and SAT without marking which points are simulated; please add a legend or state the simulated nature explicitly.
Circularity Check
No circular derivation: the paper applies the established SAM objective to PNNs and validates it with genuine hardware experiments and independent simulations; the flatness-transfer premise is an unsupported generalization, not a relabeled fit or self-citation chain.
full rationale
I walked the claimed derivation chain and found no step in which a prediction reduces, by the paper's own equations or by load-bearing self-citation, to its inputs. The core objective in Eq. (2) adds the squared gradient norm to the loss; this is a training objective, not a prediction. The SAM-type update in Eqs. (3)-(4) is explicitly attributed to Foret et al. [42], and the paper implements it as described. In the MRR experiment, training is done on a deliberately simplified digital model (all MRRs treated as identical, side effects ignored) and then the resulting currents are deployed on a real chip; the reported accuracies (97% vs 80% for standard BP, and 91% without TEC) are external hardware measurements, so no fitted quantity is renamed as a prediction. In the diffractive experiment, gradients with respect to misalignment parameters are estimated by finite differences, and the trained weights are then evaluated on the physical free-space setup; this is again a genuine out-of-sample hardware measurement, even though the tested perturbation classes overlap with those used in training. The MZI transfer study trains offline on an ideal model or online in a PAT-style simulated device and then tests on independently sampled error instances; the target devices are not used to fit any parameter. The decomposition of dL/dTheta into dL/dW times dW/dTheta in Fig. 1d is simply the chain rule and is not used to derive a circular conclusion. Self-citations, including [18] and [25], provide background and a baseline for comparison; they are not load-bearing. The main weakness, namely that flatness in the approximate digital model is assumed to persist on real hardware under arbitrary modeling error, is an empirical overgeneralization and a correctness/support concern, not a circular step: the central claims are supported by independent hardware and simulation evidence that does not reduce to the training objective itself.
Assumptions & free parameters
free parameters (3)
- Perturbation radius r (Eq. 3) =
not reported in main text
- Sharpness regularization weight alpha1 (Eq. 4) =
not reported in main text
- Finite-difference perturbation delta_theta, angle step mu, weight step rho (Eqs. 5-6) =
not reported in main text
assumptions (5)
- domain assumption Flat minima in the approximate digital training model remain low-loss regions on the real physical system.
- domain assumption Fabrication variance in MZI meshes is modeled as zero-mean Gaussian phase and beam-splitter errors with variance 0.15.
- domain assumption The one-sided finite-difference gradient dL/dtheta is an adequate estimator for updating misalignment parameters during SAM.
- domain assumption The maximum Hessian eigenvalue of the digital training model is a valid proxy for deployment robustness.
- standard math The first-order SAM gradient approximation from Foret et al. [42] is valid for the physical parameterization.
Cite this review
Pith. "Pith review of Perfecting Imperfect Physical Neural Networks with Transferable Robustness using Sharpness-Aware Training." pith.science (2026). https://pith.science/paper/4Y7W7FXJ
@misc{pith2026241112352,
author = {Pith},
title = {Pith review of: Perfecting Imperfect Physical Neural Networks with Transferable Robustness using Sharpness-Aware Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/4Y7W7FXJ}},
note = {Machine review of arXiv:2411.12352}
}
read the original abstract
AI models are essential in science and engineering, but recent advances are pushing the limits of traditional digital hardware. To address these limitations, physical neural networks (PNNs), which use physical substrates for computation, have gained increasing attention. However, developing effective training methods for PNNs remains a significant challenge. Current approaches, regardless of offline and online training, suffer from significant accuracy loss. Offline training is hindered by imprecise modeling, while online training yields device-specific models that can't be transferred to other devices due to manufacturing variances. Both methods face challenges from perturbations after deployment, such as thermal drift or alignment errors, which make trained models invalid and require retraining. Here, we address the challenges with both offline and online training through a novel technique called Sharpness-Aware Training (SAT), where we innovatively leverage the geometry of the loss landscape to tackle the problems in training physical systems. SAT enables accurate training using efficient backpropagation algorithms, even with imprecise models. PNNs trained by SAT offline even outperform those trained online, despite modeling and fabrication errors. SAT also overcomes online training limitations by enabling reliable transfer of models between devices. Finally, SAT is highly resilient to perturbations after deployment, allowing PNNs to continuously operate accurately under perturbations without retraining. We demonstrate SAT across three types of PNNs, showing it is universally applicable, regardless of whether the models are explicitly known. This work offers a transformative, efficient approach to training PNNs, addressing critical challenges in analog computing and enabling real-world deployment.
Forward citations
Cited by 1 Pith paper
-
Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization
SAM's calibration benefit is attributed to implicit entropy maximization, but the proof rests on an unstated gradient-norm assumption; CSAM shows further ECE reductions.
Reference graph
Works this paper leans on
-
[1]
Nature 521(7553), 436–444 (2015)
LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature 521(7553), 436–444 (2015)
2015
-
[2]
Nature 323(6088), 533–536 (1986)
Rumelhart, D.E., Hinton, G.E., Williams, R.J.: Learning representations by back- propagating errors. Nature 323(6088), 533–536 (1986)
1986
-
[3]
Science 382(6676), 1297–1303 (2023)
Momeni, A., Rahmani, B., Mall´ ejac, M., Del Hougne, P., Fleury, R.: Backpropagation-free training of deep physical neural networks. Science 382(6676), 1297–1303 (2023)
2023
-
[4]
Nature 601(7894), 549–555 (2022)
Wright, L.G., Onodera, T., Stein, M.M., Wang, T., Schachter, D.T., Hu, Z., McMahon, P.L.: Deep physical neural networks trained with backpropagation. Nature 601(7894), 549–555 (2022)
work page 2022
-
[5]
Nature 588(7836), 39–47 (2020)
Wetzstein, G., Ozcan, A., Gigan, S., Fan, S., Englund, D., Soljaˇ ci´ c, M., Denz, C., Miller, D.A., Psaltis, D.: Inference in artificial intelligence with deep optics and photonics. Nature 588(7836), 39–47 (2020)
work page 2020
-
[6]
Nature563(7730), 230–234 (2018) 20
Romera, M., Talatchian, P., Tsunegi, S., Abreu Araujo, F., Cros, V., Bortolotti, P., Trastoy, J., Yakushiji, K., Fukushima, A., Kubota, H.,et al.: Vowel recognition with four coupled spin-torque nano-oscillators. Nature563(7730), 230–234 (2018) 20
work page 2018
-
[7]
Nature 577(7790), 341–345 (2020)
Chen, T., Gelder, J., Ven, B., Amitonov, S.V., De Wilde, B., Ruiz Euler, H.-C., Broersma, H., Bobbert, P.A., Zwanenburg, F.A., Wiel, W.G.: Classification with a disordered dopant-atom network in silicon. Nature 577(7790), 341–345 (2020)
work page 2020
-
[8]
Nature Electronics 3(7), 360–370 (2020)
Grollier, J., Querlioz, D., Camsari, K., Everschor-Sitte, K., Fukami, S., Stiles, M.D.: Neuromorphic spintronics. Nature Electronics 3(7), 360–370 (2020)
work page 2020
Show all 49 references
-
[9]
Science 384(6692), 202–209 (2024)
Xu, Z., Zhou, T., Ma, M., Deng, C., Dai, Q., Fang, L.: Large-scale photonic chiplet taichi empowers 160-tops/w artificial general intelligence. Science 384(6692), 202–209 (2024)
2024
-
[10]
Nature Photonics, 1–9 (2024)
Xia, F., Kim, K., Eliezer, Y., Han, S., Shaughnessy, L., Gigan, S., Cao, H.: Non- linear optical encoding enabled by recurrent linear scattering. Nature Photonics, 1–9 (2024)
2024
-
[11]
Nature Photonics, 1–7 (2024)
Yildirim, M., Dinc, N.U., Oguz, I., Psaltis, D., Moser, C.: Nonlinear processing with linear optics. Nature Photonics, 1–7 (2024)
2024
-
[12]
Light: Science & Applications 13(1), 263 (2024)
Fu, T., Zhang, J., Sun, R., Huang, Y., Xu, W., Yang, S., Zhu, Z., Chen, H.: Optical neural networks: progress and challenges. Light: Science & Applications 13(1), 263 (2024)
2024
-
[13]
Nature 623(7985), 48–57 (2023)
Chen, Y., Nazhamaiti, M., Xu, H., Meng, Y., Zhou, T., Li, G., Fan, J., Wei, Q., Wu, J., Qiao, F., et al.: All-analog photoelectronic chip for high-speed vision tasks. Nature 623(7985), 48–57 (2023)
2023
-
[14]
Nature Photonics, 1–12 (2023)
Youngblood, N., R ´ ıos Ocampo, C.A., Pernice, W.H., Bhaskaran, H.: Integrated optical memristors. Nature Photonics, 1–12 (2023)
2023
-
[15]
Nature Photonics 17(8), 723–730 (2023)
Chen, Z., Sludds, A., Davis III, R., Christen, I., Bernstein, L., Ateshian, L., Heuser, T., Heermeier, N., Lott, J.A., Reitzenstein, S., et al.: Deep learning with coherent vcsel neural networks. Nature Photonics 17(8), 723–730 (2023)
2023
-
[16]
Nature 606(7914), 501–506 (2022)
Ashtiani, F., Geers, A.J., Aflatouni, F.: An on-chip photonic deep neural network for image classification. Nature 606(7914), 501–506 (2022)
2022
-
[17]
Light: Science & Applications 11(1), 30 (2022)
Zhou, H., Dong, J., Cheng, J., Dong, W., Huang, C., Shen, Y., Zhang, Q., Gu, M., Qian, C., Chen, H., et al.: Photonic matrix multiplication lights up photonic accelerator and beyond. Light: Science & Applications 11(1), 30 (2022)
2022
-
[18]
Nature Electronics 4(11), 837–844 (2021)
Huang, C., Fujisawa, S., Lima, T.F., Tait, A.N., Blow, E.C., Tian, Y., Bilodeau, S., Jha, A., Yaman, F., Peng, H.-T., et al.: A silicon photonic–electronic neural network for fibre nonlinearity compensation. Nature Electronics 4(11), 837–844 (2021)
2021
-
[19]
Nature 589(7840), 52–58 (2021)
Feldmann, J., Youngblood, N., Karpov, M., Gehring, H., Li, X., Stappers, M., Le Gallo, M., Fu, X., Lukashchuk, A., Raja, A.S., et al.: Parallel convolutional 21 processing using an integrated photonic tensor core. Nature 589(7840), 52–58 (2021)
2021
-
[20]
Nature Photonics 15(2), 102–114 (2021)
Shastri, B.J., Tait, A.N., Lima, T., Pernice, W.H., Bhaskaran, H., Wright, C.D., Prucnal, P.R.: Photonics for artificial intelligence and neuromorphic computing. Nature Photonics 15(2), 102–114 (2021)
2021
-
[21]
Nature Photonics 15(5), 367–373 (2021)
Zhou, T., Lin, X., Wu, J., Chen, Y., Xie, H., Li, Y., Fan, J., Wu, H., Fang, L., Dai, Q.: Large-scale neuromorphic optoelectronic computing with a reconfigurable diffractive processing unit. Nature Photonics 15(5), 367–373 (2021)
2021
-
[22]
Science 361(6406), 1004–1008 (2018)
Lin, X., Rivenson, Y., Yardimci, N.T., Veli, M., Luo, Y., Jarrahi, M., Ozcan, A.: All-optical machine learning using diffractive deep neural networks. Science 361(6406), 1004–1008 (2018)
2018
-
[23]
Nature Photonics 11(7), 441–446 (2017)
Shen, Y., Harris, N.C., Skirlo, S., Prabhu, M., Baehr-Jones, T., Hochberg, M., Sun, X., Zhao, S., Larochelle, H., Englund, D., et al.: Deep learning with coherent nanophotonic circuits. Nature Photonics 11(7), 441–446 (2017)
2017
-
[24]
Science Advances 9(28), 3436 (2023)
Vadlamani, S.K., Englund, D., Hamerly, R.: Transferable learning on analog hardware. Science Advances 9(28), 3436 (2023)
2023
-
[25]
Optica 11(8), 1039– 1049 (2024)
Xu, T., Zhang, W., Zhang, J., Luo, Z., Xiao, Q., Wang, B., Luo, M., Xu, X., Shastri, B.J., Prucnal, P.R., et al.: Control-free and efficient integrated photonic neural networks via hardware-aware training and pruning. Optica 11(8), 1039– 1049 (2024)
2024
-
[26]
Nature Machine Intelligence 5(10), 1119–1129 (2023)
Zheng, Z., Duan, Z., Chen, H., Yang, R., Gao, S., Zhang, H., Xiong, H., Lin, X.: Dual adaptive training of photonic neural networks. Nature Machine Intelligence 5(10), 1119–1129 (2023)
2023
-
[27]
Optica 9(7), 803–811 (2022)
Spall, J., Guo, X., Lvovsky, A.I.: Hybrid training of optical neural networks. Optica 9(7), 803–811 (2022)
2022
-
[28]
Laser & Photonics Reviews 18(4), 2300445 (2024)
Zhan, Y., Zhang, H., Lin, H., Chin, L.K., Cai, H., Karim, M.F., Poenar, D.P., Jiang, X., Mak, M.-W., Kwek, L.C., et al.: Physics-aware analytic-gradient train- ing of photonic neural networks. Laser & Photonics Reviews 18(4), 2300445 (2024)
2024
-
[29]
Nature Communications 14(1), 2535 (2023)
Huo, Y., Bao, H., Peng, Y., Gao, C., Hua, W., Yang, Q., Li, H., Wang, R., Yoon, S.-E.: Optical neural network via loose neuron array and functional learning. Nature Communications 14(1), 2535 (2023)
2023
-
[30]
Nature Communications 15(1), 6189 (2024) 22
Cheng, J., Huang, C., Zhang, J., Wu, B., Zhang, W., Liu, X., Zhang, J., Tang, Y., Zhou, H., Zhang, Q., et al.: Multimodal deep learning using on-chip diffrac- tive optics with in situ training capability. Nature Communications 15(1), 6189 (2024) 22
2024
-
[31]
Opto- Electronic Advances 7(4), 230182–1 (2024)
Wan, Y., Liu, X., Wu, G., Yang, M., Yan, G., Zhang, Y., Wang, J.: Efficient stochastic parallel gradient descent training for on-chip optical processor. Opto- Electronic Advances 7(4), 230182–1 (2024)
2024
-
[32]
arXiv preprint arXiv:2208.01623 (2022)
Bandyopadhyay, S., Sludds, A., Krastanov, S., Hamerly, R., Harris, N., Bunandar, D., Streshinsky, M., Hochberg, M., Englund, D.: Single chip photonic deep neural network with accelerated training. arXiv preprint arXiv:2208.01623 (2022)
2022 arXiv
-
[33]
Nature 632(8024), 280–286 (2024)
Xue, Z., Zhou, T., Xu, Z., Yu, S., Dai, Q., Fang, L.: Fully forward mode training for optical neural networks. Nature 632(8024), 280–286 (2024)
2024
-
[34]
Science 380(6643), 398–404 (2023)
Pai, S., Sun, Z., Hughes, T.W., Park, T., Bartlett, B., Williamson, I.A., Minkov, M., Milanizadeh, M., Abebe, N., Morichetti, F., et al.: Experimentally realized in situ backpropagation for deep learning in photonic neural networks. Science 380(6643), 398–404 (2023)
2023
-
[35]
Nature Physics 20(9), 1434–1440 (2024)
Wanjura, C.C., Marquardt, F.: Fully nonlinear neuromorphic computing with linear wave scattering. Nature Physics 20(9), 1434–1440 (2024)
2024
-
[36]
Optica5(7), 864–871 (2018)
Hughes, T.W., Minkov, M., Shi, Y., Fan, S.: Training of photonic neural networks through in situ backpropagation and gradient measurement. Optica5(7), 864–871 (2018)
2018
-
[37]
Photonics Research 8(6), 940–953 (2020)
Zhou, T., Fang, L., Yan, T., Wu, J., Li, Y., Fan, J., Wu, H., Lin, X., Dai, Q.: In situ optical backpropagation training of diffractive optical neural networks. Photonics Research 8(6), 940–953 (2020)
2020
-
[38]
Nature Communications 6(1), 6729 (2015)
Hermans, M., Burm, M., Van Vaerenbergh, T., Dambre, J., Bienstman, P.: Train- able hardware for dynamical computing using error backpropagation through physical media. Nature Communications 6(1), 6729 (2015)
2015
-
[39]
Nature Communications 13(1), 7847 (2022)
Nakajima, M., Inoue, K., Tanaka, K., Kuniyoshi, Y., Hashimoto, T., Nakajima, K.: Physical deep learning with biologically inspired training method: gradient- free approach for physical hardware. Nature Communications 13(1), 7847 (2022)
2022
-
[40]
Optica9(12), 1323– 1332 (2022)
Filipovich, M.J., Guo, Z., Al-Qadasi, M., Marquez, B.A., Morison, H.D., Sorger, V.J., Prucnal, P.R., Shekhar, S., Shastri, B.J.: Silicon photonic architecture for training deep neural networks with direct feedback alignment. Optica9(12), 1323– 1332 (2022)
2022
-
[41]
Advances in neural information processing systems 31 (2018)
Li, H., Xu, Z., Taylor, G., Studer, C., Goldstein, T.: Visualizing the loss landscape of neural nets. Advances in neural information processing systems 31 (2018)
2018
-
[42]
In: International Conference on Learning Representations (2021)
Foret, P., Kleiner, A., Mobahi, H., Neyshabur, B.: Sharpness-aware minimization for efficiently improving generalization. In: International Conference on Learning Representations (2021). https://openreview.net/forum?id=6Tm1mposlrM 23
2021
-
[43]
Proceedings of the IEEE 86(11), 2278–2324 (1998)
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P.: Gradient-based learning applied to document recognition. Proceedings of the IEEE 86(11), 2278–2324 (1998)
1998
-
[44]
In: 2020 IEEE International Conference on Big Data (Big Data), pp
Yao, Z., Gholami, A., Keutzer, K., Mahoney, M.W.: Pyhessian: Neural networks through the lens of the hessian. In: 2020 IEEE International Conference on Big Data (Big Data), pp. 581–590 (2020). IEEE
2020
-
[45]
Nature Communications 13(1), 123 (2022)
Wang, T., Ma, S.-Y., Wright, L.G., Onodera, T., Richard, B.C., McMahon, P.L.: An optical neural network using less than 1 photon per multiplication. Nature Communications 13(1), 123 (2022)
2022
-
[46]
Optica 3(12), 1460– 1465 (2016)
Clements, W.R., Humphreys, P.C., Metcalf, B.J., Kolthammer, W.S., Walmsley, I.A.: Optimal design for universal multiport interferometers. Optica 3(12), 1460– 1465 (2016)
2016
-
[47]
Optica 8(10), 1247–1255 (2021)
Bandyopadhyay, S., Hamerly, R., Englund, D.: Hardware error correction for programmable photonics. Optica 8(10), 1247–1255 (2021)
2021
-
[48]
Advances in Neural Information Processing Systems 35, 23439–23451 (2022)
Du, J., Zhou, D., Feng, J., Tan, V., Zhou, J.T.: Sharpness-aware training for free. Advances in Neural Information Processing Systems 35, 23439–23451 (2022)
2022
-
[49]
Physical Review Applied 11(6), 064044 (2019) 24
Pai, S., Bartlett, B., Solgaard, O., Miller, D.A.: Matrix optimization on universal unitary photonic devices. Physical Review Applied 11(6), 064044 (2019) 24
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.