REVIEW 2 major objections 7 minor 38 references
Inverse design of the transmission matrix in a random system using Reinforcement Learning
T0 review · 2 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that reinforcement learning can inverse-design the transmission matrix of a random scattering system, demonstrating rank-1, exceptional-point, and degenerate transmission matrices in a 2D billiard cavity.
desk verdict A plausible new RL application to transmission-matrix inverse design, but the scale-free cost functions leave the actual transmission strength unreported, which is the main gap to fix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the $2\times2$ transmission matrix $T$ of a two-port billiard, computed with an FDTD solver, and the central mechanism is the PPO policy network that proposes small scatterer displacements. The state is the normalized coordinate list $[x_1,y_1,\ldots,x_n,y_n]$, actions are normalized increments scaled by $0.005$, and each step's reward is the negative cost. Three scale-free costs steer the matrix: for rank-1, cost $=1-\tau_1/\sum_i\tau_i$; for degenerate eigenvalues, cost $=((t_{11}-t_{22})^2+4t_{12}t_{21})/(\sum_{ij}|t_{ij}|^2+\epsilon)$; for degenerate transmission eigenvalues, cost $=|\tau_1-\tau_2|/(\sum_{ij}|t_{ij}|^2+\epsilon)$, with $\epsilon=10^{-8}$. Over 1024-step episodes with a mixed reset strategy that returns to the best-known configuration or starts fresh, the agent improves its episode reward and reaches configurations that satisfy the target.
What would settle it
Run the released code from ten random seeds and count how often each seed reaches the stated cost threshold for the rank-1 target within a fixed episode budget; if most seeds fail, or if a greedy local search using the same number of FDTD evaluations matches the performance, the paper's claim that RL reliably solves this inverse-design problem is falsified.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the singularities and spectral structure of the transmission matrix of a random system are controllable by treating scatterer positions as the action space of a PPO agent. The agent's reward is the negative of a scale-free cost function, so the optimization cannot cheat by shrinking the overall matrix. The converged configurations show a rank-1 transmission matrix with a lowest transmission eigenvalue of $1.06\times10^{-7}$ and $2\pi$ phase winding of $\det(T)$ around the transmission-zero frequency; degenerate eigenvalues with eigenvector coalescence $C=1$ and asymmetry factor $A=1$ for the Jordan-type mode basis; and degenerate transmission eigenvalues with eigenchannel participation number $N_{\mathrm{eff}}=2$ at the target frequency. The paper concludes that RL is a general methodology for inverse design of scattering matrices in random media.
Load-bearing premise
The load-bearing premise is that PPO, with this state/action parameterization, these cost functions, and 1024-step episodes, can navigate the highly non-convex and oscillatory objective landscape to a good solution within a reasonable number of episodes; the paper shows single successful examples but no success rate, seed variation, or comparison to other optimizers.
Editorial extensions
If this is right
- A rank-1 transmission matrix acts as a fixed-ratio power splitter: any input wavefront produces the same output speckle pattern up to a scalar, and an input orthogonal to the feature vector gives a transmissionless mode.
- Degenerate eigenvalues realize an exceptional point at the target frequency, giving unidirectional mode conversion and enhanced sensitivity to perturbations that scale as a fractional power of the perturbation.
- Degenerate transmission eigenvalues make the transmitted power independent of the input speckle pattern, since $T$ becomes proportional to a unitary matrix and $N_{\mathrm{eff}}=N$.
- The RL loop can replace the transmission matrix with any linear response matrix, so the same framework can target focusing inside a random medium or a prescribed density-of-states spectrum.
- Because the reward is scale-free and no surrogate model is used, the method is aimed directly at the singularities of the transmission matrix, which auxiliary predictors handle poorly.
Reading between the lines
- The paper leaves open whether the same PPO loop beats a simple hill-climbing search with the same number of FDTD calls; if it does not, the benefit of RL would be exploratory rather than computational.
- The reported asymmetry exponent of 0.9, versus 0.5 predicted by first-order perturbation theory, suggests that scatterer displacements couple nonlinearly to the transmission-matrix perturbation; deriving that coupling map could turn the empirical fit into a predictive design rule.
- The rank-1 and degenerate-transmission targets are global properties of $T$ at one frequency; the same cost logic could be extended to bandwidth-averaged costs to design devices that hold their behavior over a frequency window.
- The framework's practical value in experiments will depend on whether the optimized scatterer positions remain effective under fabrication disorder; adding a robustness term to the cost would be a direct extension.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Using a 2D billiard with two input/output waveguides and up to 20 movable dielectric cylinders, the author applies Proximal Policy Optimization (PPO) to position the cylinders so that the simulated transmission matrix attains one of three target properties: rank-1 structure with a transmission zero, a degenerate eigenvalue (exceptional point with coalescing eigenvectors), or degenerate transmission eigenvalues. The agent observes the normalized scatterer coordinates, applies incremental coordinate shifts, and receives a reward equal to the negative of a scale-free cost function computed from FDTD (Meep). For each target, a single converged design is shown and characterized with eigenchannel, phase-winding, coalescence, or participation-number diagnostics.
Significance. If the demonstrations are robust, the paper would provide a conceptually simple and general RL-based route to inverse-design scattering-matrix singularities in disordered media, complementing existing local-optimization and supervised-learning approaches. The external FDTD simulator is an appropriate ground truth, the code is open-sourced, and the physical diagnostics (phase winding, eigenvector coalescence, perturbation scaling) are independent of the training reward, which is a genuine strength. The principal scientific risk is that none of the designs is characterized by an absolute transmission scale, leaving the possibility that the claimed targets are realized at negligible transmission; the lack of repeated runs further limits the evidence for the methodology's reliability.
major comments (2)
- [Table 1] The claim that a scale-free cost prevents the RL from driving the matrix scale toward zero is not correct. All three costs in Table 1 are also minimized by a zero or arbitrarily weak transmission matrix: the rank-1 cost vanishes for any rank-1 matrix (including tau_1 -> 0), the degenerate-eigenvalue cost vanishes for nilpotent matrices of arbitrarily small norm, and the degenerate-transmission-eigenvalue cost vanishes for T=0 (where the denominator is epsilon). The manuscript reports only the smallest transmission eigenvalue (1.06e-7) in Fig. 2b and never reports the dominant eigenvalue, the matrix norm, or the total transmitted power. Please report the absolute scale of the optimized transmission matrices, or add an explicit scale constraint, so that the designed systems can be distinguished from vacuous near-zero-transmission configurations.
- [Figures 2-4] Each target is demonstrated only once, and the text's phrase 'multiple converged results' in the Results section refers to one example per target. No seed variation, success rates, or error bars are provided, and the training curve in Fig. 1c is described only as 'typical.' Without these, the claim that PPO reliably navigates the non-convex landscape to these inverse-design targets is not statistically supported, nor is there any comparison with simple baselines such as random search or evolutionary optimization. Please provide multiple independent runs (varying initial conditions and PPO seeds) and report the distribution of outcomes, including failure cases and computational cost.
minor comments (7)
- [Eq. (1)] The determinant formula is garbled in the rendered text; please rewrite it as det(T(f)) proportional to a product over transmission zeros divided by a product over poles, and define all symbols including the relationship of M and N to the number of channels and resonant modes.
- [Table 1] The cost-function formulas are not legible as typeset; please write them explicitly in LaTeX notation, define tau_i, t_ij, epsilon, and state clearly that epsilon is a fixed regularization constant.
- [Fig. 1c] The description of the upper training panel is ambiguous: it says the panel plots the maximum Reward for every 128 steps but also mentions a running max and a -log10(-Reward) transformation; please clarify the exact plotted quantity and axis labels.
- [Degenerate eigenvalues] The definition of eigenvector coalescence C is garbled; it should be written as a product over pairs of normalized eigenvectors of (1 - |v_i . v_j|^2) with an appropriate normalization, and the normalization constant should be stated.
- [Rank-1 TM] The paragraph discussing a second-order transmission zero admits a negative result (phase winding number 0 rather than 4pi, attributed to the statistical rarity of double-zero eigenvalues) without quantitative support. Please report the achieved eigenvalue magnitudes and the number of training attempts, or state that this is an anecdotal observation.
- [Methods] The manuscript lacks specific PPO hyperparameters, Meep simulation resolution, FDTD run times, and the scatterer boundary-handling rules beyond the 'last object wins' mention; since the code is open-sourced this is not blocking, but a short methods paragraph would improve reproducibility.
- [Throughout] There are several typographical errors and notation inconsistencies, including 'frequncy', 'propogating', 'orthorgonal', and 'choosen'; a careful proofreading pass is recommended.
Circularity Check
No significant circularity: the design loop is PPO minimizing directly-defined TM costs against Meep FDTD ground truth, and the only self-citation is not load-bearing.
full rationale
The paper's derivation chain is self-contained and not circular. The demonstrated TMs are produced by PPO optimizing scatterer positions, with the Meep FDTD solver as an external ground truth; the cost functions in Table 1 are direct measures of the target (rank-1, degenerate eigenvalue, degenerate transmission eigenvalue), so minimizing them to find scatterer configurations is the inverse-design task itself, not a hidden prediction. The theory checks (2π phase winding of det(T), eigenvector coalescence, and the perturbation exponent 0.5 vs fitted 0.9) are independent of the cost and are not fitted to the RL results. The only self-citation, ref. [3] (Kang and Genack), supports the TM-zero phase-winding interpretation, but Fig. 2c independently verifies the winding, so the citation is not load-bearing. Two non-circular concerns remain: no multi-seed statistics or success rate are given, and the scale-free costs near Table 1 (e.g., cost = 1 - τ1/Στi for the rank-1 target) are also satisfied by arbitrarily small transmission matrices, while absolute transmitted power is not reported; these are correctness and verifiability gaps, not circular reductions of the derivation to its inputs.
Assumptions & free parameters
free parameters (2)
- PPO algorithm hyperparameters =
gamma=0.999, episode length L=1024, action scale 0.005, reset probability 0.7
- Cost function regularization epsilon =
1e-8
assumptions (4)
- domain assumption The 2x2 transmission matrix extracted from Meep FDTD fully characterizes the scattering system at 15 GHz.
- standard math The determinant of the transmission matrix obeys the Heidelberg rational-function model (Eq. 1).
- standard math Eigenvalue and transmission eigenvalue decompositions of the TM are the correct descriptors for the designed phenomena.
- domain assumption The 'last object wins' overlap rule in Meep is an acceptable representation of overlapping scatterers.
Cite this review
Pith. "Pith review of Inverse design of the transmission matrix in a random system using Reinforcement Learning." pith.science (2026). https://pith.science/paper/F3E7HCA2
@misc{pith2026250613057,
author = {Pith},
title = {Pith review of: Inverse design of the transmission matrix in a random system using Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/F3E7HCA2}},
note = {Machine review of arXiv:2506.13057}
}
read the original abstract
This work presents an approach to the inverse design of scattering systems by modifying the transmission matrix using reinforcement learning. We utilize Proximal Policy Optimization to navigate the highly non-convex landscape of the object function to achieve three types of transmission matrices: (1) Fixed-ratio power conversion and zero-transmission mode in rank-1 matrices, (2) exceptional points with degenerate eigenvalues and unidirectional mode conversion, and (3) uniform channel participation is enforced when transmission eigenvalues are degenerate.
Reference graph
Works this paper leans on
-
[1]
Longhi, $\mathcal{PT}$-symmetric laser absorber, Phys
S. Longhi, $\mathcal{PT}$-symmetric laser absorber, Phys. Rev. A 82, 031801 (2010)
work page 2010
-
[2]
Miri and A
M.-A. Miri and A. Alù, Exceptional points in optics and photonics, Science 363, eaar7709 (2019)
2019
-
[3]
Y. Kang and A. Z. Genack, Transmission zeros with topological symmetry in complex systems, Phys. Rev. B 103, L100201 (2021)
work page 2021
-
[4]
Y. D. Chong, L. Ge, H. Cao, and A. D. Stone, Coherent Perfect Absorbers: Time-Reversed Lasers, Phys. Rev. Lett. 105, 053901 (2010)
2010
-
[5]
W. R. Sweeney, C. W. Hsu, and A. D. Stone, Theory of reflectionless scattering modes, Phys. Rev. A 102, 063511 (2020)
work page 2020
- [6]
-
[7]
J. Erb, N. Shaibe, R. Calvo, D. P. Lathrop, T. M. Antonsen, T. Kottos, and S. M. Anlage, Topology and manipulation of scattering singularities in complex non-Hermitian systems: Two-channel case, Phys. Rev. Res. 7, 023090 (2025)
work page 2025
-
[8]
Silver et al., Mastering the game of Go without human knowledge, Nature 550, 354 (2017)
D. Silver et al., Mastering the game of Go without human knowledge, Nature 550, 354 (2017)
work page 2017
Show all 38 references
-
[9]
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction (A Bradford Book, Cambridge, MA, USA, 2018)
2018
-
[10]
AlMahamid and K
F. AlMahamid and K. Grolinger, Reinforcement Learning Algorithms: An Overview and Classification, in 2021 IEEE Canadian Conference on Electrical and Computer Engineering (CCECE) (2021), pp. 1–7
2021
-
[11]
A. K. Shakya, G. Pillai, and S. Chakrabarty, Reinforcement learning algorithms: A brief survey, Expert Systems with Applications 231, 120495 (2023)
2023
-
[12]
L. Deng, Y. Xu, and Y. Liu, Hybrid inverse design of photonic structures by combining optimization methods with neural networks, Photonics and Nanostructures - Fundamentals and Applications 52, 101073 (2022)
2022
-
[13]
Molesky, Z
S. Molesky, Z. Lin, A. Y. Piggott, W. Jin, J. Vuckovid, and A. W. Rodriguez, Inverse design in nanophotonics, Nature Photon 12, 659 (2018)
2018
-
[14]
A. Y. Piggott, J. Lu, K. G. Lagoudakis, J. Petykiewicz, T. M. Babinec, and J. Vučkovid, Inverse design and demonstration of a compact and broadband on-chip wavelength demultiplexer, Nature Photon 9, 374 (2015)
2015
-
[15]
D. Liu, Y. Tan, E. Khoram, and Z. Yu, Training Deep Neural Networks for the Inverse Design of Nanophotonic Structures, ACS Photonics 5, 1365 (2018)
2018
-
[16]
Z. Li, W. Liu, D. Ma, S. Yu, H. Cheng, D.-Y. Choi, J. Tian, and S. Chen, Inverse Design of Few-Layer Metasurfaces Empowered by the Matrix Theory of Multilayer Optics, Phys. Rev. Appl. 17, 024008 (2022)
2022
-
[17]
W. Ji, J. Chang, H.-X. Xu, J. R. Gao, S. Gröblacher, H. P. Urbach, and A. J. L. Adam, Recent advances in metasurface design and quantum optics applications with machine learning, physics-informed neural networks, and topology optimization methods, Light Sci Appl 12, 169 (2023)
2023
-
[19]
An et al., A Deep Learning Approach for Objective-Driven All-Dielectric Metasurface Design, ACS Photonics 6, 3196 (2019)
S. An et al., A Deep Learning Approach for Objective-Driven All-Dielectric Metasurface Design, ACS Photonics 6, 3196 (2019)
2019
-
[20]
S. So, T. Badloe, J. Noh, J. Bravo-Abad, and J. Rho, Deep learning enabled inverse design in nanophotonics, Nanophotonics 9, 1041 (2020)
2020
-
[21]
W. Ma, Z. Liu, Z. A. Kudyshev, A. Boltasseva, W. Cai, and Y. Liu, Deep learning for the design of photonic structures, Nat. Photonics 15, 77 (2021)
2021
-
[22]
R. Li, C. Zhang, W. Xie, Y. Gong, F. Ding, H. Dai, Z. Chen, F. Yin, and Z. Zhang, Deep reinforcement learning empowers automated inverse design and optimization of photonic crystals for nanoscale laser cavities, Nanophotonics 12, 319 (2023)
2023
-
[23]
Sajedian, T
I. Sajedian, T. Badloe, and J. Rho, Optimisation of colour generation from dielectric nanostructures using reinforcement learning, Opt. Express, OE 27, 5874 (2019)
2019
-
[24]
Sajedian, H
I. Sajedian, H. Lee, and J. Rho, Double-deep Q-learning to increase the efficiency of metasurface holograms, Sci Rep 9, 10899 (2019)
2019
-
[25]
Hooten, R
S. Hooten, R. G. Beausoleil, and T. V. Vaerenbergh, Inverse design of grating couplers using the policy gradient method from reinforcement learning, Nanophotonics 10, 3843 (2021)
2021
-
[26]
C. Sun, E. Kaiser, S. L. Brunton, and J. Nathan Kutz, Deep reinforcement learning for optical systems: A case study of mode-locked lasers, Mach. Learn.: Sci. Technol. 1, 045013 (2020)
2020
-
[27]
R. J. Williams, Simple statistical gradient-following algorithms for connectionist reinforcement learning, Mach Learn 8, 229 (1992)
1992
-
[28]
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, Playing Atari with Deep Reinforcement Learning, arXiv:1312.5602
-
[29]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal Policy Optimization Algorithms, arXiv:1707.06347
-
[30]
van Hasselt, A
H. van Hasselt, A. Guez, and D. Silver, Deep Reinforcement Learning with Double Q- Learning, arXiv:1509.06461
-
[31]
J. J. M. Verbaarschot, H. A. Weidenmüller, and M. R. Zirnbauer, Grassmann integration in stochastic quantum physics: The case of compound-nucleus scattering, Physics Reports 129, 367 (1985)
1985
-
[32]
Rotter, Effective Hamiltonian and unitarity of the S matrix, Phys
I. Rotter, Effective Hamiltonian and unitarity of the S matrix, Phys. Rev. E 68, 016211 (2003)
2003
-
[33]
Y. V. Fyodorov and H.-J. Sommers, Statistics of resonance poles, phase shifts and time delays in quantum chaotic scattering: Random matrix approach for systems with broken time-reversal invariance, Journal of Mathematical Physics 38, 1918 (1997)
1997
-
[34]
A. F. Oskooi, D. Roundy, M. Ibanescu, P. Bermel, J. D. Joannopoulos, and S. G. Johnson, Meep: A flexible free-software package for electromagnetic simulations by the FDTD method, Computer Physics Communications 181, 687 (2010)
2010
-
[35]
T. P. Dussauge, W. J. Sung, O. J. Pinon Fischer, and D. N. Mavris, A reinforcement learning approach to airfoil shape optimization, Sci Rep 13, 9753 (2023)
2023
-
[36]
Feng, Y.-L
L. Feng, Y.-L. Xu, W. S. Fegadolli, M.-H. Lu, J. E. B. Oliveira, V. R. Almeida, Y.-F. Chen, and A. Scherer, Experimental demonstration of a unidirectional reflectionless parity-time metamaterial at optical frequencies, Nature Mater 12, 108 (2013)
2013
-
[37]
M. Davy, Z. Shi, and A. Z. Genack, Focusing through random media: Eigenchannel participation number and intensity correlation, Phys. Rev. B 85, 035105 (2012)
2012
-
[38]
Cheng and A
X. Cheng and A. Z. Genack, Focusing and energy deposition inside random media, Optics Letters, Vol. 39, Issue 21, Pp. 6324-6327 (2014)
2014
-
[39]
M. Davy, Z. Shi, J. Wang, X. Cheng, and A. Z. Genack, Transmission Eigenchannels and the Densities of States of Random Media, Phys. Rev. Lett. 114, 033901 (2015)
2015
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.