REVIEW 3 major objections 4 minor 44 references
Diffractive neural networks for mode-sorting with flexible detection regions
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Training a mode-sorter's detection regions along with its phase plates roughly doubles efficiency at matched crosstalk in simulation, and the gain persists in experiment.
desk verdict Good idea, honest experiments, but the flexible-vs-fixed comparison is confounded by an unoptimized baseline; the advantage is real but likely smaller than claimed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the flexible detection region: a trainable, non-overlapping partition {D_j} of the output plane, into which the network directs each input mode. Instead of prescribing exact output field shapes, the network maximises the diagonal of the intensity matrix I_ij = ∫_{D_j} |Ψ_output^(i)|² dx dy through the weighted loss α·Losseff + (1−α)·Lossxtalk, with Losseff = −(1/n) Σ_i I_ii and Lossxtalk = (1/n) Σ_i (1 − I_ii / Σ_j I_ij). The hyperparameter α sets the efficiency-crosstalk trade-off, and the regions are re-optimised against experimentally measured outputs to compensate SLM imperfections.
What would settle it
An apples-to-apples benchmark settles it: train one sorter with flexible regions and another with fixed regions whose positions and sizes are also optimised (by grid search or joint gradient training, then frozen), using identical phase-plate budgets and 25, 100, and 210 modes across at least two modal bases. If the optimised fixed-region design matches the flexible one in efficiency at equal crosstalk, the central advantage claim would be refuted; if the gap persists, it is confirmed. A purely experimental check: send a free-space communication signal through both sorters and compare channel
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a mode-sorter's output detection regions should be part of the trainable set of parameters, not a fixed prescription. For tasks that only measure modal intensities, the network need not produce a prescribed output field shape; it only needs to concentrate each input mode's light within its own region of the detection plane. The authors optimise the phase plates and the non-overlapping region set {D_j} jointly by backpropagation, using a loss that weights efficiency against crosstalk via a hyperparameter α. The result is a substantially better efficiency-crosstalk trade-off than the fixed-region baseline—58.5% versus 30% efficiency at 29% crosst
Load-bearing premise
The claim rests on a comparison in which the fixed-region baseline consists of circular regions whose positions and sizes are not themselves optimised; a well-optimised fixed-region design could plausibly close much of the efficiency gap.
Editorial extensions
If this is right
- Mode-sorters for intensity-measuring applications—free-space communication, imaging, endoscopy, passive superresolution—can collect roughly twice as much light per mode at the same crosstalk without adding phase plates.
- The hyperparameter α gives a tunable efficiency-crosstalk trade-off, so the same trained hardware can be configured for high-efficiency state discrimination or for low-crosstalk imaging.
- Re-optimising detection regions against the measured output fields recovers performance lost to experimental imperfections, without redesigning the phase plates; the single-plate, 4-mode sorter then beats even the simulated fixed-region design.
- Enforcing the HG symmetry of the phase plates in the network parametrisation keeps simulated performance nearly unchanged while making experimental alignment much easier.
- Because the hardware is just one SLM and one mirror, the sorter is easily reproducible in an ordinary optics lab, and fabricated phase plates can replace the SLM where its cavity effect and pixel crosstalk limit performance.
Reading between the lines
- The method is basis-agnostic: the training loop consumes labelled input modes and never assumes Hermite-Gaussian structure, so the same gain should transfer to LG, OAM, Zernike, or arbitrary speckle bases—a direct next test.
- How much of the advantage is intrinsic to region flexibility depends on the baseline: a fairer benchmark would optimise fixed-region positions and sizes with comparable effort, and comparing on 100+ modes or non-orthogonal states would bound the effect.
- The design principle reaches beyond optics: any wave-based processor whose readout integrates intensity over a detector partition has that partition as a legitimate trainable parameter, so acoustic, microwave, or terahertz implementations could adopt the same trick.
- A promising extension is a hybrid scheme that trains flexible regions while also matching the field where a few output channels couple into fibres, bridging the intensity-measurement and fibre-coupling regimes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes training diffractive neural network mode-sorters with flexible output detection regions that are jointly optimized with the phase plates. It presents simulations and experiments for 1-, 2-, and 3-plate sorters sorting up to 25 Hermite-Gaussian modes, and reports that flexible regions improve the efficiency-crosstalk trade-off compared to fixed circular detection regions (e.g., 58.5% vs 30% efficiency at 29% crosstalk in simulation). The authors argue that this approach outperforms traditional MPLC mode-sorting methods.
Significance. The idea of including detection-region geometry in the trainable parameter set is simple, general, and potentially useful; if the comparison is fair, the reported efficiency gain is substantial. The manuscript demonstrates the effect in both simulation and experiment, and the use of an open-source differentiable simulator (TorchOptics) supports reproducibility. The explicit Pareto-style trade-off via the hyperparameter alpha is a useful design element. However, the strength of the central claim depends on the fairness of the fixed-detection baseline and on the evaluation protocol; both need strengthening before the 'outperforms' statement can be accepted.
major comments (3)
- [§3.2, Fig. 2] The central comparison uses a fixed-detection baseline that is not optimized. The text states that the fixed regions are 'circular at fixed locations, with the radii and center positions chosen such that the total area is similar to the total area of the flexible detection regions.' No optimization of these region parameters is reported. The flexible detector, by contrast, is trained jointly via Eqs. (4)-(7). The reported improvement (30% to 58.5% efficiency at 29% crosstalk) therefore conflates shape flexibility with optimization of region placement. Please optimize the fixed-region centers and radii with the same loss, or compare against another optimized fixed-region sorter (e.g., WMM), and provide the resulting curve. Without this, the 'outperforms' claim is not established.
- [§3.3, Fig. 6] The experimental evaluation re-optimizes the flexible detection regions against the measured output fields before computing performance ('we re-optimize the detection regions accounting for the experimentally measured outputs', Fig. 5). If the same measurements are then used to report efficiency and crosstalk, this gives the flexible detector an advantage that the fixed detector does not receive. Please evaluate on held-out measurements, or at least provide a split-data comparison and quantify the optimism. The manuscript should also state explicitly whether the fixed-region results in Fig. 6 were re-optimized in the same way.
- [Abstract and §4] The abstract concludes that the approach 'outperforms traditional mode-sorting methods', but the only direct comparator presented is a DONN with unoptimized fixed circular detection. No WMM-trained MPLC, log-polar sorter, or other established mode-sorter is benchmarked. This claim is broader than the evidence. I recommend either adding such a benchmark or narrowing the conclusion to 'outperforms this particular fixed-detection baseline', pending the outcome of the first comment.
minor comments (4)
- [§3.2 / Fig. 6] The efficiency definition in the Fig. 6 caption (ratio to output intensity with all phase plates set to 0) differs from the efficiency used in simulation via Eq. (5). Please align the definitions or explain why the experimental metric is equivalent.
- [Global] Typographical issues: 'charactrization' in §3.1; 'Fountaineet al.' missing space; 'a∼ 20% crosstalk' spacing. The inset in Fig. 2(a) for alpha≈0 is difficult to read and should be enlarged.
- [§3.3 / Fig. 5] The Fig. 5 caption is terse; it would help to state explicitly which panels correspond to simulation and which to experiment, and to define the purple and orange circles referenced in the text.
- [Availability] No data-availability or code-availability statement is included. Given the reproducibility-oriented claims, providing trained phase-plate profiles and detection-region parameters (or a link to code) would be valuable.
Circularity Check
No significant circularity: the paper optimizes a differentiable loss (Eqs. 4-7) and reports the optimized values, which is standard optimization. Minor concerns: experimental detection regions are re-optimized on the same measured outputs used for evaluation, and the fixed-detection baseline is unoptimized; these weaken the comparison but do not make the derivation circular.
-
fitted input called prediction
[Section 3.3, 'Experimental Results'; Fig. 6]
"Experimental imperfections cause the mode output fields to deviate from their theoretically predicted shapes. To address this, we re-optimize the detection regions accounting for the experimentally measured outputs, as illustrated in Fig. 5."
The experimental efficiency and crosstalk reported in Fig. 6 are computed using detection regions D_j that were re-optimized on the very same measured output fields used for evaluation. Since efficiency (Eq. 5) and crosstalk (Eq. 6) are defined through the region integrals I_ij in Eq. (4), re-optimizing the regions directly maximizes the reported efficiency on the evaluation data. The fixed-region comparison is not re-optimized in the same way, so the flexible-vs-fixed advantage shown in the experiment is partly an in-sample fitting artifact rather than an independent, out-of-sample performance gain.
full rationale
This paper is a standard optimization demonstration: the authors define a differentiable loss (Eqs. 5-7) and train phase plates and detection regions by backpropagation, then report the resulting efficiency and crosstalk. Reporting the optimized training objective is not circular; it is the normal way to present a trained optical system. The central comparison (flexible vs. fixed detection regions) in simulation trains both phase-plate sets with the same loss, so the advantage of flexible regions reflects the inclusion of region parameters in the training set, which matches the paper's stated claim. The only self-citations (TorchOptics [38], prior superresolution works [10,11]) are tool/context citations and are not load-bearing. One caveat deserves note: in the experimental section, the detection regions are re-optimized on the same experimentally measured outputs used to compute the reported efficiency/crosstalk (Fig. 6), and the fixed-region baseline is not given the same re-optimization, so the experimental flexible-vs-fixed advantage is partly an in-sample fitting effect. This weakens the strength of the 'outperforms' claim but does not make the derivation circular, especially because the simulation independently shows the same improvement. The unoptimized fixed baseline (circular regions with area matching, no region-parameter training) is a comparison-quality issue, not a circular-reasoning issue. Overall: no significant circularity; score 2.
Assumptions & free parameters
free parameters (3)
- phase plate phase profiles (200x200 pixels per plate, times number of plates) =
optimized by gradient descent; not reported in text
- detection region geometry (positions, sizes, shapes) =
optimized by gradient descent; not reported, only visualized in Figures 4(a) and 5
- loss weight alpha =
varied between 0 and 1
assumptions (3)
- domain assumption Scalar diffraction theory accurately models the free-space propagation between phase plates.
- domain assumption The SLM phase plates are ideal phase-only modulators in the first diffraction order.
- domain assumption Input HG modes are generated with high fidelity (97.8% mean fidelity).
Cite this review
Pith. "Pith review of Diffractive neural networks for mode-sorting with flexible detection regions." pith.science (2026). https://pith.science/paper/UNY2Z65F
@misc{pith2026250820058,
author = {Pith},
title = {Pith review of: Diffractive neural networks for mode-sorting with flexible detection regions},
year = {2026},
howpublished = {\url{https://pith.science/paper/UNY2Z65F}},
note = {Machine review of arXiv:2508.20058}
}
read the original abstract
Mode-sorting is a procedure that decomposes a light field into a basis of transverse modes, directing each mode into a separate spatial location, allowing the constituent mode intensities to be measured simultaneously. We demonstrate a mode-sorter based on a diffractive optical neural network and show that it is advantageous to include the output detection regions into the trainable set of parameters of that network. This approach outperforms traditional mode-sorting methods, achieving higher efficiency for the same crosstalk levels.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Z. Zhu, M. Janasik, A. Fyffe,et al., “Compensation-free high-dimensional free-space optical communication using turbulence-resilient vector beams,” Nat. Commun.12(2021)
work page 2021
-
[2]
Communication with spatially modulated light through turbulent air across vienna,
M. Krenn, R. Fickler, M. Fink,et al., “Communication with spatially modulated light through turbulent air across vienna,” New J. Phys.16, 113028 (2014)
work page 2014
-
[3]
Exact solution to simultaneous intensity and phase encryption with a single phase-only hologram,
E. Bolduc, N. Bent, E. Santamato,et al., “Exact solution to simultaneous intensity and phase encryption with a single phase-only hologram,” Opt. Lett.38, 3546–3549 (2013)
2013
-
[4]
Space-division multiplexing for optical fiber communications,
B. J. Puttnam, G. Rademacher, and R. S. Luís, “Space-division multiplexing for optical fiber communications,” Optica 8, 1186 (2021)
work page 2021
-
[5]
Quantum theory of superresolution for two incoherent optical point sources,
M. Tsang, R. Nair, and X.-M. Lu, “Quantum theory of superresolution for two incoherent optical point sources,” Phys. Rev. X6, 031033 (2016)
2016
-
[6]
How to build the “optical inverse
U. G. B¯utait˙e, H. Kupianskyi, T. Čižmár, and D. B. Phillips, “How to build the “optical inverse” of a multimode fibre,” Intell. Comput. (2022)
work page 2022
-
[7]
Orbitalangularmomentum-mediatedmachinelearningforhigh-accuracymode-feature encoding,
X.Fang, X.Hu, B.Li, et al., “Orbitalangularmomentum-mediatedmachinelearningforhigh-accuracymode-feature encoding,” Light Sci. Appl.13 (2024)
work page 2024
-
[8]
Z. Dutton, R. Kerviche, A. Ashok, and S. Guha, “Attaining the quantum limit of superresolution in imaging an object’s length via predetection spatial-mode sorting,” Phys. Rev. A99 (2019)
work page 2019
Show all 44 references
-
[9]
Superresolution linear optical imaging in the far field,
A. A. Pushkina, G. Maltese, J. I. Costa-Filho,et al., “Superresolution linear optical imaging in the far field,” Phys. Rev. Lett.127, 253602 (2021)
2021
-
[10]
Passive superresolution imaging of incoherent objects,
J. Frank, A. Duplinskiy, K. Bearne, and A. I. Lvovsky, “Passive superresolution imaging of incoherent objects,” Optica 10, 1147–1152 (2023)
2023
-
[11]
Tsang’s resolution enhancement method for imaging with focused illumination,
A. Duplinskiy, J. Frank, K. Bearne, and A. I. Lvovsky, “Tsang’s resolution enhancement method for imaging with focused illumination,” Light. Sci. Appl.14, 159 (2025)
2025
-
[12]
Efficient sorting of orbital angular momentum states of light,
G. C. Berkhout, M. P. Lavery, J. Courtial,et al., “Efficient sorting of orbital angular momentum states of light,” Phys. Rev. Lett.105(2010)
2010
-
[13]
Sorting photons by radial quantum number,
Y. Zhou, M. Mirhosseini, D. Fu,et al., “Sorting photons by radial quantum number,” Phys. Rev. Lett.119 (2017)
2017
-
[14]
Realization of a scalable laguerre–gaussian mode sorter based on a robust radial mode sorter,
D. Fu, Y. Zhou, R. Qi,et al., “Realization of a scalable laguerre–gaussian mode sorter based on a robust radial mode sorter,” Opt. Express26, 33057 (2018)
2018
-
[15]
Hermite–gaussian mode sorter,
Y. Zhou, J. Zhao, Z. Shi,et al., “Hermite–gaussian mode sorter,” Opt. Lett.43, 5263 (2018)
2018
-
[16]
Programmable unitary spatial mode manipulation,
J.-F. Morizur, L. Nicholls, P. Jian,et al., “Programmable unitary spatial mode manipulation,” Tech. rep. (2010)
2010
-
[17]
Efficient and mode selective spatial mode multiplexer based on multi-plane light conversion,
G. Labroille, B. Denolle, P. Jian,et al., “Efficient and mode selective spatial mode multiplexer based on multi-plane light conversion,” Opt. Express22, 15599–15607 (2014)
2014
-
[18]
Laguerre-gaussian mode sorter,
N. K. Fontaine, R. Ryf, H. Chen,et al., “Laguerre-gaussian mode sorter,” Nat. Commun.10 (2019)
2019
-
[19]
High-dimensional spatial mode sorting and optical circuit design using multi-plane light conversion,
H. Kupianskyi, S. A. R. Horsley, and D. B. Phillips, “High-dimensional spatial mode sorting and optical circuit design using multi-plane light conversion,” APL Photonics8(2023)
2023
-
[20]
Simultaneously sorting overlapping quantum states of light,
S. Goel, M. Tyler, F. Zhu,et al., “Simultaneously sorting overlapping quantum states of light,” Phys. Rev. Lett.130 (2023)
2023
-
[21]
Modal beam splitter: Determination of the transversal components of an electromagnetic light field,
M. Mazilu, T. Vettenburg, M. Ploschner,et al., “Modal beam splitter: Determination of the transversal components of an electromagnetic light field,” Sci. Reports7 (2017)
2017
-
[22]
Inverse design of gradient-index volume multimode converters,
N. Barré and A. Jesacher, “Inverse design of gradient-index volume multimode converters,” Opt. Express30, 10573 (2022)
2022
-
[23]
Full characterization of the transmission properties of a multi-plane light converter,
P. Boucher, A. Goetschy, G. Sorelli,et al., “Full characterization of the transmission properties of a multi-plane light converter,” Phys. Rev. Res.3 (2021)
2021
-
[24]
Characterization and applications of spatial mode multiplexers based on multi-plane light conversion,
G. Labroille, N. Barré, O. Pinel,et al., “Characterization and applications of spatial mode multiplexers based on multi-plane light conversion,” Opt. Fiber Technol.35, 93–99 (2017)
2017
-
[25]
Performance optimization of multi-plane light conversion (mplc) mode multiplexer by error tolerance analysis,
J. Fang, J. Bu, J. Li,et al., “Performance optimization of multi-plane light conversion (mplc) mode multiplexer by error tolerance analysis,” Opt. Express29, 37852 (2021)
2021
-
[26]
Optical circuit design based on a wavefront-matching method,
T. Hashimoto, T. Saida, I. Ogawa,et al., “Optical circuit design based on a wavefront-matching method,” Opt. Lett. 30, 2620–2622 (2005)
2005
-
[27]
Full-field mode sorter using two optimized phase transformations for high-dimensional quantum cryptography,
R. Fickler, F. Bouchard, E. Giese,et al., “Full-field mode sorter using two optimized phase transformations for high-dimensional quantum cryptography,” J. Opt. (United Kingdom)22 (2020)
2020
-
[28]
All-optical machine learning using diffractive deep neural networks,
X. Lin, Y. Rivenson, N. T. Yardimci,et al., “All-optical machine learning using diffractive deep neural networks,” Science 361, 1004–1008 (2018)
2018
-
[29]
Review of diffractive deep neural networks,
Y. Sun, M. Dong, M. Yu,et al., “Review of diffractive deep neural networks,” J. Opt. Soc. Am. B40, 2951–2961 (2023)
2023
-
[30]
Diffractive deep neural networks: Theories, optimization, and applications,
H. Chen, S. Lou, Q. Wang,et al., “Diffractive deep neural networks: Theories, optimization, and applications,” Appl. Phys. Rev.11 (2024)
2024
-
[31]
Wavefront matching method as a deep neural network and mutual use of their techniques,
T. Hashimoto, “Wavefront matching method as a deep neural network and mutual use of their techniques,” Opt. Commun. 498 (2021)
2021
-
[32]
All-optical signal processing of vortex beams with diffractive deep neural networks,
“All-optical signal processing of vortex beams with diffractive deep neural networks,” Phys. Rev. Appl.15 (2021)
2021
-
[33]
A physical neural network training approach toward multi-plane light conversion design,
Z. Zhu, J. H. Doerr, G. Li, and S. Pang, “A physical neural network training approach toward multi-plane light conversion design,” (2023)
2023
-
[34]
Broadband, low-crosstalk, and massive-channels oam modes de/multiplexing based on optical diffraction neural network,
Z. Liu, S. Gao, Z. Lai,et al., “Broadband, low-crosstalk, and massive-channels oam modes de/multiplexing based on optical diffraction neural network,” Laser Photonics Rev.17 (2023)
2023
-
[35]
Newopticalwaveguidedesignbasedonwavefrontmatching method,
Y.Sakamaki,T.Saida,T.Hashimoto,andH.Takahashi,“Newopticalwaveguidedesignbasedonwavefrontmatching method,” J. Light. Technol.25, 3511–3518 (2007)
2007
-
[36]
Opticalorbitalangularmomentummultiplexingcommunicationviainversely-designed multiphase plane light conversion,
J.Fang,J.Li,A.Kong, et al.,“Opticalorbitalangularmomentummultiplexingcommunicationviainversely-designed multiphase plane light conversion,” Photonics Res.10, 2015 (2022)
2015
-
[37]
Phase determination for image and diffraction plane pictures in the electron microscope,
R. W. Gerchberg and W. O. Saxton, “Phase determination for image and diffraction plane pictures in the electron microscope,” Optik34, 275–284 (1971)
1971
-
[38]
Torchoptics: An open-source python library for differentiable fourier optics simulations,
M. J. Filipovich and A. I. Lvovsky, “Torchoptics: An open-source python library for differentiable fourier optics simulations,” arXiv preprint arXiv:2411.18591 (2024). 9 pages, 6 figures
2024 arXiv
-
[39]
Spatial filtering for zero-order and twin-image elimination in digital off-axis holography,
E. Cuche, P. Marquet, and C. Depeursinge, “Spatial filtering for zero-order and twin-image elimination in digital off-axis holography,” Appl. Opt.39, 4070–4075 (2000)
2000
-
[40]
Comprehensivemodelandperformanceoptimization of phase-only spatial light modulators,
A.A.Pushkina,J.I.Costa-Filho,G.Maltese,andA.I.Lvovsky,“Comprehensivemodelandperformanceoptimization of phase-only spatial light modulators,” Meas. Sci. Technol.31 (2020)
2020
-
[41]
Simplifying tailored generation of complex structured femtosecond pulses with easily fabricated phase plates,
P. Veselá, J. Junek, R. Doleček,et al., “Simplifying tailored generation of complex structured femtosecond pulses with easily fabricated phase plates,” Opt. Express32, 24756 (2024)
2024
-
[42]
Phase-contrast microscopy,
C. R. Burch and J. P. P. Stock, “Phase-contrast microscopy,” J. Sci. Instrum.19, 71 (1942)
1942
-
[43]
Moderndark-fieldmicroscopyandthehistoryofitsdevelopment,
S.H.Gage,“Moderndark-fieldmicroscopyandthehistoryofitsdevelopment,”Trans.Am.Microsc.Soc. 39,95–141 (1920)
1920
-
[44]
Orbital angular momentum: origins, behavior and applications,
A. M. Yao and M. J. Padgett, “Orbital angular momentum: origins, behavior and applications,” Adv. Opt. Photon.3, 161–204 (2011)
2011
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.