{"id":"7e4c2d6b-6701-4911-84ca-9ddbc78ece5a","arxiv_id":"2508.20058","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A diffractive neural network mode-sorter with trainable output detection regions achieves higher efficiency at equal crosstalk than fixed-region sorters.","lead":"This paper shows that a light-beam sorter built from programmable phase plates works better when the output detection regions are also chosen automatically during training. The method achieves higher efficiency at the same crosstalk level than sorters with fixed output regions, in simulation and experiment.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed-detection baseline in Figs. 2/6 is an unoptimized circular-region placement; the claimed flexible-region advantage may be largely due to region optimization rather than shape flexibility.","rationale":"The reader's conditional verdict already identifies the same load-bearing assumption: the fixed-detection baseline is not optimized. My stress-test agrees; the strongest claim depends entirely on a fair comparison. The abstract and conclusion describe the flexible-region result as outperforming traditional mode-sorting, but the only direct evidence is a simulation/experiment against an ad-hoc circular fixed-region baseline. Because the flexible regions are learned while the fixed regions are not, the result is consistent with an alternative explanation: most of the gain comes from simply optimizing region placement (centers/radii) rather than from the shape flexibility or joint training. The proposed concrete test—optimizing circular regions with the same loss—directly separates these explanations. I recommend keeping the reader's CONDITIONAL verdict: the paper is interesting and the simulation-experiment agreement is a real strength, but the advertised claim needs this baseline check before it can be accepted as stated. The re-optimization of regions on the measured outputs used for the reported metrics is an additional concern, but it is secondary and can be addressed separately.","tokens_in":7419,"tokens_out":4940,"duration_ms":54602,"concrete_test":"Rerun the same TorchOptics training pipeline with detection regions constrained to circles, but optimize their centers and radii with the same loss function Eq. (7) and the same α sweep as the flexible-region training. Keep the total area and number of regions matched. If the best optimized circular-region sorter reaches an efficiency close to 58.5% at 29% crosstalk (within ~10% relative), then the flexible-region advantage is not established; if it remains near 30%, the advantage is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the comparison between flexible detection regions and a fixed-detection baseline. In §3.2 the fixed-region data are generated with circular regions at fixed locations, with radii and centers chosen so as to match the total area of the flexible regions. No optimization of these region parameters is described, and no use of the same loss Eq. (7) over region placement is reported. The flexible regions, by contrast, are optimized jointly with the phase plates via Eqs. (4)-(7). Thus the comparison conflates two changes: (i) the shape of the regions (circular fixed vs. arbitrary learned) and (ii) the optimization of region placement. A fixed circular layout whose centers and radii are optimized could plausibly recover a substantial part of the 30%→58.5% gap, which would weaken the abstract's 'outperforms' claim. The claim is also broader than the evidence because no direct comparison to WMM-based MPLC sorters is provided, but the unoptimized baseline is the more immediate threat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes training diffractive neural network mode-sorters with flexible output detection regions that are jointly optimized with the phase plates. It presents simulations and experiments for 1-, 2-, and 3-plate sorters sorting up to 25 Hermite-Gaussian modes, and reports that flexible regions improve the efficiency-crosstalk trade-off compared to fixed circular detection regions (e.g., 58.5% vs 30% efficiency at 29% crosstalk in simulation). The authors argue that this approach outperforms traditional MPLC mode-sorting methods.","tokens_in":7618,"tokens_out":5498,"duration_ms":58702,"significance":"The idea of including detection-region geometry in the trainable parameter set is simple, general, and potentially useful; if the comparison is fair, the reported efficiency gain is substantial. The manuscript demonstrates the effect in both simulation and experiment, and the use of an open-source differentiable simulator (TorchOptics) supports reproducibility. The explicit Pareto-style trade-off via the hyperparameter alpha is a useful design element. However, the strength of the central claim depends on the fairness of the fixed-detection baseline and on the evaluation protocol; both need strengthening before the 'outperforms' statement can be accepted.","major_comments":[{"comment":"The central comparison uses a fixed-detection baseline that is not optimized. The text states that the fixed regions are 'circular at fixed locations, with the radii and center positions chosen such that the total area is similar to the total area of the flexible detection regions.' No optimization of these region parameters is reported. The flexible detector, by contrast, is trained jointly via Eqs. (4)-(7). The reported improvement (30% to 58.5% efficiency at 29% crosstalk) therefore conflates shape flexibility with optimization of region placement. Please optimize the fixed-region centers and radii with the same loss, or compare against another optimized fixed-region sorter (e.g., WMM), and provide the resulting curve. Without this, the 'outperforms' claim is not established.","section":"§3.2, Fig. 2"},{"comment":"The experimental evaluation re-optimizes the flexible detection regions against the measured output fields before computing performance ('we re-optimize the detection regions accounting for the experimentally measured outputs', Fig. 5). If the same measurements are then used to report efficiency and crosstalk, this gives the flexible detector an advantage that the fixed detector does not receive. Please evaluate on held-out measurements, or at least provide a split-data comparison and quantify the optimism. The manuscript should also state explicitly whether the fixed-region results in Fig. 6 were re-optimized in the same way.","section":"§3.3, Fig. 6"},{"comment":"The abstract concludes that the approach 'outperforms traditional mode-sorting methods', but the only direct comparator presented is a DONN with unoptimized fixed circular detection. No WMM-trained MPLC, log-polar sorter, or other established mode-sorter is benchmarked. This claim is broader than the evidence. I recommend either adding such a benchmark or narrowing the conclusion to 'outperforms this particular fixed-detection baseline', pending the outcome of the first comment.","section":"Abstract and §4"}],"minor_comments":[{"comment":"The efficiency definition in the Fig. 6 caption (ratio to output intensity with all phase plates set to 0) differs from the efficiency used in simulation via Eq. (5). Please align the definitions or explain why the experimental metric is equivalent.","section":"§3.2 / Fig. 6"},{"comment":"Typographical issues: 'charactrization' in §3.1; 'Fountaineet al.' missing space; 'a∼ 20% crosstalk' spacing. The inset in Fig. 2(a) for alpha≈0 is difficult to read and should be enlarged.","section":"Global"},{"comment":"The Fig. 5 caption is terse; it would help to state explicitly which panels correspond to simulation and which to experiment, and to define the purple and orange circles referenced in the text.","section":"§3.3 / Fig. 5"},{"comment":"No data-availability or code-availability statement is included. Given the reproducibility-oriented claims, providing trained phase-plate profiles and detection-region parameters (or a link to code) would be valuable.","section":"Availability"}],"recommendation":"major_revision","confidential_remarks":"The central idea is plausible and the experiments are nontrivial, but the headline superiority claim is currently supported only by a weak baseline and a potentially optimistic evaluation protocol. These issues are fixable within the scope of the manuscript, so I recommend major revision rather than rejection. The lack of comparison with WMM or another established mode-sorter is a secondary but relevant concern for a general optics journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: the paper trains detection regions as part of the network — that's new and useful. The experiments are careful and the simulation-experiment agreement is decent. My main hesitation is the comparison: the fixed-detection baseline uses circular regions placed arbitrarily (only total area matched), so the claim that flexible regions outperform fixed ones conflates shape flexibility with the act of optimizing the regions at all. A fair test would train circular regions with the same loss. Also the detection regions are re-optimized on the same measured outputs used to compute the performance numbers; a held-out set would be better.\n\nWhat's good: the training scheme is straightforward and clearly explained. The loss weighting (Eq. 7) gives a tunable trade-off. The experimental results are honest — they show the gap to simulation and attribute it to SLM imperfections. They use open-source TorchOptics. The idea is a meaningful extension of prior MPLC/DONN training.\n\nSoft spots in proportion: The baseline issue is the main one. It's not fatal because the paper does show that adding detection regions to the trainable set improves over their baseline, but the magnitude of the improvement (30%→58.5% efficiency) is likely overestimated. It would be surprising if optimizing circular regions didn't recover a large part of that gap. The re-optimization issue is minor but should be addressed with a validation set. The abstract's 'outperforms traditional mode-sorting methods' is broader than the evidence, since no WMM comparison is made.\n\nWho this is for: people working on MPLC/DONN mode-sorters, and anyone using intensity-based mode demultiplexing. It's a solid incremental contribution. I'd suggest accepting for peer review with a request to add an optimized fixed-region baseline and a held-out evaluation.","headline":"Good idea, honest experiments, but the flexible-vs-fixed comparison is confounded by an unoptimized baseline; the advantage is real but likely smaller than claimed.","tokens_in":8109,"tokens_out":1978,"would_cite":true,"duration_ms":22437,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training a mode-sorter's detection regions along with its phase plates roughly doubles efficiency at matched crosstalk in simulation, and the gain persists in experiment.","keywords":["mode-sorting","diffractive optical neural networks","multi-plane light conversion","trainable detection regions","Hermite-Gaussian modes","efficiency-crosstalk trade-off","backpropagation training","spatial mode demultiplexing"],"falsifier":"An apples-to-apples benchmark settles it: train one sorter with flexible regions and another with fixed regions whose positions and sizes are also optimised (by grid search or joint gradient training, then frozen), using identical phase-plate budgets and 25, 100, and 210 modes across at least two modal bases. If the optimised fixed-region design matches the flexible one in efficiency at equal crosstalk, the central advantage claim would be refuted; if the gap persists, it is confirmed. A purely experimental check: send a free-space communication signal through both sorters and compare channel","tokens_in":7295,"feed_emoji":"🔦","tokens_out":17171,"duration_ms":154839,"temperature":0.7,"pith_summary":"Mode-sorting is the task of decomposing a light field into transverse spatial modes and steering each mode to its own spot on a detector, so all mode intensities can be read at once. The paper shows that a diffractive optical neural network—an SLM-based multi-plane light converter—can be trained to do this better if the boundaries of the detector regions are themselves trainable parameters, optimised by backpropagation together with the phase plates. Previous methods fix the output field or the detection geometry in advance; the authors argue that for intensity-measuring tasks only the partition of the detector matters, and giving the network freedom over that partition roughly doubles efficiency at equal crosstalk (58.5% versus 30% for 25 Hermite-Gaussian modes at 29% crosstalk in simulation). They confirm the trend in experiments on sorters handling 4, 9, 16, and 25 modes, and show that re-optimising the regions against measured outputs compensates experimental imperfections. If the claim holds, mode-sorters for communication, imaging, and quantum applications become simpler to build and more efficient at the same level of crosstalk.","feed_headline":"Trainable detection regions double mode-sorter efficiency","feed_subtitle":"At equal crosstalk, the trainable-region design collects 58.5% of the light versus 30% for fixed regions.","key_machinery":"The central object is the flexible detection region: a trainable, non-overlapping partition {D_j} of the output plane, into which the network directs each input mode. Instead of prescribing exact output field shapes, the network maximises the diagonal of the intensity matrix I_ij = ∫_{D_j} |Ψ_output^(i)|² dx dy through the weighted loss α·Losseff + (1−α)·Lossxtalk, with Losseff = −(1/n) Σ_i I_ii and Lossxtalk = (1/n) Σ_i (1 − I_ii / Σ_j I_ij). The hyperparameter α sets the efficiency-crosstalk trade-off, and the regions are re-optimised against experimentally measured outputs to compensate SLM imperfections.","core_discovery":"On the paper's own terms, the discovery is that a mode-sorter's output detection regions should be part of the trainable set of parameters, not a fixed prescription. For tasks that only measure modal intensities, the network need not produce a prescribed output field shape; it only needs to concentrate each input mode's light within its own region of the detection plane. The authors optimise the phase plates and the non-overlapping region set {D_j} jointly by backpropagation, using a loss that weights efficiency against crosstalk via a hyperparameter α. The result is a substantially better efficiency-crosstalk trade-off than the fixed-region baseline—58.5% versus 30% efficiency at 29% crosst","pith_inferences":["The method is basis-agnostic: the training loop consumes labelled input modes and never assumes Hermite-Gaussian structure, so the same gain should transfer to LG, OAM, Zernike, or arbitrary speckle bases—a direct next test.","How much of the advantage is intrinsic to region flexibility depends on the baseline: a fairer benchmark would optimise fixed-region positions and sizes with comparable effort, and comparing on 100+ modes or non-orthogonal states would bound the effect.","The design principle reaches beyond optics: any wave-based processor whose readout integrates intensity over a detector partition has that partition as a legitimate trainable parameter, so acoustic, microwave, or terahertz implementations could adopt the same trick.","A promising extension is a hybrid scheme that trains flexible regions while also matching the field where a few output channels couple into fibres, bridging the intensity-measurement and fibre-coupling regimes."],"forward_implications":["Mode-sorters for intensity-measuring applications—free-space communication, imaging, endoscopy, passive superresolution—can collect roughly twice as much light per mode at the same crosstalk without adding phase plates.","The hyperparameter α gives a tunable efficiency-crosstalk trade-off, so the same trained hardware can be configured for high-efficiency state discrimination or for low-crosstalk imaging.","Re-optimising detection regions against the measured output fields recovers performance lost to experimental imperfections, without redesigning the phase plates; the single-plate, 4-mode sorter then beats even the simulated fixed-region design.","Enforcing the HG symmetry of the phase plates in the network parametrisation keeps simulated performance nearly unchanged while making experimental alignment much easier.","Because the hardware is just one SLM and one mirror, the sorter is easily reproducible in an ordinary optics lab, and fabricated phase plates can replace the SLM where its cavity effect and pixel crosstalk limit performance."],"supporting_citations":[{"why":"Supplies the wavefront matching method, the standard MPLC training algorithm that the flexible-region approach is benchmarked against.","marker":"[26]"},{"why":"Introduces multi-plane light converters, the programmable phase-plate hardware platform that the trained sorter implements.","marker":"[16]"},{"why":"Sets the MPLC mode-sorting state of the art (210 HG modes into fibres at 19% crosstalk) that motivates the search for better training.","marker":"[18]"},{"why":"Shows that earlier gradient-descent training of MPLCs found no significant advantage over WMM, the premise the flexible-region result departs from.","marker":"[19]"},{"why":"Argues that wavefront matching is interpretable as a gradient-descent variant, motivating backpropagation training of the DONN.","marker":"[31]"},{"why":"Provides the experimental precedent of OAM mode de/multiplexing with an optical diffraction neural network trained by backpropagation.","marker":"[34]"},{"why":"Supplies the differentiable Fourier-optics simulation library used to train and test the mode-sorters numerically.","marker":"[38]"}],"fun_headline_variants":["Trainable detection regions double mode-sorter efficiency","Mode-sorter with trainable regions: efficiency doubles","Trainable output regions yield 2x mode-sorting efficiency","Jointly training detection regions doubles mode-sorter efficiency","Same crosstalk, double efficiency with trainable regions"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim rests on a comparison in which the fixed-region baseline consists of circular regions whose positions and sizes are not themselves optimised; a well-optimised fixed-region design could plausibly close much of the efficiency gap.","fun_headline_variants_meta":{"raw":{"variants":["Trainable detection regions double mode-sorter efficiency","Mode-sorter with trainable regions: efficiency doubles","Trainable output regions yield 2x mode-sorting efficiency","Jointly training detection regions doubles mode-sorter efficiency","Same crosstalk, double efficiency with trainable regions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001001,"raw_usage":{"total_tokens":4003,"prompt_tokens":605,"completion_tokens":3398,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":349,"completion_tokens_details":{"reasoning_tokens":3334}},"tokens_in":349,"tokens_out":3398,"duration_ms":25753,"temperature":1.0,"reasoning_tokens":3334,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:13:11.088328+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An apples-to-apples benchmark settles it: train one sorter with flexible regions and another with fixed regions whose positions and sizes are also optimised (by grid search or joint gradient training, then frozen), using identical phase-plate budgets and 25, 100, and 210 modes across at least two modal bases. If the optimised fixed-region design matches the flexible one in efficiency at equal crosstalk, the central advantage claim would be refuted; if the gap persists, it is confirmed. A purely experimental check: send a free-space communication signal through both sorters and compare channel","supporting_citations":[{"cited_title":"Optical circuit design based on a wavefront-matching method,","cited_arxiv_id":null,"evidence_quote":"Supplies the wavefront matching method, the standard MPLC training algorithm that the flexible-region approach is benchmarked against."},{"cited_title":"Programmable unitary spatial mode manipulation,","cited_arxiv_id":null,"evidence_quote":"Introduces multi-plane light converters, the programmable phase-plate hardware platform that the trained sorter implements."},{"cited_title":"Laguerre-gaussian mode sorter,","cited_arxiv_id":null,"evidence_quote":"Sets the MPLC mode-sorting state of the art (210 HG modes into fibres at 19% crosstalk) that motivates the search for better training."},{"cited_title":"High-dimensional spatial mode sorting and optical circuit design using multi-plane light conversion,","cited_arxiv_id":null,"evidence_quote":"Shows that earlier gradient-descent training of MPLCs found no significant advantage over WMM, the premise the flexible-region result departs from."},{"cited_title":"Wavefront matching method as a deep neural network and mutual use of their techniques,","cited_arxiv_id":null,"evidence_quote":"Argues that wavefront matching is interpretable as a gradient-descent variant, motivating backpropagation training of the DONN."},{"cited_title":"Broadband, low-crosstalk, and massive-channels oam modes de/multiplexing based on optical diffraction neural network,","cited_arxiv_id":null,"evidence_quote":"Provides the experimental precedent of OAM mode de/multiplexing with an optical diffraction neural network trained by backpropagation."}],"review_version":1}