Pith. sign in

REVIEW 4 major objections 4 minor 31 references

Trainable dynamical masking for readout-free optical computing

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that trainable dynamical masks on a telecom Mach-Zehnder modulator can take over the readout layer's role, letting an optical computing machine work with a single readout weight.

desk verdict Trainable temporal masks can plausibly replace a high-dimensional readout in a telecom-based optical ELM, but the main claim rests on undocumented FFT simulation details. read the letter →

arxiv 2505.23464 v1 pith:Q4NNIVKR submitted 2025-05-29 physics.optics

classification physics.optics
keywords trainabledynamicalmaskextremelearningmachineopticalcomputingMach-Zehndermodulatorreadout-freetimeseriespredictiondispersiontelecom-gradedevices
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a physical extreme-learning-machine computer built from standard telecom parts—a Mach-Zehnder modulator, a dispersive medium, and a photodiode—can shift the learning burden from a traditional readout layer onto two time-varying masks imposed on the modulator's drive voltage. The central result is that shrinking the readout to a single weight barely hurts performance on the tested regression and chaotic time-series tasks, because the trained masks effectively absorb the readout's function. This matters because it replaces a large, hard-to-implement readout with mask waveforms that can in principle be updated on a high-speed telecom modulator. If the behavior holds in hardware, it would simplify the training interface for analog optical computers.

What carries the argument

The central object is the trainable dynamical mask: a time-varying modulation of the Mach-Zehnder modulator's drive voltage, with two components $k_{\mathrm{mask}}(t)$ and $b_{\mathrm{mask}}(t)$, each holding a set of samples per input symbol. It is trainable because the samples are adjusted by gradient descent together with any readout weights. The mask exploits the modulator's cosine transfer function to create a nonlinear high-dimensional feature space, and the dispersive medium's quadratic phase response mixes neighboring symbols so that information lost when the readout is small can be recovered. The photodiode's square-law detection completes the chain by converting the optical field to a measured current. Together these elements make the readout layer optional.

What would settle it

Take the same configuration (six input features, mask size 8 per symbol, dispersion $D=0.4$ ns/nm, readout size 1), implement it on a physical 45 GHz modulator, fiber, and photodiode, train the masks, and compare test loss to the simulated readout-free results; if the real loss is substantially worse, or the system requires a readout larger than one to recover, the transfer-of-function claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the function of the trainable readout layer is transferred to the trainable masks. The input is encoded as $V_1(t)=V_\pi(k_{\mathrm{mask}}(t)X_{\mathrm{data}}(t)+b_{\mathrm{mask}}(t))$, where $k_{\mathrm{mask}}$ and $b_{\mathrm{mask}}$ are multiplicative and additive masks sampled multiple times per symbol; the modulated light then passes through a dispersive medium and a square-law photodiode. The authors show for six input features that test loss stays nearly flat as readout size is reduced from 48 to 1 when masks are trained and dispersion is present, while random fixed masks or missing dispersion degrade sharply. They interpret dispersion as the mechanism that mixes information across symbols when the readout is smaller than the number of input features. The same readout-free configuration produces accurate predictions on a chaotic time series for horizons up to 500 steps.

Load-bearing premise

The load-bearing premise is that the simulated chain—an ideal cosine Mach-Zehnder modulator, a lossless purely dispersive medium, and a square-law photodiode—faithfully represents real off-the-shelf telecom hardware, and that the masks can actually be updated at 45 GHz; the paper provides no experiment.

Editorial extensions

If this is right

  • Readout-free operation: the system can produce a prediction from a single photocurrent sample, with the masks carrying the task-specific transformation.
  • Hardware simplification: a high-dimensional readout is replaced by two mask waveforms, so the physical trainable interface is the modulator drive rather than an array of output weights.
  • Dispersion dependence: dispersion is not an imperfection but a functional mixing layer, and it becomes decisive when the readout has fewer elements than input features.
  • Parameter efficiency: on the yacht regression benchmark the model uses at most 96 trainable parameters and outperforms a linear regression baseline, with performance better than the reported 870-parameter multilayer perceptron.
  • Benchmark validation: the readout-free model predicts a chaotic time series up to 500 steps, with median accuracy improving over earlier reservoir-computing results on the same task.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the readout-free effect generalizes beyond the two benchmarks, the same masking scheme could be applied to other differentiable nonlinear optical elements, such as semiconductor optical amplifiers or nonlinear fibers, turning them into trainable feature maps without changing their internal parameters.
  • Editorial inference: a natural experimental test would be to implement the mask updates in hardware on a real 45 GHz modulator and compare the trained waveforms to the simulated ones; the paper currently provides numerical evidence only.
  • Editorial inference: because the masks are trained jointly with the readout, the scheme suggests an architecture where all trainable parameters live in the time domain at the modulator, potentially allowing online training at the full signal rate rather than through a separate digital readout stage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript proposes an optical extreme learning machine architecture composed of a Mach-Zehnder modulator driven by trainable temporal masks, a dispersive medium, and a photodiode, with an optional readout layer that can be reduced to a single weight. The authors simulate two benchmarks—yacht hydrodynamics regression and Mackey-Glass time-series prediction—and report that when the masks and dispersion are present, reducing the readout to one element degrades performance only minimally (Fig. 3b), implying that the readout function is transferred to the masks. The paper is entirely numerical, based on a chain of idealized component models in Eqs. (1)–(4).

Significance. If the readout-free plateau is robust, the proposal offers a practical simplification for physical ELM/reservoir computing: a single photocurrent sample could replace a high-dimensional trained readout, using off-the-shelf telecom components. The paper has strengths: the Fig. 3b study is repeated 25 times and reports medians and percentiles; the masks and readout are trained end-to-end on a training set and evaluated on held-out data, so there is no derivation masquerading as prediction. However, the central claim rests on unstated simulation conventions (zero-padding, readout-sample timing, photodiode response) that are not specified, and the headline regression result is a single run.

major comments (4)
  1. [Section II, Eq. (3); Section IV, Fig. 3b] The dispersion operation is implemented in the frequency domain as a pure phase filter, which in a finite FFT window corresponds to circular convolution unless the sequence is zero-padded. The manuscript never states whether zero-padding or guard intervals are used, nor where in the output window the single readout sample is taken. If the last sample of the M-symbol window is read while the circular tail from the first symbol wraps into it, the apparent aggregation of all input features into one photocurrent sample is a periodic-boundary artifact rather than a causal effect of the telecom chain. Because the readout-free plateau in Fig. 3b is the paper's central claim, the authors must specify the simulation convention and repeat the experiment with causal dispersion (e.g., zero-padded FFT or time-domain convolution) and with a well-defined readout timing.
  2. [Section IV, Fig. 2] The headline R2=0.997 for the readout-free regression is reported for a single training run with no error bars or cross-validation. Given that Fig. 3b shows 25-run medians and percentiles, the authors should provide the same repeated-run statistics for Fig. 2 to demonstrate that the result is not seed-dependent. In addition, the comparison with the multilayer perceptron from Ref. [29] uses an external protocol with an unknown data split and normalization; the authors should state whether their 70/30 split and normalization are compatible with the benchmark, or provide a direct reimplementation of the MLP under identical conditions.
  3. [Section III, Eq. (5); Section IV] The number of trainable parameters in the masks is ambiguous. The text says kmask(t) and bmask(t) each have size Nmask samples per symbol and concludes there are M×Nmask independent values, but a time mask of length Nmask that is applied per symbol usually implies Nmask trainable parameters (the same pattern repeated for every symbol). Later, the yacht example states "mask’s size is eight elements per symbol" and "total length (for six symbols) of 48 independent trainable parameters," which indicates M×Nmask independent values but contradicts the size-Nmask description. This ambiguity affects the claim that the model has an order-of-magnitude fewer parameters than the MLP (96 vs 870), and the stated maximum of 96 parameters is inconsistent with 48+48+2=98. The authors should clarify the mask structure and reconcile the parameter counts.
  4. [Section II, Eq. (4); Section IV] The photodiode impulse response H(τ) in Eq. (4) is never specified, and no absolute symbol rate or readout sampling times are given. The timing of the single readout sample relative to the dispersed impulse response is essential for the readout-free claim: a photodiode with finite bandwidth or a non-δ impulse response will mix neighboring symbols differently than an ideal detector. The authors should state H(τ), the symbol rate, the number of samples per symbol, and the readout sample position used in the simulations, and confirm that the single-sample readout is robust to realistic photodiode bandwidth.
minor comments (4)
  1. [Section V, Fig. 4b] The metric is labeled NMSE in the figure and text, but the comparison values from Refs. [1] and [21] are reported as NRMSE84 and NMSE300, respectively. Please standardize the notation and define the normalization used.
  2. [Section V] The Mackey-Glass prediction procedure is under-specified; state whether the predictions are recursive (feeding predicted values back) or one-step, and describe how the 20 input symbols are constructed from the time series.
  3. [Fig. 3b caption] The caption text "for varying readout sizes 8" is unclear; it should read something like "for varying readout size; the mask has 8 samples per symbol."
  4. [Section VI] There is a typo: "high-dimentional" should be "high-dimensional."

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the readout-free plateau is an end-to-end trained and held-out-tested numerical result, not a fitted parameter renamed as a prediction or a self-citation-derived theorem.

full rationale

Walking the claimed derivation chain, the central result (Fig. 3b) is empirical rather than derivational: the paper trains "the masks and readout layer using 70% of data and keeping other 30% for the test" and reports 25 runs per configuration, so the readout-free test loss is a held-out evaluation, not a quantity forced by the model definition. Equations (2)-(4) are standard component models (MZM cosine transfer, dispersive-medium phase factor, photodiode convolution) cited to external references [24]-[27]; no equation defines the claimed readout-free performance in terms of itself or of the trained masks. The phrase "the function of trainable readout is effectively transferred to the masks" is a post-hoc interpretation of the simulation curve, and the paper supports it with controlled comparisons (non-trainable masks and dispersion removed). The authors' own prior work [20,21] appears only for context and as a Mackey-Glass benchmark; it is not invoked as a uniqueness theorem or as the justification for the readout-free claim, so removing those citations would not collapse the argument. The skeptic's concern about FFT circular convolution, missing zero-padding or guard intervals, unspecified photodiode impulse response H(τ), and readout-sample timing is a legitimate reproducibility and physical-realism risk that the paper should document, but it is not circularity: it does not exhibit an equation equal to its input by construction, nor a fitted parameter renamed as a prediction. The paper also explicitly limits itself to "numerical modeling", so the absence of a hardware experiment is a stated scope limitation rather than a hidden circular step.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The scheme rests on standard telecom device models (MZM cosine, fiber dispersion, square-law photodetection) and on several hand-chosen hyperparameters (mask size, dispersion, initial scaling, bandwidth cutoff). No new physical entities are introduced. The main risk is the gap between ideal-component models and real devices, which is not experimentally validated.

free parameters (6)
  • Mask size Nmask = 8
    Set by hand to 8 samples per symbol for both tasks; determines the feature-space dimension and the number of trainable mask parameters.
  • Accumulated dispersion D = 0.4 ns/nm
    Chosen by hand for both tasks; the paper shows dispersion is critical when the readout is small but does not optimize D.
  • Initial mask scaling factor = 5
    Random initial masks are scaled by 5 to place the operating point in a moderately nonlinear region of the MZM curve; an ad hoc training choice.
  • 45 GHz bandwidth cutoff = 45 GHz
    Chosen to mimic realistic device bandwidth; the exact filter response is not specified.
  • Training epochs for yacht regression = 20000
    Used for the yacht task; no early stopping or epoch count is given for the Mackey-Glass task.
  • Data split and normalization = 70/30; zero mean, max absolute value one
    Normalization for the optical model differs from the zero mean/variance-one normalization used for the linear baseline, complicating direct comparison.
assumptions (5)
  • domain assumption MZM amplitude transfer is the ideal cosine law, Eq. (2), with push-pull driving V2 = -V1.
    The entire nonlinear feature map relies on this transfer function; real MZMs have extinction ratio, chirp, and voltage-dependent loss which are neglected.
  • domain assumption Dispersion medium is linear second-order dispersion, Eq. (3): E_out(ω) = E_in(ω) exp(-i ω^2 A / 2).
    Neglects fiber loss, higher-order dispersion, self-phase modulation, and noise; real propagation adds impairments.
  • domain assumption Photodiode is an ideal square-law detector with impulse response H(τ), Eq. (4).
    Neglects detector noise, saturation, and bandwidth limitations beyond the unspecified H(τ).
  • domain assumption Gradient descent through the simulated chain in TensorFlow corresponds to a trainable physical implementation.
    No experimental demonstration shows that masks can be updated physically at the modeled speed; hardware-in-the-loop training is not addressed.
  • ad hoc to paper Random mask initialization scaled by 5 is sufficient for convergence.
    A heuristic unique to this paper; the central performance depends on it, as stated in Section III.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Trainable dynamical masking for readout-free optical computing." pith.science (2026). https://pith.science/paper/Q4NNIVKR

@misc{pith2026250523464,
  author       = {Pith},
  title        = {Pith review of: Trainable dynamical masking for readout-free optical computing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q4NNIVKR}},
  note         = {Machine review of arXiv:2505.23464}
}
read the original abstract

Nonlinear systems, transforming an input signal into a high-dimensional output feature space, can be used for non-conventional computing. This approach, however, requires a change of system parameters during training rather than coefficients in a software program. We propose here to use available off-the-shelf high-speed optical communication devices and technologies to implement a trainable dynamical mask in addition to or even instead of the traditional readout layer for extreme learning machine-based computing. The computational potential of the proposed approach is demonstrated with both regression and time series prediction tasks.

Figures

Figures reproduced from arXiv: 2505.23464 by the authors.

Figure 1
Figure 1. The principal scheme of the computing sys [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (a) Training and test loss and (b) prediction [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. (a) The distribution of the argument π∆V 2Vπ of the trained model and corresponding to it the output of MZM transfer function, Eq. (2). (b) The dependence of the model’s effectiveness on the number of elements for varying readout sizes 8 [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: (a) The Mackey-Glass time series prediction [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

31 extracted references · 26 canonical work pages

  1. [20]

    Optical neuromorphic computing via temporal up-sampling and trainable encoding on a telecom device platform

    E. Manuylovich, D. Stoliarov, D. Saad, and S. K. Tu- ritsyn, Optical computing using telecom device platform for nonlinear mapping to high-dimensional feature space, arXiv:2501.03429 (2025)

  2. [29]

    Gerritsma, R

    J. Gerritsma, R. Onnink, and A. Versluis, Yacht Hydro- dynamics, UCI Machine Learning Repository (1981)

  3. [1]

    echo state

    H. Jaeger, The “echo state” approach to analysing and training recurrent neural networks – with an Erratum note, GMD Technical Report 148 (German National Re- search Center for Information Technology, Bonn, Ger- many, 2001)

  4. [2]

    Maass, T

    W. Maass, T. Natschl¨ ager, and H. Markram, Real-time computing without stable states: A new framework for neural computation based on perturbations, Neural Com- putation 14, 2531 (2002)

  5. [3]

    Huang, Q.-Y

    G.-B. Huang, Q.-Y. Zhu, and C.-K. Siew, Extreme learn- ing machine: Theory and applications, Neurocomputing 70, 489 (2006)

  6. [4]

    Jaeger, B

    H. Jaeger, B. Noheda, and W. G. van der Wiel, To- ward a formal theory for computing machines made out of whatever physics offers, Nature Communications14, 4911 (2023)

  7. [5]

    Tanaka, T

    G. Tanaka, T. Yamane, J. B. H´ eroux, R. Nakane, N. Kanazawa, S. Takeda, H. Numata, D. Nakano, and A. Hirose, Recent advances in physical reservoir comput- ing: A review, Neural Networks115, 100–123 (2019)

  8. [6]

    L. G. Wright, T. Onodera, M. M. Stein, T. Wang, D. T. Schachter, Z.Hu,andP.L.McMahon,Deepphysicalneu- ral networks trained with backpropagation, Nature601, 549–555 (2022). 5

Show all 31 references
  1. [7]

    L. G. Wright, Training physical systems like neural net- works, inPhotonic Computing: From Materials and De- vices to Systems and Applications, Vol. PC13113 (2024) p. PC131130C

  2. [8]

    R. P. Feynman, Simulating physics with computers, In- ternational journal of theoretical physics21, 467 (1982)

  3. [9]

    T. W. Hughes, I. A. D. Williamson, M. Minkov, and S. Fan, Wave physics as an analog recurrent neural net- work, Science Advances5, eaay6946 (2019)

  4. [10]

    Marcucci, D

    G. Marcucci, D. Pierangeli, and C. Conti, Theory of neuromorphic computing by waves: machine learning by rogue waves, dispersive shocks, and solitons, Physical Re- view Letters 125, 093901 (2020)

  5. [11]

    T. Zhou, F. Scalzo, and B. Jalali, Nonlinear Schr¨ odinger Kernel for Hardware Acceleration of Machine Learning, Journal of Lightwave Technology40, 1308 (2022)

  6. [12]

    A. M. I. Muda and U. Te˘ gin, Optical computing with su- percontinuum generation in photonic crystal fibers, Opt. Express 33, 7852 (2025)

  7. [13]

    D. R. Solli and B. Jalali, Analog optical computing, Na- ture Photonics 9, 704 (2015)

  8. [14]

    P. L. McMahon, The physics of optical computing, Na- ture Reviews Physics5, 717 (2023)

  9. [15]

    B. Kia, J. F. Lindner, and W. L. Ditto, Nonlinear dynam- ics as an engine of computation, Philosophical Transac- tions of the Royal Society A: Mathematical, Physical and Engineering Sciences 375, 20160222 (2017)

  10. [16]

    Oguz, J.-L

    I. Oguz, J.-L. Hsieh, N. U. Dinc, U. Te˘ gin, M. Yildirim, C. Gigli, C. Moser, and D. Psaltis, Programming non- linear propagation for efficient optical learning machines, Advanced Photonics6, 016002 (2024)

  11. [17]

    Gigan, Imaging and computing with disorder, Nature Physics 18, 980 (2022)

    S. Gigan, Imaging and computing with disorder, Nature Physics 18, 980 (2022)

  12. [18]

    X. Xu, M. Tan, B. Corcoran, J. Wu, A. Boes, T. G. Nguyen, S. T. Chu, B. E. Little, D. G. Hicks, R. Moran- dotti, et al., 11 TOPS photonic convolutional accelerator for optical neural networks, Nature589, 44 (2021)

  13. [19]

    N. M. Estakhri, B. Edwards, and N. Engheta, Inverse- designed metastructures that solve equations, Science 363, 1333 (2019)

  14. [21]

    are depicted. VI. DISCUSSION AND CONCLUSIONS We designed and demonstrated via numerical model- ing a new physics-based ELM computational scheme us- ing commercially available high-speed optical communi- cation components and technologies. The main novel el- ements of the propo...

  15. [22]

    Manuylovich, A

    E. Manuylovich, A. Bednyakova, D. Ivoilov, I. Terekhov, and S. Turitsyn, SOA-based reservoir computing using upsampling, Optics Letters49, 5827 (2024)

  16. [23]

    Appeltant, G

    L. Appeltant, G. Van der Sande, J. Danckaert, and I. Fis- cher, Constructing optimized binary masks for reservoir computing with delay systems, Scientific reports4, 3629 (2014)

  17. [24]

    Kuriki, J

    Y. Kuriki, J. Nakayama, K. Takano, and A. Uchida, Im- pact of input mask signals on delay-based photonic reser- voir computing with semiconductor lasers, Optics express 26, 5777 (2018)

  18. [25]

    P. J. Winzer and R.-J. Essiambre, Advanced optical mod- ulation formats, inOptical Fiber Telecommunications VB (Elsevier, 2008) pp. 23–93

  19. [26]

    R. H. d. Souza, O. L. Coutinho, J. E. B. Oliveira, A. A. Ferreira, and J. A. J. Ribeiro, An analytical solution for fiber optic links with photonic-assisted millimeter wave upconversion due to MZM nonlinearities, Journal of Mi- crowaves, Optoelectronics and Electromagnetic App...

  20. [27]

    G. P. Agrawal,Nonlinear fiber optics, 5th ed. (Academic press, 2012)

  21. [28]

    Bowers and C

    J. Bowers and C. Burrus, Ultrawide-band long- wavelength p-i-n photodetectors, Journal of Lightwave Technology 5, 1339 (1987)

  22. [30]

    BaressiˇSegota, N

    S. BaressiˇSegota, N. Andeli´ c, J. Kudl´ aˇ cek, and R.ˇCep, Artificial neural network for predicting values of resid- uary resistance per unit weight of displacement, Pomorski zbornik 57, 9 (2019)

  23. [31]

    M. C. Mackey and L. Glass, Oscillation and Chaos in Physiological Control Systems, Science197, 287 (1977)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.