Pith. sign in

REVIEW 3 major objections 4 minor 46 references

Soft-Demapping for Short Reach Optical Communication: A Comparison of Deep Neural Networks and Volterra Series

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A small neural network can replace a 5th-order Volterra equalizer for short-reach coherent links, matching its achievable rate with 65% fewer multipliers and gaining 0.35 dB OSNR at equal complexity.

desk verdict A workmanlike experimental comparison whose headline numbers are plausible as system-level claims but rest on an unverified max-log assumption that should be checked before citing them. read the letter →

arxiv 2501.05979 v1 pith:VWEJID6K submitted 2025-01-10 eess.SP

classification eess.SP
keywords soft-decisiondemappingVolterraseriesdeepneuralnetworkequalizershort-reachopticalcommunicationcoherent64QAMbitwiseequivocationlossachievableratenonlinearequalization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to show that a compact deep neural network that outputs soft bit values directly can replace a 5th-order Volterra nonlinear equalizer followed by a max-log soft demapper in short-reach coherent optical systems. On back-to-back 92 GBd dual-polarization 64QAM measurements, the neural equalizer matches the Volterra's achievable rate with 65% fewer multipliers and improves OSNR by 0.35 dB at equal multiplier count. The key enabler is a bitwise loss function that is equivalent to binary cross-entropy and that, when training succeeds, makes the network output the logarithmic a-posteriori probability ratio, which maximizes the generalized mutual information. If correct, this gives receiver designers a concrete complexity budget for nonlinear compensation and a principled way to train soft demappers.

What carries the argument

The central object is the bitwise equivocation loss (13), equivalent to binary cross-entropy, used to train both the neural equalizer and the Volterra equalizer plus demapper. The paper proves this loss is optimal: a trained network using it learns the APP ratio log(P_{B|Y}(0|y)/P_{B|Y}(1|y)). The comparison is carried by a Taylor expansion (gradients and Hessians, (7)-(11)) that converts the trained neural network into Volterra kernels, letting the paper interpret which nonlinear orders matter and showing the neural network can adjust kernels per soft output while the Volterra cannot.

What would settle it

Replace the max-log soft demapper after the Volterra equalizer with an exact a-posteriori-probability demapper at the same complexity, or measure the achievable-rate gap between max-log and exact demapping; if the gap is comparable to 0.35 dB, the claimed equal-complexity OSNR gain would shrink or vanish.

Watch

Extended reading notes

Core claim

On coherent 92 GBd dual-polarization 64QAM back-to-back measurements where component nonlinearities dominate, a fully connected neural network trained with the bitwise equivocation loss achieves the same achievable rate as a jointly optimized 5th-order Volterra equalizer with a max-log soft demapper while requiring 65% fewer multipliers; at equal multiplier count it gains 0.35 dB in OSNR at the 15%-overhead FEC limit. The paper also proves that this loss is equivalent to binary cross-entropy and trains the network to output the logarithmic a-posteriori probability ratio, i.e., the optimal soft bit, minimizing the conditional entropy and thereby maximizing the achievable rate (12).

Load-bearing premise

The comparison presumes that the max-log approximation used for the Volterra soft demapper loses essentially nothing in the signal-to-noise range tested, so the neural network's gain is measured against a near-ideal Volterra demapper.

Editorial extensions

If this is right

  • At a fixed FEC overhead, the SDNNE with hard-tanh activations reaches Volterra-level nonlinear compensation at 136 multipliers versus 385 for the sparse VNLE, pointing to a low-complexity regime where the neural architecture degrades more gracefully.
  • Because the bitwise equivocation loss equals binary cross-entropy, standard supervised-learning pipelines can be used directly to train rate-maximizing soft demappers for coherent receivers.
  • The observation that 2nd, 3rd, and 5th order kernels dominate, while 4th, 6th, and 7th add no gain, identifies odd-order component nonlinearities as the target any equalizer must capture in this setup.
  • Training can be validated by checking that the minimizing s in the achievable-rate expression (12) equals 1 for the network's outputs, giving a practical stopping rule.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the max-log demapper assumption is relaxed at lower OSNR, the 0.35 dB claim would need re-benchmarking; one would expect the gap to shrink when the Volterra side uses an exact APP demapper.
  • The dominance of odd-order kernels suggests a memoryless or weakly memory odd-power nonlinearity model might explain most of the component distortion; that hypothesis is testable by fitting a static nonlinearity and comparing residuals.
  • The same bitwise-loss training could be applied to other equalizer structures, including Volterra variants, so the complexity comparison is between architectures, not between loss functions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a soft deep neural network equalizer (SDNNE) that performs joint nonlinear equalization and bitwise soft demapping for short-reach coherent optical systems. The SDNNE is trained with a bitwise equivocation loss (13), which the authors show to be equivalent to binary cross-entropy and to learn the a-posteriori probability ratio when training succeeds. On 92 GBd dual-polarization 64QAM back-to-back measurements, the SDNNE is compared with a 5th-order Volterra nonlinear equalizer (VNLE) followed by a max-log soft demapper with optimized slopes (MLA). The central claims are that at equal achievable rate, the SDNNE requires 65% fewer multipliers, and at equal multiplier count it provides a 0.35 dB OSNR gain at the assumed FEC limits. The paper also presents a Volterra-kernel analysis of the trained SDNNE and a complexity model in terms of hardware multipliers.

Significance. If the experimental results hold, the paper offers a practical complexity reduction for nonlinear equalization in coherent short-reach links, backed by explicit complexity formulas and a well-defined training objective. The equivalence of the bitwise equivocation loss to binary cross-entropy and its information-theoretic interpretation are useful theoretical contributions. The main claims are falsifiable and the experimental methodology is largely reproducible from the described setup. However, the central quantitative comparison depends on an unverified assumption about the max-log demapper, and the reported gains are not accompanied by confidence intervals or a separate architecture-selection campaign, which tempers the strength of the conclusions.

major comments (3)
  1. [Sec. II-E, Eq. (12)-(13)] The assertion that the max-log approximation (MLA) 'incurs virtually no loss in the SNR ranges considered in this work' is unsupported by any measurement, simulation, or bound in the manuscript. This is load-bearing because both headline results—the 65% multiplier reduction and the 0.35 dB OSNR gain—compare the SDNNE against a VNLE followed by MLA. If the MLA is lossy relative to the exact APP demapper, the reported gains partly reflect a weaker demapping baseline rather than a superior equalizer. The only supporting evidence cited (a 0.002 bits/QAM improvement from bitwise training over MSE training) validates the loss function, not the optimality of MLA. The authors should report the minimizing s in Eq. (12) for both the VNLE+MLA and the SDNNE at the tested OSNR values; if the optimal s for the VNLE differs materially from 1, that is direct evidence of demapper miscalibration. Alternatively, a simulation or measurement comparison against a VNLE with exact APP demapping would resolve the issue.
  2. [Sec. III-B and IV-C] The quantitative claims (0.35 dB OSNR gain, 65% multiplier reduction) are based on the achievable-rate curves in Figs. 7 and 14, but no confidence intervals or repeated-capture statistics are provided. The SDNNE architectures (17|26|25|3 and subsequently 17|16|10|3) and the VNLE memory configurations are selected on the same measurement campaign, and the reported gains are the best among the evaluated configurations. This creates a risk of selection bias: the headline numbers may overstate the expected gain on new captures. The authors should provide error bars or repeated independent measurements, and ideally perform architecture selection on a separate training campaign or use a nested validation procedure.
  3. [Sec. II-C and III-C] The Volterra-kernel comparison between the SDNNE and the VNLE is performed by Taylor expansion of the trained tanh SDNNE at y0 = 0 (Eqs. (7)-(11)). The input signal is power-normalized but not zero-mean at the operating point, so the kernels at y0 = 0 may not be representative of the network's local behavior at typical input values. This does not invalidate the central complexity/performance comparison, but it undermines the interpretability claim in Sec. III-C that the extracted kernels reflect the equalizer's actual operation. The authors should justify the choice y0 = 0 or evaluate the expansion at the mean of the received signal.
minor comments (4)
  1. [Fig. 7 vs Fig. 8] The VNLE 4th-order configuration is labeled '17:17:11:3' in Fig. 7 but '17:17:11:5' in Fig. 8; the authors should reconcile this discrepancy.
  2. [Appendix A, Eqs. (20)-(21)] The derivation switches from natural logarithm in Eq. (20) to base-2 logarithm in Eq. (21) without comment; the equivalence holds only up to a constant scaling factor, which is immaterial for optimization but should be stated explicitly.
  3. [References] Reference [36] is cited in Fig. 7 as the source of the 15% OH FEC limit, while the text cites [35, Table 9.1]; also, reference [38] is cited in Fig. 7 for the 20% OH limit, but [38] is a channel-estimation paper, not a code-limit reference. Please verify the citations.
  4. [Throughout] There are several typographical issues, e.g., 'V olterra' in the title and headers, 'multiplers' in the Fig. 12 caption, and inconsistent use of 'optNtaps' versus 'optStruc' in figure labels. A thorough proofread is recommended.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central comparison is an externally measured back-to-back benchmark, and the loss-optimality appendix is a self-contained derivation.

full rationale

The paper's headline claims (65% multiplier reduction at equal rate, 0.35 dB OSNR gain at equal complexity) are experimental results on coherent 92 GBd DP-64QAM back-to-back captures, evaluated with the same bitwise loss (13) and the same achievable-rate metric (12) for both the VNLE+MLA and the SDNNE. Neither claim is obtained by fitting a parameter to the metric it then predicts; both equalizer types are trained on the same measured payload and validated on held-out frames. The Appendix (A) shows by elementary algebra that (13) equals binary cross-entropy and, via the information inequality, that the minimizer is the APP ratio; this is a self-contained textbook derivation, not an import of the conclusion. The self-citations [12], [29], and [30] provide prior concepts or the standard GMI formula, but the central comparison and the loss-optimality proof do not reduce to them. The unverified Sec. II-E assertion that the max-log approximation (MLA) 'incurs virtually no loss in the SNR ranges considered' is a correctness/soundness caveat for the VNLE baseline, not a circular step: if MLA were lossy, the 0.35 dB and 65% numbers would be partially attributable to baseline weakness, but this is an experimental-design risk, not an equivalence-by-construction. Likewise, the untested condition that the minimizing s in (12) equals 1 is a missing sanity check, not a circular reduction. Accordingly, no circular step is identified.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The ledger captures the design choices and assumptions behind the comparison. No new physical entities are introduced. The main claims are empirical, so circularity is low, but the reported numbers depend on architecture and tap choices fit to the same measurement campaign, and on domain assumptions about Volterra modeling, MLA losslessness, GMI-to-FEC mapping, and the activation-pattern capacity heuristic. The loss-optimality appendix is standard logistic regression theory.

free parameters (5)
  • SDNNE architecture (layer sizes and memory taps) = 17|26|25|3 (initial), 17|16|10|3 (final), plus smaller variants
    Chosen by comparing activation-pattern bounds and measured achievable rate on the same BtB captures (Figs. 5, 12-14); not derived from a general rule.
  • VNLE memory-length configuration per polynomial order = 17:17:11:3:3 for 5th order; other tap counts varied
    Tap counts are optimized on the measurement data for each architecture (Sec. III-B, Figs. 7-8); the reported 65% and 0.35 dB results depend on these specific tap choices.
  • MLA soft-demapper slopes = not reported
    The VNLE baseline uses a max-log soft demapper whose piecewise-linear slopes are trained jointly with the VNLE (Sec. II-E); final slope values are not given, so the baseline cannot be rebuilt exactly.
  • Training hyperparameters for gradient descent = not reported
    ADAM is named (Sec. II-B) and validation is shown up to 1000 epochs, but learning rate, batch size, initialization, epochs, and regularization are not listed, so the trained networks cannot be reproduced.
  • Sparsity and pruning targets = VNLE ~25% pruning; SDNNE 20-45% pruning
    Chosen empirically to reduce multipliers while maintaining rate (Sec. IV-C, Fig. 13); affects the equal-complexity comparison points.
assumptions (6)
  • domain assumption All relevant component and channel impairments can be described by a causal or non-causal time-invariant Volterra series with finite fading memory of order up to 5.
    Used in Sec. II-A to justify the VNLE baseline and the kernel comparison; references [16], [18].
  • ad hoc to paper The max-log approximation with optimized slopes incurs virtually no loss at the SNR and OSNR values tested.
    Stated in Sec. II-E; this makes the VNLE baseline comparable to the SDNNE but is not separately validated on this hardware.
  • domain assumption The generalized mutual information in Eq. (12), evaluated at the 15% and 20% FEC code limits, predicts post-FEC performance.
    Sec. III-B converts achievable-rate curves into OSNR gains at FEC limits without running end-to-end FEC decoding; assumes the chosen code limits from [36], [38] are appropriate.
  • ad hoc to paper The activation-pattern upper bound (14) is a valid guide for selecting a sufficiently expressive SDNNE architecture.
    Sec. II-F uses this bound to label optimal designs in Fig. 5 and to choose 17|26|25|3 and 17|16|10|3; the bound concerns linear regions, not equalization loss, so it is a heuristic.
  • domain assumption Channel and component nonlinearities are static enough that one offline training pass per device suffices.
    Sec. IV-A argues aging and temperature drift are slow, so real-time training complexity is excluded from the comparison.
  • ad hoc to paper Expanding the trained tanh SDNNE as a Volterra series at y0=0 yields kernels representative of its behavior at the operating point.
    Sec. II-C and Sec. III-C compare kernels using (7), (8), (11) evaluated at y0=0; this local expansion may not capture performance away from the expansion point.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Soft-Demapping for Short Reach Optical Communication: A Comparison of Deep Neural Networks and Volterra Series." pith.science (2026). https://pith.science/paper/VWEJID6K

@misc{pith2026250105979,
  author       = {Pith},
  title        = {Pith review of: Soft-Demapping for Short Reach Optical Communication: A Comparison of Deep Neural Networks and Volterra Series},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VWEJID6K}},
  note         = {Machine review of arXiv:2501.05979}
}
read the original abstract

In optical fiber communication, optical and electrical components introduce nonlinearities, which require effective compensation to attain highest data rates. In particular, in short reach communication, components are the dominant source of nonlinearities. Volterra series are a popular countermeasure for receiver-side equalization of nonlinear component impairments and their memory effects. However, Volterra equalizer architectures are generally very complex. This article investigates soft deep neural network (DNN) architectures as an alternative for nonlinear equalization and soft-decision demapping. On coherent 92 GBd dual polarization 64QAM back-to-back measurements performance and complexity is experimentally evaluated. The proposed bit-wise soft DNN equalizer (SDNNE) is compared to a 5th order Volterra equalizer at a 15 % overhead forward error correction (FEC) limit. At equal performance, the computational complexity is reduced by 65 %. At equal complexity, the performance is improved by 0.35 dB gain in optical signal-to-noise-ratio (OSNR).

Figures

Figures reproduced from arXiv: 2501.05979 by the authors.

Figure 1
Figure 1. Single artificial neuron. inverse is prone to large numerical errors. This property calls for a trade-off between matrix conditioning and the number of kernels for maximum performance. In order to avoid the problem of ill-conditioned entirely, abovementioned iterative approaches are applied in this paper. In particular, we use gradient descent and consider two objective functions. First, we follow the standard appro… view at source ↗
Figure 3
Figure 3. Channel with a (c) VNLE accompanied by a Soft-decision (SD) [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 2
Figure 2. Channel with a (a) VNLE or (b) DNNE accompanied by a Hard [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figures from the paper (10 more)
Figure 5
Figure 5. Figure 5: Number of activation patterns of a SDNNEs with different memory [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 4
Figure 4. Figure 4: Deep Neural Network Equalizer for separate in phase (I) and [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 6
Figure 6. Figure 6: Back-to-Back offline measurement setup including Tx and Rx DSP architectures with two options a) 4x Volterra Equalizers (VNLE) and b) 4x Soft [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: 92-Gbaud 64QAM performance in terms of achievable rate versus [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 9
Figure 9. Figure 9: The linear kernels, or rather the linear finite impulse responses are [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 10
Figure 10. Figure 10: Comparison of the unrolled extracted third order nonlinear kernels, [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 11
Figure 11. Figure 11: Average achievable rate performance gain in bits/symbol related [PITH_FULL_IMAGE:figures/full_fig_p007_11.png]
Figure 12
Figure 12. Figure 12: Nonlinear compensation gains in bits/symbol related to number of [PITH_FULL_IMAGE:figures/full_fig_p008_12.png]
Figure 14
Figure 14. Figure 14: OSNR nonlinear compensation gains in dB related to number of [PITH_FULL_IMAGE:figures/full_fig_p009_14.png]
Figure 15
Figure 15. Figure 15: Bitwise soft-demapping as logistic regression. [PITH_FULL_IMAGE:figures/full_fig_p010_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 41 canonical work pages

  1. [1]

    Replacing the soft-decision FEC limit paradigm in the design of optical communi- cation systems,

    A. Alvarado, E. Agrell, D. Lavery, R. Maher, and P. Bayvel, “Replacing the soft-decision FEC limit paradigm in the design of optical communi- cation systems,” Journal of Lightwave Technology , vol. 33, no. 20, pp. 4338–4352, 2015

  2. [2]

    Bit-interleaved coded modula- tion,

    G. Caire, G. Taricco, and E. Biglieri, “Bit-interleaved coded modula- tion,” IEEE Trans. Inf. Theory , vol. 44, no. 3, pp. 927–946, 1998

  3. [3]

    Lin and D

    S. Lin and D. J. Costello, Error control coding, 2nd ed. Pearson Prentice Hall, 2004

  4. [4]

    Szczecinski and A

    L. Szczecinski and A. Alvarado, Bit-interleaved coded modulation: fundamentals, analysis and design . John Wiley & Sons, 2015

  5. [5]

    Experimental estima- tion of optical nonlinear memory channel conditional distribution using deep neural networks,

    R. Rios-M ¨uller, J. M. Estar ´an, and J. Renaudier, “Experimental estima- tion of optical nonlinear memory channel conditional distribution using deep neural networks,” in Optical Fiber Communication Conference . Optical Society of America, 2017, pp. W2A–51

  6. [6]

    Nonlinear equalizer based on neural networks for PAM-4 signal transmission using DML,

    A. G. Reza and J.-K. K. Rhee, “Nonlinear equalizer based on neural networks for PAM-4 signal transmission using DML,” IEEE Photonics Technology Letters, vol. 30, no. 15, pp. 1416–1419, 2018

  7. [7]

    100Gbps IM/DD transmission over 25km SSMF using 20G-class DML and PIN enabled by machine learning,

    P. Li, L. Yi, L. Xue, and W. Hu, “100Gbps IM/DD transmission over 25km SSMF using 20G-class DML and PIN enabled by machine learning,” in Optical Fiber Communication Conference. Optical Society of America, 2018, pp. W2A–46

  8. [8]

    Fiber nonlinearity equalization with multi-label deep learning scalable to high- order DP-QAM,

    T. Koike-Akino, D. S. Millar, K. Parsons, and K. Kojima, “Fiber nonlinearity equalization with multi-label deep learning scalable to high- order DP-QAM,” in Signal Processing in Photonic Communications . Optical Society of America, 2018, pp. SpM4G–1

Show all 46 references
  1. [9]

    Evolution from 8QAM live traffic to PS 64-QAM with neural-network based nonlinearity compensa- tion on 11000 km open subsea cable,

    V . Kamalov, L. Jovanovski, V . Vusirikala, S. Zhang, F. Yaman, K. Naka- mura, T. Inoue, E. Mateo, and Y . Inada, “Evolution from 8QAM live traffic to PS 64-QAM with neural-network based nonlinearity compensa- tion on 11000 km open subsea cable,” in Optical Fiber Communication...

  2. [10]

    Field and lab ex- perimental demonstration of nonlinear impairment compensation using neural networks,

    S. Zhang, F. Yaman, K. Nakamura, T. Inoue, V . Kamalov, L. Jovanovski, V . Vusirikala, E. Mateo, Y . Inada, and T. Wang, “Field and lab ex- perimental demonstration of nonlinear impairment compensation using neural networks,” Nature communications, vol. 10, no. 1, pp. 1–8, 2019

  3. [11]

    Equalization performance and complexity analysis of dynamic deep neural networks in long haul transmission systems,

    O. Sidelnikov, A. Redyuk, and S. Sygletos, “Equalization performance and complexity analysis of dynamic deep neural networks in long haul transmission systems,” Optics Express , vol. 26, no. 25, pp. 32 765– 32 776, 2018

  4. [12]

    Neural network-based soft-demapping for nonlinear channels,

    M. Schaedler, S. Calabr `o, F. Pittal `a, C. Bluemm, M. Kuschnerov, and S. Pachnicke, “Neural network-based soft-demapping for nonlinear channels,” in 2020 Optical Fiber Communications Conference and Exhibition (OFC). IEEE, 2020, pp. 1–3. 11

  5. [13]

    Single carrier vs. OFDM for coherent 600Gb/s data centre interconnects with nonlinear equalization,

    C. Bluemm, M. Schaedler, M. Kuschnerov, F. Pittal `a, and C. Xie, “Single carrier vs. OFDM for coherent 600Gb/s data centre interconnects with nonlinear equalization,” in 2019 Optical Fiber Communications Conference and Exhibition (OFC) . IEEE, 2019, pp. 1–3

  6. [14]

    Compensation schemes for transmitter-and receiver-based pattern-dependent distor- tion,

    A. Rezania, J. Cartledge, A. Bakhshali, and W.-Y . Chan, “Compensation schemes for transmitter-and receiver-based pattern-dependent distor- tion,” IEEE Photonics Technology Letters , vol. 28, no. 22, pp. 2641– 2644, 2016

  7. [15]

    V olterra equalization for nonlinearities in optical fiber communications,

    J. C. Cartledge, “V olterra equalization for nonlinearities in optical fiber communications,” in Signal Processing in Photonic Communications . Optical Society of America, 2017, pp. SpTu2F–2

  8. [16]

    Theory of pth-order inverses of nonlinear systems,

    M. Schetzen, “Theory of pth-order inverses of nonlinear systems,” IEEE Transactions on Circuits and Systems, vol. 23, no. 5, pp. 285–291, 1976

  9. [17]

    Neural network modeling of nonlinear systems based on volterra series extension of a linear model,

    D. I. Soloway and J. T. Bialasiewicz, “Neural network modeling of nonlinear systems based on volterra series extension of a linear model,” in Proc. IEEE International Symposium on Intelligent Control (ISIC) , 1992, pp. 7–12

  10. [18]

    Guan, FPGA-based Digital Convolution for Wireless Applications

    L. Guan, FPGA-based Digital Convolution for Wireless Applications . Springer, 2017

  11. [19]

    Complexity reduction of volterra nonlinear equalization for optical short-reach IM/DD systems,

    T. Wettlin, S. Pachnicke, T. Rahman, J. Wei, S. Calabro, and N. Sto- janovic, “Complexity reduction of volterra nonlinear equalization for optical short-reach IM/DD systems,” in Photonic Networks; 21th ITG- Symposium. VDE, 2020, pp. 1–6

  12. [20]

    Zaknich, Principles of adaptive filters and self-learning systems

    A. Zaknich, Principles of adaptive filters and self-learning systems . Leipzig, Germany: Springer Science & Business Media, 2005

  13. [21]

    Nonlinear electrical equalization for dif- ferent modulation formats with optical filtering,

    C. Xia and W. Rosenkranz, “Nonlinear electrical equalization for dif- ferent modulation formats with optical filtering,” Journal of Lightwave Technology, vol. 25, no. 4, pp. 996–1001, 2007

  14. [22]

    Calculating the singular values and pseudo- inverse of a matrix,

    G. Golub and W. Kahan, “Calculating the singular values and pseudo- inverse of a matrix,” Journal of the Society for Industrial and Applied Mathematics, Series B: Numerical Analysis , vol. 2, no. 2, pp. 205–224, 1965

  15. [23]

    F. M. Ghannouchi, O. Hammi, and M. Helaoui, Behavioral modeling and predistortion of wideband wireless transmitters . John Wiley & Sons, 2015

  16. [24]

    Orthogonal polynomials for complex Gaus- sian processes,

    R. Raich and G. T. Zhou, “Orthogonal polynomials for complex Gaus- sian processes,” IEEE transactions on signal processing, vol. 52, no. 10, pp. 2788–2797, 2004

  17. [25]

    C. M. Bishop, Pattern recognition and machine learning . Springer, 2006

  18. [26]

    Theory of the backpropagation neural network,

    R. Hecht-Nielsen, “Theory of the backpropagation neural network,” in Neural networks for perception . Elsevier, 1992, pp. 65–93

  19. [27]

    Goodfellow, Y

    I. Goodfellow, Y . Bengio, and A. Courville, Deep learning. MIT press, 2016

  20. [28]

    Adam: A method for stochastic optimization,

    D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. Int. Conference on Learning Representations , 2015

  21. [29]

    Equalizing nonlinearities with memory effects: V olterra series vs. deep neural networks,

    C. Bluemm, M. Schaedler, S. Calabr `o, G. Charlet, C. Xie, F. Pittal `a, and M. Kuschnerov, “Equalizing nonlinearities with memory effects: V olterra series vs. deep neural networks,” in ECOC. IET, 2019, pp. 1–4

  22. [30]

    Probabilistic shaping and forward error correction for fiber-optic communication systems,

    G. B ¨ocherer, P. Schulte, and F. Steiner, “Probabilistic shaping and forward error correction for fiber-optic communication systems,”Journal of Lightwave Technology, vol. 37, no. 2, pp. 230–244, 2019

  23. [31]

    On the number of linear regions of deep neural networks,

    G. F. Montufar, R. Pascanu, K. Cho, and Y . Bengio, “On the number of linear regions of deep neural networks,” in Advances in neural information processing systems , 2014, pp. 2924–2932

  24. [32]

    Universal approximation bounds for superpositions of a sigmoidal function,

    A. R. Barron, “Universal approximation bounds for superpositions of a sigmoidal function,” IEEE Transactions on Information theory , vol. 39, no. 3, pp. 930–945, 1993

  25. [33]

    On the number of linear regions of deep neural networks,

    G. Mont ´ufar, R. Pascanu, K. Cho, and Y . Bengio, “On the number of linear regions of deep neural networks,” in Proc. Int. Conf. on Neural Information Processing, 2014, pp. 2924–2932

  26. [34]

    Notes on the number of linear regions of deep neural networks,

    G. Mont ´ufar, “Notes on the number of linear regions of deep neural networks,” Int. Conf. on Sampling Theory and Applications (SampTA) , 2017

  27. [35]

    Jia and L

    Z. Jia and L. A. Campos, Eds., Coherent Optics for Access Networks . CRC Press, 2020

  28. [36]

    Open ROADM MSA 3.01 W-Port Digital Specification (200G-400G),

    “Open ROADM MSA 3.01 W-Port Digital Specification (200G-400G),”

  29. [37]

    Implementation of 64QAM at 42.66 GBaud using 1.5 samples per symbol DAC and demonstration of up to 300 km fiber transmission,

    F. Buchali, A. Klekamp, L. Schmalen, and T. Drenski, “Implementation of 64QAM at 42.66 GBaud using 1.5 samples per symbol DAC and demonstration of up to 300 km fiber transmission,” in Optical Fiber Communication Conference . Optical Society of America, 2014, pp. M2A–1

  30. [38]

    Training-aided frequency-domain channel estimation and equalization for single-carrier coherent optical transmission systems,

    F. Pittala, I. Slim, A. Mezghani, and J. A. Nossek, “Training-aided frequency-domain channel estimation and equalization for single-carrier coherent optical transmission systems,” Journal of Lightwave Technol- ogy, vol. 32, no. 24, pp. 4849–4863, 2014

  31. [39]

    Design and implementation of a neural network based predistorter for enhanced mobile broadband,

    C. Tarver, A. Balatsoukas-Stimming, and J. R. Cavallaro, “Design and implementation of a neural network based predistorter for enhanced mobile broadband,” in 2019 IEEE International Workshop on Signal Processing Systems (SiPS) . IEEE, 2019, pp. 296–301

  32. [40]

    The CORDIC algorithm: new results for fast VLSI implementation,

    J. Duprat and J.-M. Muller, “The CORDIC algorithm: new results for fast VLSI implementation,” IEEE Transactions on Computers , vol. 42, no. 2, pp. 168–178, 1993

  33. [41]

    Regression shrinkage and selection via the lasso,

    R. Tibshirani, “Regression shrinkage and selection via the lasso,” Jour- nal of the Royal Statistical Society: Series B (Methodological) , vol. 58, no. 1, pp. 267–288, 1996

  34. [42]

    To prune, or not to prune: exploring the efficacy of pruning for model compression,

    M. Zhu and S. Gupta, “To prune, or not to prune: exploring the efficacy of pruning for model compression,” Int. Conf. on Learning Representations, 2018

  35. [43]

    K. P. Murphy, Machine learning: a probabilistic perspective . MIT press, 2012

  36. [44]

    G ´eron, Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow: Concepts, tools, and techniques to build intelligent systems

    A. G ´eron, Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow: Concepts, tools, and techniques to build intelligent systems. Sevastopol, CA, USA: O’Reilly Media, 2019

  37. [45]

    T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed. Hoboken, NJ, USA: John Wiley & Sons, Inc., 2006

  38. [2019]

    Available: http://openroadm.org/download.html

    [Online]. Available: http://openroadm.org/download.html

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.