{"id":"7874b4ce-2180-4d2e-ac10-9a90424477ab","arxiv_id":"2501.05979","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A bitwise-trained deep neural network equalizer matches a 5th-order Volterra equalizer at 65% fewer multipliers and gains 0.35 dB OSNR at equal complexity in 92 GBd 64QAM back-to-back measurements.","lead":"The paper compares a bitwise-trained neural network soft-demapper with a Volterra equalizer on 92 GBd dual-polarization 64QAM back-to-back optical measurements. It reports equal performance with 65% fewer multipliers, or a 0.35 dB OSNR gain at equal complexity.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central comparison rests on the unverified claim (Sec. II-E) that max-log soft-demapping for the VNLE is virtually lossless; if false, the SDNNE's 0.35 dB / 65% gains partly reflect a weaker baseline.","rationale":"The reader identifies the same weakest assumption I find most load-bearing: the unverified max-log optimality of the VNLE soft-demapper. The strongest claim is an experimental complexity/performance comparison, and that comparison is only meaningful if the VNLE baseline includes a demapper close to the APP-optimal one. The manuscript's own Sec. II-E states the MLA assumption without quantification. I found no internal contradiction in the derivation of the bitwise loss in Appendix A; the proof is sound under the stated convergence and capacity assumptions, but it also depends on the unverified s=1 condition in Eq. (12), so it does not rescue the baseline concern. Other limitations, such as the single measurement campaign, lack of error bars, and architecture selection on the same data, are real but secondary: they affect generalizability rather than specifically explaining the claimed gap. The proposed re-benchmark with an exact APP demapper is the single check that would determine whether the 0.35 dB / 65% result is caused by the equalizer architecture or by the weaker VNLE demapping baseline. I therefore keep the reader's CONDITIONAL verdict.","tokens_in":17475,"tokens_out":10355,"duration_ms":106632,"concrete_test":"Recompute Fig. 14 with the VNLE soft-demapper replaced by a near-exact APP demapper: estimate the conditional density of the equalized VNLE output given each transmitted 64QAM symbol from the held-out frames, compute exact LLRs, and re-optimize VNLE kernels under loss (13). If the 0.35 dB gap at equal multipliers and the 65% equal-performance point persist, the max-log assumption is not load-bearing; if they shrink materially, the central claim is inflated. Also report the minimizing s in Eq. (12) for both the VNLE and SDNNE; s near 1 for both would support calibration.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result compares a VNLE plus a max-log soft-demapper (MLA) against the SDNNE, which is a universal function approximator trained with loss (13). The paper asserts in Sec. II-E that MLA 'incurs virtually no loss in the SNR ranges considered in this work,' but no measurement or bound is provided. If MLA is lossy at 64QAM / 33.8 dB OSNR, the 0.35 dB equal-complexity gain and the 65% equal-performance multiplier reduction are not attributable to the SDNNE equalizer itself; they may partly measure the difference between a piecewise-linear demapper and a learned nonlinear one. The only supporting evidence, a 0.002 bits/QAM gain of bitwise-trained VNLE over MSE-trained VNLE, tests the loss function, not the optimality of MLA. The paper also never reports the minimizing s in Eq. (12) for either scheme; if optimal s differs materially from 1 for the VNLE's LLRs, that is direct evidence of demapper miscalibration. This is a correctness risk in the central comparison, not merely a missing detail.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a soft deep neural network equalizer (SDNNE) that performs joint nonlinear equalization and bitwise soft demapping for short-reach coherent optical systems. The SDNNE is trained with a bitwise equivocation loss (13), which the authors show to be equivalent to binary cross-entropy and to learn the a-posteriori probability ratio when training succeeds. On 92 GBd dual-polarization 64QAM back-to-back measurements, the SDNNE is compared with a 5th-order Volterra nonlinear equalizer (VNLE) followed by a max-log soft demapper with optimized slopes (MLA). The central claims are that at equal achievable rate, the SDNNE requires 65% fewer multipliers, and at equal multiplier count it provides a 0.35 dB OSNR gain at the assumed FEC limits. The paper also presents a Volterra-kernel analysis of the trained SDNNE and a complexity model in terms of hardware multipliers.","tokens_in":17745,"tokens_out":4439,"duration_ms":43249,"significance":"If the experimental results hold, the paper offers a practical complexity reduction for nonlinear equalization in coherent short-reach links, backed by explicit complexity formulas and a well-defined training objective. The equivalence of the bitwise equivocation loss to binary cross-entropy and its information-theoretic interpretation are useful theoretical contributions. The main claims are falsifiable and the experimental methodology is largely reproducible from the described setup. However, the central quantitative comparison depends on an unverified assumption about the max-log demapper, and the reported gains are not accompanied by confidence intervals or a separate architecture-selection campaign, which tempers the strength of the conclusions.","major_comments":[{"comment":"The assertion that the max-log approximation (MLA) 'incurs virtually no loss in the SNR ranges considered in this work' is unsupported by any measurement, simulation, or bound in the manuscript. This is load-bearing because both headline results—the 65% multiplier reduction and the 0.35 dB OSNR gain—compare the SDNNE against a VNLE followed by MLA. If the MLA is lossy relative to the exact APP demapper, the reported gains partly reflect a weaker demapping baseline rather than a superior equalizer. The only supporting evidence cited (a 0.002 bits/QAM improvement from bitwise training over MSE training) validates the loss function, not the optimality of MLA. The authors should report the minimizing s in Eq. (12) for both the VNLE+MLA and the SDNNE at the tested OSNR values; if the optimal s for the VNLE differs materially from 1, that is direct evidence of demapper miscalibration. Alternatively, a simulation or measurement comparison against a VNLE with exact APP demapping would resolve the issue.","section":"Sec. II-E, Eq. (12)-(13)"},{"comment":"The quantitative claims (0.35 dB OSNR gain, 65% multiplier reduction) are based on the achievable-rate curves in Figs. 7 and 14, but no confidence intervals or repeated-capture statistics are provided. The SDNNE architectures (17|26|25|3 and subsequently 17|16|10|3) and the VNLE memory configurations are selected on the same measurement campaign, and the reported gains are the best among the evaluated configurations. This creates a risk of selection bias: the headline numbers may overstate the expected gain on new captures. The authors should provide error bars or repeated independent measurements, and ideally perform architecture selection on a separate training campaign or use a nested validation procedure.","section":"Sec. III-B and IV-C"},{"comment":"The Volterra-kernel comparison between the SDNNE and the VNLE is performed by Taylor expansion of the trained tanh SDNNE at y0 = 0 (Eqs. (7)-(11)). The input signal is power-normalized but not zero-mean at the operating point, so the kernels at y0 = 0 may not be representative of the network's local behavior at typical input values. This does not invalidate the central complexity/performance comparison, but it undermines the interpretability claim in Sec. III-C that the extracted kernels reflect the equalizer's actual operation. The authors should justify the choice y0 = 0 or evaluate the expansion at the mean of the received signal.","section":"Sec. II-C and III-C"}],"minor_comments":[{"comment":"The VNLE 4th-order configuration is labeled '17:17:11:3' in Fig. 7 but '17:17:11:5' in Fig. 8; the authors should reconcile this discrepancy.","section":"Fig. 7 vs Fig. 8"},{"comment":"The derivation switches from natural logarithm in Eq. (20) to base-2 logarithm in Eq. (21) without comment; the equivalence holds only up to a constant scaling factor, which is immaterial for optimization but should be stated explicitly.","section":"Appendix A, Eqs. (20)-(21)"},{"comment":"Reference [36] is cited in Fig. 7 as the source of the 15% OH FEC limit, while the text cites [35, Table 9.1]; also, reference [38] is cited in Fig. 7 for the 20% OH limit, but [38] is a channel-estimation paper, not a code-limit reference. Please verify the citations.","section":"References"},{"comment":"There are several typographical issues, e.g., 'V olterra' in the title and headers, 'multiplers' in the Fig. 12 caption, and inconsistent use of 'optNtaps' versus 'optStruc' in figure labels. A thorough proofread is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a preprint of an older manuscript (submitted October 2020, accepted January 2021) posted to arXiv in 2025. The authors should confirm with the journal that this is an original submission and not a duplicate publication. The central concern is the unverified max-log assumption; if the authors can provide the requested s-value analysis or an exact-APP comparison, the paper would be considerably stronger."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a workmanlike experimental comparison, and the 65% multiplier reduction and 0.35 dB OSNR gain are plausible as system-level claims. The weakest link is the unsupported \"max-log is virtually lossless\" assertion; if that fails, part of the reported gain is just better demapping, not better equalization.\n\nWhat is new: a soft DNN equalizer trained with bitwise equivocation, compared against a jointly optimized Volterra equalizer plus max-log soft demapper on real 92 GBd DP-64QAM back-to-back captures. The complexity model counts multipliers explicitly for both sides, and the kernel-extraction diagnostic is a nice way to see what the DNN actually learned. The activation-function comparison (I-tanh vs H-tanh vs ReLU) is practically useful. The appendix showing equivalence to logistic regression is textbook, but correctly done, so no harm.\n\nThe central experimental comparison is on measured data and both systems are trained with the same bitwise loss, so there is no obvious circularity. The numbers in Figs. 7, 12, and 14 support the stated trends.\n\nSoft spots, in order of importance:\n1. The paper never validates the MLA \"virtually no loss\" claim. For 64QAM at these SNRs max-log is usually close but not exact, and the only evidence given (0.002 bits/QAM gain from bitwise vs MSE training) speaks to the loss function, not the demapper gap. The stress-test note is on target. I would not call it fatal: the paper's practical claim is that one learned system can replace VNLE+MLA, and that can still be true if MLA is imperfect. But the 65% and 0.35 dB numbers should be read as system-level, not equalizer-only.\n2. No error bars or shared data/code. The architecture search and final evaluation use the same measurement campaign, so overfitting to the test set is possible.\n3. The s=1 check described in Sec. II-E is never reported. That is a cheap, useful diagnostic and its absence is a real omission.\n4. The reference list stops in 2020; the header shows acceptance in 2021, but the 2025 posting makes the literature comparison stale.\n\nWho this is for: applied optical DSP researchers comparing ML equalizers against Volterra baselines. They will get a useful benchmark and a clear complexity framework. The theory reader will not learn anything new.\n\nRecommendation: send it to peer review, yes. A good referee should push for the MLA validation, error bars, and the s=1 check, but the paper deserves the scrutiny.","headline":"A workmanlike experimental comparison whose headline numbers are plausible as system-level claims but rest on an unverified max-log assumption that should be checked before citing them.","tokens_in":18299,"tokens_out":3249,"would_cite":true,"duration_ms":33331,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A small neural network can replace a 5th-order Volterra equalizer for short-reach coherent links, matching its achievable rate with 65% fewer multipliers and gaining 0.35 dB OSNR at equal complexity.","keywords":["soft-decision demapping","Volterra series","deep neural network equalizer","short-reach optical communication","coherent 64QAM","bitwise equivocation loss","achievable rate","nonlinear equalization"],"falsifier":"Replace the max-log soft demapper after the Volterra equalizer with an exact a-posteriori-probability demapper at the same complexity, or measure the achievable-rate gap between max-log and exact demapping; if the gap is comparable to 0.35 dB, the claimed equal-complexity OSNR gain would shrink or vanish.","tokens_in":17197,"feed_emoji":"📡","tokens_out":4452,"duration_ms":38868,"temperature":0.7,"pith_summary":"The paper tries to show that a compact deep neural network that outputs soft bit values directly can replace a 5th-order Volterra nonlinear equalizer followed by a max-log soft demapper in short-reach coherent optical systems. On back-to-back 92 GBd dual-polarization 64QAM measurements, the neural equalizer matches the Volterra's achievable rate with 65% fewer multipliers and improves OSNR by 0.35 dB at equal multiplier count. The key enabler is a bitwise loss function that is equivalent to binary cross-entropy and that, when training succeeds, makes the network output the logarithmic a-posteriori probability ratio, which maximizes the generalized mutual information. If correct, this gives receiver designers a concrete complexity budget for nonlinear compensation and a principled way to train soft demappers.","feed_headline":"Neural net equalizer matches Volterra at 65% lower multiplier cost","feed_subtitle":"92 GBd 64QAM back-to-back test: same achievable rate with 136 multipliers vs 385, plus 0.35 dB at equal cost.","key_machinery":"The central object is the bitwise equivocation loss (13), equivalent to binary cross-entropy, used to train both the neural equalizer and the Volterra equalizer plus demapper. The paper proves this loss is optimal: a trained network using it learns the APP ratio log(P_{B|Y}(0|y)/P_{B|Y}(1|y)). The comparison is carried by a Taylor expansion (gradients and Hessians, (7)-(11)) that converts the trained neural network into Volterra kernels, letting the paper interpret which nonlinear orders matter and showing the neural network can adjust kernels per soft output while the Volterra cannot.","core_discovery":"On coherent 92 GBd dual-polarization 64QAM back-to-back measurements where component nonlinearities dominate, a fully connected neural network trained with the bitwise equivocation loss achieves the same achievable rate as a jointly optimized 5th-order Volterra equalizer with a max-log soft demapper while requiring 65% fewer multipliers; at equal multiplier count it gains 0.35 dB in OSNR at the 15%-overhead FEC limit. The paper also proves that this loss is equivalent to binary cross-entropy and trains the network to output the logarithmic a-posteriori probability ratio, i.e., the optimal soft bit, minimizing the conditional entropy and thereby maximizing the achievable rate (12).","pith_inferences":["If the max-log demapper assumption is relaxed at lower OSNR, the 0.35 dB claim would need re-benchmarking; one would expect the gap to shrink when the Volterra side uses an exact APP demapper.","The dominance of odd-order kernels suggests a memoryless or weakly memory odd-power nonlinearity model might explain most of the component distortion; that hypothesis is testable by fitting a static nonlinearity and comparing residuals.","The same bitwise-loss training could be applied to other equalizer structures, including Volterra variants, so the complexity comparison is between architectures, not between loss functions."],"forward_implications":["At a fixed FEC overhead, the SDNNE with hard-tanh activations reaches Volterra-level nonlinear compensation at 136 multipliers versus 385 for the sparse VNLE, pointing to a low-complexity regime where the neural architecture degrades more gracefully.","Because the bitwise equivocation loss equals binary cross-entropy, standard supervised-learning pipelines can be used directly to train rate-maximizing soft demappers for coherent receivers.","The observation that 2nd, 3rd, and 5th order kernels dominate, while 4th, 6th, and 7th add no gain, identifies odd-order component nonlinearities as the target any equalizer must capture in this setup.","Training can be validated by checking that the minimizing s in the achievable-rate expression (12) equals 1 for the network's outputs, giving a practical stopping rule."],"supporting_citations":[{"why":"Establishes the one-to-one replacement of a Volterra nonlinear equalizer by a deep neural network equalizer trained on MSE, which this paper extends to soft outputs.","marker":"[29]"},{"why":"Provides the achievable-rate (GMI) expression (12) used both as the evaluation metric and as the basis for the bitwise equivocation loss.","marker":"[30]"},{"why":"Supplies the Taylor-expansion method for converting a trained neural network into Volterra kernels, enabling the kernel-level comparison.","marker":"[17]"},{"why":"Defines the multiplier-count model for Volterra equalizer complexity that the paper adapts in (15).","marker":"[39]"},{"why":"Provides the gradual pruning algorithm used to generate the sparse VNLE and SDNNE complexity curves.","marker":"[42]"},{"why":"Documents the max-log approximation (MLA) that the Volterra soft demapper uses, whose negligible loss is a load-bearing assumption.","marker":"[4]"}],"fun_headline_variants":["Neural net equalizer beats Volterra at 65% lower cost","DNN equalizer: same rate, 65% fewer multipliers","Soft DNN demapper cuts equalizer complexity by 65%","Neural net equalizer matches Volterra with 65% less compute","Deep learning equalizer: 65% cheaper, 0.35 dB better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison presumes that the max-log approximation used for the Volterra soft demapper loses essentially nothing in the signal-to-noise range tested, so the neural network's gain is measured against a near-ideal Volterra demapper.","fun_headline_variants_meta":{"raw":{"variants":["Neural net equalizer beats Volterra at 65% lower cost","DNN equalizer: same rate, 65% fewer multipliers","Soft DNN demapper cuts equalizer complexity by 65%","Neural net equalizer matches Volterra with 65% less compute","Deep learning equalizer: 65% cheaper, 0.35 dB better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00055,"raw_usage":{"total_tokens":2593,"prompt_tokens":879,"completion_tokens":1714,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1618}},"tokens_in":495,"tokens_out":1714,"duration_ms":11845,"temperature":1.0,"reasoning_tokens":1618,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:07:17.843725+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the max-log soft demapper after the Volterra equalizer with an exact a-posteriori-probability demapper at the same complexity, or measure the achievable-rate gap between max-log and exact demapping; if the gap is comparable to 0.35 dB, the claimed equal-complexity OSNR gain would shrink or vanish.","supporting_citations":[{"cited_title":"Equalizing nonlinearities with memory effects: V olterra series vs. deep neural networks,","cited_arxiv_id":null,"evidence_quote":"Establishes the one-to-one replacement of a Volterra nonlinear equalizer by a deep neural network equalizer trained on MSE, which this paper extends to soft outputs."},{"cited_title":"Probabilistic shaping and forward error correction for fiber-optic communication systems,","cited_arxiv_id":null,"evidence_quote":"Provides the achievable-rate (GMI) expression (12) used both as the evaluation metric and as the basis for the bitwise equivocation loss."},{"cited_title":"Neural network modeling of nonlinear systems based on volterra series extension of a linear model,","cited_arxiv_id":null,"evidence_quote":"Supplies the Taylor-expansion method for converting a trained neural network into Volterra kernels, enabling the kernel-level comparison."},{"cited_title":"Design and implementation of a neural network based predistorter for enhanced mobile broadband,","cited_arxiv_id":null,"evidence_quote":"Defines the multiplier-count model for Volterra equalizer complexity that the paper adapts in (15)."},{"cited_title":"To prune, or not to prune: exploring the efficacy of pruning for model compression,","cited_arxiv_id":null,"evidence_quote":"Provides the gradual pruning algorithm used to generate the sparse VNLE and SDNNE complexity curves."}],"review_version":1}