REVIEW 4 major objections 4 minor 30 references
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read RADE, a neural codec for HF radio, sends intelligible speech at SNRs 4–13 dB lower than analog SSB.
desk verdict A real engineering advance with code and field tests, but the headline dB gains are measured with Whisper WER, not human ears, so the claimed size of the advantage is not yet established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a joint source–channel autoencoder that passes information through continuously valued QAM symbols. Encoder stacks of 1D convolutions and gated recurrent units map 20-dimensional vocoder features $f$ (18 Bark-scale cepstral coefficients, pitch period, voicing) into an 80-dimensional latent vector $z$ every 40 ms; the decoder is a symmetric stack with gated linear units that returns four feature frames. OFDM places three consecutive latent vectors into one 120 ms frame with pilot symbols for phase and magnitude equalisation. Training applies a $c_{\tanh}(x)=\tanh(|x|)e^{j\arg x}$ bottleneck to the time-domain signal to enforce the peak-power limit, and multiplies each frequency-domain symbol by a real fading magnitude $h_c=|H(e^{j\omega_c})|$ derived from the two-path Watterson model, so the autoencoder learns to compensate for both amplifier saturation and multipath notches in a single optimisation.
What would settle it
A direct test would be a double-blind listening test in which naive transcribers hear RADE and SSB audio at matched SNRs from the same simulated or over-the-air channels. If human word error rates do not reproduce roughly 4 dB at link closure and 13 dB at good quality, the ASR-based claim is overstated. A second check is hardware measurement of the RADE waveform's actual PAPR and adjacent-channel spectrum, which would confirm or refute the assumed 7 dB peak-power advantage over SSB.
Extended reading notes
Core claim
The discovery is that a voice link with no intermediate bitstream can be trained end to end for the HF channel. The autoencoder's latent vector $z$ is placed directly onto complex QAM symbols $q$; the receiver's equalised symbols $\hat{q}$ are mapped straight back to vocoder features $\hat{f}$ for synthesis. Because of this, the network can learn a nonlinear mapping that spreads speech information across symbols in a way that survives additive noise and frequency-selective fading. The training signal includes a $c_{\tanh}$ magnitude bottleneck in the time domain, which drives the waveform's peak-to-average power ratio below 1 dB, and the multipath channel is simulated by a two-path Watterson model applied in the frequency domain. The result, measured by ASR word error rate, is speech that clearly outperforms analog SSB at equal receiver SNR, with a graceful quality–SNR trade-off instead of a digital cliff.
Load-bearing premise
The advantage over SSB is measured by the word error rate of a machine speech recognizer, and the paper's numbers assume that recognizer tracks what a human listener would understand; informal listening and a field demonstration are the only human evidence.
Editorial extensions
If this is right
- At equal receiver SNR, RADE gives lower word error than SSB, so the same transmitter could close a voice link at roughly 4 dB lower SNR and reach good quality 13 dB lower.
- Because the waveform has below 1 dB PAPR, a peak-power-limited transmitter can run RADE at up to about 7 dB higher mean power than a typical SSB signal, extending range.
- Unlike FEC-based digital voice systems that stop working below a threshold, RADE degrades gradually, so partial communication remains possible as SNR falls.
- The 120 ms one-way delay is compatible with push-to-talk radio, and the system works with practical acquisition and equalisation stages built from classical DSP.
- The system also shows qualitative robustness to channel impairments that were not part of the training set, such as impulse noise.
Reading between the lines
- If the ASR result holds up in human listening tests, RADE could make HF voice usable by non-specialist operators at much lower SNR, effectively increasing the coverage area of an existing transmitter without raising peak power.
- The ctanh-bottleneck training trick decouples PAPR control from multipath equalisation, an approach that could be transferred to conventional OFDM systems to reduce PAPR without explicit peak-reduction algorithms.
- Because the system has no bitstream, the same architecture might be retrained for low-rate data or telemetry over HF as easily as for speech, with the QAM-symbol representation acting as a learned physical-layer code.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents RADE, a neural autoencoder that transmits speech over HF radio by mapping vocoder features directly to continuous QAM symbols carried on OFDM subcarriers. The encoder/decoder pair is trained end-to-end with a ctanh power-amplifier saturation bottleneck and a frequency-domain Watterson multipath model, replacing the classical separation of source coding, channel coding, and modulation. The authors report that RADE matches or exceeds analog SSB and the digital FreeDV 700D system in Whisper-based word error rate (WER) over simulated AWGN and multipath-poor (MPP) channels, with claimed gains of 4 dB at the 30% WER link-closure threshold and 13 dB at the 5% WER good-quality threshold. They also provide an over-the-air demonstration using amateur radio transmitters and KiwiSDR receivers, and release source code and audio samples.
Significance. If the reported gains hold, RADE would be a practically important advance for HF voice: it replaces the classical modulation/FEC chain with a learned analog joint source-channel code, achieves graceful degradation instead of a digital cliff, and operates at low PAPR with only 120 ms algorithmic delay. The paper is commendably concrete: it evaluates on held-out Librispeech utterances, includes a clean-feature FARGAN control, provides source code and audio samples, and reports complexity figures. The central risk is that the dB-level claims rest entirely on an unvalidated ASR proxy for intelligibility and on a PAPR advantage that is never measured on the final transmitted waveform; the over-the-air test is anecdotal. These issues are fixable within the manuscript's scope, so the work is promising but needs revision.
major comments (4)
- [Section 4, Fig. 6] The central quantitative claim—4 dB improvement at 30% WER and 13 dB at 5% WER—rests entirely on Whisper WER as a proxy for intelligibility, and this proxy is not validated for the two very different distortion types being compared. SSB outputs at low SNR are band-limited, Hilbert-compressed, and corrupted by colored and multipath noise, while RADE outputs are synthesized by FARGAN from decoded features; Whisper is trained on natural speech and may be systematically more tolerant of one distortion than the other. The clean-feature FARGAN control shows only that the vocoder is transparent to Whisper, not to human listeners. I request either a listening study (e.g., DRT/MRT or a standardized subjective intelligibility test) or a clearly argued calibration of WER to human intelligibility for both systems; without this, the "clearly surpasses" claim is not supported.
- [Section 4, Fig. 6] The WER curves are single-run evaluations on 500 Librispeech samples, and no confidence intervals, error bars, or significance tests are reported. The 4 dB and 13 dB margins could be within sampling noise, especially near the 5% WER threshold where a few misrecognized utterances can shift the operating point. Please report bootstrap intervals over utterances or multiple draws, and state precisely how the SNR on the horizontal axis (labeled "SNR3k") is defined and how it relates to Eq/N0 used in training.
- [Section 3 and Section 4] The paper claims a PAPR of less than 1 dB and then uses this to argue for an additional 7 dB peak-power advantage over SSB, but the PAPR of the complete transmitted waveform—after pilot insertion, cyclic prefix, OFDM framing, and the ctanh bottleneck—is never measured. The <1 dB figure appears to come from the training-time signal, not from the final waveform that is actually transmitted. Please report the measured PAPR (e.g., complementary CDF at 0.01% probability) of the final RADE waveform and of the SSB compressor used in the comparison; the 7 dB advantage is load-bearing for the practical SNR comparison.
- [Section 3, Eqs. (4)-(5)] Training applies the multipath channel as frequency-domain magnitude-only fading and explicitly assumes that phase equalization and ISI removal are performed by the classical DSP receiver, while the equalization penalty is separately estimated at 2 dB in Section 2. It is not shown that the trained encoder/decoder remains robust to the residual phase error and inter-carrier/inter-symbol interference that the LS equalizer actually leaves at low SNR. Please provide an end-to-end evaluation that includes the full receiver chain and quantifies the residual equalization error at the operating SNR, or state explicitly that the Section 4 simulations already include these effects and show the comparison.
minor comments (4)
- [Section 3] There is a typo in "minimse" (should be "minimise"), and the signal-to-noise notation Eq/N0 appears as "E q/N0" in several places; use consistent math formatting.
- [Figure 6] The axis label "SNR3k" is not defined in the text; please define the 3000 Hz noise bandwidth reference and state how it is computed for the SSB and RADE signals.
- [Section 4.2] The over-the-air demonstration is a valuable existence proof but is described as informal; please label it as anecdotal and avoid relying on it for the quantitative dB claims, or provide scored WER or listening-test results from the recorded over-the-air samples.
- [Section 5] The claim that RADE shows "robustness to channel impairments we did not train for, e.g. impulse noise" is not supported by any presented experiment; either add data or remove the claim.
Circularity Check
No circularity: RADE's central dB advantage is measured against SSB on held-out LibriSpeech via Whisper ASR, independent of the training objective.
full rationale
The central quantitative claim (4 dB gain at 30% WER and 13 dB at 5% WER over SSB, for both AWGN and MPP channels) is obtained by passing 500 held-out LibriSpeech samples through RADE, SSB, and FreeDV 700D simulations and measuring Whisper word error rate. The paper explicitly states that the LibriSpeech speech and Watterson model data used for evaluation were not part of the training set, so the reported improvement is an external, held-out measurement rather than a fitted prediction. The training objective is reconstruction loss L(f, f_hat) on vocoder features, while the evaluation metric is an external ASR word error rate; these are not the same quantity and the result is not forced by construction. The architecture is derived from the authors' earlier DRED and FARGAN work, but the load-bearing radio-performance claim does not reduce to those citations: the FARGAN-clean condition is an empirical control shown in the paper as evidence that the chosen classical features are not the limiting factor, and the PAPR claim follows from the ctanh bottleneck explicitly included in the training loop. The reliance on Whisper ASR as a proxy for human intelligibility, and the absence of formal listening tests, are threats to external validity rather than circularity under the review rules. No step in the derivation chain equates the input with the output, and no fitted parameter is renamed as a prediction.
Assumptions & free parameters
assumptions (4)
- domain assumption The Watterson two-path channel model (Eq. 4) is representative of real HF multipath propagation.
- domain assumption Phase equalisation and ISI removal are assumed perfect during training (Section 3).
- ad hoc to paper The ctanh nonlinearity (Eq. 3) models power amplifier saturation.
- domain assumption Whisper ASR word error rate is a valid proxy for human intelligibility over HF channels.
Cite this review
Pith. "Pith review of RADE: A Neural Codec for Transmitting Speech over HF Radio Channels." pith.science (2026). https://pith.science/paper/WN5R3WAN
@misc{pith2026250506671,
author = {Pith},
title = {Pith review of: RADE: A Neural Codec for Transmitting Speech over HF Radio Channels},
year = {2026},
howpublished = {\url{https://pith.science/paper/WN5R3WAN}},
note = {Machine review of arXiv:2505.06671}
}
read the original abstract
Speech compression is commonly used to send voice over radio channels in applications such as mobile telephony and two-way push-to-talk (PTT) radio. In classical systems, the speech codec is combined with forward error correction, modulation and radio hardware. In this paper we describe an autoencoder that replaces many of the traditional signal processing elements with a neural network. The encoder takes a vocoder feature set (short term spectrum, pitch, voicing), and produces discrete time, but continuously valued quadrature amplitude modulation (QAM) symbols. We use orthogonal frequency domain multiplexing (OFDM) to send and receive these symbols over high frequency (HF) radio channels. The decoder converts received QAM symbols to vocoder features suitable for synthesis. The autoencoder has been trained to be robust to additive Gaussian noise and multipath channel impairments while simultaneously maintaining a Peak To Average Power Ratio (PAPR) of less than 1 dB. Over simulated and real world HF radio channels we have achieved output speech intelligibility that clearly surpasses existing analog and digital radio systems over a range of SNRs.
Reference graph
Works this paper leans on
-
[1]
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
INTRODUCTION High-frequency (HF) push-to-talk (PTT) radio has the benefit of operating without infrastructure over ranges of several thousands of km. Applications include humanitarian, remote area, emergency and government communication when access to cellular and satellite systems cannot be guaranteed. The HF radio signal propagates from the transmitter ...
work page Pith review arXiv 1915
-
[2]
Unlike classical approaches there is no intermediate bit stream, and the QAM symbols (Fig
An autoencoder that combines the classical DSP functions of quantisation, channel coding, and modulation to generate discrete time but continuously valued (analog) QAM symbols directly from vocoder features. Unlike classical approaches there is no intermediate bit stream, and the QAM symbols (Fig. 3) emerge from the training process rather than being memb...
-
[3]
A training procedure that minimises the end-to-end distortion of vocoder features in the presence of additive Gaussian noise and frequency-selective fading, while simultaneously generating an OFDM waveform with low Peak To Average Power Ratio (PAPR). In Section 2 we describe how we have combined a neural encoder and decoder with OFDM to develop the RADE s...
-
[4]
RADE DESIGN Rather than operating directly on the time-domain signal like recent end-to-end neural codecs [10], RADE uses classical acoustic features similar to those used in LPCNet [11]. We use 20-dimensional feature vectors f that consist of 18 Bark-scale cepstral coefficients, the pitch period, and a voicing parameter. Using classical features not only...
-
[5]
OFDM performs equalisation using a single complex multiply of each symbol which allows us to efficiently represent the multipath channel fading in the ML frame-rate processing
-
[6]
The Fs = 8000 Hz sample rate processing is performed efficiently in classical DSP, with the ML processing at a much slower frame rate Rf = 25 Hz
-
[7]
Alternatively, these tasks can also be accomplished using ML [13] [14]
We perform acquisition, synchronisation, and sample rate conversion using well-known DSP techniques. Alternatively, these tasks can also be accomplished using ML [13] [14]. A disadvantage of OFDM is high peak to average power ratio (PAPR), which for a given transmitter peak power reduces the available SNR at the receiver. We have largely overcome this iss...
-
[8]
5 illustrates the configuration used for training
TRAINING Fig. 5 illustrates the configuration used for training. We use a mixed sample rate design, with most of the signal processing occurring at the subcarrier rate Rs, and selected portions at the sample rate Fs. Only the RADE encoder and decoder have trainable parameters. The bottleneck is defined as: ctanh(x) = tanh(|x|)ej arg[x] (3) which simulates...
Show all 30 references
-
[9]
Since intelligibility – more than quality – is the primary goal for HF radio, we use Automatic Speech Recognition (ASR) to evaluate the performance of the proposed system
EV ALUA TION AND RESULTS Informal listening tests on simulated and over the air samples suggest RADE significantly outperforms SSB. Since intelligibility – more than quality – is the primary goal for HF radio, we use Automatic Speech Recognition (ASR) to evaluate the performan...
-
[10]
It is robust to AWGN and multipath channel impairments, and the transmit signal has a PAPR of less than 1 dB
CONCLUSION We have combined a ML vocoder, ML autoencoder and classical DSP OFDM to build a system capable of sending speech over HF radio channels. It is robust to AWGN and multipath channel impairments, and the transmit signal has a PAPR of less than 1 dB. Unusually for HF sp...
2000
-
[11]
Early history of single-sideband transmission,
A. A. Oswald, “Early history of single-sideband transmission,” Proceedings of the IRE , vol. 44, no. 12, pp. 1676–1679, 1956
1956
-
[12]
An introduction to single- sideband communications,
J. F. Honey and D. K. Weaver, “An introduction to single- sideband communications,” Proceedings of the IRE , vol. 44, no. 12, pp. 1667–1675, 1956
1956
-
[13]
Project 25-DataOverview- NewTechStandards,
T. I. Association et al. , “Project 25-DataOverview- NewTechStandards,” ANSI/TIA-102.BAEA-A
-
[14]
An introduction to deep learning for the physical layer,
T. J. O’Shea and J. Hoydis, “An introduction to deep learning for the physical layer,” 2017. [Online]. Available: https://arxiv.org/abs/1702.00832
2017 arXiv
-
[15]
DRED: Deep REDundancy coding of speech using a rate-distortion- optimized variational autoencoder,
J.-M. Valin, J. B ¨uthe, A. Mustafa, and M. Klingbeil, “DRED: Deep REDundancy coding of speech using a rate-distortion- optimized variational autoencoder,” IEEE Journal of Selected Topics in Signal Processing , 2024
2024
-
[16]
A hybrid deep-learning approach for single channel HF-SSB speech enhancement,
Y . Chen, B. Dong, X. Zhang, P. Gao, and S. Li, “A hybrid deep-learning approach for single channel HF-SSB speech enhancement,” IEEE Wireless Communications Letters , vol. 10, no. 10, pp. 2165–2169, 2021
2021
-
[17]
Low-latency deep analog speech transmission using joint source channel coding,
M. Bokaei, J. Jensen, S. Doclo, and J. Østergaard, “Low-latency deep analog speech transmission using joint source channel coding,” IEEE Journal of Selected Topics in Signal Processing , 2025
2025
-
[18]
OFDM-guided deep joint source channel coding for wireless multipath fading channels,
M. Yang, C. Bian, and H.-S. Kim, “OFDM-guided deep joint source channel coding for wireless multipath fading channels,” IEEE Transactions on Cognitive Communications and Network- ing, vol. 8, no. 2, pp. 584–599, 2022
2022
-
[19]
Very low complexity speech synthesis using framewise autoregressive GAN (FAR- GAN) with pitch prediction,
J.-M. Valin, A. Mustafa, and J. B ¨uthe, “Very low complexity speech synthesis using framewise autoregressive GAN (FAR- GAN) with pitch prediction,” 2024
2024
-
[20]
SoundStream: An end-to-end neural audio codec,
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 30, pp. 495–507, 2021
2021
-
[21]
LPCNet: Improving neural speech synthesis through linear prediction,
J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. International Con- ference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5891–5895
2019
-
[22]
Densely connected convolutional networks,
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 4700–4708
2017
-
[23]
Deep learning based communication over the air,
S. D ¨orner, S. Cammerer, J. Hoydis, and S. Ten Brink, “Deep learning based communication over the air,” IEEE Journal of Selected Topics in Signal Processing , vol. 12, no. 1, pp. 132–143, 2017
2017
-
[24]
OFDM-autoencoder for end-to-end learning of communications systems,
A. Felix, S. Cammerer, S. D ¨orner, J. Hoydis, and S. Ten Brink, “OFDM-autoencoder for end-to-end learning of communications systems,” in 2018 IEEE 19th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC) . IEEE, 2018, pp. 1–5
2018
-
[25]
Low papr waveform design for OFDM systems based on convolutional au- toencoder,
Y . Huleihel, E. Ben-Dror, and H. H. Permuter, “Low papr waveform design for OFDM systems based on convolutional au- toencoder,” in 2020 IEEE International Conference on Advanced Networks and Telecommunications Systems (ANTS) . IEEE, 2020, pp. 1–6
2020
-
[26]
ITU-R F.1487: Testing of HF modems with bandwidths of up to about 12 kHz using ionospheric channel simulators,
“ITU-R F.1487: Testing of HF modems with bandwidths of up to about 12 kHz using ionospheric channel simulators,” 2000
2000
-
[27]
Robust speech recognition via large- scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large- scale weak supervision,” 2022. [Online]. Available: https: //arxiv.org/abs/2212.04356
2022 arXiv
-
[28]
Codec 2 Algorithm Description,
D. Rowe, “Codec 2 Algorithm Description,” https://rowetel.com/ downloads/codec2 doc.html
-
[29]
W ASPAA 2025 RADE Demonstration Speech Samples,
——, “W ASPAA 2025 RADE Demonstration Speech Samples,” https://freedv.org/waspaa-2025-rade-demonstration-speech- samples/
2025
-
[30]
RADAE (RADE) GitHub Repository,
——, “RADAE (RADE) GitHub Repository,” https://github.com/ drowe67/radae
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.