Pith. sign in

REVIEW 1 major objections 1 minor 61 references

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

T0 review · 1 major / 1 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read QC-GAN uses quaternion conformers and Hamilton products to match state-of-the-art speech enhancement quality with under one million parameters.

desk verdict QC-GAN gets solid PESQ scores at very low parameter counts by swapping in a quaternion Conformer, but the abstract leaves the efficiency mechanism untested. read the letter →

arxiv 2606.18611 v2 pith:3BDP6FVF submitted 2026-06-17 cs.SD cs.AIcs.LGstat.ML

classification cs.SDcs.AIcs.LGstat.ML
keywords speechenhancementquaternionneuralnetworksgenerativeadversarialnetworkparameterefficiencyconformerHamiltonproductperceptualquality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents QC-GAN as a speech enhancement system built around a Quaternion Conformer generator trained with MetricGAN. The core mechanism is the Hamilton product, which applies structured weight sharing across real and imaginary components to capture magnitude-phase relationships while cutting the total parameter count. On the VoiceBank+DEMAND dataset this yields a PESQ of 3.48 using only 0.89 million parameters, matching larger conventional models, and a 35-thousand-parameter version still reaches 3.23. The same model also performs well on the DNS-Challenge 3 set, showing the efficiency carries over to more realistic conditions.

What carries the argument

Quaternion Conformer generator that uses the Hamilton product to perform structured weight sharing across magnitude and phase components.

What would settle it

Train an otherwise identical real-valued Conformer GAN limited to exactly 0.89 million parameters and measure whether its PESQ on VoiceBank+DEMAND falls measurably below 3.48.

Watch

Extended reading notes

Core claim

QC-GAN combines a Quaternion Conformer generator with MetricGAN-based training. The Hamilton product encodes magnitude and phase via structured weight sharing, reducing layer parameters while preserving their interdependencies. On VoiceBank+DEMAND this produces a PESQ score of 3.48 at 0.89 million parameters and 3.23 at 35 thousand parameters, with further confirmation of generalization on DNS-Challenge 3.

Load-bearing premise

The Hamilton product successfully encodes the necessary magnitude-phase interdependencies in speech signals without degrading perceptual quality.

Editorial extensions

If this is right

  • Speech enhancement models become feasible on embedded or mobile hardware due to the halved parameter budget.
  • Metric-optimized discriminators can drive perceptual scores even when the generator operates in a reduced quaternion space.
  • The same architecture scales down to tens of thousands of parameters while still exceeding many classical enhancement baselines.
  • Performance holds on both simulated and real-world noise recordings from the DNS-Challenge 3 corpus.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Quaternion weight sharing could be tested in related audio tasks such as source separation or bandwidth extension to check for similar parameter savings.
  • Replacing standard conformer blocks with quaternion versions in other GAN or diffusion pipelines might yield efficiency gains without retraining the entire stack.
  • Direct measurement of phase-magnitude coupling strength before and after the Hamilton product would clarify how much of the efficiency comes from the algebraic structure itself.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The manuscript proposes QC-GAN, a parameter-efficient speech enhancement model that pairs a Quaternion Conformer generator with MetricGAN-based training. The core mechanism is the Hamilton product within quaternion conformer layers, which is asserted to encode magnitude-phase interdependencies via structured weight sharing and thereby halve parameter counts while preserving performance. On the VoiceBank+DEMAND dataset the model is reported to reach PESQ 3.48 at 0.89 M parameters (comparable to SOTA at less than half the size) and PESQ 3.23 at a 35 K-parameter variant; generalization is claimed on the DNS-Challenge 3 dataset.

Significance. If the efficiency attribution holds after proper validation, the work would demonstrate a practical route to high-performing speech enhancement models small enough for resource-constrained devices, while the quaternion treatment of complex speech features could motivate analogous structured representations in other audio tasks.

major comments (1)
  1. [Abstract] Abstract: the central efficiency claim—that the Hamilton product encodes magnitude-phase interdependencies via structured weight sharing and thereby reduces layer parameters while preserving performance—is presented without any supporting ablation (quaternion vs. real-valued Conformer at matched parameter count), without an explicit mapping of 2-D complex speech features onto 4-D quaternion components, and without evidence that the reported PESQ scores could not be obtained by a conventional real-valued model of identical size. This attribution is therefore unverified and load-bearing for the paper’s primary contribution.
minor comments (1)
  1. [Abstract] Abstract: concrete PESQ scores and parameter counts are stated, yet no baselines, error bars, dataset splits, or training details are supplied, preventing independent verification of the numerical claims.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the constructive feedback. The major comment identifies a valid gap in supporting evidence for the central efficiency claim, which we will address through targeted revisions including new ablations and clarifications.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central efficiency claim—that the Hamilton product encodes magnitude and phase via structured weight sharing and thereby reduces the layer parameters while preserving their interdependencies—is presented without any supporting ablation (quaternion vs. real-valued Conformer at matched parameter count), without an explicit mapping of 2-D complex speech features onto 4-D quaternion components, and without evidence that the reported PESQ scores could not be obtained by a conventional real-valued model of identical size. This attribution is therefore unverified and load-bearing for the paper’s primary contribution.

    Authors: We agree that the abstract presents the efficiency attribution without the requested direct validations, making the claim unverified as noted. In revision we will add: (1) an explicit section describing the mapping from 2-D complex spectrogram features (real/imaginary parts) to 4-D quaternion components with the associated Hamilton product formulation; (2) a controlled ablation comparing the Quaternion Conformer generator to a real-valued Conformer baseline at identical parameter count (0.89 M and 35 K) on VoiceBank+DEMAND; (3) results demonstrating that the reported PESQ values are not matched by the real-valued counterpart of the same size. These additions will be reflected in the abstract and methods. We view this as a necessary strengthening of the primary contribution. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical results on standard benchmarks with no derivation chain

full rationale

The paper proposes the QC-GAN architecture and reports PESQ scores (3.48 at 0.89M parameters on VoiceBank+DEMAND; 3.23 at 35K parameters) as direct empirical outcomes of training and evaluation on fixed public datasets. No equations, parameter-fitting steps, or derivations are described that would reduce these metrics to inputs by construction. The Hamilton-product claim is presented as an architectural motivation rather than a self-referential prediction or fitted quantity. No self-citation load-bearing steps, uniqueness theorems, or ansatz smuggling appear in the provided text. The result is therefore self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Abstract provides no explicit free parameters or invented entities; the central efficiency claim rests on the unproven domain assumption that quaternion weight sharing preserves necessary signal interdependencies for speech.

assumptions (1)
  • domain assumption Hamilton product encodes magnitude and phase via structured weight sharing while preserving interdependencies
    Stated directly in abstract as the mechanism enabling parameter reduction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement." pith.science (2026). https://pith.science/paper/3BDP6FVF

@misc{pith2026260618611,
  author       = {Pith},
  title        = {Pith review of: QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3BDP6FVF}},
  note         = {Machine review of arXiv:2606.18611}
}
read the original abstract

We propose a parameter-efficient speech enhancement framework, Quaternion Conformer GAN (QC-GAN), which combines a Quaternion Conformer generator with MetricGAN-based training. The Hamilton product encodes the magnitude and phase via structured weight sharing, reducing the number of layer parameters while preserving their interdependencies. A metric-learning discriminator was employed to maximize perceptual quality by optimizing the approximate perceptual evaluation scores. On the VoiceBank+DEMAND dataset, QC-GAN achieved a Perceptual Evaluation of Speech Quality (PESQ) score of 3.48 with only 0.89M parameters, delivering a performance comparable to state-of-the-art models at less than half their size. A 35K-parameter variant achieved a PESQ score of 3.23, surpassing conventional methods with significantly fewer parameters. Evaluation on the DNS-Challenge 3 dataset further confirmed generalization to real-world conditions.

Figures

Figures reproduced from arXiv: 2606.18611 by the authors.

Figure 1
Figure 1. Feature learning in real-valued vs. quaternion layers. A real-valued layer considers magnitude and phase indepen￾dently, whereas a quaternion layer couples them through the Hamilton product as an inductive bias. To overcome this trade-off, we focus on Quaternion Neu￾ral Networks (QNNs), which leverage Hamilton’s quaternion algebra [13]. In a QNN, the Hamilton product performs linear transformations through structure… view at source ↗
Figure 2
Figure 2. Overall architecture of the proposed QC-GAN, comprising a Quaternion Encoder, a QG-Dilated DenseNet encoder (b), a two-stage Quaternion Conformer bottleneck (c), a dual-branch decoder (mask and complex residual), and a metric discriminator (d). The overall generator architecture is shown in (a). transformation y = W ⊗ x can be formulated as a real-valued matrix-vector multiplication using the Hamilton matrix HW :  … view at source ↗
Figure 3
Figure 3. PESQ curves over 100 training epochs for QC-GAN, Real-NN (32K and 140K), and QC-GAN ablated variants (w/o Discriminator and w/o TSQ-Conformer). QC-GAN (Tiny, 35K), and Real-NN (140K), which had 4× the parameters, corresponding to the theoretical reduction ratio of the Hamilton product. For a fair comparison, both variants used the same four-channel input features (Eq. 19), training configura￾tion, and loss functions… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Spectrogram comparison on a VoiceBank+DEMAND test sample. Yellow dashed boxes mark the pre-speech silent region (0–0.5 s), where QC-GAN (Tiny) achieves noise suppres￾sion closest to the clean reference. to the theoretical reduction ratio of the Hamilton product, at￾tai…
Figure 5
Figure 5. Figure 5: Phase reconstruction metrics on the Voice￾Bank+DEMAND test set. Lower values indicate better phase preservation. QC-GAN outperforms both Real-NN variants across all three metrics [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Phase reconstruction for QC-GAN and Real-NN (32K): spectrograms (top) and absolute phase error weighted by clean speech amplitude (bottom). QC-GAN reduces mean weighted phase error by 14.7% (0.414→0.353 rad). 6.2. Phase Preservation Analysis The ablation study ( [PITH…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 3 canonical work pages

  1. [1]

    QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

    Introduction Speech enhancement (SE) is a critical task that aims to improve the perceptual quality of speech signals in real-world environ- ments. With recent advances in deep learning, methods operating in the time-frequency (T–F) domain have become a dominant approach, resulting in significant performance improvements. In particular, Transformers with ...

  2. [2]

    Related Work 2.1. Speech-Enhancement Models Deep learning-based speech enhancement (SE) has evolved from early convolutional neural network (CNN) [ 21] and recurrent neural network (RNN) architectures [22] to more sophisticated Transformer and Conformer models. In particular, the Con- former excels at simultaneously capturing local and global de- pendenci...

  3. [3]

    Sec- tion 3.1 introduces the fundamentals of quaternion algebra, and Section 3.2 defines the quaternion neural network layers used as building blocks

    Proposed Method This section describes the proposed QC-GAN architecture. Sec- tion 3.1 introduces the fundamentals of quaternion algebra, and Section 3.2 defines the quaternion neural network layers used as building blocks. Sections 3.3 and 3.4 present the overall QC- GAN architecture and its loss functions, respectively. Finally, Section 3.5 describes th...

  4. [4]

    Dataset 4.1.1

    Experiment 4.1. Dataset 4.1.1. V oiceBank+DEMAND We evaluated QC-GAN using the V oiceBank+DEMAND dataset [36]2, which comprises 28 speakers (11,572 utterances) for training and 2 speakers (824 utterances) for testing. Clean speech from the V oice Bank corpus [37] was mixed with noise from the DEMAND database [38] and artificial sources at SNRs of 0–15 dB ...

  5. [5]

    To isolate the contribution of quaternion algebra, we constructed a real-valued counterpart, Real-NN, by replacing all quaternion layers with standard real-valued layers

    Ablation Study To validate the effectiveness of our architectural design choices, we conducted an ablation study using QC-GAN (Tiny) config- ured on the V oiceBank+DEMAND dataset [36]. To isolate the contribution of quaternion algebra, we constructed a real-valued counterpart, Real-NN, by replacing all quaternion layers with standard real-valued layers. W...

  6. [6]

    The comparison uses QC-GAN (Tiny) and Real- NN from the ablation study

    Discussion This section examines the advantages of QC-GAN over Real- NN, focusing on the learning characteristics and phase handling that contribute to its superior parameter efficiency and improved performance. The comparison uses QC-GAN (Tiny) and Real- NN from the ablation study. 6.1. Why Quaternion Networks Achieve Parameter Effi- ciency Quaternion la...

  7. [7]

    Conclusion We proposed QC-GAN, a speech-enhancement framework that integrates Quaternion Neural Networks into a Conformer ar- chitecture with MetricGAN-based training. Leveraging the Hamilton product, QC-GAN achieved performance competi- tive with SoTA models with approximately half their parameters, and the Tiny model (35K) attained a PESQ of 3.23 on V o...

  8. [8]

    Acknowledgements The authors thank the anonymous reviewers for their constructive comments

Show all 61 references
  1. [9]

    The authors reviewed and edited the outputs and take full responsibility for the final content of this paper

    Generative AI Use Disclosure Anthropic Claude was used for English translation assistance, and Paperpal (by Editage) was used for grammatical proofread- ing. The authors reviewed and edited the outputs and take full responsibility for the final content of this paper

  2. [10]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” inAdvances in Neural Information Processing Systems, vol. 30, 2017

  3. [11]

    Conformer: Convolution- augmented transformer for speech recognition,

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution- augmented transformer for speech recognition,” inProc. Inter- speech, 2020, pp. 5036–5040

  4. [12]

    An investigation of incorporating mamba for speech enhancement,

    R. Chao, W.-H. Cheng, M. L. Quatra, S. M. Siniscalchi, C.-H. H. Yang, S.-W. Fu, and Y . Tsao, “An investigation of incorporating mamba for speech enhancement,” in2024 IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 302–308

  5. [13]

    CMGAN: Conformer-based Metric GAN for Speech Enhancement,

    R. Cao, S. Abdulatif, and B. Yang, “CMGAN: Conformer-based Metric GAN for Speech Enhancement,” inInterspeech 2022, 2022, pp. 936–940

  6. [14]

    Mp-senet: A speech enhancement model with parallel denoising of magnitude and phase spectra,

    Y .-X. Lu, Y . Ai, and Z.-H. Ling, “Mp-senet: A speech enhancement model with parallel denoising of magnitude and phase spectra,” inINTERSPEECH 2023. ISCA, Aug. 2023, pp. 3834–3838. [Online]. Available: http://dx.doi.org/10.21437/ Interspeech.2023-1441

  7. [15]

    Towards model compression for deep learn- ing based speech enhancement,

    K. Tan and D. Wang, “Towards model compression for deep learn- ing based speech enhancement,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1785–1794, 2021

  8. [16]

    Lsenet: A lightweight spec- tral enhancement network for high-quality speech processing on resource-constrained platforms,

    H. Il Koh, S. Na, and M. N. Kim, “Lsenet: A lightweight spec- tral enhancement network for high-quality speech processing on resource-constrained platforms,”IEEE Access, vol. 13, pp. 116 934– 116 943, 2025

  9. [17]

    Deepfilternet: A low complexity speech enhancement framework for full-band audio based on deep filtering,

    H. Schr ¨oter, A. N. Escalante, T. Rosenkranz, and A. K. Maier, “Deepfilternet: A low complexity speech enhancement framework for full-band audio based on deep filtering,”ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. ...

  10. [18]

    DeepFilterNet: Perceptually motivated real-time speech enhance- ment,

    H. Schr ¨oter, T. Rosenkranz, A. N. Escalante-B., and A. Maier, “DeepFilterNet: Perceptually motivated real-time speech enhance- ment,” inINTERSPEECH, 2023

  11. [19]

    Deep- filternet2: Towards real-time speech enhancement on embedded devices for full-band audio,

    H. Schr¨oter, A. Maier, A. Escalante-B, and T. Rosenkranz, “Deep- filternet2: Towards real-time speech enhancement on embedded devices for full-band audio,” in2022 International Workshop on Acoustic Signal Enhancement (IWAENC), 2022, pp. 1–5

  12. [20]

    The importance of phase in speech enhancement,

    K. Paliwal, K. W ´ojcicki, and B. Shannon, “The importance of phase in speech enhancement,”Speech Communication, vol. 53, no. 4, pp. 465–494, 2011. [Online]. Available: https://www. sciencedirect.com/science/article/pii/S0167639310002086

  13. [21]

    Percep- tual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

    A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Percep- tual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in2001 IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 2...

  14. [22]

    On quaternions, or on a new system of imaginar- ies in algebra,

    W. R. Hamilton, “On quaternions, or on a new system of imaginar- ies in algebra,” 1847

  15. [23]

    A quaternary version of the back-propagation algorithm,

    T. Nitta, “A quaternary version of the back-propagation algorithm,” inProceedings of ICNN’95 - International Conference on Neural Networks, vol. 5, 1995, pp. 2753–2756

  16. [24]

    Quaternion convolutional neural networks,

    X. Zhu, Y . Xu, H. Xu, and C. Chen, “Quaternion convolutional neural networks,” inProceedings of the European Conference on Computer Vision (ECCV), September 2018

  17. [25]

    Quaternion approximate networks for enhanced image classification and oriented object detection,

    B. Grant and P. Wang, “Quaternion approximate networks for enhanced image classification and oriented object detection,” in 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025, pp. 4240–4247

  18. [26]

    Transformers at a fraction,

    A. Mukhopadhyay, R. B. Joshi, N. Tiwari, and S. Mishra, “Transformers at a fraction,” inNorthern Lights Deep Learning Conference 2025, 2024. [Online]. Available: https://openreview. net/forum?id=1U0kkt7ymn

  19. [27]

    Learning characteristics of reverse quaternion neural network,

    S. Yamauchi, T. Nitta, and T. Ohnishi, “Learning characteristics of reverse quaternion neural network,” in2025 International Joint Conference on Neural Networks (IJCNN), 2025, pp. 1–8

  20. [28]

    Quaternion convolutional neural networks for end-to-end automatic speech recognition,

    T. Parcollet, Y . Zhang, M. Morchid, C. Trabelsi, G. Linar `es, R. De Mori, and Y . Bengio, “Quaternion convolutional neural networks for end-to-end automatic speech recognition,” 06 2018

  21. [29]

    Quaternion recurrent neural networks,

    T. Parcollet, M. Ravanelli, M. Morchid, G. Linar `es, C. Trabelsi, R. D. Mori, and Y . Bengio, “Quaternion recurrent neural networks,” inInternational Conference on Learning Representations,

  22. [30]

    Available: https://openreview.net/forum?id= ByMHvs0cFQ

    [Online]. Available: https://openreview.net/forum?id= ByMHvs0cFQ

  23. [31]

    A fully convolutional neural network for speech enhancement,

    S. R. Park and J. W. Lee, “A fully convolutional neural network for speech enhancement,”Interspeech 2017, 08 2017. [Online]. Available: https://cir.nii.ac.jp/crid/1360011145594860160

  24. [32]

    Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr,

    F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr,” inLatent V ariable Analysis and Signal Separation, E. Vincent, A. Yeredor, Z. Koldovs...

  25. [33]

    Lisennet: Lightweight sub-band and dual-path modeling for real-time speech enhancement,

    H. Yan, J. Zhang, C. Fan, Y . Zhou, and P. Liu, “Lisennet: Lightweight sub-band and dual-path modeling for real-time speech enhancement,” inICASSP 2025 - 2025 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–5

  26. [34]

    Quaternion neural networks for multi-channel distant speech recog- nition,

    X. Qiu, T. Parcollet, M. Ravanelli, N. D. Lane, and M. Morchid, “Quaternion neural networks for multi-channel distant speech recog- nition,” inInterspeech 2020, 2020, pp. 329–333

  27. [35]

    Quaternion neural networks for 3d sound source localization in reverberant environments,

    M. R. Celsi, S. Scardapane, and D. Comminiello, “Quaternion neural networks for 3d sound source localization in reverberant environments,” in2020 IEEE 30th International Workshop on Machine Learning for Signal Processing (MLSP), 2020, pp. 1–6

  28. [36]

    A short- time objective intelligibility measure for time-frequency weighted noisy speech,

    C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short- time objective intelligibility measure for time-frequency weighted noisy speech,” in2010 IEEE International Conference on Acous- tics, Speech and Signal Processing, 2010, pp. 4214–4217

  29. [37]

    MetricGAN: Gen- erative adversarial networks based black-box metric scores op- timization for speech enhancement,

    S.-W. Fu, C.-F. Liao, Y . Tsao, and S.-D. Lin, “MetricGAN: Gen- erative adversarial networks based black-box metric scores op- timization for speech enhancement,” inProceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Resea...

  30. [38]

    MetricGAN+: An improved version of MetricGAN for speech enhancement,

    S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y . Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” inInterspeech, 2021, pp. 201–205

  31. [39]

    Deep quaternion networks,

    C. J. Gaudet and A. Maida, “Deep quaternion networks,”2018 International Joint Conference on Neural Networks (IJCNN), pp. 1–8, 2017. [Online]. Available: https://api.semanticscholar.org/ CorpusID:6912552

  32. [40]

    Lightweight and efficient neural natural language processing with quaternion networks,

    Y . Tay, A. Zhang, A. T. Luu, J. Rao, S. Zhang, S. Wang, J. Fu, and S. C. Hui, “Lightweight and efficient neural natural language processing with quaternion networks,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. T...

  33. [41]

    Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,

    A. Pandey and D. Wang, “Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,” inICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 6629–6633

  34. [42]

    Wavenet: A generative model for raw audio,

    A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” in9th ISCA Work- shop on Speech Synthesis Workshop (SSW 9), 2016, p. 125

  35. [43]

    Quater- nion convolutional neural network for color image classification and forensics,

    Q. Yin, J. Wang, X. Luo, J. Zhai, S. K. Jha, and Y .-Q. Shi, “Quater- nion convolutional neural network for color image classification and forensics,”IEEE Access, vol. 7, pp. 20 293–20 301, 2019

  36. [44]

    Delving deep into recti- fiers: Surpassing human-level performance on imagenet classifica- tion,

    K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into recti- fiers: Surpassing human-level performance on imagenet classifica- tion,” in2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1026–1034

  37. [45]

    A deep learning loss function based on the perceptual evaluation of the speech quality,

    J. M. Martin-Do ˜nas, A. M. Gomez, J. A. Gonzalez, and A. M. Peinado, “A deep learning loss function based on the perceptual evaluation of the speech quality,”IEEE Signal Processing Letters, vol. 25, no. 11, pp. 1680–1684, 2018

  38. [46]

    Investigating rnn-based speech enhancement methods for noise- robust text-to-speech,

    C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating rnn-based speech enhancement methods for noise- robust text-to-speech,” in9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), 2016, pp. 146–152

  39. [47]

    The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,

    C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” inProc. Int. Conf. Oriental COCOSDA, Nov. 2013

  40. [48]

    The diverse environments multi-channel acoustic noise database (DEMAND): A database of multichannel environmental noise recordings,

    J. Thiemann, N. Ito, and E. Vincent, “The diverse environments multi-channel acoustic noise database (DEMAND): A database of multichannel environmental noise recordings,” inProceedings of Meetings on Acoustics, vol. 19, no. 1. Acoustical Society of America, 2013, p. 035081

  41. [49]

    IN- TERSPEECH 2021 deep noise suppression challenge,

    C. K. A. Reddy, H. Dubey, K. Koishida, A. Nair, V . Gopal, R. Cut- ler, S. Braun, H. Gamper, R. Aichner, and S. Srinivasan, “IN- TERSPEECH 2021 deep noise suppression challenge,” inProc. Interspeech 2021, 2021, pp. 2796–2800

  42. [50]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” inInternational Conference on Learning Representations, 2019. [Online]. Available: https://openreview.net/ forum?id=Bkg6RiCqY7

  43. [51]

    Evaluation of objective quality measures for speech enhancement,

    Y . Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,”IEEE Transactions on Audio, Speech, and Language Processing, vol. 16, no. 1, pp. 229–238, 2008

  44. [52]

    Dnsmos p.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

    C. K. Reddy, V . Gopal, and R. Cutler, “Dnsmos p.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,” inICASSP 2022 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022

  45. [53]

    Dnsmos: A non-intrusive perceptual objective speech qual- ity metric to evaluate noise suppressors,

    ——, “Dnsmos: A non-intrusive perceptual objective speech qual- ity metric to evaluate noise suppressors,” inICASSP 2021 IEEE International Conference on Acoustics, Speech and Signal Process- ing (ICASSP). IEEE, 2021, pp. 6493–6497

  46. [54]

    SEGAN: Speech enhance- ment generative adversarial network,

    S. Pascual, A. Bonafonte, and J. Serr`a, “SEGAN: Speech enhance- ment generative adversarial network,” inProc. of Interspeech, 2017, pp. 3642–3646

  47. [55]

    Dpt-fsnet: Dual-path transformer based full-band and sub-band fusion network for speech enhancement,

    F. Dang, H. Chen, and P. Zhang, “Dpt-fsnet: Dual-path transformer based full-band and sub-band fusion network for speech enhancement,”ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6857–6861, 2021. [Online]. Available:...

  48. [56]

    A hybrid DSP/deep learning approach to real-time full-band speech enhancement,

    J.-M. Valin, “A hybrid DSP/deep learning approach to real-time full-band speech enhancement,” inProc. MMSP, 2018, pp. 1–5

  49. [57]

    First coarse, fine afterward: A lightweight two-stage complex approach for monaural speech enhancement,

    F. Dang, H. Chen, Q. Hu, P. Zhang, and Y . Yan, “First coarse, fine afterward: A lightweight two-stage complex approach for monaural speech enhancement,”Speech Communication, vol. 146, pp. 32–44, 2023

  50. [58]

    FSPEN: An ultra-lightweight network for real time speech en- hancement,

    L. Yang, W. Liu, R. Meng, G. Lee, S. Baek, and H.-G. Moon, “FSPEN: An ultra-lightweight network for real time speech en- hancement,” inProc. ICASSP, 2024, pp. 10 671–10 675

  51. [59]

    Data augmentation and loss normalization for deep noise suppression,

    S. Braun and I. Tashev, “Data augmentation and loss normalization for deep noise suppression,” 2020. [Online]. Available: https://arxiv.org/abs/2008.06412

  52. [60]

    Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement,

    Y . Hu, Y . Liu, S. Lv, M. Xing, S. Zhang, Y . Fu, J. Wu, B. Zhang, and L. Xie, “Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement,” inInterspeech, 2020. [Online]. Available: https://api.semanticscholar.org/CorpusID: 220936516

  53. [61]

    Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses,

    Y . Ai and Z.-H. Ling, “Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.