Pith. sign in

REVIEW 3 major objections 5 minor 51 references

From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A two-stage speech enhancer that first cleans continuous features, then generates discrete tokens, claims to beat single-paradigm models.

desk verdict Two-stage continuous-then-discrete GSE with a RootLM/BranchLM codec hierarchy is a real architectural contribution, but private fine-tuning data and missing significance tests keep the 'surpasses' claim from being proven. read the letter →

arxiv 2507.19062 v1 pith:73UJPIK3 submitted 2025-07-25 cs.SD eess.AS

classification cs.SDeess.AS
keywords generalspeechenhancementhierarchicallanguagemodelsneuralaudiocodecresidualvectorquantizationcompounddistortionsrestorationpacketlossconcealmentcross-domaincollaborativeoptimization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces OmniGSE, a general speech enhancement system aimed at speech degraded by several distortions at once: background noise, reverberation, limited bandwidth, clipping, and packet loss. Its central claim is that no single paradigm is enough, because discriminative models are precise on regression tasks like denoising while generative models are flexible on reconstruction tasks like declipping, so the two should be combined. OmniGSE does this in two stages: it first enhances the continuous features produced by a neural audio codec encoder, then uses a hierarchical language model to generate the codec's discrete tokens. The paper reports that this architecture beats both discriminative and generative baselines across multiple benchmark test sets, with the largest advantage on compound distortions. A sympathetic reader would care because real-world recordings usually suffer several degradation types together, and most existing methods are built for only one.

What carries the argument

The load-bearing object is the hierarchical language model built on the codec's residual vector quantization: a RootLM models shared acoustic content, while per-level BranchLMs predict each codebook level conditioned on the RootLM and on the previous level, which is what lets the system regenerate missing spectral or temporal content while keeping acoustic consistency. The supporting mechanism is the channel-split NAC-RoFormer, which lowers the computational cost of attending to 1024-dimensional codec features by grouping channels and alternating temporal and cross-group attention.

What would settle it

Train OmniGSE on public data only and evaluate on a held-out compound-distortion set whose distortion recipe differs from the training recipe; if its advantage over single-paradigm baselines disappears, the central claim of general superiority is not supported.

Watch

Extended reading notes

Core claim

Stage I uses a channel-split network with dual-path rotary-position attention to map the codec encoder's features for distorted speech toward the clean-speech features produced by a teacher codec encoder, and the codec encoder itself is fine-tuned on distorted input at the same time. Stage II takes those enhanced pre-quantized features as conditioning and autoregressively predicts the residual vector quantization tokens of clean speech. The hierarchical language model consists of one RootLM, which captures acoustic features shared across codebook levels, and a separate BranchLM per level, which captures the progressive relationship from one codebook level to the next; the levels are trained with teacher forcing under a cross-entropy loss. The discovery, stated on the paper's terms, is that this continuous-to-discrete collaboration resolves the precision-versus-flexibility tradeoff, and that the hierarchical LM reduces inter-level prediction conflicts that a single shared LM would introduce.

Load-bearing premise

The central claim assumes the performance gains come from the two-stage architecture itself, not from the private high-fidelity fine-tuning data or from the hand-chosen distortion mix aligning with the test sets.

Editorial extensions

If this is right

  • A single OmniGSE model could replace separate denoising, dereverberation, bandwidth extension, declipping, and packet-loss concealment modules.
  • Because the first stage improves the conditioning features, the second stage can focus on regenerating missing content instead of suppressing noise, so the two stages compound.
  • The paper's ablation says using separate BranchLMs avoids the pattern conflicts that a single multi-level LM produces.
  • The framework reports its restoration quality without extra self-supervised semantic features, which lowers conditioning cost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the RootLM/BranchLM split is a general recipe for any residual-vector-quantization codec, so it could transfer to other neural codecs and to low-bitrate speech synthesis.
  • Editorial inference: the architecture implies that discriminative feature cleaning can act as a universal conditioning front-end for generative audio models, not just for speech enhancement.
  • Editorial inference: the paper does not report how sensitive the results are to the hand-chosen distortion probabilities; varying them per deployment is a direct test of the framework's real-world generality.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes OmniGSE, a two-stage general speech enhancement framework that combines a continuous feature enhancement stage (channel-split NAC-RoFormer) with a discrete token generation stage based on a hierarchical language model (RootLM plus BranchLMs). The authors evaluate on DNS2020 denoising/dereverberation, Voicefixer SR/GSR, and Interspeech 2022 PLC benchmarks, and they report ablations of the two stages, the NAC-RoFormer, the hierarchical LM, encoder fine-tuning, and teacher forcing. The central claim is that this cross-domain collaborative design surpasses existing single-paradigm discriminative and generative methods, with particular strength in compound distortion scenarios.

Significance. The architectural idea is coherent: a discriminative pre-enhancement stage feeds high-SNR continuous features to a generative LM-based token predictor, and the hierarchical RootLM/BranchLM design explicitly models RVQ inter-codebook dependencies. The ablations in Table 5 provide useful evidence that each component contributes to the final performance. The breadth of the benchmark coverage (denoising, dereverberation, super-resolution, restoration, packet loss) is a strength. However, the central comparative claim is currently not fully supported because the evaluation uses a private fine-tuning dataset unavailable to the baselines, reports only point estimates without confidence intervals, and the listening tests omit the most relevant generative baselines. If these evaluation gaps are addressed, the framework would be a solid contribution to general speech enhancement.

major comments (3)
  1. [Sec. 4.1, Tables 1-4] The comparison is confounded by the private fine-tuning data. Section 4.1 states that the authors 'fine-tuned our model on a private high-fidelity speech dataset,' but the baselines in Tables 1-4 were not given this data. Additionally, the training distortion recipe in Sec. 3.1 (noise 100%, reverb 50%, other distortions equally likely) may be aligned with the test-set conditions, and the baselines were not trained with this exact recipe. Therefore the reported gains cannot be attributed solely to the proposed architecture. I request a controlled experiment: (i) train OmniGSE without the private data; (ii) fine-tune at least one representative generative baseline (e.g., MaskSR or AnyEnhance) on the same private data and distortion pipeline; and (iii) report the performance difference to isolate the contribution of the private data.
  2. [Tables 1-3] No confidence intervals or significance tests are reported, and the claimed superiority is not consistent across all metrics. For example, in Table 2 OmniGSEfb has lower SBS (0.930 vs 0.941) and SIM (0.935 vs 0.943) than AnyEnhance, and in Table 3 OmniGSEfb has lower NISQA (4.293 vs 4.335) than MaskSR. Some OVRL differences in Table 1 are as small as 0.026 (3.444 vs 3.418). Without variance estimates or statistical tests, the statement in the abstract that OmniGSE 'surpasses existing models across multiple benchmarks' is overstated, and the specific claim of excelling in compound distortions is not uniformly supported by the GSR results.
  3. [Figures 3-4] The subjective listening tests compare OmniGSE only against discriminative baselines (FullSubNet, VoiceFixer, TF-GridNet) and omit the strongest generative competitors (MaskSR, AnyEnhance, LLaSE-G1) that are the direct point of comparison for the generative stage. Figure 3 and Figure 4 also lack details on the number of listeners, the number of stimuli, and any statistical analysis. Please extend the listening test to include at least the strongest generative baselines and report listener counts and significance testing.
minor comments (5)
  1. [Sec. 2.2] The word 'knowlege' should be 'knowledge' in the sentence about discrete codebooks encapsulating rich prior knowledge.
  2. [Sec. 4.1] The description of the private high-fidelity dataset is too vague; please report its size, duration, and whether it was used only for fine-tuning after the main training or as part of the main training mixture, since this materially affects the comparison.
  3. [Sec. 4.2] The full-band and wideband models contain roughly 0.97B and 1.23B parameters, respectively, yet Sec. 2.2 claims reduced computational cost. Please report inference time (e.g., RTF) and memory usage to substantiate the efficiency claim.
  4. [Figure 5] The x-axis label 'SNR' is ambiguous; please define whether this is signal-domain SNR or a feature-space measure, and additionally report the final objective scores for each conditioning feature rather than only the feature SNR.
  5. [Table 1] Some baseline entries have missing values (e.g., VoiceFixer SBS and SIM, SELM and GenSE NISQA/SBS/SIM); please indicate whether these are not reported or not applicable, and consider supplementing the missing metrics for a complete comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical system built on external benchmarks, with no derivation step that reduces to its own inputs.

full rationale

The central claim, that OmniGSE surpasses existing models on general speech enhancement benchmarks, is an empirical claim supported by comparisons on public test sets (Interspeech 2020 DNS, Voicefixer SR/GSR, Interspeech 2022 PLC). The training objectives are Lemb = MSE(Fenh, Ftea) in Eq. (6) and the code cross-entropy loss in Eq. (7), both of which supervise the model toward teacher-NAC targets for clean speech. These are standard training losses, not quantities used to derive benchmark scores. No parameter is fitted to the test metrics, and no metric used in Tables 1-4 appears as an input to the model. The paper does not invoke a uniqueness theorem or an ansatz from the authors' prior work; references to DAC, WavLM, DNSMOS, and NISQA are external, and no cited result is loaded with the paper's own conclusion. The disclosed use of a private high-fidelity speech dataset in Sec. 4.1 is a legitimate evaluation-fairness concern: the baselines were not retrained on that data, so the reported advantages could in principle reflect data rather than architecture. However, that is a controlled-comparison issue, not circularity under the defined patterns; it does not make the benchmark claim true by construction. The distortion simulation recipe (noise at 100%, reverb at 50%, other distortions equally likely) likewise matches common evaluation practice and is not a fitted parameter renamed as a prediction. Overall, no step of the paper's argument reduces, by definition or by self-citation, to its own inputs, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is an empirical system paper, so the ledger is mostly free of fitted constants. The distortion probabilities in Sec. 3.1 are hand-chosen design choices rather than numbers fitted to benchmarks, so they are not listed as free parameters. The main external inputs are the pretrained DAC, public datasets, and a private fine-tuning dataset. The listed axioms are the load-bearing domain assumptions about the codec representation, the teacher targets, the evaluation metrics, and the realism of the synthetic distortions.

assumptions (4)
  • domain assumption The pre-trained DAC codebooks provide a sufficient discrete representation of clean speech such that codec reconstruction error is small relative to enhancement gains.
    Stage II trains LMs to predict DAC tokens from a teacher network; if DAC cannot represent clean speech well, token prediction cannot yield high-quality output. Invoked in Sec. 3.4.
  • domain assumption The teacher NAC's clean-speech continuous embeddings are a valid regression target for stage I.
    Stage I optimizes MSE against teacher embeddings from clean speech; if the teacher embeddings carry a bias, the enhanced features inherit it. Sec. 3.3, Eq. 6.
  • domain assumption DNSMOS, NISQA, PLCMOS, SBS, and SIM, along with the collected MOS, are reliable proxies for perceptual quality and speaker fidelity.
    All conclusions rest on these metrics; no listening test is fully described and no significance testing is reported. Sec. 4.3.
  • domain assumption The synthetic distortion pipeline (sequential noise, reverb, then one of clipping, super-resolution, or packet loss) is representative of real-world compound distortions.
    Training data is generated with fixed probabilities; if real distortions combine differently, the model may not generalize. Sec. 3.1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models." pith.science (2026). https://pith.science/paper/73UJPIK3

@misc{pith2026250719062,
  author       = {Pith},
  title        = {Pith review of: From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/73UJPIK3}},
  note         = {Machine review of arXiv:2507.19062}
}
read the original abstract

This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications.

Figures

Figures reproduced from arXiv: 2507.19062 by the authors.

Figure 1
Figure 1. Workflow of the proposed OmniGSE framework. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Topology of the hierarchical language model. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Violin plots of NMOS and SMOS scores for various methods on the Interspeech 2020 Challenge blind test set. The [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Violin plots of NMOS and SMOS scores for various methods on the Voicefixer SR and Voicefixer GSR test sets. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Results and SNR of different input features on the [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 36 canonical work pages

  1. [1]

    Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, and Yang Zhang. 2021. Hi-Fi Multi-Speaker English TTS Dataset. In Interspeech. ISCA, 2776–2780

  2. [2]

    Sebastian Braun and Ivan Tashev. 2020. Data Augmentation and Loss Normaliza- tion for Deep Noise Suppression. In SPECOM (Lecture Notes in Computer Science, Vol. 12335). Springer, 79–86

  3. [3]

    Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. 2022. WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing. IEEE J. Sel. Top. Signal Process. 16,...

  4. [4]

    Sanyuan Chen, Chengyi Wang, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al . 2025. Neural codec language models are zero-shot text to speech synthesizers. IEEE Transactions on Audio, Speech and Language Processing (2025)

  5. [5]

    Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2023. High Fidelity Neural Audio Compression. Trans. Mach. Learn. Res. 2023 (2023)

  6. [6]

    Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi. 2020. Real Time Speech Enhancement in the Waveform Domain. In INTERSPEECH. ISCA, 3291–3295

  7. [7]

    Lorenz Diener, Marju Purin, Sten Sootla, Ando Saabas, Robert Aichner, and Ross Cutler. 2023. PLCMOS - A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms. In INTERSPEECH. ISCA, 2533–2537

  8. [8]

    Lorenz Diener, Sten Sootla, Solomiya Branets, Ando Saabas, Robert Aichner, and Ross Cutler. 2022. INTERSPEECH 2022 Audio Deep Packet Loss Concealment Challenge. In INTERSPEECH. ISCA, 580–584

Show all 51 references
  1. [9]

    Harishchandra Dubey, Ashkan Aazami, Vishak Gopal, Babak Naderi, Sebastian Braun, Ross Cutler, Alex Ju, Mehdi Zohourian, Min Tang, Mehrsa Golestaneh, et al. 2024. Icassp 2023 deep noise suppression challenge. IEEE Open Journal of Signal Processing (2024)

  2. [10]

    Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra

  3. [11]

    Xiang Hao, Xiangdong Su, Radu Horaud, and Xiaofei Li. 2021. Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement. In ICASSP. IEEE, 6633–6637

  4. [12]

    Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE ACM Trans. Audio Speech Lang. Process. 29 (2021), 3451–3460

  5. [13]

    Shengpeng Ji, Ziyue Jiang, Xize Cheng, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Ruiqi Li, Ziang Zhang, Xiaoda Yang, Rongjie Huang, Yidi Jiang, Qian Chen, Siqi Zheng, Wen Wang, and Zhou Zhao. 2024. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio L...

  6. [14]

    Boyi Kang, Xinfa Zhu, Zihan Zhang, Zhen Ye, Mingshuai Liu, Ziqian Wang, Yike Zhu, Guobin Ma, Jun Chen, Longshuai Xiao, et al. 2025. LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement. arXiv preprint arXiv:2503.00493 (2025)

  7. [15]

    Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2023. High-Fidelity Audio Compression with Improved RVQGAN. In NeurIPS

  8. [16]

    Haoyang Li, Jia Qi Yip, Tianyu Fan, and Eng Siong Chng. 2025. Speech En- hancement Using Continuous Embeddings of Neural Audio Codec. CoRR abs/2502.16240 (2025)

  9. [17]

    Nan Li, Xiguang Zheng, Chen Zhang, Liang Guo, and Bing Yu. 2022. End-to-End Multi-Loss Training for Low Delay Packet Loss Concealment. In INTERSPEECH. ISCA, 585–589

  10. [18]

    Xu Li, Qirui Wang, and Xiaoyu Liu. 2024. MaskSR: Masked Language Model for Full-band Speech Restoration. CoRR abs/2406.02092 (2024)

  11. [19]

    Baiyun Liu, Qi Song, Mingxue Yang, Wuwen Yuan, and Tianbao Wang. 2022. PLCNet: Real-time Packet Loss Concealment with Semi-supervised Generative Adversarial Network. In INTERSPEECH. ISCA, 575–579

  12. [20]

    Plumbley

    Haohe Liu, Ke Chen, Qiao Tian, Wenwu Wang, and Mark D. Plumbley. 2024. Audiosr: Versatile Audio Super-Resolution at Scale. In ICASSP. IEEE, 1076–1080

  13. [21]

    Haohe Liu, Xubo Liu, Qiuqiang Kong, Qiao Tian, Yan Zhao, DeLiang Wang, Chuanzeng Huang, and Yuxuan Wang. 2022. VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration. In INTERSPEECH. ISCA, 4232–4236

  14. [22]

    Ziyin Liu, Tilman Hartwig, and Masahito Ueda. 2020. Neural Networks Fail to Learn Periodic Functions and How to Fix It. In NeurIPS

  15. [23]

    Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In ICLR (Poster). OpenReview.net

  16. [24]

    Wei Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, and Yun-Ning Hung. 2024. Music Source Separation With Band-Split Rope Transformer. In ICASSP. IEEE, 481–485

  17. [25]

    Ye-Xin Lu, Yang Ai, and Zhen-Hua Ling. 2023. MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra. InINTERSPEECH. ISCA, 3834–3838

  18. [26]

    Gabriel Mittag, Babak Naderi, Assmaa Chehadi, and Sebastian Möller. 2021. NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets. In Interspeech. ISCA, 2127–2131

  19. [27]

    Chandan K. A. Reddy, Vishak Gopal, and Ross Cutler. 2022. Dnsmos P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors. In ICASSP. IEEE, 886–890

  20. [28]

    Chandan K. A. Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, Puneet Rana, Sriram Srinivasan, and Johannes Gehrke. 2020. The INTERSPEECH 2020 Deep Noise Suppression Challeng...

  21. [29]

    Takaaki Saeki, Soumi Maiti, Shinnosuke Takamichi, Shinji Watanabe, and Hiroshi Saruwatari. 2024. SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics. CoRR abs/2401.16812 (2024)

  22. [30]

    Jianlin Su, Murtadha H. M. Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. RoFormer: Enhanced transformer with Rotary Position Embedding. Neurocomputing 568 (2024), 127063

  23. [31]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. LLaMA: Open and Efficient Foundation ...

  24. [32]

    Ter- riberry, Michael Klingbeil, Paris Smaragdis, and Arvindh Krishnaswamy

    Jean-Marc Valin, Ahmed Mustafa, Christopher Montgomery, Timothy B. Ter- riberry, Michael Klingbeil, Paris Smaragdis, and Arvindh Krishnaswamy. 2022. Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model. In INTERSPEECH. ISCA, 570–574

  25. [33]

    Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al. 2017. CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit. University of Edinburgh. The Centre for Speech Technology Research (CSTR) 6 (2017), 15

  26. [34]

    Zhong-Qiu Wang, Samuele Cornell, Shukjae Choi, Younglo Lee, Byeong-Yeol Kim, and Shinji Watanabe. 2023. TF-GridNet: Integrating Full- and Sub-Band Modeling for Speech Separation. IEEE ACM Trans. Audio Speech Lang. Process. 31 (2023), 3221–3236

  27. [35]

    Ziqian Wang, Xinfa Zhu, Zihan Zhang, Yuanjun Lv, Ning Jiang, Guoqing Zhao, and Lei Xie. 2024. SELM: Speech Enhancement using Discrete Tokens and Language Models. In ICASSP. IEEE, 11561–11565

  28. [36]

    Gordon Wichern, Joe Antognini, Michael Flynn, Licheng Richard Zhu, Emmett McQuinn, Dwight Crow, Ethan Manilow, and Jonathan Le Roux. 2019. WHAM!: Extending Speech Separation to Noisy Environments. In INTERSPEECH. ISCA, 1368–1372

  29. [37]

    Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. 2022. Vision Transformer with Deformable Attention. In CVPR. IEEE, 4784–4793

  30. [38]

    Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2024. BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec. CoRR abs/2409.05377 (2024)

  31. [39]

    Haici Yang, Jiaqi Su, Minje Kim, and Zeyu Jin. 2024. Genhancer: High-fidelity speech enhancement via generative modeling on discrete codec tokens. In Proc. Interspeech 2024. 1170–1174

  32. [40]

    Jixun Yao, Hexin Liu, Chen Chen, Yuchen Hu, Chng Eng Siong, and Lei Xie. 2025. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling. CoRR abs/2502.02942 (2025)

  33. [41]

    Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang, Xu Tan, Jiahe Lei, Yi Peng, Haohe Liu, Yizhu Jin, Zheqi Dai, Hongzhan Lin, Jianyi Chen, Xingjian Du, Liu- meng Xue, Yunlin Chen, Zhifei Li, Lei Xie, Qiuqiang Kong, Yike Guo, and Wei Xue

  34. [42]

    Jia Qi Yip, Shengkui Zhao, Dianwen Ng, Eng Siong Chng, and Bin Ma. 2024. Towards audio codec-based speech separation. arXiv preprint arXiv:2406.12434 (2024)

  35. [43]

    Jianwei Yu and Yi Luo. 2023. Efficient Monaural Speech Enhancement with Universal Sample Rate Band-Split RNN. In ICASSP. IEEE, 1–5

  36. [44]

    Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big Bird: Transformers for Longer Sequences. InNeurIPS

  37. [45]

    Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

    Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019. LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech. In INTERSPEECH. ISCA, 1526–1530

  38. [46]

    Junan Zhang, Jing Yang, Zihao Fang, Yuancheng Wang, Zehua Zhang, Zhuo Wang, Fan Fan, and Zhizheng Wu. 2025. AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement. CoRR abs/2501.15417 (2025)

  39. [47]

    Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Anurag Kumar, Jan Pirklbauer, Marvin Sach, Shinji Watanabe, Tim Fingscheidt, and Yanmin Qian. 2024. URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement. Co...

  40. [48]

    Zihan Zhang, Jiayao Sun, Xianjun Xia, Chuanzeng Huang, Yijian Xiao, and Lei Xie. 2024. Bs-Plcnet: Band-Split Packet Loss Concealment Network with Multi- Task Learning Framework and Multi-Discriminators. In ICASSP Workshops. IEEE, 23–24

  41. [49]

    Watcharasupat, and Woon-Seng Gan

    Shengkui Zhao, Bin Ma, Karn N. Watcharasupat, and Woon-Seng Gan. 2022. FR- CRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement. In ICASSP. IEEE, 9281–9285

  42. [2022]

    IEEE ACM Trans

    FSD50K: An Open Dataset of Human-Labeled Sound Events. IEEE ACM Trans. Audio Speech Lang. Process. 30 (2022), 829–852

  43. [2025]

    CoRR abs/2502.04128 (2025)

    Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis. CoRR abs/2502.04128 (2025)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.