Pith. sign in

Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Deep learning-based speech enhancement (SE) models have achieved impressive performance in the past decade. Numerous advanced architectures have been designed to deliver state-of-the-art performance; however, their scalability potential remains unrevealed. Meanwhile, the majority of research focuses on small-sized datasets with restricted diversity, leading to a plateau in performance improvement. In this paper, we aim to provide new insights for addressing the above issues by exploring the scalability of SE models in terms of architectures, model sizes, compute budgets, and dataset sizes. Our investigation involves several popular SE architectures and speech data from different domains. Experiments reveal both similarities and distinctions between the scaling effects in SE and other tasks such as speech recognition. These findings further provide insights into the under-explored SE directions, e.g., larger-scale multi-domain corpora and efficiently scalable architectures.

citation-role summary

background 1

citation-polarity summary

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

support 1

representative citing papers

Training-Free Multi-Step Audio Source Separation

cs.SD · 2025-05-26 · conditional · novelty 6.0

Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.

citing papers explorer

Showing 1 of 1 citing paper.

  • Training-Free Multi-Step Audio Source Separation cs.SD · 2025-05-26 · conditional · none · ref 55 · internal anchor

    Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.