Pith. sign in

Scaling Speech Enhancement in Unseen Environments with Noise Embeddings

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We address the problem of speech enhancement generalisation to unseen environments by performing two manipulations. First, we embed an additional recording from the environment alone, and use this embedding to alter activations in the main enhancement subnetwork. Second, we scale the number of noise environments present at training time to 16,784 different environments. Experiment results show that both manipulations reduce word error rates of a pretrained speech recognition system and improve enhancement quality according to a number of performance measures. Specifically, our best model reduces the word error rate from 34.04% on noisy speech to 15.46% on the enhanced speech. Enhanced audio samples can be found in https://speechenhancement.page.link/samples.

citation-role summary

background 1

citation-polarity summary

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Training-Free Multi-Step Audio Source Separation

cs.SD · 2025-05-26 · conditional · novelty 6.0

Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.

citing papers explorer

Showing 1 of 1 citing paper.

  • Training-Free Multi-Step Audio Source Separation cs.SD · 2025-05-26 · conditional · none · ref 15 · internal anchor

    Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.