Pith. sign in

REVIEW 2 cited by

URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04660 v1 pith:IP6DO2KN submitted 2024-06-07 eess.AS cs.SD

classification eess.AScs.SD
keywords dataspeechsub-taskschallengedifferentenhancementevaluationexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To fill this gap and promote research toward universal SE, we establish a new SE challenge, named URGENT, to focus on the universality, robustness, and generalizability of SE. We aim to extend the SE definition to cover different sub-tasks to explore the limits of SE models, starting from denoising, dereverberation, bandwidth extension, and declipping. A novel framework is proposed to unify all these sub-tasks in a single model, allowing the use of all existing SE approaches. We collected public speech and noise data from different domains to construct diverse evaluation data. Finally, we discuss the insights gained from our preliminary baseline experiments based on both generative and discriminative SE methods with 12 curated metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Model as Loss: A Self-Consistent Training Paradigm

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Using the model's own encoder as a feature loss improves perceptual quality and iterative stability of a speech enhancement model compared with a WavLM-based loss.

  2. Training-Free Multi-Step Audio Source Separation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Iteratively remixing and re-separating the input mixture, with the best blend chosen by a quality metric, improves pretrained one-step audio separation models without any retraining.

Pith tools