Pith. sign in

REVIEW 1 cited by

Model selection for deep audio source separation via clustering analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12626 v2 pith:VGKC6UBR submitted 2019-10-23 eess.AS cs.LGcs.SDstat.ML

classification eess.AScs.LGcs.SDstat.ML
keywords modelaudiomixtureseparationdeepensemblegivenmixtures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio source separation is the process of separating a mixture (e.g. a pop band recording) into isolated sounds from individual sources (e.g. just the lead vocals). Deep learning models are the state-of-the-art in source separation, given that the mixture to be separated is similar to the mixtures the deep model was trained on. This requires the end user to know enough about each model's training to select the correct model for a given audio mixture. In this work, we automate selection of the appropriate model for an audio mixture. We present a confidence measure that does not require ground truth to estimate separation quality, given a deep model and audio mixture. We use this confidence measure to automatically select the model output with the best predicted separation quality. We compare our confidence-based ensemble approach to using individual models with no selection, to an oracle that always selects the best model and to a random model selector. Results show our confidence-based ensemble significantly outperforms the random ensemble over general mixtures and approaches oracle performance for music mixtures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models

    astro-ph.IM 2025-06 conditional novelty 5.0 of 10

    Whale and bird vocalizations, converted to abstract tokens and fed to GPT-2, trigger more structured language-like output than white noise, suggesting a reactivity-based screening metric for SETI.

Pith tools