Pith. sign in

REVIEW 4 cited by

Universal Source Separation with Weakly Labelled Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07447 v1 pith:KIVNCVFK submitted 2023-05-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords sourceseparationaudiosounddataseparateclassesdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Universal source separation (USS) is a fundamental research task for computational auditory scene analysis, which aims to separate mono recordings into individual source tracks. There are three potential challenges awaiting the solution to the audio source separation task. First, previous audio source separation systems mainly focus on separating one or a limited number of specific sources. There is a lack of research on building a unified system that can separate arbitrary sources via a single model. Second, most previous systems require clean source data to train a separator, while clean source data are scarce. Third, there is a lack of USS system that can automatically detect and separate active sound classes in a hierarchical level. To use large-scale weakly labeled/unlabeled audio data for audio source separation, we propose a universal audio source separation framework containing: 1) an audio tagging model trained on weakly labeled data as a query net; and 2) a conditional source separation model that takes query net outputs as conditions to separate arbitrary sound sources. We investigate various query nets, source separation models, and training strategies and propose a hierarchical USS strategy to automatically detect and separate sound classes from the AudioSet ontology. By solely leveraging the weakly labelled AudioSet, our USS system is successful in separating a wide variety of sound classes, including sound event separation, music source separation, and speech enhancement. The USS system achieves an average signal-to-distortion ratio improvement (SDRi) of 5.57 dB over 527 sound classes of AudioSet; 10.57 dB on the DCASE 2018 Task 2 dataset; 8.12 dB on the MUSDB18 dataset; an SDRi of 7.28 dB on the Slakh2100 dataset; and an SSNR of 9.00 dB on the voicebank-demand dataset. We release the source code at https://github.com/bytedance/uss

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis

    eess.AS 2025-09 unverdicted novelty 7.0 of 10

    DeepASA unifies source separation, dereverberation, SED, classification, and DoAE via object-oriented processing, chain-of-inference, and temporal coherence matching, reporting SOTA on ASA2, MC-FUSS, and STARSS23.

  2. MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    MAGE unifies text, visual, and audio-conditioned music generation and editing in one flow-based latent model with dynamic modality masking and cross-gated control.

  3. Class-Aware Permutation-Invariant Signal-to-Distortion Ratio for Semantic Segmentation of Sound Scene with Same-Class Sources

    eess.AS 2026-01 conditional novelty 6.0 of 10

    A class-aware permutation-invariant SDR loss and matching metric let label-queried source separation and DCASE-style S5 evaluation handle mixtures with multiple same-class sources.

  4. MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

    cs.SD 2026-04 unverdicted novelty 5.0 of 10

    A shared continuous-latent flow model generates music from text/vision or extracts a target source from a mixture via visual-audio alignment, gated modulation, and dynamic modality masking.

Pith tools