Pith. sign in

REVIEW 2 cited by

Text-Driven Separation of Arbitrary Sounds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.05738 v1 pith:SO32EQ47 submitted 2022-04-12 eess.AS cs.SD

classification eess.AScs.SD
keywords audiomodelsoundwordssourceapproacharbitraryclipconditioned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is achieved by combining two distinct models. The first model, SoundWords, is trained to jointly embed both an audio clip and its textual description to the same embedding in a shared representation. The second model, SoundFilter, takes a mixed source audio clip as an input and separates it based on a conditioning vector from the shared text-audio representation defined by SoundWords, making the model agnostic to the conditioning modality. Evaluating on multiple datasets, we show that our approach can achieve an SI-SDR of 9.1 dB for mixtures of two arbitrary sounds when conditioned on text and 10.1 dB when conditioned on audio. We also show that SoundWords is effective at learning co-embeddings and that our multi-modal training approach improves the performance of SoundFilter.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fine-grained Soundscape Control for Augmented Hearing

    cs.SD 2026-02 conditional novelty 6.0 of 10

    Aurchestra enables real-time, on-device per-class sound extraction and volume control for up to five simultaneous sound classes on hearables.

  2. SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A U-Net conditioned on target direction and a text embedding extracts the target ambisonic sound field from mixtures, outperforming beamforming baselines and showing the largest semantic benefit when a secondary sourc...

Pith tools