Inter-speaker relative cues, expressed as text differences like louder, faster, or female, guide target speech extraction and outperform random cue subsets on a new five-language dataset.
Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In daily life, we encounter a variety of sounds, both desirable and undesirable, with limited control over their presence and volume. Our work introduces "Listen, Chat, and Remix" (LCR), a novel multimodal sound remixer that controls each sound source in a mixture based on user-provided text instructions. LCR distinguishes itself with a user-friendly text interface and its unique ability to remix multiple sound sources simultaneously within a mixture, without needing to separate them. Users input open-vocabulary text prompts, which are interpreted by a large language model to create a semantic filter for remixing the sound mixture. The system then decomposes the mixture into its components, applies the semantic filter, and reassembles filtered components back to the desired output. We developed a 160-hour dataset with over 100k mixtures, including speech and various audio sources, along with text prompts for diverse remixing tasks including extraction, removal, and volume control of single or multiple sources. Our experiments demonstrate significant improvements in signal quality across all remixing tasks and robust performance in zero-shot scenarios with varying numbers and types of sound sources. An audio demo is available at: https://listenchatremix.github.io/demo.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
Inter-speaker relative cues, expressed as text differences like louder, faster, or female, guide target speech extraction and outperform random cue subsets on a new five-language dataset.