Pith. sign in

REVIEW 3 cited by

Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07284 v4 pith:LGQHN66O submitted 2023-10-11 eess.AS cs.CL

classification eess.AScs.CL
keywords cuesextractionspeakercocktailpartyaloneconcernsdemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans can easily isolate a single speaker from a complex acoustic environment, a capability referred to as the "Cocktail Party Effect." However, replicating this ability has been a significant challenge in the field of target speaker extraction (TSE). Traditional TSE approaches predominantly rely on voiceprints, which raise privacy concerns and face issues related to the quality and availability of enrollment samples, as well as intra-speaker variability. To address these issues, this work introduces a novel text-guided TSE paradigm named LLM-TSE. In this paradigm, a state-of-the-art large language model, LLaMA 2, processes typed text input from users to extract semantic cues. We demonstrate that textual descriptions alone can effectively serve as cues for extraction, thus addressing privacy concerns and reducing dependency on voiceprints. Furthermore, our approach offers flexibility by allowing the user to specify the extraction or suppression of a speaker and enhances robustness against intra-speaker variability by incorporating context-dependent textual information. Experimental results show competitive performance with text-based cues alone and demonstrate the effectiveness of using text as a task selector. Additionally, they achieve a new state-of-the-art when combining text-based cues with pre-registered cues. This work represents the first integration of LLMs with TSE, potentially establishing a new benchmark in solving the cocktail party problem and expanding the scope of TSE applications by providing a versatile, privacy-conscious solution.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Inter-speaker relative cues, expressed as text differences like louder, faster, or female, guide target speech extraction and outperform random cue subsets on a new five-language dataset.

  2. Detect, Attend and Extract: Keyword Guided Target Speaker Extraction

    eess.AS 2026-02 conditional novelty 5.0 of 10

    Keyword-guided target speaker extraction (DAE-TSE) uses a few words spoken by the target to detect, localize, and extract that speaker's full utterance from a two-speaker mixture.

  3. M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

    eess.AS 2025-05 conditional novelty 5.0 of 10

    M3ANet aligns EEG and speech representations with InfoNCE contrastive learning and encodes speech with multi-scale convolutions plus GroupMamba, improving brain-assisted target speaker extraction on three datasets.

Pith tools