Pith. sign in

REVIEW 6 cited by

TSELM: Target Speaker Extraction using Discrete Tokens and Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07841 v3 pith:HV2IICLO submitted 2024-09-12 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords tokenstselmmodelslanguageresultsspeakertargetaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate target speaker information. Language models are employed to capture the sequence dependencies, while a scalable HiFi-GAN is used to reconstruct the audio from the tokens. By applying a cross-entropy loss, TSELM models the probability distribution of output tokens, thus converting the complex regression problem of audio generation into a classification task. Experimental results show that TSELM achieves excellent results in speech quality and comparable results in speech intelligibility.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

    cs.SD 2025-10 conditional novelty 6.0 of 10

    A 63M-parameter decoder-only LM, UniSE, unifies speech restoration, target speaker extraction, and speech separation by generating BiCodec discrete tokens under task-specific prompts.

  2. FlowTSE: Target Speaker Extraction with Flow Matching

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Conditional flow matching on mel-spectrograms with a phase-conditioned vocoder matches or beats published TSE baselines on Libri2Mix.

  3. GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

    eess.AS 2025-12 conditional novelty 5.0 of 10

    A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.

  4. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

    cs.SD 2025-02 conditional novelty 5.0 of 10

    A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.

  5. Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

    eess.AS 2025-01 conditional novelty 5.0 of 10

    A generative target-speech extractor trained jointly with a Whisper-based transcript-prediction loss achieves better intelligibility and competitive quality on Libri2Mix and WSJ0-2mix than discrete-token and mask-base...

  6. Overview of the Amphion Toolkit (v0.2)

    cs.SD 2025-01 conditional novelty 4.0 of 10

    Amphion v0.2 is an open-source toolkit for audio, music, and speech generation, adding a 101K-hour multilingual dataset, processing pipelines, and pretrained models.

Pith tools