Pith. sign in

REVIEW 3 cited by

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15799 v1 pith:HHBY2SZJ submitted 2024-09-24 eess.AS cs.SD

classification eess.AScs.SD
keywords speakertargetwesepspeechtoolkitapplicationsdataextraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to its potential for various applications such as user-customized interfaces and hearing aids, or as a crutial front-end processing technologies for subsequential tasks such as speech recognition and speaker recongtion. However, there are currently few open-source toolkits or available pre-trained models for off-the-shelf usage. In this work, we introduce WeSep, a toolkit designed for research and practical applications in TSE. WeSep is featured with flexible target speaker modeling, scalable data management, effective on-the-fly data simulation, structured recipes and deployment support. The toolkit is publicly avaliable at \url{https://github.com/wenet-e2e/WeSep.}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A cascaded pipeline of audio compression, latent diffusion extraction, and generative correction achieves state-of-the-art target speech extraction quality and intelligibility on Libri2Mix and out-of-domain data.

  2. The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

    eess.AS 2026-07 conditional novelty 4.0 of 10

    A cascaded smart-glasses TSA-ASR system with a dominant-speaker overlap fallback achieved 7.10% tcpCER on two-person dialogues and 34.04% on multi-party meetings, ranking second on the meeting track.

  3. Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling

    cs.SD 2025-07 conditional novelty 4.0 of 10

    A centroid-based speaker consistency loss plus conditional loss suppression improves target speaker extraction quality and speaker similarity across multiple backbone models.

Pith tools