Pith. sign in

REVIEW 1 cited by

X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12766 v1 pith:VAGHI45R submitted 2020-10-24 eess.AS

classification eess.AS
keywords speechspeakerx-tasnettargetaudiosoutputresultsspeakers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extracting the speech of a target speaker from mixed audios, based on a reference speech from the target speaker, is a challenging yet powerful technology in speech processing. Recent studies of speaker-independent speech separation, such as TasNet, have shown promising results by applying deep neural networks over the time-domain waveform. Such separation neural network does not directly generate reliable and accurate output when target speakers are specified, because of the necessary prior on the number of speakers and the lack of robustness when dealing with audios with absent speakers. In this paper, we break these limitations by introducing a new speaker-aware speech masking method, called X-TaSNet. Our proposal adopts new strategies, including a distortion-based loss and corresponding alternating training scheme, to better address the robustness issue. X-TaSNet significantly enhances the extracted speech quality, doubling SDRi and SI-SNRi of the output speech audio over state-of-the-art voice filtering approach. X-TaSNet also improves the reliability of the results by improving the accuracy of speaker identity in the output audio to 95.4%, such that it returns silent audios in most cases when the target speaker is absent. These results demonstrate X-TaSNet moves one solid step towards more practical applications of speaker extraction technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion

    cs.SD 2024-11 conditional novelty 4.0 of 10

    X-CrossNet applies the CrossNet separation backbone to target speaker extraction with cross-attention speaker embedding fusion, reporting small improvements on WSJ0-2mix and WHAMR!.

Pith tools