Pith. sign in

REVIEW 1 cited by

Exploring Voice Conversion based Data Augmentation in Text-Dependent Speaker Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.10710 v1 pith:6FNNX3FV submitted 2020-11-21 cs.SD eess.AS

classification cs.SDeess.AS
keywords datatext-dependentconversionspeakertrainingverificationvoiceaugmentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we focus on improving the performance of the text-dependent speaker verification system in the scenario of limited training data. The speaker verification system deep learning based text-dependent generally needs a large scale text-dependent training data set which could be labor and cost expensive, especially for customized new wake-up words. In recent studies, voice conversion systems that can generate high quality synthesized speech of seen and unseen speakers have been proposed. Inspired by those works, we adopt two different voice conversion methods as well as the very simple re-sampling approach to generate new text-dependent speech samples for data augmentation purposes. Experimental results show that the proposed method significantly improves the Equal Error Rare performance from 6.51% to 4.51% in the scenario of limited training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Fusing zero-shot TTS-generated speech embeddings with original short-utterance embeddings reduces speaker-verification EER by 10-16% on VoxCeleb1 without retraining.

Pith tools