Pith. sign in

REVIEW 1 cited by

An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.03648 v2 pith:YBZLYVR5 submitted 2020-08-09 eess.AS cs.SD

classification eess.AScs.SD
keywords conversionvoicespeakerspeechchallengesdeepidentitylearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech processing techniques, such as speech analysis, spectral conversion, prosody conversion, speaker characterization, and vocoding. With the recent advances in theory and practice, we are now able to produce human-like voice quality with high speaker similarity. In this paper, we provide a comprehensive overview of the state-of-the-art of voice conversion techniques and their performance evaluation methods from the statistical approaches to deep learning, and discuss their promise and limitations. We will also report the recent Voice Conversion Challenges (VCC), the performance of the current state of technology, and provide a summary of the available resources for voice conversion research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

    cs.SD 2025-01 reject novelty 5.0 of 10

    Stepback trains a voice converter with two decoders and a self-destructive loss to separate speaker identity from linguistic content, but the preprint contains no reported evaluation results.

Pith tools