Pith. sign in

REVIEW 1 cited by

Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.02812 v2 pith:5NB5BQBE submitted 2018-04-09 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords speakervoiceconversionlatenttargetdataparallelrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for each target speaker. In this paper, we propose an adversarial learning framework for voice conversion, with which a single model can be trained to convert the voice to many different speakers, all without parallel data, by separating the speaker characteristics from the linguistic content in speech signals. An autoencoder is first trained to extract speaker-independent latent representations and speaker embedding separately using another auxiliary speaker classifier to regularize the latent representation. The decoder then takes the speaker-independent latent representation and the target speaker embedding as the input to generate the voice of the target speaker with the linguistic content of the source utterance. The quality of decoder output is further improved by patching with the residual signal produced by another pair of generator and discriminator. A target speaker set size of 20 was tested in the preliminary experiments, and very good voice quality was obtained. Conventional voice conversion metrics are reported. We also show that the speaker information has been properly reduced from the latent representations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

    cs.SD 2025-01 reject novelty 5.0 of 10

    Stepback trains a voice converter with two decoders and a self-destructive loss to separate speaker identity from linguistic content, but the preprint contains no reported evaluation results.

Pith tools