Pith. sign in

REVIEW 1 cited by

Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.00198 v5 pith:AQDWY4QR submitted 2020-02-01 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords conversionemotionaldatadifferentprosodyspeechtransformmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content. Many studies require parallel speech data between different emotional patterns, which is not practical in real life. Moreover, they often model the conversion of fundamental frequency (F0) with a simple linear transform. As F0 is a key aspect of intonation that is hierarchical in nature, we believe that it is more adequate to model F0 in different temporal scales by using wavelet transform. We propose a CycleGAN network to find an optimal pseudo pair from non-parallel training data by learning forward and inverse mappings simultaneously using adversarial and cycle-consistency losses. We also study the use of continuous wavelet transform (CWT) to decompose F0 into ten temporal scales, that describes speech prosody at different time resolution, for effective F0 conversion. Experimental results show that our proposed framework outperforms the baselines both in objective and subjective evaluations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

    eess.AS 2025-05 conditional novelty 6.0 of 10

    EmoCorrector retrieves emotional speech samples matching the edited text and uses them to post-correct the emotion of TSE output, supported by the new synthetic ECD-TSE dataset.

Pith tools