Pith. sign in

REVIEW 2 cited by

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.07025 v3 pith:FDCAXRP3 submitted 2020-05-13 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords conversionemotionalemotionspeaker-independentvoiceanyoneconvertframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried out under the assumption that emotion is speaker-dependent. We consider that there is a common code between speakers for emotional expression in a spoken language, therefore, a speaker-independent mapping between emotional states is possible. In this paper, we propose a speaker-independent emotional voice conversion framework, that can convert anyone's emotion without the need for parallel data. We propose a VAW-GAN based encoder-decoder structure to learn the spectrum and prosody mapping. We perform prosody conversion by using continuous wavelet transform (CWT) to model the temporal dependencies. We also investigate the use of F0 as an additional input to the decoder to improve emotion conversion performance. Experiments show that the proposed speaker-independent framework achieves competitive results for both seen and unseen speakers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion

    cs.SD 2025-06 conditional novelty 6.0 of 10

    TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthet...

  2. Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Maestro-EVC independently controls content, speaker, and emotion in voice conversion using separate references and explicit prosody modeling, outperforming StyleVC and ZEST on emotion similarity and prosody.

Pith tools