Pith. sign in

REVIEW 2 cited by

Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.10326 v2 pith:FRN7CKOE submitted 2021-10-20 eess.AS cs.SD

classification eess.AScs.SD
keywords emotionalstyleidentityspeakerconversionexpressivespeakersstylevc
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional style for different speakers. Inspired by the recent success of speaker disentanglement with variational autoencoder (VAE), we propose an any-to-any expressive voice conversion framework, that is called StyleVC. StyleVC is designed to disentangle linguistic content, speaker identity, pitch, and emotional style information. We study the use of style encoder to model emotional style explicitly. At run-time, StyleVC converts both speaker identity and emotional style for arbitrary speakers. Experiments validate the effectiveness of our proposed framework in both objective and subjective evaluations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Fast-VGAN is a lightweight GAN-based voice converter that explicitly controls F0, phoneme timing, and intensity, achieving near-perfect intelligibility and competitive speaker similarity on a small test set.

  2. Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Maestro-EVC independently controls content, speaker, and emotion in voice conversion using separate references and explicit prosody modeling, outperforming StyleVC and ZEST on emotion similarity and prosody.

Pith tools