Pith. sign in

REVIEW 1 cited by

Emotional Prosody Control for Speech Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.04730 v1 pith:ZL2FLOYG submitted 2021-11-07 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords speechemotioncontrolsystemtextcontinuousemotionalprosody
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from prosody sequences in training data or transferred from a source style. We propose a text to speech(TTS) system, where a user can choose the emotion of generated speech from a continuous and meaningful emotion space (Arousal-Valence space). The proposed TTS system can generate speech from the text in any speaker's style, with fine control of emotion. We show that the system works on emotion unseen during training and can scale to previously unseen speakers given his/her speech sample. Our work expands the horizon of the state-of-the-art FastSpeech2 backbone to a multi-speaker setting and gives it much-coveted continuous (and interpretable) affective control, without any observable degradation in the quality of the synthesized speech.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Fast-VGAN is a lightweight GAN-based voice converter that explicitly controls F0, phoneme timing, and intensity, achieving near-perfect intelligibility and competitive speaker similarity on a small test set.

Pith tools