Pith. sign in

REVIEW 1 cited by

Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.08352 v1 pith:DB77RA3K submitted 2021-06-15 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords acousticmodelspeechtextfeaturespredictedvariationcontrol
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data is to provide acoustic information as an additional learning signal. When generating speech, modifying this acoustic information enables multiple distinct renditions of a text to be produced. Since much of the unexplained variation is in the prosody, we propose a model that generates speech explicitly conditioned on the three primary acoustic correlates of prosody: $F_{0}$, energy and duration. The model is flexible about how the values of these features are specified: they can be externally provided, or predicted from text, or predicted then subsequently modified. Compared to a model that employs a variational auto-encoder to learn unsupervised latent features, our model provides more interpretable, temporally-precise, and disentangled control. When automatically predicting the acoustic features from text, it generates speech that is more natural than that from a Tacotron 2 model with reference encoder. Subsequent human-in-the-loop modification of the predicted acoustic features can significantly further increase naturalness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Counterfactual gradient edits to a pretrained TTS model's encoder activations can control prosody and correct mispronunciations at inference time, at least on Tacotron 2.

Pith tools