Pith. sign in

REVIEW 2 cited by

Emphasis control for parallel neural TTS

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.03012 v2 pith:DQBJTBIV submitted 2021-10-06 eess.AS cs.CL

classification eess.AScs.CL
keywords emphasiscontrollatentneuralparallelspaceableduration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent parallel neural text-to-speech (TTS) synthesis methods are able to generate speech with high fidelity while maintaining high performance. However, these systems often lack control over the output prosody, thus restricting the semantic information conveyable for a given text. This paper proposes a hierarchical parallel neural TTS system for prosodic emphasis control by learning a latent space that directly corresponds to a change in emphasis. Three candidate features for the latent space are compared: 1) Variance of pitch and duration within words in a sentence, 2) Wavelet-based feature computed from pitch, energy, and duration, and 3) Learned combination of the two aforementioned approaches. At inference time, word-level prosodic emphasis is achieved by increasing the feature values of the latent space for the given words. Experiments show that all the proposed methods are able to achieve the perception of increased emphasis with little loss in overall quality. Moreover, emphasized utterances were preferred in a pairwise comparison test over the non-emphasized utterances, indicating promise for real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving French Synthetic Speech Quality via SSML Prosody Control

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Two fine-tuned LLMs predict SSML prosody tags that raise French TTS naturalness from a 3.20 to 3.87 MOS.

  2. Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts

    cs.CL 2025-09 conditional novelty 4.0 of 10

    AI voice interviewers are already fit for closed-ended surveys and partially for open-ended interviews, but transcription, emotion, and probing weaknesses limit qualitative use.

Pith tools