Pith. sign in

REVIEW 7 cited by

StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15439 v2 pith:QHQXLYSN submitted 2022-05-30 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords speechmodelparalleldiverseemotionalgenerativemodelsmonotonic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text-to-Speech (TTS) has recently seen great progress in synthesizing high-quality speech owing to the rapid development of parallel TTS systems, but producing speech with naturalistic prosodic variations, speaking styles and emotional tones remains challenging. Moreover, since duration and speech are generated separately, parallel TTS models still have problems finding the best monotonic alignments that are crucial for naturalistic speech synthesis. Here, we propose StyleTTS, a style-based generative model for parallel TTS that can synthesize diverse speech with natural prosody from a reference speech utterance. With novel Transferable Monotonic Aligner (TMA) and duration-invariant data augmentation schemes, our method significantly outperforms state-of-the-art models on both single and multi-speaker datasets in subjective tests of speech naturalness and speaker similarity. Through self-supervised learning of the speaking styles, our model can synthesize speech with the same prosodic and emotional tone as any given reference speech without the need for explicitly labeling these categories.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  2. RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching

    eess.AS 2025-06 conditional novelty 6.0 of 10

    RapFlow-TTS applies consistency flow matching to TTS and, together with adversarial and scheduling techniques, matches the naturalness of slower ODE-based TTS at only two synthesis steps.

  3. TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

    eess.AS 2025-05 conditional novelty 6.0 of 10

    TCSinger 2 generates zero-shot singing voices in nine languages with style transfer from audio prompts and multi-level style control from natural language prompts, using blurred boundary encoders, contrastive prompt a...

  4. Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Continuous autoregressive text-to-speech with a Gaussian-mixture codec matches or beats a discrete-codec VALL-E baseline with a fraction of the language model parameters.

  5. DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A sketch-conditioned diffusion TTS model that turns coarse user-drawn pitch and energy trends into natural, precisely controlled speech.

  6. ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis

    cs.CL 2024-12 conditional novelty 6.0 of 10

    ProsodyFM adds a phrase-break encoder, a break-duration predictor, and a bank of intonation-shape tokens to a flow-matching TTS backbone, improving prosody and intelligibility without explicit prosodic labels.

  7. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

Pith tools