Pith. sign in

REVIEW 2 cited by

Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.11129 v2 pith:TWTD3DMT submitted 2020-05-22 eess.AS cs.SD

classification eess.AScs.SD
keywords modelgenerativeglow-ttsmodelsmonotonicparallelspeechalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, text-to-speech (TTS) models such as FastSpeech and ParaNet have been proposed to generate mel-spectrograms from text in parallel. Despite the advantage, the parallel TTS models cannot be trained without guidance from autoregressive TTS models as their external aligners. In this work, we propose Glow-TTS, a flow-based generative model for parallel TTS that does not require any external aligner. By combining the properties of flows and dynamic programming, the proposed model searches for the most probable monotonic alignment between text and the latent representation of speech on its own. We demonstrate that enforcing hard monotonic alignments enables robust TTS, which generalizes to long utterances, and employing generative flows enables fast, diverse, and controllable speech synthesis. Glow-TTS obtains an order-of-magnitude speed-up over the autoregressive model, Tacotron 2, at synthesis with comparable speech quality. We further show that our model can be easily extended to a multi-speaker setting.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ALAS: An Automatic Latent Alignment Score for Audio Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ALAS is a reference-based score for audio-text alignment in speech LLMs, computed from frozen hidden states and a Whisper-derived alignment path, with no training or fitted classifier.

  2. Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages

    eess.AS 2025-07 conditional novelty 4.0 of 10

    In a five-language zero-shot TTS study, no single duration prediction strategy dominates: speaker-prompted durations help some languages, infilling durations help others, and results vary by metric.

Pith tools