Pith. sign in

REVIEW 1 cited by

Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.10288 v2 pith:N4F2ZOOB submitted 2019-10-23 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords attentionmechanismslocation-relativetextutterancesabilityadditivealignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text. We show that these failures can be addressed using simple location-relative attention mechanisms that do away with content-based query/key comparisons. We compare two families of attention mechanisms: location-relative GMM-based mechanisms and additive energy-based mechanisms. We suggest simple modifications to GMM-based attention that allow it to align quickly and consistently during training, and introduce a new location-relative attention mechanism to the additive energy-based family, called Dynamic Convolution Attention (DCA). We compare the various mechanisms in terms of alignment speed and consistency during training, naturalness, and ability to generalize to long utterances, and conclude that GMM attention and DCA can generalize to very long utterances, while preserving naturalness for shorter, in-domain utterances.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

    cs.MM 2025-05 conditional novelty 5.0 of 10

    A movie dubbing system that uses a vision-language model to extract scene type and speaker attributes from silent video, then feeds those attributes as extra conditions into a diffusion-based speech generator, support...

Pith tools