PS-TTS and PS-Comet TTS use isochrony via language model paraphrasing plus phonetic synchronization with DTW on vowel distances to achieve better lip-sync and semantic preservation in automated dubbing than standard TTS or voice actors on tested language pairs.
arXiv:2203.11389 (2022)
3 Pith papers cite this work. Polarity classification is still indexing.
fields
eess.AS 3years
2026 3verdicts
UNVERDICTED 3representative citing papers
DNSMOS-C adds MOS-guided triplet contrastive loss to DNSMOS Pro for improved correlation, out-of-domain generalization, and emergent low-dimensional quality ordering in embeddings via unified training.
Voice range indicates TTS model capability with VITS highest, Glow-TTS best at soft phonation, and CPPs of 7-8 dB marking natural quality while values over 10 dB sound robotic.
citing papers explorer
-
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
PS-TTS and PS-Comet TTS use isochrony via language model paraphrasing plus phonetic synchronization with DTW on vowel distances to achieve better lip-sync and semantic preservation in automated dubbing than standard TTS or voice actors on tested language pairs.
-
DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning
DNSMOS-C adds MOS-guided triplet contrastive loss to DNSMOS Pro for improved correlation, out-of-domain generalization, and emergent low-dimensional quality ordering in embeddings via unified training.
-
Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment
Voice range indicates TTS model capability with VITS highest, Glow-TTS best at soft phonation, and CPPs of 7-8 dB marking natural quality while values over 10 dB sound robotic.