Pith. sign in

REVIEW 19 cited by

A Survey on Neural Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.15561 v3 pith:E4SNF44Q submitted 2021-06-29 eess.AS cs.CLcs.LGcs.MMcs.SD

classification eess.AScs.CLcs.LGcs.MMcs.SD
keywords speechneuralresearchsurveytextfutureincludingindustry
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the development of deep learning and artificial intelligence, neural network-based TTS has significantly improved the quality of synthesized speech in recent years. In this paper, we conduct a comprehensive survey on neural TTS, aiming to provide a good understanding of current research and future trends. We focus on the key components in neural TTS, including text analysis, acoustic models and vocoders, and several advanced topics, including fast TTS, low-resource TTS, robust TTS, expressive TTS, and adaptive TTS, etc. We further summarize resources related to TTS (e.g., datasets, opensource implementations) and discuss future research directions. This survey can serve both academic researchers and industry practitioners working on TTS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    Bagpiper-TTS uses natural language prompts and intent reasoning to derive rich captions that guide a single model for universal speech synthesis across classical TTS, multi-talker, singing, and role-play tasks.

  2. N\"ushuVoice: Reviving the Voice of Endangered N\"ushu with Pitch-Aware Text-to-Speech

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    NüshuVoice releases the first sentence-level Nüshu TTS dataset and shows that an F0-conditioned VITS model using five-level pitch notation outperforms baselines on spectral fidelity, pitch accuracy, and intelligibility.

  3. Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages

    eess.AS 2026-04 unverdicted novelty 7.0 of 10

    Introduces the Indic-CodecFake dataset for Indic codec deepfakes and SATYAM, a novel hyperbolic ALM that outperforms baselines through dual-stage semantic-prosodic fusion using Bhattacharya distance.

  4. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

    cs.CL 2023-01 unverdicted novelty 7.0 of 10

    VALL-E is a neural codec language model trained on 60K hours of speech that performs zero-shot TTS, synthesizing natural speech that matches an unseen speaker's voice, emotion, and environment from a 3-second prompt.

  5. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.

  6. DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    DisSpeech synthesizes controllable stuttered Mandarin speech via discrete tokens and stuttering event labels to augment ASR datasets, improving recognition to 4.19% CER on stuttered tasks with minimal impact on fluent speech.

  7. Asymmetric Phase Coding Audio Watermarking

    cs.CR 2026-05 unverdicted novelty 6.0 of 10

    APC embeds compact Ed25519 signatures into audio phase data with error correction to achieve 97.5-98.3% cryptographic verification under eight attack types at mean PESQ 3.02.

  8. Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative

    cs.SD 2026-03 accept novelty 6.0 of 10

    RuASD is a comprehensive Russian speech anti-spoofing dataset featuring 37 synthesis systems and a robustness evaluation pipeline for real-world channel distortions.

  9. NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

    cs.SD 2025-08 unverdicted novelty 6.0 of 10

    A Mandarin speech pipeline that tags non-word vocalizations (laughter, breath, filled pauses) for ASR and controls their generation in TTS, backed by a claimed 573-hour, 174,179-utterance word-level annotated corpus.

  10. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  11. Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

    cs.CL 2026-08 conditional novelty 5.0 of 10

    No single Urdu TTS system wins across all metrics, and the system listeners liked most, Google Gemini, was the furthest from reference audio on objective acoustic scores.

  12. Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A lightweight face adapter plus soft-tuning aligns face embeddings to a frozen StyleTTS 2 style space, yielding natural zero-shot face-to-speech and language-agnostic transfer to Spanish.

  13. Position: Towards Responsible Evaluation for Text-to-Speech

    eess.AS 2025-10 conditional novelty 5.0 of 10

    A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.

  14. Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model

    eess.AS 2025-08 conditional novelty 5.0 of 10

    Speaker-conditioned phrasing with phoneme-level PLMs (MP BERT) improves pause prediction from F0.5 0.3719 to 0.4991, and a few-shot adapter generalizes to unseen speakers.

  15. One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

    eess.AS 2026-04 unverdicted novelty 4.0 of 10

    A system based on OmniVoice with multi-model ensemble distillation for fine-tuning shows consistent gains in intelligibility metrics while keeping speaker similarity for cross-lingual scientific speech.

  16. XR-CareerAssist: An Immersive Platform for Personalised Career Guidance Leveraging Extended Reality and Multimodal AI

    cs.CE 2026-04 unverdicted novelty 4.0 of 10

    XR-CareerAssist fuses XR and five AI modules into a Unity-based immersive platform for multilingual, personalized career guidance via 3D avatars and dynamic Sankey diagrams, reporting 78.3% user satisfaction in a 23-p...

  17. Marco-Voice Technical Report

    cs.CL 2025-08 reject novelty 4.0 of 10

    Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...

  18. One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

    eess.AS 2026-04 unverdicted novelty 3.0 of 10

    Authors submit a cross-lingual voice cloning system to IWSLT 2026 using OmniVoice fine-tuned on ensemble-distilled synthetic data, reporting gains in WER, CER, and speaker similarity for scientific texts in three languages.

  19. Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment

    eess.AS 2026-04 unverdicted novelty 3.0 of 10

    Voice range indicates TTS model capability with VITS highest, Glow-TTS best at soft phonation, and CPPs of 7-8 dB marking natural quality while values over 10 dB sound robotic.

Pith tools