REVIEW 10 cited by
LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces a new speech dataset called ``LibriTTS-R'' designed for text-to-speech (TTS) use. It is derived by applying speech restoration to the LibriTTS corpus, which consists of 585 hours of speech data at 24 kHz sampling rate from 2,456 speakers and the corresponding texts. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the sound quality improved. Experimental results show that the LibriTTS-R ground-truth samples showed significantly improved sound quality compared to those in LibriTTS. In addition, neural end-to-end TTS trained with LibriTTS-R achieved speech naturalness on par with that of the ground-truth samples. The corpus is freely available for download from \url{http://www.openslr.org/141/}.
Forward citations
Cited by 10 Pith papers
-
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
Evaluating voice anonymization at low false-positive rates reveals much stronger membership inference and attribute leakage than Equal Error Rate reports.
-
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.
-
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
SpeechRefiner, a conformer-based conditional flow matching model, improves SIGMOS perceptual quality scores on speech processed by various front-ends, including unseen systems.
-
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthet...
-
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Authors propose ESS-CLAP and RA-CLAP, contrastive speech-text models for emotional speaking style retrieval, evaluated on PromptSpeech, TextrolSpeech, and SpeechCraft.
-
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
A zero-shot TTS method that splits reference audio into ASR word segments, encodes local styles, and merges them via self-attention improves intelligibility and speaker similarity on unseen voices.
-
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
A three-stage training pipeline aligns face and text encoders to a pretrained speech-encoder space, then trains VITS on speech embeddings, and reports gains over single-modal baselines.
-
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
RT-VC converts a speaker's voice to a new target voice in real time on a CPU with 61.4ms latency, matching the quality of the current SOTA StreamVC.
-
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.
-
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
A coarse-to-fine multi-step prediction of hierarchical emotion labels in TTS yields marginal quality gains over the authors' single-step baseline.
Discussion (0). Sign in to comment.