Pith. sign in

REVIEW 10 cited by

LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18802 v1 pith:IT3N7G2W submitted 2023-05-30 eess.AS cs.SD

classification eess.AScs.SD
keywords libritts-rspeechcorpuslibrittssamplesground-truthimprovedquality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces a new speech dataset called ``LibriTTS-R'' designed for text-to-speech (TTS) use. It is derived by applying speech restoration to the LibriTTS corpus, which consists of 585 hours of speech data at 24 kHz sampling rate from 2,456 speakers and the corresponding texts. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the sound quality improved. Experimental results show that the LibriTTS-R ground-truth samples showed significantly improved sound quality compared to those in LibriTTS. In addition, neural end-to-end TTS trained with LibriTTS-R achieved speech naturalness on par with that of the ground-truth samples. The corpus is freely available for download from \url{http://www.openslr.org/141/}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks

    cs.CR 2025-09 conditional novelty 6.0 of 10

    Evaluating voice anonymization at low false-positive rates reveals much stronger membership inference and attribute leakage than Equal Error Rate reports.

  2. Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.

  3. SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms

    eess.AS 2025-06 conditional novelty 6.0 of 10

    SpeechRefiner, a conformer-based conditional flow matching model, improves SIGMOS perceptual quality scores on speech processed by various front-ends, including unseen systems.

  4. In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion

    cs.SD 2025-06 conditional novelty 6.0 of 10

    TES-VC can change both the speaker's voice and the acoustic environment of an audio clip from text prompts while preserving the words, using retrieval of known timbre embeddings and latent diffusion trained on synthet...

  5. RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Authors propose ESS-CLAP and RA-CLAP, contrastive speech-text models for emotional speaking style retrieval, evaluated on PromptSpeech, TextrolSpeech, and SpeechCraft.

  6. GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A zero-shot TTS method that splits reference audio into ASR word segments, encodes local styles, and merges them via self-attention improves intelligibility and speaker similarity on unseen voices.

  7. A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A three-stage training pipeline aligns face and text encoders to a pretrained speech-encoder space, then trains VITS on speech embeddings, and reports gains over single-modal baselines.

  8. RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding

    eess.AS 2025-06 conditional novelty 5.0 of 10

    RT-VC converts a speaker's voice to a new target voice in real time on a CPU with 61.4ms latency, matching the quality of the current SOTA StreamVC.

  9. CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

    cs.SD 2025-05 reject novelty 5.0 of 10

    A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.

  10. Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis

    cs.SD 2025-07 conditional novelty 4.0 of 10

    A coarse-to-fine multi-step prediction of hierarchical emotion labels in TTS yields marginal quality gains over the authors' single-step baseline.

Pith tools