Pith. sign in

REVIEW 11 cited by

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05763 v3 pith:MCVQJPPX submitted 2024-06-09 eess.AS

classification eess.AS
keywords corpuswenetspeech4ttsdatasegmentsystemsbenchmarklargemandarin
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus derived from the open-sourced WenetSpeech dataset. Tailored for the text-to-speech tasks, we refined WenetSpeech by adjusting segment boundaries, enhancing the audio quality, and eliminating speaker mixing within each segment. Following a more accurate transcription process and quality-based data filtering process, the obtained WenetSpeech4TTS corpus contains $12,800$ hours of paired audio-text data. Furthermore, we have created subsets of varying sizes, categorized by segment quality scores to allow for TTS model training and fine-tuning. VALL-E and NaturalSpeech 2 systems are trained and fine-tuned on these subsets to validate the usability of WenetSpeech4TTS, establishing baselines on benchmark for fair comparison of TTS systems. The corpus and corresponding benchmarks are publicly available on huggingface.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

    eess.AS 2026-08 conditional novelty 7.0 of 10

    SemBridge supervises autoregressive states with discrete semantic tokens during training, improving content fidelity of continuous-latent speech generation without changing inference.

  2. Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A jointly trained speech tokenizer and flow-matching decoder achieve strong zero-shot TTS and voice conversion on English and Mandarin, with WER below ground truth in several tests.

  3. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  4. FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    FlowTTS-GRPO fine-tunes open-source flow-matching TTS models with multi-objective online RL via ODE-to-SDE conversion, improving speaker similarity and quality on CosyVoice 3.0 and F5-TTS.

  5. MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MFLA adds finite look-ahead attention plus a CIF-based token counter to Whisper, enabling streaming recognition with a wait-k latency-quality trade-off.

  6. Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper releases Teochew-Wild, the first publicly available Teochew speech corpus with orthographic and pinyin annotations, and shows it supports ASR and TTS training.

  7. TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

    eess.AS 2024-12 conditional novelty 6.0 of 10

    TouchASP trains a single elastic mixture-of-experts ASR model on 1M hours of partly pseudo-labeled audio and reports SpeechIO CER dropping from 4.98% to 2.45% while adding multi-task perception.

  8. Adaptive Duration Model for Text Speech Alignment

    cs.SD 2025-07 conditional novelty 5.0 of 10

    DurFormer, an adaptive duration prediction model with speed, scene, and semantic conditioning, improves phoneme-level duration accuracy and lowers word error rate in Mandarin text-to-speech.

  9. UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

    cs.SD 2025-05 conditional novelty 5.0 of 10

    The authors propose DistilCodec, a 32,768-code single-codebook audio codec, and UniTTS, a Qwen2.5-7B TTS model trained with audio, text, and cross-modal autoregressive tasks on interleaved prompts.

  10. DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue

    cs.CL 2025-04 conditional novelty 5.0 of 10

    A three-agent script-writer, synthesizer, and critic pipeline generates a bilingual multi-party speech dataset whose quality matches manually assembled datasets in TTS training benchmarks.

  11. The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024

    cs.SD 2024-12 conditional novelty 4.0 of 10

    A codec language model with delay-pattern decoding, classifier-free guidance, and spontaneous-data fine-tuning achieved the top naturalness score in the CoVoC 2024 challenge.

Pith tools