Pith. sign in

REVIEW 23 cited by

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.11567 v2 pith:R4ZFCRJR submitted 2020-10-22 cs.SD eess.AS

classification cs.SDeess.AS
keywords multi-speakercorpussystemsynthesisaishell-3mandarinsimilarityspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present AISHELL-3, a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems. The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in Chinese character-level and pinyin-level are provided along with the recordings. We present a baseline system that uses AISHELL-3 for multi-speaker Madarin speech synthesis. The multi-speaker speech synthesis system is an extension on Tacotron-2 where a speaker verification model and a corresponding loss regarding voice similarity are incorporated as the feedback constraint. We aim to use the presented corpus to build a robust synthesis model that is able to achieve zero-shot voice cloning. The system trained on this dataset also generalizes well on speakers that are never seen in the training process. Objective evaluation results from our experiments show that the proposed multi-speaker synthesis system achieves high voice similarity concerning both speaker embedding similarity and equal error rate measurement. The dataset, baseline system code and generated samples are available online.

Discussion (0). Sign in to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

    eess.AS 2026-06 unverdicted novelty 7.0 of 10

    SpeechEditBench provides seven atomic editing tasks, compositional multi-operation instructions, and an anchor-based protocol yielding target success, preservation success, and joint success metrics; evaluations show ...

  2. AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    AffectCodec is an emotion-guided neural speech codec that preserves emotional cues during quantization while maintaining semantic fidelity and prosodic naturalness.

  3. VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    VITA-QinYu is the first expressive end-to-end spoken language model supporting role-playing and singing alongside conversation, trained on 15.8K hours of data and outperforming prior models on expressiveness and conve...

  4. SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation

    cs.CL 2026-01 unverdicted novelty 7.0 of 10

    SpeechMedAssist adapts SpeechLMs for medical consultations via two-stage training (text knowledge injection then limited speech re-alignment) using 10k synthesized samples and outperforms baselines in effectiveness an...

  5. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  6. DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    DisSpeech synthesizes controllable stuttered Mandarin speech via discrete tokens and stuttering event labels to augment ASR datasets, improving recognition to 4.19% CER on stuttered tasks with minimal impact on fluent speech.

  7. UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    UR-BERT scales multilingual TTS encoders to 495 languages via Romanization unification and speech token prediction, outperforming baselines with better generalization.

  8. Benchmarking Neural Speech Compression from a Rate-Distortion Perspective

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    ECC integrates hyperprior side information, channel-wise context, latent residual prediction, temporal modeling, and entropy skip into a learned entropy model, yielding 39.9% and 76.3% average BD-rate reductions on Vi...

  9. CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    CleanCodec reframes audio tokenization as a selective information bottleneck to encode only perceptually important features at 12.5 tokens per second, outperforming prior codecs in efficiency, speaker similarity, and ...

  10. Aliasing-Free Neural Audio Synthesis

    cs.SD 2025-12 conditional novelty 6.0 of 10

    Pupu-Vocoder and Pupu-Codec integrate differentiable anti-aliasing into neural audio models to eliminate aliasing artifacts from non-linear activations and upsampling, yielding better results on music and singing voice.

  11. Aliasing-Free Neural Audio Synthesis

    cs.SD 2025-12 conditional novelty 6.0 of 10

    Pupu-Vocoder and Pupu-Codec use a closed-form anti-aliased SnakeBeta activation and resampling-based upsampling to reduce aliasing and improve singing, music, and audio synthesis.

  12. CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

    cs.SD 2025-09 unverdicted novelty 6.0 of 10

    CoMelSinger introduces a discrete token-based zero-shot SVS framework on MaskGCT with coarse-to-fine contrastive learning and an SVT module to improve melody control and reduce prosody leakage.

  13. DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DeCodec learns a single neural codec that disentangles speech, background sound, semantic content, and paralinguistic style into orthogonal quantized streams, enabling reconstruction, enhancement, voice conversion, AS...

  14. SwiftF0: Fast and Accurate Monophonic Pitch Detection

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SwiftF0 estimates monophonic pitch from a compact STFT-CNN, reporting better accuracy than CREPE under 10 dB noise at 42x lower CPU cost, alongside a new synthetic speech dataset and a six-component evaluation metric.

  15. UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling

    eess.AS 2025-08 conditional novelty 6.0 of 10

    UniFlow unifies four speech front-end tasks in one continuous-latent generative model with task-ID conditioning and reports competitive, but not uniformly superior, benchmark scores.

  16. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  17. ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    ContextCodec uses a dual-branch encoder with CLIP-style contrastive training on phoneme-aligned context features plus autoregressive refinement to improve quality-intelligibility at bitrates down to 500 bps.

  18. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 accept novelty 5.0 of 10

    MLLM-enabled video translation is usefully framed as three roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than a cascade of ASR, MT, TTS, and lip-sync.

  19. Schr\"odinger Bridge Mamba for One-Step Speech Enhancement

    cs.SD 2025-10 conditional novelty 5.0 of 10

    A Mamba-based speech enhancer trained with Schrödinger Bridge objectives produces strong denoising and dereverberation in one inference step with a low real-time factor.

  20. Kimi-Audio Technical Report

    eess.AS 2025-04 unverdicted novelty 5.0 of 10

    Kimi-Audio is an open-source audio foundation model that achieves state-of-the-art results on speech recognition, audio understanding, question answering, and conversation after pre-training on more than 13 million ho...

  21. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

  22. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  23. AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

    cs.SD 2026-04 unverdicted novelty 3.0 of 10

    AT-ADD introduces standardized tracks and datasets for evaluating audio deepfake detectors on speech under real-world conditions and on diverse unknown audio types to promote generalization beyond speech-centric methods.

Pith tools