Pith. sign in

REVIEW 1 cited by

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18584 v3 pith:G34LGQL4 submitted 2024-09-27 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechdatasetchildrenmodelsmandarinyoungagedchildmandarin
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research. The dataset is now open-source and freely available for all academic purposes on https://github.com/flageval-baai/ChildMandarin.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EchoVoices: Preserving Generational Voices and Memories for Seniors and Children

    cs.SD 2025-07 conditional novelty 5.0 of 10

    EchoVoices combines k-NN-enhanced Whisper ASR, two-stage VITS TTS, and an LLM/RAG persona agent to create digital personas for seniors and children, reporting CER and similarity gains on two Mandarin datasets.

Pith tools