Pith. sign in

REVIEW 3 cited by

Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10507 v1 pith:VFPOGT3D submitted 2024-06-15 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords finetuningmodelsspeechchildsfmsvariousbenchmarkdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech foundation models (SFMs) have achieved state-of-the-art results for various speech tasks in supervised (e.g. Whisper) or self-supervised systems (e.g. WavLM). However, the performance of SFMs for child ASR has not been systematically studied. In addition, there is no benchmark for child ASR with standard evaluations, making the comparisons of novel ideas difficult. In this paper, we initiate and present a comprehensive benchmark on several child speech databases based on various SFMs (Whisper, Wav2vec2.0, HuBERT, and WavLM). Moreover, we investigate finetuning strategies by comparing various data augmentation and parameter-efficient finetuning (PEFT) methods. We observe that the behaviors of these methods are different when the model size increases. For example, PEFT matches the performance of full finetuning for large models but worse for small models. To stabilize finetuning using augmented data, we propose a perturbation invariant finetuning (PIF) loss as a regularization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Multi-model ASR consensus (BEACON) curates 413 h of CHILDES with corrected timestamps; the 283 h ASR subset yields up to 19.5% relative WER reduction on four held-out child benchmarks.

  2. SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SimClass is a new 391-hour simulated classroom speech dataset with game-engine babble noise; ASR fine-tuning on it beats Librispeech and TEDLIUM on real classroom test sets.

  3. Robust fine-tuning of speech recognition models via model merging: application to disordered speech

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Merging multiple fine-tuned Whisper models reduces word error rate on dysarthric speech by 12-16% relative to standard fine-tuning, with gains on long audio and low-data settings.

Pith tools