A LoRA-tuned Whisper encoder with a factorized tier/family speaker token and a temporal smoothing loss improves multi-tier audio tagging for daylong infant recordings.
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Speech foundation models, trained on vast datasets, have opened unique opportunities in addressing challenging low-resource speech understanding, such as child speech. In this work, we explore the capabilities of speech foundation models on child-adult speaker diarization. We show that exemplary foundation models can achieve 39.5% and 62.3% relative reductions in Diarization Error Rate and Speaker Confusion Rate, respectively, compared to previous speaker diarization methods. In addition, we benchmark and evaluate the speaker diarization results of the speech foundation models with varying the input audio window size, speaker demographics, and training data ratio. Our results highlight promising pathways for understanding and adopting speech foundation models to facilitate child speech understanding.
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
A LoRA-tuned Whisper encoder with a factorized tier/family speaker token and a temporal smoothing loss improves multi-tier audio tagging for daylong infant recordings.