A LoRA-tuned Whisper encoder with a factorized tier/family speaker token and a temporal smoothing loss improves multi-tier audio tagging for daylong infant recordings.
Ssast: Self- supervised audio spectrogram transformer,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
A LoRA-tuned Whisper encoder with a factorized tier/family speaker token and a temporal smoothing loss improves multi-tier audio tagging for daylong infant recordings.