REVIEW 1 cited by
Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pre-training objective is designed for speech recognition. In this paper, we study which factor leads to the success of self-supervised learning on speaker-related tasks, e.g. speaker verification (SV), through a series of carefully designed experiments. Our empirical results on the Voxceleb-1 dataset suggest that the benefit of SSL to SV task is from a combination of mask speech prediction loss, data scale, and model size, while the SSL quantizer has a minor impact. We further employ the integrated gradients attribution method and loss landscape visualization to understand the effectiveness of self-supervised learning for speaker recognition performance.
Forward citations
Cited by 1 Pith paper
-
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
Layer-wise probing of MusicFM and MuQ shows acoustic-to-semantic feature progression across layers, and single-layer selection often outperforms all-layer aggregation on MIR tasks.
Discussion (0). Sign in to comment.