On the NIST SRE24 audio track, a ResNet152 pre-trained on 8kHz/GSM-augmented VoxBlink2 and fine-tuned on telephone speech with 40s segments achieved the best EER and Cprimary among the tested frontends.
Analysis of ABC Frontend Audio Systems for the NIST-SRE24
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of telephone recordings for training (fixed) or adding publicly available data (open condition). Under these constraints, we develop the best possible speaker embedding extractors for the pre-dominant conversational telephone speech (CTS) domain. We explored architectures based on ResNet with different pooling mechanisms, recently introduced ReDimNet architecture, as well as a system based on the XLS-R model, which represents the family of large pre-trained self-supervised models. In open condition, we train on VoxBlink2 dataset, containing 110 thousand speakers across multiple languages. We observed a good performance and robustness of VoxBlink-trained models, and our experiments show practical recipes for developing state-of-the-art frontends for speaker recognition.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Analysis of ABC Frontend Audio Systems for the NIST-SRE24
On the NIST SRE24 audio track, a ResNet152 pre-trained on 8kHz/GSM-augmented VoxBlink2 and fine-tuned on telephone speech with 40s segments achieved the best EER and Cprimary among the tested frontends.