On the NIST SRE24 audio track, a ResNet152 pre-trained on 8kHz/GSM-augmented VoxBlink2 and fine-tuned on telephone speech with 40s segments achieved the best EER and Cprimary among the tested frontends.
Training data and augmentations For the fixed condition, we used the NIST CTS Superset [4] to train the embedding extractors
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Analysis of ABC Frontend Audio Systems for the NIST-SRE24
On the NIST SRE24 audio track, a ResNet152 pre-trained on 8kHz/GSM-augmented VoxBlink2 and fine-tuned on telephone speech with 40s segments achieved the best EER and Cprimary among the tested frontends.