A new hybrid spoofed-audio benchmark is claimed to show that fine-tuning on it reaches 97%+ accuracy, but the reported numbers are internally inconsistent.
A Comparison of Differential Performance Metrics for the Evaluation of Automatic Speaker Verification Fairness
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
When decisions are made and when personal data is treated by automated processes, there is an expectation of fairness -- that members of different demographic groups receive equitable treatment. This expectation applies to biometric systems such as automatic speaker verification (ASV). We present a comparison of three candidate fairness metrics and extend previous work performed for face recognition, by examining differential performance across a range of different ASV operating points. Results show that the Gini Aggregation Rate for Biometric Equitability (GARBE) is the only one which meets three functional fairness measure criteria. Furthermore, a comprehensive evaluation of the fairness and verification performance of five state-of-the-art ASV systems is also presented. Our findings reveal a nuanced trade-off between fairness and verification accuracy underscoring the complex interplay between system design, demographic inclusiveness, and verification reliability.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
A new hybrid spoofed-audio benchmark is claimed to show that fine-tuning on it reaches 97%+ accuracy, but the reported numbers are internally inconsistent.