Using character-duration sequences from WhisperX, a transformer identifies speakers with balanced accuracy of 0.39 on LibriSpeech but only 0.03 on VoxCeleb1, and fusion with x-vectors does not improve accuracy.
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, we investigate the impact of speech temporal dynamics in application to automatic speaker verification and speaker voice anonymization tasks. We propose several metrics to perform automatic speaker verification based only on phoneme durations. Experimental results demonstrate that phoneme durations leak some speaker information and can reveal speaker identity from both original and anonymized speech. Thus, this work emphasizes the importance of taking into account the speaker's speech rate and, more importantly, the speaker's phonetic duration characteristics, as well as the need to modify them in order to develop anonymization systems with strong privacy protection capacity.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Rhythm Features for Speaker Identification
Using character-duration sequences from WhisperX, a transformer identifies speakers with balanced accuracy of 0.39 on LibriSpeech but only 0.03 on VoxCeleb1, and fusion with x-vectors does not improve accuracy.