A context-dependent encoding of phoneme durations identifies speakers far better than average-duration vectors and remains effective on anonymized speech without retraining on anonymized data.
Voice Conversion-based Privacy through Adversarial Information Hiding
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion that allows controlling the leakage of identity-bearing information using adversarial information hiding. This enables a deliberate trade-off between maintaining source-speech characteristics and modification of speaker identity. As such, the approach improves on voice-conversion techniques like CycleGAN and StarGAN, which were not designed for privacy, meaning that converted speech may leak personal information in unpredictable ways. Our approach is also more flexible than ASR-TTS voice conversion pipelines, which by design discard all prosodic information linked to textual content. Evaluations show that the proposed system successfully modifies perceived speaker identity whilst well maintaining source lexical content.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
A context-dependent encoding of phoneme durations identifies speakers far better than average-duration vectors and remains effective on anonymized speech without retraining on anonymized data.