A dual anonymization pipeline that uses an LLM to replace sensitive entities and a prompt-driven TTS to generate speech with a new, unrelated voice.
People are poorly equipped to detect AI-powered voice clones
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
As generative artificial intelligence (AI) continues its ballistic trajectory, everything from text to audio, image, and video generation continues to improve at mimicking human-generated content. Through a series of perceptual studies, we report on the realism of AI-generated voices in terms of identity matching and naturalness. We find human participants cannot consistently identify recordings of AI-generated voices. Specifically, participants perceived the identity of an AI-voice to be the same as its real counterpart approximately 80% of the time, and correctly identified a voice as AI generated only about 60% of the time.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
SecureSpeech: Prompt-based Speaker and Content Protection
A dual anonymization pipeline that uses an LLM to replace sensitive entities and a prompt-driven TTS to generate speech with a new, unrelated voice.