Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.
Comparison of the ITU-t p.85 standard to other methods for the evaluation of text-to-speech systems,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection
Phoneme-aligned Grad-CAM on a WavLM-CNN detector reveals significant attack- and speaker-dependent importance of vowels, fricatives and pauses for spoof vs bona-fide decisions on ASVspoof 5.