REVIEW 4 cited by
WavLM model ensemble for audio deepfake detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Audio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples, thus enabling their use in malicious activities. In this paper we address the issue of audio deepfake detection as it was set in the ASVspoof5 challenge. First, we benchmark ten types of pretrained representations and show that the self-supervised representations stemming from the wav2vec2 and wavLM families perform best. Of the two, wavLM is better when restricting the pretraining data to LibriSpeech, as required by the challenge rules. To further improve performance, we finetune the wavLM model for the deepfake detection task. We extend the ASVspoof5 dataset with samples from other deepfake detection datasets and apply data augmentation. Our final challenge submission consists of a late fusion combination of four models and achieves an equal error rate of 6.56% and 17.08% on the two evaluation sets.
Forward citations
Cited by 4 Pith papers
-
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
A four-backbone SSL ensemble achieves near-perfect deepfake detection in the ImageCLEF 2026 track, while the same team's generated audio ranks first in the generation track, revealing an OR-versus-AND asymmetry betwee...
-
HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition
Fusing wav2vec2 and SoundStream features in hyperbolic space improves speech emotion recognition accuracy on CREMA-D and Emo-DB, according to reported results.
-
Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems
A systematic review of deepfake detection finds a pervasive lack of adversarial robustness evaluation across all modalities and calls for resilient, modality-agnostic detectors.
-
Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
A gender-separated countermeasure built from probability-mass-function time embeddings improves tandem spoofing-robust speaker verification on ASVspoof2019, but only when thresholds are tuned on the evaluation set.
Discussion (0). Continue with ORCID to comment.