REVIEW 5 cited by
Investigating self-supervised front ends for speech spoofing countermeasures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised speech model is a rapid progressing research topic, and many pre-trained models have been released and used in various down stream tasks. For speech anti-spoofing, most countermeasures (CMs) use signal processing algorithms to extract acoustic features for classification. In this study, we use pre-trained self-supervised speech models as the front end of spoofing CMs. We investigated different back end architectures to be combined with the self-supervised front end, the effectiveness of fine-tuning the front end, and the performance of using different pre-trained self-supervised models. Our findings showed that, when a good pre-trained front end was fine-tuned with either a shallow or a deep neural network-based back end on the ASVspoof 2019 logical access (LA) training set, the resulting CM not only achieved a low EER score on the 2019 LA test set but also significantly outperformed the baseline on the ASVspoof 2015, 2021 LA, and 2021 deepfake test sets. A sub-band analysis further demonstrated that the CM mainly used the information in a specific frequency band to discriminate the bona fide and spoofed trials across the test sets.
Forward citations
Cited by 5 Pith papers
-
Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
A zero-shot open-set speech deepfake source tracing framework using adapted SSL-AASIST embeddings and AAM loss achieves EER of 16.43% in OOD trials with cosine scoring, outperforming few-shot alternatives.
-
Emoanti: audio anti-deepfake with refined emotion-guided representations
EmoAnti fine-tunes Wav2Vec2 on emotion recognition and refines the resulting features with a convolutional residual extractor, achieving low EER on ASVspoof LA but worse performance on DF than its own no-finetuning baseline.
-
Generalizable Audio Spoofing Detection using Non-Semantic Representations
Frozen non-semantic TRILLson embeddings with a lightweight backend beat prior spoofing detectors on out-of-domain datasets while staying competitive in-domain.
-
Multi-Granularity Adaptive Time-Frequency Attention Framework for Audio Deepfake Detection under Real-World Communication Degradations
A multi-granularity adaptive attention model for audio deepfake detection remains accurate across six speech codecs and five packet-loss levels, reportedly outperforming baselines.
-
Segment Transformer: AI-Generated Music Detection via Music Structural Analysis
A two-stage transformer framework classifies AI-generated music from short clips and beat-segmented full tracks, reporting 99.9% accuracy on SONICS without releasing code or ablations.
Discussion (0). Sign in to comment.