REVIEW 2 cited by
Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rise of AI-driven generative models has enabled the creation of highly realistic speech deepfakes - synthetic audio signals that can imitate target speakers' voices - raising critical security concerns. Existing methods for detecting speech deepfakes primarily rely on supervised learning, which suffers from two critical limitations: limited generalization to unseen synthesis techniques and a lack of explainability. In this paper, we address these issues by introducing a novel interpretable one-class detection framework, which reframes speech deepfake detection as an anomaly detection task. Our model is trained exclusively on real speech to characterize its distribution, enabling the classification of out-of-distribution samples as synthetically generated. Additionally, our framework produces interpretable anomaly maps during inference, highlighting anomalous regions across both time and frequency domains. This is done through a Student-Teacher Feature Pyramid Matching system, enhanced with Discrepancy Scaling to improve generalization capabilities across unseen data distributions. Extensive evaluations demonstrate the superior performance of our approach compared to the considered baselines, validating the effectiveness of framing speech deepfake detection as an anomaly detection problem.
Forward citations
Cited by 2 Pith papers
-
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
Phoneme-level decomposition of speech improves person-of-interest deepfake detection by making decisions robust to noise and compression while localizing artifacts to specific phonemes.
-
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
Segment-level EER overstates deployment readiness for partial fake speech localizers, which drop from 7.6% to above 40% EER on out-of-domain test sets.
Discussion (0). Continue with ORCID to comment.