A naturalness-based curriculum and per-sample dynamic temperature reduce speech deepfake detection EER on ASVspoof 2021 DF from 2.45% to 1.88%.
Naturalness-Aware Curriculum Learning with Dynamic Temperature for Speech Deepfake Detection
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed speech. This study proposes naturalness-aware curriculum learning, a novel training framework that leverages speech naturalness to enhance the robustness and generalization of SDD. This approach measures sample difficulty using both ground-truth labels and mean opinion scores, and adjusts the training schedule to progressively introduce more challenging samples. To further improve generalization, a dynamic temperature scaling method based on speech naturalness is incorporated into the training process. A 23% relative reduction in the EER was achieved in the experiments on the ASVspoof 2021 DF dataset, without modifying the model architecture. Ablation studies confirmed the effectiveness of naturalness-aware training strategies for SDD tasks.
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Naturalness-Aware Curriculum Learning with Dynamic Temperature for Speech Deepfake Detection
A naturalness-based curriculum and per-sample dynamic temperature reduce speech deepfake detection EER on ASVspoof 2021 DF from 2.45% to 1.88%.