Pith. sign in

REVIEW 20 cited by

Audio Deepfake Detection: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14970 v1 pith:FRFOETAL submitted 2023-08-29 cs.SD eess.AS

Audio Deepfake Detection: A Survey

classification cs.SD eess.AS
keywords detectiondeepfakeaudiosurveydatasetsdevelopmentsevaluationfeatures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Audio deepfake detection is an emerging active topic. A growing number of literatures have aimed to study deepfake detection algorithms and achieved effective performance, the problem of which is far from being solved. Although there are some review literatures, there has been no comprehensive survey that provides researchers with a systematic overview of these developments with a unified evaluation. Accordingly, in this survey paper, we first highlight the key differences across various types of deepfake audio, then outline and analyse competitions, datasets, features, classifications, and evaluation of state-of-the-art approaches. For each aspect, the basic techniques, advanced developments and major challenges are discussed. In addition, we perform a unified comparison of representative features and classifiers on ASVspoof 2021, ADD 2023 and In-the-Wild datasets for audio deepfake detection, respectively. The survey shows that future research should address the lack of large scale datasets in the wild, poor generalization of existing detection methods to unknown fake attacks, as well as interpretability of detection results.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Probing-Guided Layer Selection from Self-Supervised Speech Models for Generalizable Audio Deepfake Detection

    cs.SD 2026-06 unverdicted novelty 7.0

    Probing-guided selection of depth zones from frozen SSL speech models yields compact classifiers with 28% relative EER improvement on cross-domain deepfake detection tasks.

  2. MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

    cs.SD 2026-05 unverdicted novelty 7.0

    MixFake is a new benchmark for mixed-authenticity audio and a multi-stream prompt tuning method achieves 0.95% EER foreground and 7.72% absolute gain in complex background deepfake detection.

  3. Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

    cs.SD 2026-05 unverdicted novelty 7.0

    A new dataset, iterative coarse-to-fine localization framework, and segment-level IoU F1 metric tackle the open problem of detecting multiple unknown word-level inpainted regions in speech.

  4. Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages

    eess.AS 2026-04 unverdicted novelty 7.0

    Introduces the Indic-CodecFake dataset for Indic codec deepfakes and SATYAM, a novel hyperbolic ALM that outperforms baselines through dual-stage semantic-prosodic fusion using Bhattacharya distance.

  5. ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics

    cs.SD 2026-04 unverdicted novelty 7.0

    ArtifactNet extracts codec residuals from spectrograms with a 4M-parameter network to detect AI music at F1=0.9829 and 1.49% FPR on unseen tracks from 22 generators, outperforming larger baselines.

  6. Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

    cs.CV 2026-04 conditional novelty 7.0

    Introduces the LDD task, ListenForge dataset built from five listening head generation methods, and MANet model that detects listening forgeries via motion inconsistencies guided by audio semantics.

  7. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  8. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

    cs.SD 2026-07 conditional novelty 6.0

    A new Identity Sensitivity Score flags misclassified audio deepfake detections with AUC up to 0.954, but its claim to isolate speaker-identity behavior from plain confidence is not yet controlled.

  9. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

    cs.SD 2026-07 conditional novelty 6.0

    Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.

  10. DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

    cs.MM 2026-06 unverdicted novelty 6.0

    DetectZoo is a unified toolkit providing reference implementations of 61 detectors, native loaders for 22 benchmark datasets, and a standardized evaluation pipeline for AI-generated content detection across text, audi...

  11. Asymmetric Phase Coding Audio Watermarking

    cs.CR 2026-05 unverdicted novelty 6.0

    APC embeds compact Ed25519 signatures into audio phase data with error correction to achieve 97.5-98.3% cryptographic verification under eight attack types at mean PESQ 3.02.

  12. Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

    cs.SD 2026-05 unverdicted novelty 6.0

    Phoneme-level analysis using self-supervised embeddings identifies higher divergence in complex vowels and fricatives for emotional voice conversion deepfakes, enabling more interpretable detection across emotions.

  13. Split and Conquer Partial Deepfake Speech

    cs.SD 2026-04 unverdicted novelty 6.0

    A two-stage boundary detection plus segment classification method with multi-length training achieves state-of-the-art results for detecting and localizing partial deepfakes on PartialSpoof and Half-Truth benchmarks.

  14. Synthetic Trust Attacks: Modeling How Generative AI Manipulates Human Decisions in Social Engineering Fraud

    cs.CR 2026-04 unverdicted novelty 6.0

    The paper proposes Synthetic Trust Attacks (STAs) as a formal threat model with an eight-stage attack chain (STAM) that shifts defense focus from detecting synthetic media to protecting human decision processes in soc...

  15. AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features

    cs.HC 2025-09 unverdicted novelty 6.0

    The AuthGlass dataset and proposed multi-modal models achieve state-of-the-art results on voice liveness detection and user authentication for smart glasses.

  16. A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

    cs.SD 2026-05 unverdicted novelty 5.0

    A survey of Large Audio Language Models that establishes a taxonomy of trustworthiness vulnerabilities and proposes a Defense-in-Depth roadmap for audio intelligence.

  17. Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis

    cs.SD 2026-03 unverdicted novelty 5.0

    Fairness metrics uncover gender disparities in audio deepfake detection error distributions that standard Equal Error Rate metrics obscure.

  18. Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing

    eess.AS 2025-09 unverdicted novelty 5.0

    A zero-shot open-set speech deepfake source tracing framework using adapted SSL-AASIST embeddings and AAM loss achieves EER of 16.43% in OOD trials with cosine scoring, outperforming few-shot alternatives.

  19. Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection

    cs.SD 2026-06 unverdicted novelty 4.0

    Dual-granularity orthogonal disentanglement framework achieves EERs of 1.35%, 7.88%, and 21.58% on ASVspoof 2019 LA, ASVspoof 2021 DF, and In-the-Wild datasets, outperforming gradient reversal by 2.60% on cross-datase...

  20. Classical Machine Learning Baselines for Deepfake Audio Detection on the Fake-or-Real Dataset

    eess.AS 2026-04 unverdicted novelty 3.0

    RBF SVM achieves ~93% accuracy and ~7% EER on deepfake audio detection using prosodic and spectral features from the FoR dataset at 44.1 kHz and 16 kHz sampling rates.