Pith. sign in

REVIEW 11 cited by

ADD 2023: the Second Audio Deepfake Detection Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13774 v1 pith:AOPINLM3 submitted 2023-05-23 cs.SD eess.AS

ADD 2023: the Second Audio Deepfake Detection Challenge

classification cs.SD eess.AS
keywords audiodeepfakefakedetectionchallengeevaluationgameincludes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the world to build new innovative technologies that can further accelerate and foster research on detecting and analyzing deepfake speech utterances. Different from previous challenges (e.g. ADD 2022), ADD 2023 focuses on surpassing the constraints of binary real/fake classification, and actually localizing the manipulated intervals in a partially fake speech as well as pinpointing the source responsible for generating any fake audio. Furthermore, ADD 2023 includes more rounds of evaluation for the fake audio game sub-challenge. The ADD 2023 challenge includes three subchallenges: audio fake game (FG), manipulation region location (RL) and deepfake algorithm recognition (AR). This paper describes the datasets, evaluation metrics, and protocols. Some findings are also reported in audio deepfake detection tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

    cs.SD 2026-05 unverdicted novelty 7.0

    MixFake is a new benchmark for mixed-authenticity audio and a multi-stream prompt tuning method achieves 0.95% EER foreground and 7.72% absolute gain in complex background deepfake detection.

  2. Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages

    eess.AS 2026-04 unverdicted novelty 7.0

    Introduces the Indic-CodecFake dataset for Indic codec deepfakes and SATYAM, a novel hyperbolic ALM that outperforms baselines through dual-stage semantic-prosodic fusion using Bhattacharya distance.

  3. Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

    cs.CV 2026-04 conditional novelty 7.0

    Introduces the LDD task, ListenForge dataset built from five listening head generation methods, and MANet model that detects listening forgeries via motion inconsistencies guided by audio semantics.

  4. GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

    cs.SD 2026-06 unverdicted novelty 6.0

    GRIDEX is the first pipeline to generate grid-grounded structured forensic explanations for deepfake spectrogram anomalies using a two-stage SFT plus GRPO training process on vision-language models.

  5. Ethical and Technical Limits of Deepfake Speech Datasets

    cs.SD 2026-06 unverdicted novelty 6.0

    Audit of 39 deepfake speech datasets shows most lack demographic metadata making fairness checks infeasible and reveals substantial overlap in bona fide sources that undermines cross-dataset generalization claims.

  6. Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

    cs.SD 2026-06 unverdicted novelty 6.0

    A gated fusion of XLSR-53 and CORES features with energy margin and diversity losses reaches 97.6% ID accuracy and reduces FPR95 by 83.5% relative to the Interspeech 2025 baseline on MLAAD.

  7. RTCFake: Speech Deepfake Detection in Real-Time Communication

    cs.SD 2026-04 unverdicted novelty 6.0

    RTCFake is the first large-scale dataset of real-time communication speech deepfakes paired with offline versions, paired with a phoneme-guided consistency learning method that improves cross-platform and noise-robust...

  8. ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks

    eess.AS 2026-04 unverdicted novelty 6.0

    ProSDD learns speaker-conditioned prosodic variation from real speech via supervised masked prediction and jointly optimizes it with spoof detection, cutting EER substantially on ASVspoof 2024 and emotional datasets.

  9. MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

    cs.SD 2024-01 unverdicted novelty 6.0

    MLAAD provides a large-scale multi-language synthetic audio dataset for training and evaluating audio anti-spoofing models, showing better training performance than InTheWild and FakeOrReal and alternating superiority...

  10. Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis

    cs.SD 2026-06 unverdicted novelty 4.0

    Experiments indicate that diversity of attack generation methods improves cross-dataset generalization in speech anti-spoofing more than increasing training data scale.

  11. SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing

    cs.SD 2026-06 conditional novelty 3.0

    SpAArSIST sparsifies AASIST by swapping learned pooling for explicit magnitude-based scoring and mean aggregation, cutting compute 20.7% and improving In-the-Wild EER to 2.82%.