Pith. sign in

REVIEW 3 cited by

CLAD: Robust Audio Deepfake Detection Against Manipulation Attacks with Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15854 v1 pith:5EKYXYMV submitted 2024-04-24 cs.CR cs.LGcs.SDeess.AS

classification cs.CRcs.LGcs.SDeess.AS
keywords detectionaudiocladattacksdeepfakemanipulationrobustnesscontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing prevalence of audio deepfakes poses significant security threats, necessitating robust detection methods. While existing detection systems exhibit promise, their robustness against malicious audio manipulations remains underexplored. To bridge the gap, we undertake the first comprehensive study of the susceptibility of the most widely adopted audio deepfake detectors to manipulation attacks. Surprisingly, even manipulations like volume control can significantly bypass detection without affecting human perception. To address this, we propose CLAD (Contrastive Learning-based Audio deepfake Detector) to enhance the robustness against manipulation attacks. The key idea is to incorporate contrastive learning to minimize the variations introduced by manipulations, therefore enhancing detection robustness. Additionally, we incorporate a length loss, aiming to improve the detection accuracy by clustering real audios more closely in the feature space. We comprehensively evaluated the most widely adopted audio deepfake detection models and our proposed CLAD against various manipulation attacks. The detection models exhibited vulnerabilities, with FAR rising to 36.69%, 31.23%, and 51.28% under volume control, fading, and noise injection, respectively. CLAD enhanced robustness, reducing the FAR to 0.81% under noise injection and consistently maintaining an FAR below 1.63% across all tests. Our source code and documentation are available in the artifact repository (https://github.com/CLAD23/CLAD).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Small semantic-preserving changes to transcripts, passed through text-to-speech, significantly reduce the accuracy of both open-source and commercial audio anti-spoofing detectors.

  2. SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks

    cs.SD 2025-07 conditional novelty 5.0 of 10

    The paper claims that a collaborative defense generator plus triplet learning keeps audio deepfake detectors at roughly 98 percent accuracy against GAN-based anti-forensic attacks.

  3. Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A systematic review of deepfake detection finds a pervasive lack of adversarial robustness evaluation across all modalities and calls for resilient, modality-agnostic detectors.

Pith tools