Pith. sign in

REVIEW 1 cited by

What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.13887 v2 pith:5E357ERQ submitted 2025-01-23 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords audioanalyzedatasetsdeepfakeexplanationsimportancelargelimited
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adding explanations to audio deepfake detection (ADD) models will boost their real-world application by providing insight on the decision making process. In this paper, we propose a relevancy-based explainable AI (XAI) method to analyze the predictions of transformer-based ADD models. We compare against standard Grad-CAM and SHAP-based methods, using quantitative faithfulness metrics as well as a partial spoof test, to comprehensively analyze the relative importance of different temporal regions in an audio. We consider large datasets, unlike previous works where only limited utterances are studied, and find that the XAI methods differ in their explanations. The proposed relevancy-based XAI method performs the best overall on a variety of metrics. Further investigation on the relative importance of speech/non-speech, phonetic content, and voice onsets/offsets suggest that the XAI results obtained from analyzing limited utterances don't necessarily hold when evaluated on large datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks

    cs.SD 2025-07 conditional novelty 5.0 of 10

    The paper claims that a collaborative defense generator plus triplet learning keeps audio deepfake detectors at roughly 98 percent accuracy against GAN-based anti-forensic attacks.

Pith tools