Pith. sign in

REVIEW 3 cited by

Are audio DeepFake detection models polyglots?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.17924 v2 pith:KSXHE6VM submitted 2024-12-23 cs.SD eess.AS

classification cs.SDeess.AS
keywords detectionaudioadaptationbenchmarkdatasetsdeepfakeefficacyenglish
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since the majority of audio DeepFake (DF) detection methods are trained on English-centric datasets, their applicability to non-English languages remains largely unexplored. In this work, we present a benchmark for the multilingual audio DF detection challenge by evaluating various adaptation strategies. Our experiments focus on analyzing models trained on English benchmark datasets, as well as intra-linguistic (same-language) and cross-linguistic adaptation approaches. Our results indicate considerable variations in detection efficacy, highlighting the difficulties of multilingual settings. We show that limiting the dataset to English negatively impacts the efficacy, while stressing the importance of the data in the target language.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tell me Habibi, is it Real or Fake?

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.

  2. Multilingual Source Tracing of Speech Deepfakes: A First Benchmark

    eess.AS 2025-08 conditional novelty 6.0 of 10

    The first multilingual source-tracing benchmark for speech deepfakes, showing LFCC-ECAPA-TDNN generalizes best across languages.

  3. Multi-level SSL Feature Gating for Audio Deepfake Detection

    cs.SD 2025-09 conditional novelty 5.0 of 10

    An XLS-R based audio deepfake detector combining gated multi-kernel convolutions with a CKA dissimilarity loss reports top EERs on 19LA, 21DF, and In-The-Wild benchmarks.

Pith tools