Pith. sign in

REVIEW 1 cited by

Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.06922 v1 pith:ULQTPHOE submitted 2024-08-13 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords asvspoof5dataasvspoofaudioaugmentationdeepfakefeaturesfrequency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ASVspoof5, the fifth edition of the ASVspoof series, is one of the largest global audio security challenges. It aims to advance the development of countermeasure (CM) to discriminate bonafide and spoofed speech utterances. In this paper, we focus on addressing the problem of open-domain audio deepfake detection, which corresponds directly to the ASVspoof5 Track1 open condition. At first, we comprehensively investigate various CM on ASVspoof5, including data expansion, data augmentation, and self-supervised learning (SSL) features. Due to the high-frequency gaps characteristic of the ASVspoof5 dataset, we introduce Frequency Mask, a data augmentation method that masks specific frequency bands to improve CM robustness. Combining various scale of temporal information with multiple SSL features, our experiments achieved a minDCF of 0.0158 and an EER of 0.55% on the ASVspoof 5 Track 1 evaluation progress set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tandem spoofing-robust automatic speaker verification based on time-domain embeddings

    eess.AS 2024-12 conditional novelty 3.0 of 10

    A gender-separated countermeasure built from probability-mass-function time embeddings improves tandem spoofing-robust speaker verification on ASVspoof2019, but only when thresholds are tuned on the evaluation set.

Pith tools