Pith. sign in

REVIEW 1 cited by

Self-supervised Transformer for Deepfake Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.01265 v1 pith:QE3RCYBB submitted 2022-03-02 cs.CV

classification cs.CV
keywords methoddeepfakedetectionself-superviseddatafeaturefeaturesgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The fast evolution and widespread of deepfake techniques in real-world scenarios require stronger generalization abilities of face forgery detectors. Some works capture the features that are unrelated to method-specific artifacts, such as clues of blending boundary, accumulated up-sampling, to strengthen the generalization ability. However, the effectiveness of these methods can be easily corrupted by post-processing operations such as compression. Inspired by transfer learning, neural networks pre-trained on other large-scale face-related tasks may provide useful features for deepfake detection. For example, lip movement has been proved to be a kind of robust and good-transferring highlevel semantic feature, which can be learned from the lipreading task. However, the existing method pre-trains the lip feature extraction model in a supervised manner, which requires plenty of human resources in data annotation and increases the difficulty of obtaining training data. In this paper, we propose a self-supervised transformer based audio-visual contrastive learning method. The proposed method learns mouth motion representations by encouraging the paired video and audio representations to be close while unpaired ones to be diverse. After pre-training with our method, the model will then be partially fine-tuned for deepfake detection task. Extensive experiments show that our self-supervised method performs comparably or even better than the supervised pre-training counterpart.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A 0.48M-parameter single-stream network with iterative audio-visual fusion outperforms larger two-stream baselines on DF-TIMIT, FakeAVCeleb, and DFDC deepfake detection benchmarks.

Pith tools