Pith. sign in

REVIEW 7 cited by

Deepfake Video Detection Using Convolutional Vision Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.11126 v3 pith:IJ7QQ4IP submitted 2021-02-22 cs.CV

classification cs.CV
keywords convolutionaltransformervisiondetectionfeaturesvideoachievedalter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid advancement of deep learning models that can generate and synthesis hyper-realistic videos known as Deepfakes and their ease of access to the general public have raised concern from all concerned bodies to their possible malicious intent use. Deep learning techniques can now generate faces, swap faces between two subjects in a video, alter facial expressions, change gender, and alter facial features, to list a few. These powerful video manipulation methods have potential use in many fields. However, they also pose a looming threat to everyone if used for harmful purposes such as identity theft, phishing, and scam. In this work, we propose a Convolutional Vision Transformer for the detection of Deepfakes. The Convolutional Vision Transformer has two components: Convolutional Neural Network (CNN) and Vision Transformer (ViT). The CNN extracts learnable features while the ViT takes in the learned features as input and categorizes them using an attention mechanism. We trained our model on the DeepFake Detection Challenge Dataset (DFDC) and have achieved 91.5 percent accuracy, an AUC value of 0.91, and a loss value of 0.32. Our contribution is that we have added a CNN module to the ViT architecture and have achieved a competitive result on the DFDC dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Replacing the global-pooling readout of a frozen video backbone with a velocity-gated, per-channel-magnitude readout improves AI-generated video detection cross-generator accuracy by several AUC points.

  2. CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.

  3. Revisiting Simple Baselines for In-The-Wild Deepfake Detection

    cs.CV 2025-09 conditional novelty 4.0 of 10

    Finetuned CLIP-pretrained ConvNeXt-base and ViT-b32 classifiers reach 81% accuracy on Deepfake-Eval-2024, within noise of the leading commercial detector's 82%.

  4. SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A spatial-frequency feature fusion network with attention achieves improved accuracy on remote sensing image forgery detection and introduces a stable-diffusion-based benchmark.

  5. Enhancing Abnormality Identification: Robust Out-of-Distribution Strategies for Deepfake Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    The paper introduces a deepfake OOD detector that combines reconstruction residual, latent encoding, and softmax confidence, and shows strong results only when real OOD samples are available for training.

  6. Enhancing Deepfake Detection using SE Block Attention with CNN

    cs.CV 2025-06 conditional novelty 3.0 of 10

    Adding SE attention blocks to a small CNN raises deepfake detection accuracy from 91.13% to 94.14% on the StyleGAN subset of DFFD, but the result rests on a single run with questionable comparisons.

  7. Face Deepfakes -- A Comprehensive Review

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A review of face deepfake generation and detection finds that off-the-shelf deepfake tools such as Wav2Lip and SimSwap achieve high attack success rates against lightweight face recognition models.

Pith tools