REVIEW 7 cited by
Deepfake Video Detection Using Convolutional Vision Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rapid advancement of deep learning models that can generate and synthesis hyper-realistic videos known as Deepfakes and their ease of access to the general public have raised concern from all concerned bodies to their possible malicious intent use. Deep learning techniques can now generate faces, swap faces between two subjects in a video, alter facial expressions, change gender, and alter facial features, to list a few. These powerful video manipulation methods have potential use in many fields. However, they also pose a looming threat to everyone if used for harmful purposes such as identity theft, phishing, and scam. In this work, we propose a Convolutional Vision Transformer for the detection of Deepfakes. The Convolutional Vision Transformer has two components: Convolutional Neural Network (CNN) and Vision Transformer (ViT). The CNN extracts learnable features while the ViT takes in the learned features as input and categorizes them using an attention mechanism. We trained our model on the DeepFake Detection Challenge Dataset (DFDC) and have achieved 91.5 percent accuracy, an AUC value of 0.91, and a loss value of 0.32. Our contribution is that we have added a CNN module to the ViT architecture and have achieved a competitive result on the DFDC dataset.
Forward citations
Cited by 7 Pith papers
-
Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection
Replacing the global-pooling readout of a frozen video backbone with a velocity-gated, per-channel-magnitude readout improves AI-generated video detection cross-generator accuracy by several AUC points.
-
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
CAD combines cross-modal lip-speech alignment with per-modality artifact distillation and reports 99.96% AUC on IDForge-v2, with strong cross-dataset results.
-
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
Finetuned CLIP-pretrained ConvNeXt-base and ViT-b32 classifiers reach 81% accuracy on Deepfake-Eval-2024, within noise of the leading commercial detector's 82%.
-
SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection
A spatial-frequency feature fusion network with attention achieves improved accuracy on remote sensing image forgery detection and introduces a stable-diffusion-based benchmark.
-
Enhancing Abnormality Identification: Robust Out-of-Distribution Strategies for Deepfake Detection
The paper introduces a deepfake OOD detector that combines reconstruction residual, latent encoding, and softmax confidence, and shows strong results only when real OOD samples are available for training.
-
Enhancing Deepfake Detection using SE Block Attention with CNN
Adding SE attention blocks to a small CNN raises deepfake detection accuracy from 91.13% to 94.14% on the StyleGAN subset of DFFD, but the result rests on a single run with questionable comparisons.
-
Face Deepfakes -- A Comprehensive Review
A review of face deepfake generation and detection finds that off-the-shelf deepfake tools such as Wav2Lip and SimSwap achieve high attack success rates against lightweight face recognition models.
Discussion (0). Continue with ORCID to comment.