Pith. sign in

REVIEW 25 cited by

FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.05080 v4 pith:77SMH234 submitted 2021-08-11 cs.CV cs.MMcs.SDeess.AS

FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

classification cs.CV cs.MMcs.SDeess.AS
keywords deepfakedatasetaudiodevelopmultimodalpersonvideovideos
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a person's identity in a video by replacing his/her face with another person's face. Recently, a new problem of generating synthesized human voice of a person is emerging, where AI-based deep learning models can synthesize any person's voice requiring just a few seconds of audio. With the emerging threat of impersonation attacks using deepfake audios and videos, a new generation of deepfake detectors is needed to focus on both video and audio collectively. To develop a competent deepfake detector, a large amount of high-quality data is typically required to capture real-world (or practical) scenarios. Existing deepfake datasets either contain deepfake videos or audios, which are racially biased as well. As a result, it is critical to develop a high-quality video and audio deepfake dataset that can be used to detect both audio and video deepfakes simultaneously. To fill this gap, we propose a novel Audio-Video Deepfake dataset, FakeAVCeleb, which contains not only deepfake videos but also respective synthesized lip-synced fake audios. We generate this dataset using the most popular deepfake generation methods. We selected real YouTube videos of celebrities with four ethnic backgrounds to develop a more realistic multimodal dataset that addresses racial bias, and further help develop multimodal deepfake detectors. We performed several experiments using state-of-the-art detection methods to evaluate our deepfake dataset and demonstrate the challenges and usefulness of our multimodal Audio-Video deepfake dataset.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Toward Calibrated, Fair, and accurate Deepfake Detection

    cs.LG 2026-06 unverdicted novelty 7.0

    Face-Feature Tuning is a label-free logit remapping method that reduces FPR/TPR gaps across groups in deepfake detection while preserving overall accuracy.

  2. LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization

    cs.CV 2026-04 unverdicted novelty 7.0

    LAVA is a layered audio-visual watermarking system using cross-modal fusion and calibration-aware alignment to achieve robust deepfake tamper detection and localization under compression and asynchrony.

  3. BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling

    cs.CV 2026-04 unverdicted novelty 7.0

    A landmark-based detector that flags lip-sync deepfakes by elevated kinematic variance in velocity, acceleration, and jerk over short windows, independent of pixels or audio.

  4. VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?

    cs.CV 2025-12 unverdicted novelty 7.0

    VideoASMR-Bench shows state-of-the-art VLMs fail to reliably detect AI-generated ASMR videos from real ones, though humans can still identify the fakes relatively easily.

  5. MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

    cs.CV 2025-11 conditional novelty 7.0

    MVAD is the first comprehensive benchmark dataset for AI-generated multimodal video-audio detection, with three realistic forgery patterns, high-quality outputs from state-of-the-art models, and diversity across visua...

  6. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  7. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  8. Detecting AI-Generated Video: A Vision-Language Dual-View Survey

    cs.CV 2026-07 conditional novelty 6.0

    AIGC-V detection should be treated as factual fidelity verification and organized by a four-layer vision-language dual-view taxonomy spanning cues, motion, cross-modal consistency, and world-level reasoning.

  9. SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

    cs.SD 2026-07 conditional novelty 6.0

    SynSFX provides a multi-generator sound-effect deepfake corpus showing speech detectors fail, joint training mitigates forgetting, but generalization to unseen generators remains poor due to artifact overfitting.

  10. The Alpha Blending Hypothesis: Compositing Shortcut in Deepfake Detection

    cs.CV 2026-05 unverdicted novelty 6.0

    Deepfake detectors act as alpha blending searchers; training solely on self-blended real images yields top cross-dataset generalization on 15 datasets without using synthetic deepfakes.

  11. Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

    cs.CV 2026-05 unverdicted novelty 6.0

    A training-free dual-system framework refines anomaly score ordering on uncertain samples from self-supervised talking head forgery detectors to improve detection performance.

  12. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 unverdicted novelty 6.0

    The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...

  13. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 conditional novelty 6.0

    Four-class audio-visual DeepFake detectors misclassify authentic but semantically mismatched audio-video pairs; five-class training plus ImageBind similarity improves models that can learn cross-modal semantics.

  14. BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling

    cs.CV 2026-04 unverdicted novelty 6.0

    Lip-sync deepfakes can be detected zero-shot across generators and languages by measuring elevated velocity/acceleration/jerk variance in perioral landmark trajectories that real speech biomechanics forbid.

  15. BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling

    cs.CV 2026-04 unverdicted novelty 6.0

    BioLip detects lip-sync deepfakes via temporal lip jitter, a measurable elevation in lip position variance caused by generative models violating biomechanical articulation constraints.

  16. AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

    cs.CV 2026-04 unverdicted novelty 6.0

    AIFIND stabilizes incremental face forgery detection by aligning volatile features to invariant semantic anchors from low-level artifacts using attention and harmonization modules.

  17. Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes

    cs.MM 2026-04 unverdicted novelty 6.0

    AVPF generates self-created audio-visual pseudo-fakes from real samples to train deepfake detectors that generalize better, with reported average gains up to 7.4%.

  18. Deepfake Detection that Generalizes Across Benchmarks

    cs.CV 2025-08 accept novelty 6.0

    GenD achieves state-of-the-art average cross-dataset AUROC in deepfake detection by parameter-efficient adaptation of a foundational vision encoder with hyperspherical manifold enforcement via L2 normalization and met...

  19. Orthogonal Subspace Decomposition for Generalizable AI-Generated Image Detection

    cs.CV 2024-11 unverdicted novelty 6.0

    Orthogonal subspace decomposition via SVD on vision foundation model features preserves high-rank pre-trained knowledge by freezing principal components and adapting residuals, reducing overfitting for better generali...

  20. LoCC: Detection and Localization of Lip-Syncing Deepfakes via Counterfactual Frame Consistency

    cs.CV 2026-06 unverdicted novelty 5.0

    LoCC detects and localizes lip-syncing deepfakes at frame and segment levels by measuring inconsistencies between each frame and a counterfactual estimate from temporal neighbors via teacher-student learning, outperfo...

  21. The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive

    cs.CR 2026-05 unverdicted novelty 5.0

    Deepfake research prepared for a public-figure catastrophe that did not occur, leaving dominant real harms like NCII and voice scams under-defended.

  22. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

    cs.CV 2026-05 unverdicted novelty 5.0

    Omni-Fake delivers a unified multimodal deepfake benchmark dataset and RL-driven detector that reports gains in accuracy, cross-modal generalization, and explainability over prior baselines.

  23. Deepfake News Detection: A Multimodal Framework Integrating LipNet, DeepSpeech and ResNET for Enhanced Audio-Visual Analysis

    cs.CR 2026-07 reject novelty 4.0

    Using LipNet, DeepSpeech2, BlazeFace, and ResNet18 features with Random Forest, the authors report 94% accuracy on FakeAVCeleb audio features, but evaluation flaws make that result unsupported.

  24. EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

    cs.AI 2026-05 unverdicted novelty 4.0

    Emo-Boost augments low-level deepfake detectors with intra- and inter-modal emotion consistency checks to raise cross-manipulation generalization AUC by 2.1% on FakeAVCeleb.

  25. Ensemble Deep Learning Approaches for AI-Altered Video Detection

    cs.CV 2026-07 conditional novelty 2.0

    An ensemble of AASIST, EfficientNet, XceptionNet, and MesoNet achieves ~70% cross-dataset deepfake detection accuracy, with voting-based fusion slightly outperforming score-based fusion.