Pith. sign in

REVIEW 7 cited by

Evolving from Single-modal to Multi-modal Facial Deepfake Detection: Progress and Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06965 v4 pith:3XMWQD5D submitted 2024-06-11 cs.CV

classification cs.CV
keywords detectiondeepfakechallengesmulti-modalsingle-modaladvancementsemergingfacial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As synthetic media, including video, audio, and text, become increasingly indistinguishable from real content, the risks of misinformation, identity fraud, and social manipulation escalate. This survey traces the evolution of deepfake detection from early single-modal methods to sophisticated multi-modal approaches that integrate audio-visual and text-visual cues. We present a structured taxonomy of detection techniques and analyze the transition from GAN-based to diffusion model-driven deepfakes, which introduce new challenges due to their heightened realism and robustness against detection. Unlike prior surveys that primarily focus on single-modal detection or earlier deepfake techniques, this work provides the most comprehensive study to date, encompassing the latest advancements in multi-modal deepfake detection, generalization challenges, proactive defense mechanisms, and emerging datasets specifically designed to support new interpretability and reasoning tasks. We further explore the role of Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) in strengthening detection robustness against increasingly sophisticated deepfake attacks. By systematically categorizing existing methods and identifying emerging research directions, this survey serves as a foundation for future advancements in combating AI-generated facial forgeries. A curated list of all related papers can be found at \href{https://github.com/qiqitao77/Comprehensive-Advances-in-Deepfake-Detection-Spanning-Diverse-Modalities}{https://github.com/qiqitao77/Awesome-Comprehensive-Deepfake-Detection}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...

  2. SEED: A Benchmark Dataset for Sequential Facial Attribute Editing with Diffusion Models

    cs.CV 2025-05 reject novelty 6.0 of 10

    SEED is a 91,526-image benchmark of diffusion-generated sequential facial edits with sequence, mask, and prompt annotations, and FAITH adds DWT high-frequency cues to a transformer for edit-sequence detection.

  3. ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ForenX detects AI-generated images with MLLMs guided by a forensic prompt and trained on a new explanation dataset, ForgReason.

  4. Evaluating Deepfake Detectors in the Wild

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Modern deepfake detectors, evaluated on a new 500,000 image in-the-wild style benchmark built with SimSwap and Inswapper, mostly fail to generalize and degrade under simple image manipulations.

  5. ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors

    cs.CV 2026-07 conditional novelty 4.0 of 10

    An agentic VLM/LLM orchestration of five adversarial primitives raises no-query transfer attack success on deepfake detectors to 44.3% on AADD-LQ ViT-B/16, up from 39.6% for ARMOR and 19.6% for AutoAttack-PGD.

  6. Visual Language Models as Zero-Shot Deepfake Detectors

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Zero-shot VLMs scored by normalized yes/no token probabilities beat most trained deepfake detectors on a new SimSwap dataset, and a lightly fine-tuned InstructBLIP is near-perfect on DFDC-P.

  7. Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A systematic review of deepfake detection finds a pervasive lack of adversarial robustness evaluation across all modalities and calls for resilient, modality-agnostic detectors.

Pith tools