Pith. sign in

REVIEW 3 cited by

AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.15308 v2 pith:MSTEOYQO submitted 2023-11-26 cs.CV

AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset

classification cs.CV
keywords datasetdeepfakeaudio-visuallocalizationmanipulationsav-deepfake1mmethodsproposed
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality deepfake images and videos, only a few works address the problem of the localization of small segments of audio-visual manipulations embedded in real videos. In this research, we emulate the process of such content generation and propose the AV-Deepfake1M dataset. The dataset contains content-driven (i) video manipulations, (ii) audio manipulations, and (iii) audio-visual manipulations for more than 2K subjects resulting in a total of more than 1M videos. The paper provides a thorough description of the proposed data generation pipeline accompanied by a rigorous analysis of the quality of the generated data. The comprehensive benchmark of the proposed dataset utilizing state-of-the-art deepfake detection and localization methods indicates a significant drop in performance compared to previous datasets. The proposed dataset will play a vital role in building the next-generation deepfake localization methods. The dataset and associated code are available at https://github.com/ControlNet/AV-Deepfake1M .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

    cs.SD 2026-05 unverdicted novelty 7.0

    A new dataset, iterative coarse-to-fine localization framework, and segment-level IoU F1 metric tackle the open problem of detecting multiple unknown word-level inpainted regions in speech.

  2. MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

    cs.SD 2024-01 unverdicted novelty 6.0

    MLAAD provides a large-scale multi-language synthetic audio dataset for training and evaluating audio anti-spoofing models, showing better training performance than InTheWild and FakeOrReal and alternating superiority...

  3. A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators

    cs.SD 2026-03 unverdicted novelty 4.0

    Balancing diverse bonafide resources and AI generators in training data is the key to building general deepfake speech detection models.