FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces

Andreas R\"ossler; Christian Riess; Davide Cozzolino; Justus Thies; Luisa Verdoliva; Matthias Nie{\ss}ner

arxiv: 1803.09179 · v1 · pith:7763I6KAnew · submitted 2018-03-24 · 💻 cs.CV

FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces

Andreas R\"ossler , Davide Cozzolino , Luisa Verdoliva , Christian Riess , Justus Thies , Matthias Nie{\ss}ner This is my paper

classification 💻 cs.CV

keywords videosdatasetfaceintroducevideobeencompresseddatasets

0 comments

read the original abstract

With recent advances in computer vision and graphics, it is now possible to generate videos with extremely realistic synthetic faces, even in real time. Countless applications are possible, some of which raise a legitimate alarm, calling for reliable detectors of fake videos. In fact, distinguishing between original and manipulated video can be a challenge for humans and computers alike, especially when the videos are compressed or have low resolution, as it often happens on social networks. Research on the detection of face manipulations has been seriously hampered by the lack of adequate datasets. To this end, we introduce a novel face manipulation dataset of about half a million edited images (from over 1000 videos). The manipulations have been generated with a state-of-the-art face editing approach. It exceeds all existing video manipulation datasets by at least an order of magnitude. Using our new dataset, we introduce benchmarks for classical image forensic tasks, including classification and segmentation, considering videos compressed at various quality levels. In addition, we introduce a benchmark evaluation for creating indistinguishable forgeries with known ground truth; for instance with generative refinement models.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
cs.CV 2026-03 unverdicted novelty 7.0

FrameDiT proposes Matrix Attention for DiTs to achieve SOTA video generation with improved temporal coherence and efficiency comparable to local factorized attention.
The DeepSpeak Dataset
cs.CV 2024-08 unverdicted novelty 7.0

DeepSpeak provides over 100 hours of consented, identity-matched real and modern deepfake audiovisual content focused on talking heads, with evaluations showing existing detectors fail to generalize without retraining.
Latte: Latent Diffusion Transformer for Video Generation
cs.CV 2024-01 unverdicted novelty 6.0

Latte achieves state-of-the-art video generation on FaceForensics, SkyTimelapse, UCF101, and Taichi-HD by using a latent diffusion transformer with four efficient spatial-temporal decomposition variants and best-pract...
Hiding Faces in Plain Sight: Disrupting AI Face Synthesis with Adversarial Perturbations
cs.CV 2019-06 unverdicted novelty 6.0

Adversarial perturbations disrupt DNN-based face detectors under white-box, gray-box, and black-box settings to sabotage training data for AI face synthesis.
We Need No Pixels: Video Manipulation Detection Using Stream Descriptors
cs.LG 2019-06 unverdicted novelty 6.0

Video forgeries are detectable via binary classification on multimedia stream descriptors without pixel analysis.