Pith. sign in

REVIEW 4 major objections 5 minor 96 references

MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read MFFI is the first face forgery dataset to combine 50 forgery methods, varied scenes, diverse authentic faces, and transmission degradation into one 1,024,000-image benchmark.

desk verdict Useful, checkable dataset resource with real coverage breadth, but the cross-domain generalization and difficulty-gradient claims outrun the experiments. read the letter →

arxiv 2509.05592 v1 pith:LR4POLSO submitted 2025-09-06 cs.CV

classification cs.CV
keywords DeepfakedetectionFaceforgerydatasetReal-worldbenchmarkmethoddiversityImagedegradationCross-domaingeneralizationswappingDiffusionmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MFFI is a face forgery image dataset built to close the gap between laboratory deepfake benchmarks and the forgeries that actually circulate online. The paper claims that prior datasets fall short on four independent axes: coverage of modern forgery methods, variety of facial scenes, diversity of authentic (real) face data, and simulation of transmission degradation. MFFI is presented as the first dataset to handle all four axes at once, integrating 50 forgery methods and more than one million face images from four real-face sources, with a degraded test set that models internet propagation. Benchmark evaluations reported in the paper indicate that MFFI outperforms existing public datasets in scene complexity, cross-domain generalization capability, and detection difficulty gradients. If the claim is right, MFFI gives detection researchers a training and evaluation ground on which models must cope with unknown generation methods, varied faces, and messy image degradation rather than a single narrow forgery style.

What carries the argument

The load-bearing object is the dataset itself, built along four construction axes. Wider Forgery Methods contributes 50 generators, including 2025 commercial models, grouped into six forgery classes. Varied Facial Scenes applies a six-way filter so forged and real faces span ethnicities, ages, poses, occlusions, backgrounds, and lighting. Diversified Authentic Data pools real faces from four different sources. Multi-level Degradation Operations adds conventional distortions (blur, noise, sharpening, compression, geometric transforms) plus a black-box adversarial patch attack to the test set only. The evaluation machinery is a three-protocol benchmark: intra-dataset, cross-dataset, and zero-shot multimodal large-model testing, using the training configuration and metrics (ACC, AUC, EER, AP) defined by the reference benchmark implementation cited in the paper.

What would settle it

Train a small classifier on Test-D to distinguish fake from real, then check whether the degradation type or the patch location alone predicts the label near-perfectly; if it does, the Test-D difficulty gradient reflects artifacts of the degradation pipeline rather than real-world robustness. A complementary check is to build a matched set of images that passed through actual social-platform upload and download cycles and compare detector accuracy drop on that set to the drop on Test-D.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is a dataset construction and evaluation strategy rather than a new detector. It assembles 50 forgery methods across six categories (face swapping, reenactment, entire face synthesis, editing, super-resolution, and manual Photoshop), filters faces by ethnicity, age, pose, occlusion, background, and lighting, pools real images from four sources, and applies a multi-level degradation pipeline consisting of blur, noise, sharpening, compression, geometric transforms, and adversarial patches to a dedicated Test-D set. The paper reports that every tested detector loses accuracy on Test-D and that frequency-domain methods degrade most sharply, while cross-dataset evaluations indicate that models trained on MFFI transfer to unseen benchmarks and that models trained on FF++ or DF40 generalize better to MFFI than to previous test sets. These results are the evidence offered for the claim that MFFI provides superior scene complexity, cross-domain generalization capability, and detection-difficulty gradients.

Load-bearing premise

The central claim depends on the assumption that the degradation operations applied to Test-D faithfully mimic real-world image propagation and do not create shortcut cues that let a detector separate degraded fakes from degraded reals without genuine robustness.

Editorial extensions

If this is right

  • Models trained on MFFI reach comparable cross-dataset AUC on unseen benchmarks such as CDF-V1, CDF-V2, DFD, and DFDC, so the dataset functions as a general-purpose training resource rather than a style-specific one.
  • Because the frequency-domain detector SRM loses about 0.21 accuracy on Test-D while spatial detectors hold up better, real-world deployments should expect frequency-only cues to fail under transmission degradation and should design detectors that mix spatial and frequency evidence.
  • Zero-shot multimodal large language models do not beat specialized small detectors on MFFI; the paper reports overall accuracy below 0.68 for the best large models, and one model collapses toward labeling nearly all samples as fake.
  • Inclusion of 2025 commercial generators means MFFI evaluates detectors against forgery technology released after most existing benchmarks were built, covering the latest method frontier.
  • The dataset already anchors a global deepfake detection challenge with 1500 participating teams, so its utility as a community benchmark is being tested beyond the paper's own experiments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An unstated consequence is that Test-D should be audited for shortcut artifacts before being adopted as a robustness standard, because the degradation parameters and patch locations are fixed and a detector could memorize them.
  • Because MFFI contains 50 method labels, it invites a stronger 'unknown forgery' protocol than the paper reports: train on a random subset of methods and test on held-out methods to create a controlled generalization ladder.
  • The paper's binary-label limitation suggests a natural next benchmark: adding method-level and region-level annotations would let MFFI also measure localization and interpretability, not just binary detection.
  • Since video-based forgery methods are converted to frames, temporal cues are absent from the image benchmark; linking MFFI frames back to their source clips could unify image-level and video-level deepfake evaluation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces MFFI, a large-scale face forgery image dataset that integrates 50 forgery methods and 1024K samples across six categories (face swapping, reenactment, entire face synthesis, editing, super-resolution, and manual Photoshop), with a four-dimensional design targeting wider forgery methods, varied facial scenes, diversified authentic data, and multi-level degradation operations. The paper evaluates four detectors (Xception, RFM, SRM, SPSL) under intra-dataset, cross-dataset, and zero-shot MLLM protocols, and claims that MFFI outperforms existing public datasets in scene complexity, cross-domain generalization capability, and detection difficulty gradients. The dataset is publicly released and served as the basis for a Kaggle challenge.

Significance. If the four-dimensional construction and the superiority claims are substantiated, MFFI would be one of the most comprehensive and realistic face forgery benchmarks currently available, and the large number of forgery methods and real-world degradation operations would be a valuable resource for the deepfake detection community. The public release, the scale of the dataset, and the use of the DeepfakeBench protocol for experimental reproducibility are clear strengths. However, the current experiments do not provide a head-to-head comparison with baseline datasets, and key terms such as 'detection difficulty gradients' are undefined, so the significance of the paper currently rests on the dataset itself rather than on the verification of its claimed advantages.

major comments (4)
  1. [§4.4, Tables 4 and 5; Abstract and §1] The claim that MFFI outperforms existing public datasets in cross-domain generalization is not supported by the reported experiments. Table 5 reports AUC for models trained only on MFFI and tested on CDF-V1, CDF-V2, DFD, and DFDC, while Table 4 trains models on DF40-FS or FF++ and tests them on MFFI Test and Test-D; neither protocol is a matched head-to-head comparison. To substantiate the claim, the same detector architectures must be trained on MFFI and on each baseline dataset under identical protocols, with all trained models evaluated on a common set of unseen datasets. As written, the numbers in Table 5 could be typical of any training set, and the lower scores in Table 4 show only that MFFI is a difficult target, not that training on it yields better generalization.
  2. [Abstract, §1, §5, Figure 1] The term detection difficulty gradients (also phrased as gradient-based detection difficulty in §5) is never defined or quantified anywhere in the paper, and the Real-World Coefficients in the Figure 1 caption are likewise undefined. Without an operational definition and a measurement protocol, the claim that MFFI outperforms existing public datasets in detection difficulty gradients cannot be assessed. The authors should either define and measure these quantities for MFFI and the baselines, or remove the claim from the abstract and conclusions.
  3. [§1, §3.3, Table 1] The scene complexity advantage is asserted using only the qualitative checkmarks in Table 1 and the descriptions in §3.3; no quantitative distributions (e.g., ethnicity, age, pose, occlusion, background, lighting) are reported for MFFI or for the baseline datasets, and no statistical test or diversity metric is provided. Since scene complexity is one of the three headline advantages claimed in the abstract, the authors should report such distributions and a comparison with the baselines, or temper the claim.
  4. [§3.5, Table 3] The Test-D construction is not described in enough detail to rule out shortcut artifacts. The degradation operations (uniform/Gaussian blur, noise, sharpening, compression, geometric transforms, PatchAttack) are listed without parameter ranges, sampling distributions, or example images, and there is no experiment isolating whether the Test-D accuracy drops reflect realistic degradation difficulty or detector-identifiable artifacts introduced by the pipeline. To support the real-world robustness interpretation, the paper should specify the degradation protocol, show qualitative examples, and include an analysis demonstrating that the drops are not attributable to simple artifact cues.
minor comments (5)
  1. [§3.2, Table 1] The list of Face Swapping methods in §3.2 contains 16 names (SimSwap, FaceShifter, FaceFusion, FSGAN, InfoSwap, Stable-Diffusion-1.5, HiFiface, IPAdapter, MegaFS, MobileFaceSwap, FaceMerge, RAFSwap, e4s, AIM, I2G, SBI) while the text and Table 1 state 15; please reconcile the count and list.
  2. [§4.4] The sentence beginning 'we observe that SRM [41] trained FF++ (C23) [56] demonstrates superior generalizability' is grammatically incomplete and misstates the direction of the improvement; it should read that SRM trained on FF++ achieves 0.0974 higher accuracy on MFFI Test than the DF40-FS-trained SRM.
  3. [Table 2] The Test and Test-D columns have identical sample counts in every row; please state explicitly that Test-D is derived from Test by applying the degradation operations, and clarify whether the 1024K total counts each test image twice.
  4. [Figure 4] The confusion-matrix definitions in the caption (TN as correctly identified fake, TP as correctly identified real) are nonstandard; please rename or clarify to avoid ambiguity with conventional usage.
  5. [Table 1] The header 'Fake Smaples' contains a typo and should read 'Fake Samples'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: MFFI's benchmark measurements are empirical detector outputs, and its comparative claims, though under-supported, are not derived from their own inputs.

full rationale

This is a dataset-and-benchmark paper rather than a derivation from first principles, so the standard circularity patterns do not apply: there is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via self-citation. The main quantitative claims come from actual detector evaluations: intra-dataset AUC/EER/AP on MFFI Test and Test-D (Table 3), cross-dataset tests with models trained on DF40-FS or FF++(C23) and evaluated on MFFI (Table 4), and MFFI-trained models evaluated on external benchmarks CDF-V1, CDF-V2, DFD, and DFDC (Table 5). These numbers are not forced by the dataset definition; they are empirical outputs of trained models. The paper does contain self-citations, including reference [84] for the Kaggle challenge where MFFI reportedly served as the core benchmark, but these are peripheral and are not used to justify the central novelty or quality claims. The claims that MFFI 'outperforms existing public datasets' in cross-domain generalization and has 'detection difficulty gradients' are under-supported: Table 5 lacks a matched head-to-head training comparison against baseline datasets, and no difficulty-gradient metric is defined or quantified. However, under-support is a validation gap, not circularity. The Test-D performance drop is partly built into the construction because Test-D is defined as the test set with added degradations, yet the specific model rankings and degradation effects are empirical. No step in the paper reduces a predicted quantity to an input by construction, so no circularity is established.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central contribution is empirical dataset construction; it does not rest on a mathematical derivation. The listed domain assumptions are the load-bearing choices that, if invalid, would make benchmark conclusions unreliable.

free parameters (2)
  • Degradation operation settings = Not reported
    MDO (Section 3.5) applies blur, noise, color shift, compression, geometric transforms, and adversarial patches 'randomly', but no parameter ranges or probabilities are given. These choices determine Test-D difficulty and realism, so they are hand-chosen free settings on which the robustness claim rests.
  • PatchAttack patch configuration = Not reported
    PatchAttack (Section 3.5) is described as a black-box attack, but patch size, location, count, and attack budget are not specified. These affect how much Test-D performance drops.
assumptions (3)
  • domain assumption All forgery pipelines produce correctly labeled fake face images; the generated samples are representative of real-world forgeries.
    Section 3.2 lists 50/51 methods, but no per-method quality checks, visual inspection statistics, or label verification protocol are reported. Mislabeled or unrealistic samples would bias every benchmark number.
  • domain assumption Applying degradation operations only to Test-D simulates real-world propagation without creating a detectable distributional shortcut.
    Section 3.5 specifies the operation types but not intensities; the realism of this simulation is assumed rather than measured, and Test-D and Test share identical images before perturbation.
  • domain assumption The DeepfakeBench training protocol and the four chosen detectors are a neutral and sufficient evaluation standard for comparing datasets.
    Section 4.1 states all preprocessing and training follow DeepfakeBench; the paper provides no independent verification that results generalize beyond these detectors and protocol.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios." pith.science (2026). https://pith.science/paper/LR4POLSO

@misc{pith2026250905592,
  author       = {Pith},
  title        = {Pith review of: MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LR4POLSO}},
  note         = {Machine review of arXiv:2509.05592}
}
abstract

Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (\textbf{MFFI}) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates $50$ different forgery methods and contains $1024K$ image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at {https://github.com/inclusionConf/MFFI}.

Figures

Figures reproduced from arXiv: 2509.05592 by the authors.

Figure 1
Figure 1. Comparison with other dataset in diversity and real [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Examples of generated images from 6 major face forgery categories in our MFFI dataset. The categories include Face [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Fake Examples of the VFS dimension. and the background of the target. We implement 15 different meth￾ods, including SimSwap [7], FaceShifter [33], FaceFusion [10], FS￾GAN [50], InfoSwap [16], Stable-Diffusion-1.5 [55], HiFiface [72], IPAdapter [81], MegaFS [91], MobileFaceSwap [75], FaceMerge [14], RAFSwap [74], e4s [34], AIM [32], I2G [87], and SBI [58]. Face Reenactment (FR): It modifies the motion or expression i… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Zero-shot testing performance of MLLM on the MFFI Test and Test-D sets. True Negative (TN): Correctly identified [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

96 extracted references · 49 canonical work pages

  1. [1]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...

  2. [2]

    Diponkor Bala, Md Shamim Hossain, Mohammad Alamgir Hossain, Md Ibrahim Abdullah, Md Mizanur Rahman, Balachandran Manavalan, Naijie Gu, Moham- mad S Islam, and Zhangjin Huang. 2023. MonkeyNet: A robust deep convolutional neural network for monkeypox disease detection and classification.Neural Net- works161 (2023), 757–775

  3. [3]

    Chaitali Bhattacharyya, Hanxiao Wang, Feng Zhang, Sungho Kim, and Xiatian Zhu. 2024. Diffusion deepfake.arXiv preprint arXiv:2404.01579(2024)

  4. [4]

    Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, and Geor- gios Tzimiropoulos. 2023. Hyperreenact: one-shot reenactment via jointly learn- ing to refine and retarget faces. InProceedings of the IEEE/CVF International Conference on Computer Vision. 7149–7159

  5. [5]

    Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip Yu, and Lichao Sun. 2025. A survey of ai-generated content (aigc).Comput. Surveys57, 5 (2025), 1–38

  6. [6]

    Junsong Chen, YU Jincheng, GE Chongjian, Lewei Yao, Enze Xie, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2024. PixArt-alpha: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. In The Twelfth International Conference on Learning Representations

  7. [7]

    Renwang Chen, Xuanhong Chen, Bingbing Ni, and Yanhao Ge. 2020. Simswap: An efficient framework for high fidelity face swapping. InProceedings of the 28th ACM international conference on multimedia. 2003–2011

  8. [8]

    François Chollet. 2017. Xception: Deep learning with depthwise separable con- volutions. InProceedings of the IEEE conference on computer vision and pattern recognition. 1251–1258

Show all 96 references
  1. [9]

    DeepFakes. 2019. https://github.com/deepfakes/

  2. [10]

    DeepFakes. 2023. https://github.com/facefusion

  3. [11]

    Yunfeng Diao, Baiqi Wu, Ruixuan Zhang, Ajian Liu, Xingxing Wei, Meng Wang, and He Wang. 2025. TASAR: Transfer-based Attack on Skeletal Action Recogni- tion. InThe International Conference on Learning Representations (ICLR)

  4. [12]

    Yunfeng Diao, Naixin Zhai, Changtao Miao, Zitong Yu, Xingxing Wei, Xun Yang, and Meng Wang. 2024. Vulnerabilities in ai-generated image detection: The challenge of adversarial attacks.arXiv preprint arXiv:2407.20836(2024)

  5. [13]

    Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. 2020. The deepfake detection challenge (dfdc) dataset.arXiv preprint arXiv:2006.07397(2020)

  6. [14]

    Paul Dubois. 2025. Face Merge 2. https://github.com/pauldubois98/FaceMerge2

  7. [15]

    FaceSwap. 2019. https://github.com/MarekKowalski/FaceSwap

  8. [16]

    Gege Gao, Huaibo Huang, Chaoyou Fu, Zhaoyang Li, and Ran He. 2021. Infor- mation bottleneck disentanglement for identity swapping. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3404–3413

  9. [17]

    Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets.Advances in neural information processing systems27 (2014)

  10. [18]

    Yinan He, Bei Gan, Siyu Chen, Yichun Zhou, Guojun Yin, Luchuan Song, Lu Sheng, Jing Shao, and Ziwei Liu. 2021. Forgerynet: A versatile benchmark for comprehensive forgery analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4360–4369

  11. [19]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models.Advances in neural information processing systems33 (2020), 6840–6851

  12. [20]

    Fa-Ting Hong and Dan Xu. 2023. Implicit identity representation conditioned memory compensation network for talking head video generation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 23062–23072

  13. [21]

    Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. 2022. Depth-aware genera- tive adversarial network for talking head video generation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3397–3406

  14. [22]

    Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detec- tion. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2889–2898

  15. [23]

    Jimeng Team. 2025. Jimeng3.0. https://jimeng.jianying.com/ai-tool/video/ generate. Accessed: 2025-05-30

  16. [24]

    Yan Ju, Shan Jia, Jialing Cai, Haiying Guan, and Siwei Lyu. 2023. Glff: Global and local feature fusion for ai-synthesized image detection.IEEE Transactions on Multimedia26 (2023), 4073–4085

  17. [25]

    Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks.Ad- vances in neural information processing systems34 (2021), 852–863

  18. [26]

    Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator ar- chitecture for generative adversarial networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4401–4410

  19. [27]

    Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8110–8119

  20. [28]

    Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo. 2021. FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)

  21. [29]

    Kling Team. 2025. Kling1.6. https://klingai.com/global/. Accessed: 2025-05-30

  22. [30]

    Chenqi Kong, Baoliang Chen, Haoliang Li, Shiqi Wang, Anderson Rocha, and Sam Kwong. 2022. Detect and locate: Exposing face manipulation by semantic-and noise-level telltales.IEEE Transactions on Information Forensics and Security17 (2022), 1741–1756

  23. [31]

    Pavel Korshunov and Sébastien Marcel. 2018. Deepfakes: a new threat to face recognition? assessment and detection.arXiv preprint arXiv:1812.08685(2018)

  24. [32]

    Jianshu Li, Man Luo, Jian Liu, Tao Chen, Chengjie Wang, Ziwei Liu, Shuo Liu, Kewei Yang, Xuning Shao, Kang Chen, et al. 2022. Multi-Forgery Detection Chal- lenge 2022: Push the Frontier of Unconstrained and Diverse Forgery Detection. arXiv preprint arXiv:2207.13505(2022)

  25. [33]

    Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. 2019. Faceshifter: Towards high fidelity and occlusion aware face swapping.arXiv preprint arXiv:1912.13457(2019)

  26. [34]

    Maomao Li, Ge Yuan, Cairong Wang, Zhian Liu, Yong Zhang, Yongwei Nie, Jue Wang, and Dong Xu. 2023. E4S: Fine-grained Face Swapping via Editing With Regional GAN Inversion.arXiv preprint arXiv:2310.15081(2023)

  27. [35]

    Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-df: A large-scale challenging dataset for deepfake forensics. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3207–3216

  28. [36]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning

  29. [37]

    Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weiming Zhang, and Nenghai Yu. 2021. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 772–781

  30. [38]

    Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2018. Large-scale celeb- faces attributes (celeba) dataset.Retrieved August15, 2018 (2018), 11

  31. [39]

    Anwei Luo, Chenqi Kong, Jiwu Huang, Yongjian Hu, Xiangui Kang, and Alex C Kot. 2023. Beyond the prior forgery knowledge: Mining critical clues for general face forgery detection.IEEE Transactions on Information Forensics and Security 19 (2023), 1168–1182

  32. [40]

    Wuyang Luo, Su Yang, and Weishan Zhang. 2023. Reference-guided large-scale face inpainting with identity and texture control.IEEE Transactions on Circuits and Systems for Video Technology33, 10 (2023), 5498–5509

  33. [41]

    Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. 2021. Generalizing face forgery detection with high-frequency features. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16317–16326

  34. [42]

    Wang Mei and Deng Weihong. 2020. Mitigating bias in face recognition using skewness-aware reinforcement learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9322–9331

  35. [43]

    Mellis, Inc. 2024. Pika (Version 1.5). https://pika.art/. Accessed: May 30, 2025. Version 1.5 reportedly available around October 2024

  36. [44]

    Changtao Miao, Qi Chu, Weihai Li, Suichan Li, Zhentao Tan, Wanyi Zhuang, and Nenghai Yu. 2022. Learning Forgery Region-aware and ID-independent Features for Face Manipulation Detection.IEEE TBIOM4, 1 (2022), 71–84

  37. [45]

    Changtao Miao, Qi Chu, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Yue Wu, Bin Liu, Honggang Hu, and Nenghai Yu. 2023. Multi-spectral Class Center Network for Face Manipulation Detection and Localization.arXiv preprint arXiv:2305.10794 (2023)

  38. [46]

    Changtao Miao, Zichang Tan, Qi Chu, Huan Liu, Honggang Hu, and Nenghai Yu. 2023. F 2 Trans: High-Frequency Fine-Grained Transformer for Face Forgery Detection.IEEE Transactions on Information Forensics and Security18 (2023), 1039–1051

  39. [47]

    Changtao Miao, Zichang Tan, Qi Chu, Nenghai Yu, and Guodong Guo. 2022. Hier- archical frequency-assisted interactive networks for face manipulation detection. IEEE Transactions on Information Forensics and Security17 (2022), 3008–3021

  40. [48]

    Aakash Varma Nadimpalli and Ajita Rattani. 2022. On improving cross-dataset generalization of deepfake detectors. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 91–99. MM ’25, October 27–31, 2025, Dublin, Ireland Changtao Miao et al

  41. [49]

    Kartik Narayan, Harsh Agarwal, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, and Richa Singh. 2023. DF-Platter: Multi-Face Heterogeneous Deepfake Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9739–9748

  42. [50]

    Yuval Nirkin, Yosi Keller, and Tal Hassner. 2019. Fsgan: Subject agnostic face swapping and reenactment. InProceedings of the IEEE/CVF international conference on computer vision. 7184–7193

  43. [51]

    PixVerse. 2025. PixVerse AI Video Generator. https://platform.pixverse.ai/. Ac- cessed: May 30, 2025

  44. [52]

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. SDXL: Improving Latent Dif- fusion Models for High-Resolution Image Synthesis. InThe Twelfth International Conference on Learning Representations

  45. [53]

    KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar

  46. [54]

    Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. 2021. Encoding in style: a stylegan encoder for image-to-image translation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2287–2296

  47. [55]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695

  48. [56]

    Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to detect manipulated facial images. InProceedings of the IEEE/CVF international conference on computer vision. 1–11

  49. [57]

    shaoanlu. 2019. faceswap-GAN. https://github.com/shaoanlu/faceswap-GAN. Accessed: 2025-05-30

  50. [58]

    Kaede Shiohara and Toshihiko Yamasaki. 2022. Detecting deepfakes with self- blended images. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18720–18729

  51. [59]

    Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021. Motion representations for articulated animation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13653– 13662

  52. [60]

    Haixu Song, Shiyu Huang, Yinpeng Dong, and Wei-Wei Tu. 2023. Robustness and generalizability of deepfake detection: A study with diffusion models.arXiv preprint arXiv:2309.02218(2023)

  53. [61]

    Sparking Innovations Limited. 2025. Vivago AI Video Generation. https://vivago. ai/video-generation. Accessed: May 30, 2025

  54. [62]

    Zichang Tan, Zhichao Yang, Changtao Miao, and Guodong Guo. 2022. Transformer-based feature compensation and aggregation for deepfake detection. IEEE Signal Processing Letters29 (2022), 2183–2187

  55. [63]

    Hailuo AI Team. 2025. Hailuo AI: Idea to Visual - I2V-01-Live Model. https: //www.hailuoai.com

  56. [64]

    Justus Thies, Michael Zollhöfer, and Matthias Nießner. 2019. Deferred neural rendering: Image synthesis using neural textures.Acm Transactions on Graphics (TOG)38, 4 (2019), 1–12

  57. [65]

    Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. 2016. Face2face: Real-time face capture and reenactment of rgb videos. InProceedings of the IEEE conference on computer vision and pattern recognition. 2387–2395

  58. [66]

    Vidu Team. 2025. Vidu2.0. https://www.vidu.com/. Accessed: 2025-05-30

  59. [67]

    Wan Team. 2025. Wan2.1. https://wan.video/. Accessed: 2025-05-30

  60. [68]

    Chengrui Wang and Weihong Deng. 2021. Representative Forgery Mining for Fake Face Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14923–14932

  61. [69]

    Mei Wang, Weihong Deng, Jiani Hu, Xunqiang Tao, and Yaohai Huang. 2019. Racial faces in the wild: Reducing racial bias by information maximization adap- tation network. InProceedings of the ieee/cvf international conference on computer vision. 692–702

  62. [70]

    Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, Anthony Chen, Huaxia Li, Xu Tang, and Yao Hu. 2024. Instantid: Zero-shot identity-preserving generation in seconds.arXiv preprint arXiv:2401.07519(2024)

  63. [71]

    Tengfei Wang, Yong Zhang, Yanbo Fan, Jue Wang, and Qifeng Chen. 2022. High- fidelity gan inversion for image attribute editing. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11379–11388

  64. [72]

    Yuhan Wang, Xu Chen, Junwei Zhu, Wenqing Chu, Ying Tai, Chengjie Wang, Jilin Li, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2021. Hififace: 3d shape and semantic prior guided high fidelity face swapping.arXiv preprint arXiv:2106.09965 (2021)

  65. [73]

    Huawei Wei, Zejun Yang, and Zhisheng Wang. 2024. Aniportrait: Audio-driven synthesis of photorealistic portrait animation.arXiv preprint arXiv:2403.17694 (2024)

  66. [74]

    Chao Xu, Jiangning Zhang, Miao Hua, Qian He, Zili Yi, and Yong Liu. 2022. Region- aware face swapping. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7632–7641

  67. [75]

    Zhiliang Xu, Zhibin Hong, Changxing Ding, Zhen Zhu, Junyu Han, Jingtuo Liu, and Errui Ding. 2022. Mobilefaceswap: A lightweight framework for video face swapping. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 2973–2981

  68. [76]

    Zhiyuan Yan, Taiping Yao, Shen Chen, Yandan Zhao, Xinghe Fu, Junwei Zhu, Donghao Luo, Chengjie Wang, Shouhong Ding, Yunsheng Wu, et al. 2024. DF40: Toward Next-Generation Deepfake Detection. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Be...

  69. [77]

    Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu. 2023. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. InAd- vances in Neural Information Processing Systems. 4534–4565

  70. [78]

    Chenglin Yang, Adam Kortylewski, Cihang Xie, Yinzhi Cao, and Alan Yuille. 2020. Patchattack: A black-box texture-based attack with reinforcement learning. In European Conference on Computer Vision. Springer, 681–698

  71. [79]

    Shuai Yang, Liming Jiang, Ziwei Liu, and Chen Change Loy. 2022. Pastiche master: Exemplar-based high-resolution portrait style transfer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7693–7702

  72. [80]

    Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang. 2021. Gan prior embedded network for blind face restoration in the wild. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 672–681

  73. [81]

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models.arXiv preprint arXiv:2308.06721(2023)

  74. [82]

    Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z Li. 2014. Learning face representa- tion from scratch.arXiv preprint arXiv:1411.7923(2014)

  75. [83]

    Ruixuan Zhang, He Wang, Zhengyu Zhao, Zhiqing Guo, Xun Yang, Yunfeng Diao, and Meng Wang. 2025. Adversarially Robust AI-Generated Image Detection for Free: An Information Theoretic Perspective.arXiv preprint arXiv:2505.22604 (2025)

  76. [84]

    Yi Zhang, Weize Gao, Changtao Miao, Man Luo, Jianshu Li, Wenzhong Deng, Zhe Li, Bingyu Hu, Weibin Yao, Wenbo Zhou, et al. 2024. Inclusion 2024 Global Multi- media Deepfake Detection: Towards Multi-dimensional Facial Forgery Detection. arXiv preprint arXiv:2412.20833(2024)

  77. [85]

    Yaning Zhang, Zitong Yu, Tianyi Wang, Xiaobin Huang, Linlin Shen, Zan Gao, and Jianfeng Ren. 2024. Genface: A large-scale fine-grained face forgery benchmark and cross appearance-edge learning.IEEE Transactions on Information Forensics and Security(2024)

  78. [86]

    Yi Zhang, Youjun Zhao, Yuhang Wen, Zixuan Tang, Xinhua Xu, and Mengyuan Liu. 2021. Facial prior based first order motion model for micro-expression generation. InProceedings of the 29th ACM International Conference on Multimedia. 4755–4759

  79. [87]

    Tianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding, Yuanjun Xiong, and Wei Xia

  80. [88]

    Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...

  81. [89]

    Mingjian Zhu, Hanting Chen, Qiangyu Yan, Xudong Huang, Guanyu Lin, Wei Li, Zhijun Tu, Hailin Hu, Jie Hu, and Yunhe Wang. 2023. Genimage: A million-scale benchmark for detecting ai-generated image.Advances in Neural Information Processing Systems36 (2023), 77771–77782

  82. [90]

    Tianlei Zhu, Junqi Chen, Renzhe Zhu, and Gaurav Gupta. 2023. StyleGAN3: generative networks for improving the equivariance of translation and rotation. arXiv preprint arXiv:2307.03898(2023)

  83. [91]

    Yuhao Zhu, Qi Li, Jian Wang, Cheng-Zhong Xu, and Zhenan Sun. 2021. One shot face swapping on megapixels. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4834–4844

  84. [92]

    Wanyi Zhuang, Qi Chu, Zhentao Tan, Qiankun Liu, Haojie Yuan, Changtao Miao, Zixiang Luo, and Nenghai Yu. 2022. UIA-ViT: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. InComputer Vision–ECCV 2022: 17th European Conference, Tel A ...

  85. [93]

    Wanyi Zhuang, Qi Chu, Haojie Yuan, Changtao Miao, Bin Liu, and Nenghai Yu

  86. [2020]

    In Proceedings of the 28th ACM international conference on multimedia

    A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia. 484–492

  87. [2021]

    InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

    Learning Self-Consistency for Deepfake Detection. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 15023–15033

  88. [2022]

    In2022 IEEE International Conference on Multimedia and Expo (ICME)

    Towards intrinsic common discriminative features learning for face forgery detection using adversarial learning. In2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.