REVIEW 4 major objections 5 minor 96 references
MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read MFFI is the first face forgery dataset to combine 50 forgery methods, varied scenes, diverse authentic faces, and transmission degradation into one 1,024,000-image benchmark.
desk verdict Useful, checkable dataset resource with real coverage breadth, but the cross-domain generalization and difficulty-gradient claims outrun the experiments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dataset itself, built along four construction axes. Wider Forgery Methods contributes 50 generators, including 2025 commercial models, grouped into six forgery classes. Varied Facial Scenes applies a six-way filter so forged and real faces span ethnicities, ages, poses, occlusions, backgrounds, and lighting. Diversified Authentic Data pools real faces from four different sources. Multi-level Degradation Operations adds conventional distortions (blur, noise, sharpening, compression, geometric transforms) plus a black-box adversarial patch attack to the test set only. The evaluation machinery is a three-protocol benchmark: intra-dataset, cross-dataset, and zero-shot multimodal large-model testing, using the training configuration and metrics (ACC, AUC, EER, AP) defined by the reference benchmark implementation cited in the paper.
What would settle it
Train a small classifier on Test-D to distinguish fake from real, then check whether the degradation type or the patch location alone predicts the label near-perfectly; if it does, the Test-D difficulty gradient reflects artifacts of the degradation pipeline rather than real-world robustness. A complementary check is to build a matched set of images that passed through actual social-platform upload and download cycles and compare detector accuracy drop on that set to the drop on Test-D.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is a dataset construction and evaluation strategy rather than a new detector. It assembles 50 forgery methods across six categories (face swapping, reenactment, entire face synthesis, editing, super-resolution, and manual Photoshop), filters faces by ethnicity, age, pose, occlusion, background, and lighting, pools real images from four sources, and applies a multi-level degradation pipeline consisting of blur, noise, sharpening, compression, geometric transforms, and adversarial patches to a dedicated Test-D set. The paper reports that every tested detector loses accuracy on Test-D and that frequency-domain methods degrade most sharply, while cross-dataset evaluations indicate that models trained on MFFI transfer to unseen benchmarks and that models trained on FF++ or DF40 generalize better to MFFI than to previous test sets. These results are the evidence offered for the claim that MFFI provides superior scene complexity, cross-domain generalization capability, and detection-difficulty gradients.
Load-bearing premise
The central claim depends on the assumption that the degradation operations applied to Test-D faithfully mimic real-world image propagation and do not create shortcut cues that let a detector separate degraded fakes from degraded reals without genuine robustness.
Editorial extensions
If this is right
- Models trained on MFFI reach comparable cross-dataset AUC on unseen benchmarks such as CDF-V1, CDF-V2, DFD, and DFDC, so the dataset functions as a general-purpose training resource rather than a style-specific one.
- Because the frequency-domain detector SRM loses about 0.21 accuracy on Test-D while spatial detectors hold up better, real-world deployments should expect frequency-only cues to fail under transmission degradation and should design detectors that mix spatial and frequency evidence.
- Zero-shot multimodal large language models do not beat specialized small detectors on MFFI; the paper reports overall accuracy below 0.68 for the best large models, and one model collapses toward labeling nearly all samples as fake.
- Inclusion of 2025 commercial generators means MFFI evaluates detectors against forgery technology released after most existing benchmarks were built, covering the latest method frontier.
- The dataset already anchors a global deepfake detection challenge with 1500 participating teams, so its utility as a community benchmark is being tested beyond the paper's own experiments.
Reading between the lines
- An unstated consequence is that Test-D should be audited for shortcut artifacts before being adopted as a robustness standard, because the degradation parameters and patch locations are fixed and a detector could memorize them.
- Because MFFI contains 50 method labels, it invites a stronger 'unknown forgery' protocol than the paper reports: train on a random subset of methods and test on held-out methods to create a controlled generalization ladder.
- The paper's binary-label limitation suggests a natural next benchmark: adding method-level and region-level annotations would let MFFI also measure localization and interpretability, not just binary detection.
- Since video-based forgery methods are converted to frames, temporal cues are absent from the image benchmark; linking MFFI frames back to their source clips could unify image-level and video-level deepfake evaluation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces MFFI, a large-scale face forgery image dataset that integrates 50 forgery methods and 1024K samples across six categories (face swapping, reenactment, entire face synthesis, editing, super-resolution, and manual Photoshop), with a four-dimensional design targeting wider forgery methods, varied facial scenes, diversified authentic data, and multi-level degradation operations. The paper evaluates four detectors (Xception, RFM, SRM, SPSL) under intra-dataset, cross-dataset, and zero-shot MLLM protocols, and claims that MFFI outperforms existing public datasets in scene complexity, cross-domain generalization capability, and detection difficulty gradients. The dataset is publicly released and served as the basis for a Kaggle challenge.
Significance. If the four-dimensional construction and the superiority claims are substantiated, MFFI would be one of the most comprehensive and realistic face forgery benchmarks currently available, and the large number of forgery methods and real-world degradation operations would be a valuable resource for the deepfake detection community. The public release, the scale of the dataset, and the use of the DeepfakeBench protocol for experimental reproducibility are clear strengths. However, the current experiments do not provide a head-to-head comparison with baseline datasets, and key terms such as 'detection difficulty gradients' are undefined, so the significance of the paper currently rests on the dataset itself rather than on the verification of its claimed advantages.
major comments (4)
- [§4.4, Tables 4 and 5; Abstract and §1] The claim that MFFI outperforms existing public datasets in cross-domain generalization is not supported by the reported experiments. Table 5 reports AUC for models trained only on MFFI and tested on CDF-V1, CDF-V2, DFD, and DFDC, while Table 4 trains models on DF40-FS or FF++ and tests them on MFFI Test and Test-D; neither protocol is a matched head-to-head comparison. To substantiate the claim, the same detector architectures must be trained on MFFI and on each baseline dataset under identical protocols, with all trained models evaluated on a common set of unseen datasets. As written, the numbers in Table 5 could be typical of any training set, and the lower scores in Table 4 show only that MFFI is a difficult target, not that training on it yields better generalization.
- [Abstract, §1, §5, Figure 1] The term detection difficulty gradients (also phrased as gradient-based detection difficulty in §5) is never defined or quantified anywhere in the paper, and the Real-World Coefficients in the Figure 1 caption are likewise undefined. Without an operational definition and a measurement protocol, the claim that MFFI outperforms existing public datasets in detection difficulty gradients cannot be assessed. The authors should either define and measure these quantities for MFFI and the baselines, or remove the claim from the abstract and conclusions.
- [§1, §3.3, Table 1] The scene complexity advantage is asserted using only the qualitative checkmarks in Table 1 and the descriptions in §3.3; no quantitative distributions (e.g., ethnicity, age, pose, occlusion, background, lighting) are reported for MFFI or for the baseline datasets, and no statistical test or diversity metric is provided. Since scene complexity is one of the three headline advantages claimed in the abstract, the authors should report such distributions and a comparison with the baselines, or temper the claim.
- [§3.5, Table 3] The Test-D construction is not described in enough detail to rule out shortcut artifacts. The degradation operations (uniform/Gaussian blur, noise, sharpening, compression, geometric transforms, PatchAttack) are listed without parameter ranges, sampling distributions, or example images, and there is no experiment isolating whether the Test-D accuracy drops reflect realistic degradation difficulty or detector-identifiable artifacts introduced by the pipeline. To support the real-world robustness interpretation, the paper should specify the degradation protocol, show qualitative examples, and include an analysis demonstrating that the drops are not attributable to simple artifact cues.
minor comments (5)
- [§3.2, Table 1] The list of Face Swapping methods in §3.2 contains 16 names (SimSwap, FaceShifter, FaceFusion, FSGAN, InfoSwap, Stable-Diffusion-1.5, HiFiface, IPAdapter, MegaFS, MobileFaceSwap, FaceMerge, RAFSwap, e4s, AIM, I2G, SBI) while the text and Table 1 state 15; please reconcile the count and list.
- [§4.4] The sentence beginning 'we observe that SRM [41] trained FF++ (C23) [56] demonstrates superior generalizability' is grammatically incomplete and misstates the direction of the improvement; it should read that SRM trained on FF++ achieves 0.0974 higher accuracy on MFFI Test than the DF40-FS-trained SRM.
- [Table 2] The Test and Test-D columns have identical sample counts in every row; please state explicitly that Test-D is derived from Test by applying the degradation operations, and clarify whether the 1024K total counts each test image twice.
- [Figure 4] The confusion-matrix definitions in the caption (TN as correctly identified fake, TP as correctly identified real) are nonstandard; please rename or clarify to avoid ambiguity with conventional usage.
- [Table 1] The header 'Fake Smaples' contains a typo and should read 'Fake Samples'.
Circularity Check
No significant circularity: MFFI's benchmark measurements are empirical detector outputs, and its comparative claims, though under-supported, are not derived from their own inputs.
full rationale
This is a dataset-and-benchmark paper rather than a derivation from first principles, so the standard circularity patterns do not apply: there is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via self-citation. The main quantitative claims come from actual detector evaluations: intra-dataset AUC/EER/AP on MFFI Test and Test-D (Table 3), cross-dataset tests with models trained on DF40-FS or FF++(C23) and evaluated on MFFI (Table 4), and MFFI-trained models evaluated on external benchmarks CDF-V1, CDF-V2, DFD, and DFDC (Table 5). These numbers are not forced by the dataset definition; they are empirical outputs of trained models. The paper does contain self-citations, including reference [84] for the Kaggle challenge where MFFI reportedly served as the core benchmark, but these are peripheral and are not used to justify the central novelty or quality claims. The claims that MFFI 'outperforms existing public datasets' in cross-domain generalization and has 'detection difficulty gradients' are under-supported: Table 5 lacks a matched head-to-head training comparison against baseline datasets, and no difficulty-gradient metric is defined or quantified. However, under-support is a validation gap, not circularity. The Test-D performance drop is partly built into the construction because Test-D is defined as the test set with added degradations, yet the specific model rankings and degradation effects are empirical. No step in the paper reduces a predicted quantity to an input by construction, so no circularity is established.
Assumptions & free parameters
free parameters (2)
- Degradation operation settings =
Not reported
- PatchAttack patch configuration =
Not reported
assumptions (3)
- domain assumption All forgery pipelines produce correctly labeled fake face images; the generated samples are representative of real-world forgeries.
- domain assumption Applying degradation operations only to Test-D simulates real-world propagation without creating a detectable distributional shortcut.
- domain assumption The DeepfakeBench training protocol and the four chosen detectors are a neutral and sufficient evaluation standard for comparing datasets.
Cite this review
Pith. "Pith review of MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios." pith.science (2026). https://pith.science/paper/LR4POLSO
@misc{pith2026250905592,
author = {Pith},
title = {Pith review of: MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios},
year = {2026},
howpublished = {\url{https://pith.science/paper/LR4POLSO}},
note = {Machine review of arXiv:2509.05592}
}
abstract
Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (\textbf{MFFI}) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates $50$ different forgery methods and contains $1024K$ image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at {https://github.com/inclusionConf/MFFI}.
Figures
Reference graph
Works this paper leans on
-
[1]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
arXiv 2025
-
[2]
Diponkor Bala, Md Shamim Hossain, Mohammad Alamgir Hossain, Md Ibrahim Abdullah, Md Mizanur Rahman, Balachandran Manavalan, Naijie Gu, Moham- mad S Islam, and Zhangjin Huang. 2023. MonkeyNet: A robust deep convolutional neural network for monkeypox disease detection and classification.Neural Net- works161 (2023), 757–775
2023
-
[3]
Chaitali Bhattacharyya, Hanxiao Wang, Feng Zhang, Sungho Kim, and Xiatian Zhu. 2024. Diffusion deepfake.arXiv preprint arXiv:2404.01579(2024)
arXiv 2024
-
[4]
Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, and Geor- gios Tzimiropoulos. 2023. Hyperreenact: one-shot reenactment via jointly learn- ing to refine and retarget faces. InProceedings of the IEEE/CVF International Conference on Computer Vision. 7149–7159
2023
-
[5]
Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip Yu, and Lichao Sun. 2025. A survey of ai-generated content (aigc).Comput. Surveys57, 5 (2025), 1–38
2025
-
[6]
Junsong Chen, YU Jincheng, GE Chongjian, Lewei Yao, Enze Xie, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2024. PixArt-alpha: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. In The Twelfth International Conference on Learning Representations
2024
-
[7]
Renwang Chen, Xuanhong Chen, Bingbing Ni, and Yanhao Ge. 2020. Simswap: An efficient framework for high fidelity face swapping. InProceedings of the 28th ACM international conference on multimedia. 2003–2011
2020
-
[8]
François Chollet. 2017. Xception: Deep learning with depthwise separable con- volutions. InProceedings of the IEEE conference on computer vision and pattern recognition. 1251–1258
2017
Show all 96 references
-
[9]
DeepFakes. 2019. https://github.com/deepfakes/
2019
-
[10]
DeepFakes. 2023. https://github.com/facefusion
2023
-
[11]
Yunfeng Diao, Baiqi Wu, Ruixuan Zhang, Ajian Liu, Xingxing Wei, Meng Wang, and He Wang. 2025. TASAR: Transfer-based Attack on Skeletal Action Recogni- tion. InThe International Conference on Learning Representations (ICLR)
2025
-
[12]
Yunfeng Diao, Naixin Zhai, Changtao Miao, Zitong Yu, Xingxing Wei, Xun Yang, and Meng Wang. 2024. Vulnerabilities in ai-generated image detection: The challenge of adversarial attacks.arXiv preprint arXiv:2407.20836(2024)
2024
-
[13]
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. 2020. The deepfake detection challenge (dfdc) dataset.arXiv preprint arXiv:2006.07397(2020)
2020 arXiv
-
[14]
Paul Dubois. 2025. Face Merge 2. https://github.com/pauldubois98/FaceMerge2
2025
-
[15]
FaceSwap. 2019. https://github.com/MarekKowalski/FaceSwap
2019
-
[16]
Gege Gao, Huaibo Huang, Chaoyou Fu, Zhaoyang Li, and Ran He. 2021. Infor- mation bottleneck disentanglement for identity swapping. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3404–3413
2021
-
[17]
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets.Advances in neural information processing systems27 (2014)
2014
-
[18]
Yinan He, Bei Gan, Siyu Chen, Yichun Zhou, Guojun Yin, Luchuan Song, Lu Sheng, Jing Shao, and Ziwei Liu. 2021. Forgerynet: A versatile benchmark for comprehensive forgery analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4360–4369
2021
-
[19]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models.Advances in neural information processing systems33 (2020), 6840–6851
2020
-
[20]
Fa-Ting Hong and Dan Xu. 2023. Implicit identity representation conditioned memory compensation network for talking head video generation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 23062–23072
2023
-
[21]
Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. 2022. Depth-aware genera- tive adversarial network for talking head video generation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3397–3406
2022
-
[22]
Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detec- tion. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2889–2898
2020
-
[23]
Jimeng Team. 2025. Jimeng3.0. https://jimeng.jianying.com/ai-tool/video/ generate. Accessed: 2025-05-30
2025
-
[24]
Yan Ju, Shan Jia, Jialing Cai, Haiying Guan, and Siwei Lyu. 2023. Glff: Global and local feature fusion for ai-synthesized image detection.IEEE Transactions on Multimedia26 (2023), 4073–4085
2023
-
[25]
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks.Ad- vances in neural information processing systems34 (2021), 852–863
2021
-
[26]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator ar- chitecture for generative adversarial networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4401–4410
2019
-
[27]
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8110–8119
2020
-
[28]
Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo. 2021. FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
2021
-
[29]
Kling Team. 2025. Kling1.6. https://klingai.com/global/. Accessed: 2025-05-30
2025
-
[30]
Chenqi Kong, Baoliang Chen, Haoliang Li, Shiqi Wang, Anderson Rocha, and Sam Kwong. 2022. Detect and locate: Exposing face manipulation by semantic-and noise-level telltales.IEEE Transactions on Information Forensics and Security17 (2022), 1741–1756
2022
-
[31]
Pavel Korshunov and Sébastien Marcel. 2018. Deepfakes: a new threat to face recognition? assessment and detection.arXiv preprint arXiv:1812.08685(2018)
2018 arXiv
-
[32]
Jianshu Li, Man Luo, Jian Liu, Tao Chen, Chengjie Wang, Ziwei Liu, Shuo Liu, Kewei Yang, Xuning Shao, Kang Chen, et al. 2022. Multi-Forgery Detection Chal- lenge 2022: Push the Frontier of Unconstrained and Diverse Forgery Detection. arXiv preprint arXiv:2207.13505(2022)
2022 arXiv
-
[33]
Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. 2019. Faceshifter: Towards high fidelity and occlusion aware face swapping.arXiv preprint arXiv:1912.13457(2019)
2019 arXiv
-
[34]
Maomao Li, Ge Yuan, Cairong Wang, Zhian Liu, Yong Zhang, Yongwei Nie, Jue Wang, and Dong Xu. 2023. E4S: Fine-grained Face Swapping via Editing With Regional GAN Inversion.arXiv preprint arXiv:2310.15081(2023)
2023 arXiv
-
[35]
Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-df: A large-scale challenging dataset for deepfake forensics. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3207–3216
2020
-
[36]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning
2023
-
[37]
Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weiming Zhang, and Nenghai Yu. 2021. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 772–781
2021
-
[38]
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2018. Large-scale celeb- faces attributes (celeba) dataset.Retrieved August15, 2018 (2018), 11
2018
-
[39]
Anwei Luo, Chenqi Kong, Jiwu Huang, Yongjian Hu, Xiangui Kang, and Alex C Kot. 2023. Beyond the prior forgery knowledge: Mining critical clues for general face forgery detection.IEEE Transactions on Information Forensics and Security 19 (2023), 1168–1182
2023
-
[40]
Wuyang Luo, Su Yang, and Weishan Zhang. 2023. Reference-guided large-scale face inpainting with identity and texture control.IEEE Transactions on Circuits and Systems for Video Technology33, 10 (2023), 5498–5509
2023
-
[41]
Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. 2021. Generalizing face forgery detection with high-frequency features. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 16317–16326
2021
-
[42]
Wang Mei and Deng Weihong. 2020. Mitigating bias in face recognition using skewness-aware reinforcement learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9322–9331
2020
-
[43]
Mellis, Inc. 2024. Pika (Version 1.5). https://pika.art/. Accessed: May 30, 2025. Version 1.5 reportedly available around October 2024
2024
-
[44]
Changtao Miao, Qi Chu, Weihai Li, Suichan Li, Zhentao Tan, Wanyi Zhuang, and Nenghai Yu. 2022. Learning Forgery Region-aware and ID-independent Features for Face Manipulation Detection.IEEE TBIOM4, 1 (2022), 71–84
2022
-
[45]
Changtao Miao, Qi Chu, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Yue Wu, Bin Liu, Honggang Hu, and Nenghai Yu. 2023. Multi-spectral Class Center Network for Face Manipulation Detection and Localization.arXiv preprint arXiv:2305.10794 (2023)
2023 arXiv
-
[46]
Changtao Miao, Zichang Tan, Qi Chu, Huan Liu, Honggang Hu, and Nenghai Yu. 2023. F 2 Trans: High-Frequency Fine-Grained Transformer for Face Forgery Detection.IEEE Transactions on Information Forensics and Security18 (2023), 1039–1051
2023
-
[47]
Changtao Miao, Zichang Tan, Qi Chu, Nenghai Yu, and Guodong Guo. 2022. Hier- archical frequency-assisted interactive networks for face manipulation detection. IEEE Transactions on Information Forensics and Security17 (2022), 3008–3021
2022
-
[48]
Aakash Varma Nadimpalli and Ajita Rattani. 2022. On improving cross-dataset generalization of deepfake detectors. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 91–99. MM ’25, October 27–31, 2025, Dublin, Ireland Changtao Miao et al
2022
-
[49]
Kartik Narayan, Harsh Agarwal, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, and Richa Singh. 2023. DF-Platter: Multi-Face Heterogeneous Deepfake Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9739–9748
2023
-
[50]
Yuval Nirkin, Yosi Keller, and Tal Hassner. 2019. Fsgan: Subject agnostic face swapping and reenactment. InProceedings of the IEEE/CVF international conference on computer vision. 7184–7193
2019
-
[51]
PixVerse. 2025. PixVerse AI Video Generator. https://platform.pixverse.ai/. Ac- cessed: May 30, 2025
2025
-
[52]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. SDXL: Improving Latent Dif- fusion Models for High-Resolution Image Synthesis. InThe Twelfth International Conference on Learning Representations
2024
-
[53]
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar
-
[54]
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. 2021. Encoding in style: a stylegan encoder for image-to-image translation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2287–2296
2021
-
[55]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695
2022
-
[56]
Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to detect manipulated facial images. InProceedings of the IEEE/CVF international conference on computer vision. 1–11
2019
-
[57]
shaoanlu. 2019. faceswap-GAN. https://github.com/shaoanlu/faceswap-GAN. Accessed: 2025-05-30
2019
-
[58]
Kaede Shiohara and Toshihiko Yamasaki. 2022. Detecting deepfakes with self- blended images. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18720–18729
2022
-
[59]
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021. Motion representations for articulated animation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 13653– 13662
2021
-
[60]
Haixu Song, Shiyu Huang, Yinpeng Dong, and Wei-Wei Tu. 2023. Robustness and generalizability of deepfake detection: A study with diffusion models.arXiv preprint arXiv:2309.02218(2023)
2023 arXiv
-
[61]
Sparking Innovations Limited. 2025. Vivago AI Video Generation. https://vivago. ai/video-generation. Accessed: May 30, 2025
2025
-
[62]
Zichang Tan, Zhichao Yang, Changtao Miao, and Guodong Guo. 2022. Transformer-based feature compensation and aggregation for deepfake detection. IEEE Signal Processing Letters29 (2022), 2183–2187
2022
-
[63]
Hailuo AI Team. 2025. Hailuo AI: Idea to Visual - I2V-01-Live Model. https: //www.hailuoai.com
2025
-
[64]
Justus Thies, Michael Zollhöfer, and Matthias Nießner. 2019. Deferred neural rendering: Image synthesis using neural textures.Acm Transactions on Graphics (TOG)38, 4 (2019), 1–12
2019
-
[65]
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. 2016. Face2face: Real-time face capture and reenactment of rgb videos. InProceedings of the IEEE conference on computer vision and pattern recognition. 2387–2395
2016
-
[66]
Vidu Team. 2025. Vidu2.0. https://www.vidu.com/. Accessed: 2025-05-30
2025
-
[67]
Wan Team. 2025. Wan2.1. https://wan.video/. Accessed: 2025-05-30
2025
-
[68]
Chengrui Wang and Weihong Deng. 2021. Representative Forgery Mining for Fake Face Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14923–14932
2021
-
[69]
Mei Wang, Weihong Deng, Jiani Hu, Xunqiang Tao, and Yaohai Huang. 2019. Racial faces in the wild: Reducing racial bias by information maximization adap- tation network. InProceedings of the ieee/cvf international conference on computer vision. 692–702
2019
-
[70]
Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, Anthony Chen, Huaxia Li, Xu Tang, and Yao Hu. 2024. Instantid: Zero-shot identity-preserving generation in seconds.arXiv preprint arXiv:2401.07519(2024)
2024 arXiv
-
[71]
Tengfei Wang, Yong Zhang, Yanbo Fan, Jue Wang, and Qifeng Chen. 2022. High- fidelity gan inversion for image attribute editing. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11379–11388
2022
-
[72]
Yuhan Wang, Xu Chen, Junwei Zhu, Wenqing Chu, Ying Tai, Chengjie Wang, Jilin Li, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2021. Hififace: 3d shape and semantic prior guided high fidelity face swapping.arXiv preprint arXiv:2106.09965 (2021)
2021 arXiv
-
[73]
Huawei Wei, Zejun Yang, and Zhisheng Wang. 2024. Aniportrait: Audio-driven synthesis of photorealistic portrait animation.arXiv preprint arXiv:2403.17694 (2024)
2024 arXiv
-
[74]
Chao Xu, Jiangning Zhang, Miao Hua, Qian He, Zili Yi, and Yong Liu. 2022. Region- aware face swapping. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7632–7641
2022
-
[75]
Zhiliang Xu, Zhibin Hong, Changxing Ding, Zhen Zhu, Junyu Han, Jingtuo Liu, and Errui Ding. 2022. Mobilefaceswap: A lightweight framework for video face swapping. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 2973–2981
2022
-
[76]
Zhiyuan Yan, Taiping Yao, Shen Chen, Yandan Zhao, Xinghe Fu, Junwei Zhu, Donghao Luo, Chengjie Wang, Shouhong Ding, Yunsheng Wu, et al. 2024. DF40: Toward Next-Generation Deepfake Detection. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Be...
2024
-
[77]
Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu. 2023. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. InAd- vances in Neural Information Processing Systems. 4534–4565
2023
-
[78]
Chenglin Yang, Adam Kortylewski, Cihang Xie, Yinzhi Cao, and Alan Yuille. 2020. Patchattack: A black-box texture-based attack with reinforcement learning. In European Conference on Computer Vision. Springer, 681–698
2020
-
[79]
Shuai Yang, Liming Jiang, Ziwei Liu, and Chen Change Loy. 2022. Pastiche master: Exemplar-based high-resolution portrait style transfer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7693–7702
2022
-
[80]
Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang. 2021. Gan prior embedded network for blind face restoration in the wild. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 672–681
2021
-
[81]
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models.arXiv preprint arXiv:2308.06721(2023)
2023 arXiv
-
[82]
Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z Li. 2014. Learning face representa- tion from scratch.arXiv preprint arXiv:1411.7923(2014)
2014 arXiv
-
[83]
Ruixuan Zhang, He Wang, Zhengyu Zhao, Zhiqing Guo, Xun Yang, Yunfeng Diao, and Meng Wang. 2025. Adversarially Robust AI-Generated Image Detection for Free: An Information Theoretic Perspective.arXiv preprint arXiv:2505.22604 (2025)
2025 arXiv
-
[84]
Yi Zhang, Weize Gao, Changtao Miao, Man Luo, Jianshu Li, Wenzhong Deng, Zhe Li, Bingyu Hu, Weibin Yao, Wenbo Zhou, et al. 2024. Inclusion 2024 Global Multi- media Deepfake Detection: Towards Multi-dimensional Facial Forgery Detection. arXiv preprint arXiv:2412.20833(2024)
2024 arXiv
-
[85]
Yaning Zhang, Zitong Yu, Tianyi Wang, Xiaobin Huang, Linlin Shen, Zan Gao, and Jianfeng Ren. 2024. Genface: A large-scale fine-grained face forgery benchmark and cross appearance-edge learning.IEEE Transactions on Information Forensics and Security(2024)
2024
-
[86]
Yi Zhang, Youjun Zhao, Yuhang Wen, Zixuan Tang, Xinhua Xu, and Mengyuan Liu. 2021. Facial prior based first order motion model for micro-expression generation. InProceedings of the 29th ACM International Conference on Multimedia. 4755–4759
2021
-
[87]
Tianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding, Yuanjun Xiong, and Wei Xia
-
[88]
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...
2025 arXiv
-
[89]
Mingjian Zhu, Hanting Chen, Qiangyu Yan, Xudong Huang, Guanyu Lin, Wei Li, Zhijun Tu, Hailin Hu, Jie Hu, and Yunhe Wang. 2023. Genimage: A million-scale benchmark for detecting ai-generated image.Advances in Neural Information Processing Systems36 (2023), 77771–77782
2023
-
[90]
Tianlei Zhu, Junqi Chen, Renzhe Zhu, and Gaurav Gupta. 2023. StyleGAN3: generative networks for improving the equivariance of translation and rotation. arXiv preprint arXiv:2307.03898(2023)
2023 arXiv
-
[91]
Yuhao Zhu, Qi Li, Jian Wang, Cheng-Zhong Xu, and Zhenan Sun. 2021. One shot face swapping on megapixels. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4834–4844
2021
-
[92]
Wanyi Zhuang, Qi Chu, Zhentao Tan, Qiankun Liu, Haojie Yuan, Changtao Miao, Zixiang Luo, and Nenghai Yu. 2022. UIA-ViT: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. InComputer Vision–ECCV 2022: 17th European Conference, Tel A ...
2022
-
[93]
Wanyi Zhuang, Qi Chu, Haojie Yuan, Changtao Miao, Bin Liu, and Nenghai Yu
-
[2020]
In Proceedings of the 28th ACM international conference on multimedia
A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia. 484–492
-
[2021]
InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Learning Self-Consistency for Deepfake Detection. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 15023–15033
-
[2022]
In2022 IEEE International Conference on Multimedia and Expo (ICME)
Towards intrinsic common discriminative features learning for face forgery detection using adversarial learning. In2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1–6
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.