{"id":"06d46077-ab22-4cd5-9d4c-fc53f45bc838","arxiv_id":"2509.05592","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"MFFI is a 1,024,000-image face forgery benchmark spanning 50 (listed 51) forgery techniques, varied facial scenes, multiple real-data sources, and multi-level degradation operations.","lead":"This paper introduces MFFI, a large face-forgery image dataset that combines 50 forgery methods, multiple face-image sources, and simulated transmission artifacts to mimic real-world conditions. It also benchmarks several deepfake detectors on this data, reporting that the dataset is harder and more diverse than existing benchmarks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-domain generalization claim is unsupported: no matched head-to-head training comparison between MFFI and baseline datasets on common external test sets.","rationale":"The reader's weakest_assumption focused on whether Test-D degradation faithfully simulates real-world propagation, which is a valid concern about the robustness experiments. However, I find a more direct logical gap in the paper's headline comparative claim. The abstract states MFFI 'outperforms existing public datasets' in cross-domain generalization, but Table 5 only evaluates models trained on MFFI on external test sets, and Table 4 evaluates models trained on external datasets on MFFI. Neither direction permits a comparison of training sets. The missing matched comparison means the central claim of superiority is unsubstantiated, independent of any degradation realism issues. This is a load-bearing concern because the paper's practical utility is argued through these benchmark evaluations. I agree with the reader's overall CONDITIONAL verdict: the dataset resource may be valuable, but the strong claims require additional experiments (or explicit removal). My concern does not change the verdict; it reinforces the condition that matched cross-training baselines must be added. I also note the paper deserves credit for releasing the dataset and using standard DeepfakeBench protocols, and the method-count inconsistencies are secondary but should be corrected.","tokens_in":14756,"tokens_out":6694,"duration_ms":55413,"concrete_test":"Run a matched cross-dataset protocol: train Xception, RFM, SRM, and SPSL on MFFI, FF++(C23), DF40-FS, and Celeb-DF v2 (using fixed, comparable subsets) under identical DeepfakeBench settings, and evaluate every trained model on the same held-out external sets (CDF-V1, CDF-V2, DFD, DFDC). If MFFI-trained models do not achieve higher AUC than at least one baseline on most external sets, the cross-domain generalization claim should be retracted or substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 1 claim that benchmark evaluations show MFFI outperforms existing public datasets in cross-domain generalization capability. The experiments do not establish this. Table 5 trains detectors only on MFFI and tests on CDF-V1, CDF-V2, DFD, and DFDC, reporting AUC. Table 4 trains on DF40-FS and FF++(C23) and tests on MFFI Test and Test-D. Neither protocol is a head-to-head comparison: to show that MFFI improves cross-domain generalization, one must train identical models on MFFI and on each baseline dataset under the same protocol, then evaluate all models on the same external test sets. Without this, the transfer numbers in Table 5 could be typical of any training set, and the lower scores in Table 4 only show that MFFI is a difficult target, not that training on it yields better generalization. The claim 'outperforms existing public datasets' is therefore not supported by the presented experiments. Additionally, the 'detection difficulty gradient' is never defined or quantified, so that comparative claim is also unverified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces MFFI, a large-scale face forgery image dataset that integrates 50 forgery methods and 1024K samples across six categories (face swapping, reenactment, entire face synthesis, editing, super-resolution, and manual Photoshop), with a four-dimensional design targeting wider forgery methods, varied facial scenes, diversified authentic data, and multi-level degradation operations. The paper evaluates four detectors (Xception, RFM, SRM, SPSL) under intra-dataset, cross-dataset, and zero-shot MLLM protocols, and claims that MFFI outperforms existing public datasets in scene complexity, cross-domain generalization capability, and detection difficulty gradients. The dataset is publicly released and served as the basis for a Kaggle challenge.","tokens_in":15060,"tokens_out":10494,"duration_ms":78260,"significance":"If the four-dimensional construction and the superiority claims are substantiated, MFFI would be one of the most comprehensive and realistic face forgery benchmarks currently available, and the large number of forgery methods and real-world degradation operations would be a valuable resource for the deepfake detection community. The public release, the scale of the dataset, and the use of the DeepfakeBench protocol for experimental reproducibility are clear strengths. However, the current experiments do not provide a head-to-head comparison with baseline datasets, and key terms such as 'detection difficulty gradients' are undefined, so the significance of the paper currently rests on the dataset itself rather than on the verification of its claimed advantages.","major_comments":[{"comment":"The claim that MFFI outperforms existing public datasets in cross-domain generalization is not supported by the reported experiments. Table 5 reports AUC for models trained only on MFFI and tested on CDF-V1, CDF-V2, DFD, and DFDC, while Table 4 trains models on DF40-FS or FF++ and tests them on MFFI Test and Test-D; neither protocol is a matched head-to-head comparison. To substantiate the claim, the same detector architectures must be trained on MFFI and on each baseline dataset under identical protocols, with all trained models evaluated on a common set of unseen datasets. As written, the numbers in Table 5 could be typical of any training set, and the lower scores in Table 4 show only that MFFI is a difficult target, not that training on it yields better generalization.","section":"§4.4, Tables 4 and 5; Abstract and §1"},{"comment":"The term detection difficulty gradients (also phrased as gradient-based detection difficulty in §5) is never defined or quantified anywhere in the paper, and the Real-World Coefficients in the Figure 1 caption are likewise undefined. Without an operational definition and a measurement protocol, the claim that MFFI outperforms existing public datasets in detection difficulty gradients cannot be assessed. The authors should either define and measure these quantities for MFFI and the baselines, or remove the claim from the abstract and conclusions.","section":"Abstract, §1, §5, Figure 1"},{"comment":"The scene complexity advantage is asserted using only the qualitative checkmarks in Table 1 and the descriptions in §3.3; no quantitative distributions (e.g., ethnicity, age, pose, occlusion, background, lighting) are reported for MFFI or for the baseline datasets, and no statistical test or diversity metric is provided. Since scene complexity is one of the three headline advantages claimed in the abstract, the authors should report such distributions and a comparison with the baselines, or temper the claim.","section":"§1, §3.3, Table 1"},{"comment":"The Test-D construction is not described in enough detail to rule out shortcut artifacts. The degradation operations (uniform/Gaussian blur, noise, sharpening, compression, geometric transforms, PatchAttack) are listed without parameter ranges, sampling distributions, or example images, and there is no experiment isolating whether the Test-D accuracy drops reflect realistic degradation difficulty or detector-identifiable artifacts introduced by the pipeline. To support the real-world robustness interpretation, the paper should specify the degradation protocol, show qualitative examples, and include an analysis demonstrating that the drops are not attributable to simple artifact cues.","section":"§3.5, Table 3"}],"minor_comments":[{"comment":"The list of Face Swapping methods in §3.2 contains 16 names (SimSwap, FaceShifter, FaceFusion, FSGAN, InfoSwap, Stable-Diffusion-1.5, HiFiface, IPAdapter, MegaFS, MobileFaceSwap, FaceMerge, RAFSwap, e4s, AIM, I2G, SBI) while the text and Table 1 state 15; please reconcile the count and list.","section":"§3.2, Table 1"},{"comment":"The sentence beginning 'we observe that SRM [41] trained FF++ (C23) [56] demonstrates superior generalizability' is grammatically incomplete and misstates the direction of the improvement; it should read that SRM trained on FF++ achieves 0.0974 higher accuracy on MFFI Test than the DF40-FS-trained SRM.","section":"§4.4"},{"comment":"The Test and Test-D columns have identical sample counts in every row; please state explicitly that Test-D is derived from Test by applying the degradation operations, and clarify whether the 1024K total counts each test image twice.","section":"Table 2"},{"comment":"The confusion-matrix definitions in the caption (TN as correctly identified fake, TP as correctly identified real) are nonstandard; please rename or clarify to avoid ambiguity with conventional usage.","section":"Figure 4"},{"comment":"The header 'Fake Smaples' contains a typo and should read 'Fake Samples'.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself appears to be a potentially useful resource, and the public release plus the associated challenge are positive aspects. My main concern is that the paper's headline comparative claims are not backed by the experimental design; these claims should either be supported by head-to-head experiments and defined metrics, or substantially softened. I believe the issues are addressable within the manuscript's scope, which is why I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: MFFI is a genuinely useful dataset, and the four-dimensional construction (wider methods, varied scenes, diversified real data, degradation) is a real step forward for the deepfake-detection benchmark space. The abstract's comparative claims, however, are ahead of the evidence: nothing in the experiments shows that MFFI outperforms existing datasets on cross-domain generalization or 'detection difficulty gradients,' because the needed head-to-head comparison was never run.\n\nWhat is actually new and good: the dataset bundles 50 forgery techniques (including 2025 commercial generators like Wan2.1, Vidu2.0, Kling1.6, Jimeng3.0), adds Face Super-Resolution and Face Manual Photoshop as categories that DF40 and ForgeryNet don't cover, mixes four real-face sources, and applies a multi-level degradation pipeline including a PatchAttack-based adversarial component. Training code is tied to DeepfakeBench, and the dataset already anchored a Kaggle challenge. That is a concrete, checkable resource rather than a restatement of earlier datasets. The citation pattern looks normal; the self-citation to the authors' Kaggle challenge report is context, not padding.\n\nWhere it gets soft: the cross-dataset section. Table 5 only trains on MFFI and tests on CDF/DFD/DFDC; Table 4 trains on DF40-FS and FF++ and tests on MFFI. Neither is a matched comparison. To support 'outperforms existing datasets in cross-domain generalization,' you need identical models trained on MFFI and on each baseline under the same protocol, then evaluated on shared external sets. Without that, Table 4 mostly says MFFI is a harder test, and Table 5 gives absolute numbers that could come from any decent training set. The paper also never defines or quantifies 'detection difficulty gradient,' which is odd given it appears in the abstract and conclusion. Minor but worth fixing: the method count is inconsistent (50 claimed, 16 FS techniques listed while the text and figure say 15; the FMPS count also wobbles between 6 and 7). A released dataset should get its manifest right.\n\nThe degradation-shortcut worry is reasonable but, as far as I can tell, not demonstrated. Test-D is a stress test, not a claim that the pipeline perfectly simulates social-media propagation. I'd treat it as a nice robustness probe rather than a fatal flaw.\n\nBottom line: for anyone building or evaluating deepfake detectors, MFFI deserves a serious look; the resource is the contribution. The comparative claims should be either supported with head-to-head training comparisons or cut down to what the experiments actually show. I'd accept the paper for peer review, because the dataset itself is real, checkable, and useful.\n\nBest.","headline":"Useful, checkable dataset resource with real coverage breadth, but the cross-domain generalization and difficulty-gradient claims outrun the experiments.","tokens_in":15476,"tokens_out":3711,"would_cite":true,"duration_ms":33821,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MFFI is the first face forgery dataset to combine 50 forgery methods, varied scenes, diverse authentic faces, and transmission degradation into one 1,024,000-image benchmark.","keywords":["Deepfake detection","Face forgery dataset","Real-world benchmark","Forgery method diversity","Image degradation","Cross-domain generalization","Face swapping","Diffusion model forgery"],"falsifier":"Train a small classifier on Test-D to distinguish fake from real, then check whether the degradation type or the patch location alone predicts the label near-perfectly; if it does, the Test-D difficulty gradient reflects artifacts of the degradation pipeline rather than real-world robustness. A complementary check is to build a matched set of images that passed through actual social-platform upload and download cycles and compare detector accuracy drop on that set to the drop on Test-D.","tokens_in":14568,"feed_emoji":"🎭","tokens_out":14039,"duration_ms":106153,"temperature":0.7,"pith_summary":"MFFI is a face forgery image dataset built to close the gap between laboratory deepfake benchmarks and the forgeries that actually circulate online. The paper claims that prior datasets fall short on four independent axes: coverage of modern forgery methods, variety of facial scenes, diversity of authentic (real) face data, and simulation of transmission degradation. MFFI is presented as the first dataset to handle all four axes at once, integrating 50 forgery methods and more than one million face images from four real-face sources, with a degraded test set that models internet propagation. Benchmark evaluations reported in the paper indicate that MFFI outperforms existing public datasets in scene complexity, cross-domain generalization capability, and detection difficulty gradients. If the claim is right, MFFI gives detection researchers a training and evaluation ground on which models must cope with unknown generation methods, varied faces, and messy image degradation rather than a single narrow forgery style.","feed_headline":"50 forgery methods fit in one benchmark of 1M images","feed_subtitle":"It is the first benchmark to combine all four dimensions: method variety, facial scenes, authentic data, degradation.","key_machinery":"The load-bearing object is the dataset itself, built along four construction axes. Wider Forgery Methods contributes 50 generators, including 2025 commercial models, grouped into six forgery classes. Varied Facial Scenes applies a six-way filter so forged and real faces span ethnicities, ages, poses, occlusions, backgrounds, and lighting. Diversified Authentic Data pools real faces from four different sources. Multi-level Degradation Operations adds conventional distortions (blur, noise, sharpening, compression, geometric transforms) plus a black-box adversarial patch attack to the test set only. The evaluation machinery is a three-protocol benchmark: intra-dataset, cross-dataset, and zero-shot multimodal large-model testing, using the training configuration and metrics (ACC, AUC, EER, AP) defined by the reference benchmark implementation cited in the paper.","core_discovery":"On its own terms, the paper's central discovery is a dataset construction and evaluation strategy rather than a new detector. It assembles 50 forgery methods across six categories (face swapping, reenactment, entire face synthesis, editing, super-resolution, and manual Photoshop), filters faces by ethnicity, age, pose, occlusion, background, and lighting, pools real images from four sources, and applies a multi-level degradation pipeline consisting of blur, noise, sharpening, compression, geometric transforms, and adversarial patches to a dedicated Test-D set. The paper reports that every tested detector loses accuracy on Test-D and that frequency-domain methods degrade most sharply, while cross-dataset evaluations indicate that models trained on MFFI transfer to unseen benchmarks and that models trained on FF++ or DF40 generalize better to MFFI than to previous test sets. These results are the evidence offered for the claim that MFFI provides superior scene complexity, cross-domain generalization capability, and detection-difficulty gradients.","pith_inferences":["An unstated consequence is that Test-D should be audited for shortcut artifacts before being adopted as a robustness standard, because the degradation parameters and patch locations are fixed and a detector could memorize them.","Because MFFI contains 50 method labels, it invites a stronger 'unknown forgery' protocol than the paper reports: train on a random subset of methods and test on held-out methods to create a controlled generalization ladder.","The paper's binary-label limitation suggests a natural next benchmark: adding method-level and region-level annotations would let MFFI also measure localization and interpretability, not just binary detection.","Since video-based forgery methods are converted to frames, temporal cues are absent from the image benchmark; linking MFFI frames back to their source clips could unify image-level and video-level deepfake evaluation."],"forward_implications":["Models trained on MFFI reach comparable cross-dataset AUC on unseen benchmarks such as CDF-V1, CDF-V2, DFD, and DFDC, so the dataset functions as a general-purpose training resource rather than a style-specific one.","Because the frequency-domain detector SRM loses about 0.21 accuracy on Test-D while spatial detectors hold up better, real-world deployments should expect frequency-only cues to fail under transmission degradation and should design detectors that mix spatial and frequency evidence.","Zero-shot multimodal large language models do not beat specialized small detectors on MFFI; the paper reports overall accuracy below 0.68 for the best large models, and one model collapses toward labeling nearly all samples as fake.","Inclusion of 2025 commercial generators means MFFI evaluates detectors against forgery technology released after most existing benchmarks were built, covering the latest method frontier.","The dataset already anchors a global deepfake detection challenge with 1500 participating teams, so its utility as a community benchmark is being tested beyond the paper's own experiments."],"supporting_citations":[{"why":"Provides the FaceForensics++ reference, which MFFI uses as a baseline and as a cross-dataset training source in the C23 setting.","marker":"[56]"},{"why":"Supplies the Celeb-DF videos used as unseen test sets (CDF-V1, CDF-V2), against which MFFI-trained detectors are measured.","marker":"[35]"},{"why":"Gives the large-scale DFDC dataset, used as an unseen cross-dataset test set and as the multi-method baseline whose diversity MFFI extends.","marker":"[13]"},{"why":"Defines the closest prior benchmark, DF40, with 40 forgery methods, and its FS split is used for a cross-dataset training comparison.","marker":"[76]"},{"why":"Establishes the training configuration, model backbones, and evaluation metrics (ACC, AUC, EER, AP) that all experiments follow.","marker":"[77]"},{"why":"Supplies the PatchAttack black-box perturbation used to create adversarial degradation in the Test-D set.","marker":"[78]"},{"why":"Provides the CelebA image collection, one of the four real-face sources underlying the authentic-data dimension.","marker":"[38]"},{"why":"Provides the RFW dataset with balanced racial groups, grounding the ethnicity coverage in the facial-scene dimension.","marker":"[69]"},{"why":"Provides the CASIA-WebFace collection, another large real-face source used in the diversified authentic data.","marker":"[82]"}],"fun_headline_variants":["MFFI: 1M face images spanning 50 forgery methods for real-world testing","New dataset: 50 face forgery methods across 1M real-world images","Benchmark covers 50 forgery styles and 1M faces for realistic deepfake detection","MFFI dataset: 50 forgery methods, 1M images, four real-world dimensions","1M real-world face images with 50 forgery techniques in one dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the assumption that the degradation operations applied to Test-D faithfully mimic real-world image propagation and do not create shortcut cues that let a detector separate degraded fakes from degraded reals without genuine robustness.","fun_headline_variants_meta":{"raw":{"variants":["MFFI: 1M face images spanning 50 forgery methods for real-world testing","New dataset: 50 face forgery methods across 1M real-world images","Benchmark covers 50 forgery styles and 1M faces for realistic deepfake detection","MFFI dataset: 50 forgery methods, 1M images, four real-world dimensions","1M real-world face images with 50 forgery techniques in one dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00099,"raw_usage":{"total_tokens":4209,"prompt_tokens":969,"completion_tokens":3240,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":3126}},"tokens_in":585,"tokens_out":3240,"duration_ms":21576,"temperature":1.0,"reasoning_tokens":3126,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:21:57.740954+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a small classifier on Test-D to distinguish fake from real, then check whether the degradation type or the patch location alone predicts the label near-perfectly; if it does, the Test-D difficulty gradient reflects artifacts of the degradation pipeline rather than real-world robustness. A complementary check is to build a matched set of images that passed through actual social-platform upload and download cycles and compare detector accuracy drop on that set to the drop on Test-D.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the closest prior benchmark, DF40, with 40 forgery methods, and its FS split is used for a cross-dataset training comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the training configuration, model backbones, and evaluation metrics (ACC, AUC, EER, AP) that all experiments follow."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the PatchAttack black-box perturbation used to create adversarial degradation in the Test-D set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CelebA image collection, one of the four real-face sources underlying the authentic-data dimension."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the RFW dataset with balanced racial groups, grounding the ethnicity coverage in the facial-scene dimension."}],"review_version":2}