{"id":"8a2857df-5a83-44bb-a6d4-fe83c01bbb38","arxiv_id":"2506.00868","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MultiFakeVerse provides 845,286 person-centric images edited through VLM-generated instructions; state-of-the-art deepfake detectors and human observers misclassify a large fraction of them.","lead":"This paper introduces MultiFakeVerse, a dataset of 845,286 images in which AI image editors, guided by vision-language models, made subtle changes to people, objects, and scenes in real photos. The authors report that current deepfake detectors and human viewers often fail to tell the edited images from the originals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Localization masks underpinning the 'largest spatial deepfake localization dataset' claim are auto-thresholded pixel-difference maps with unreported parameters and no human validation; for GPT-Image-1 edits, differing output resolutions may invalidate pixel-wise differencing entirely.","rationale":"The central claim has two parts: a large benchmark exists, and state-of-the-art detectors and humans struggle on it. The detection results are largely independent of the localization masks, so the masks issue does not threaten the detection half. However, the paper's conclusion explicitly claims to be 'the largest image-based dataset for spatial deepfake localization.' That contribution, and the §4.0.3 localization benchmark, depend entirely on auto-computed pixel-difference masks. The reader's weakest assumption targets exactly this dependency, and I agree it is the most load-bearing concern: if the masks are invalid, a headline contribution is false and the reported IoU/AUC numbers are uninterpretable. The concern is concrete—the threshold is unreported, no human validation is performed, and GPT-Image-1's fixed output resolutions make pixel-wise differencing ill-defined without an explicit alignment step. A 200-image human annotation study would settle it. I considered the alternative concern that the 'conceptual manipulation' labels are never validated (the same VLM family that generated the data also performs the perceptual analysis), but that concern affects framing more than the benchmark claims; the detection benchmark uses only binary real/fake labels, and the human study, though small, directly supports the 'humans struggle' claim. Thus the masks are the single most load-bearing concern, and the paper should be accepted conditional on either mask validation or removal of the localization claim. The reader's CONDITIONAL verdict therefore remains appropriate.","tokens_in":12928,"tokens_out":17301,"duration_ms":159451,"concrete_test":"Take a random sample of 200 test fake images stratified by source and manipulation level. For each, show three annotators the original side-by-side and ask them to draw a polygon around the manipulated region. Compute mean IoU between the §3.2.1 auto-mask and the annotator-consensus (majority-vote) mask. If mean IoU is less than 0.4, or if the GPT-Image-1 subset shows systematically lower agreement than the Gemini subset, the localization ground truth is not validated and the §4.0.3 numbers cannot be interpreted as evidence about localization quality. The authors should also report the threshold, noise-filter parameters, and any resizing or alignment used for pixel differencing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's conclusion claims MultiFakeVerse is 'the largest image-based dataset for spatial deepfake localization.' That claim, and the localization benchmark in §4.0.3 (SIDA-13B: IoU 13.10/24.74, F1 19.92/39.40, AUC 14.06/37.53), rests entirely on masks computed in §3.2.1 by thresholding pixel-wise differences between original and edited images followed by connected-component analysis. Three problems make these masks unreliable ground truth. First, the threshold and noise-filter parameters are never reported, so the masks are not reproducible. Second, the masks are never validated against human annotations or any independent localization signal; low IoU/AUC of a detector cannot be separated from poor ground truth versus poor detector performance. Third, §2.0.3 and §3.1.2 state that GPT-Image-1 has fixed output aspect ratios and 'tends to edit in a few cases, untargeted regions due to constraints in the output aspect ratios.' Pixel-wise differencing requires identical geometry, yet no resizing, cropping, or alignment step is reported; for GPT-generated images, masks may reflect resampling artifacts rather than semantic edits. Since GPT images are a nontrivial fraction of the 758,041 fakes (§3.2.7), the localization ground truth may be corrupted for that subset. The detection claim may survive, but the localization benchmark and the 'largest localization dataset' contribution would not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MultiFakeVerse, a large-scale person-centric deepfake dataset of 845,286 images (87,245 real, 758,041 fake) created by using VLMs to propose and execute subtle semantic edits targeting emotions, narrative, and human-object interactions. The authors benchmark zero-shot and finetuned deepfake detectors (CNNSpot, TruFor, AntiFakePrompt, SIDA-13B), report a human study (61.67% accuracy), and evaluate forgery localization on SIDA-13B using masks automatically derived from pixel differences. The central claims are that the dataset is large, semantically meaningful, and hard for both detectors and humans.","tokens_in":13097,"tokens_out":6090,"duration_ms":53561,"significance":"If the dataset and its annotations are valid, MultiFakeVerse would be a valuable resource: it is substantially larger than most prior person-centric deepfake benchmarks, uses a novel VLM-driven manipulation pipeline, and provides initial evidence that current detectors and humans struggle on these subtle semantic edits. The paper also ships code and dataset links, and the zero-shot detection results (best 66.87% accuracy) are concrete falsifiable claims. However, the spatial-localization contribution currently rests on unvalidated, insufficiently documented mask generation, and the perceptual-impact analysis is self-referential; both issues need to be resolved before the dataset can serve as a reliable benchmark for localization or for claims about viewer perception.","major_comments":[{"comment":"The forgery localization ground truth is not adequately supported. Section 3.2.1 states that masks are obtained by taking the mean squared pixel difference, thresholding to remove noise, and running connected-component analysis, but the threshold value, noise-filter parameters, and any alignment/resizing step are never reported. For GPT-Image-1 edits, Section 3.1.2 notes that the model has fixed output aspect ratios (1024×1024, 1536×1024, 1024×1536), and that it 'tends to edit in a few cases, untargeted regions'; pixel-wise differencing is only meaningful if the real and edited images are geometrically aligned, and no such preprocessing is described. Since the localization benchmark in Table 4 (SIDA-13B IoU 13.10/24.74, F1 19.92/39.40, AUC 14.06/37.53) and the conclusion's claim of 'the largest image-based dataset for spatial deepfake localization' depend entirely on these masks, the authors must either release the exact mask-generation code and parameters, validate the masks on a human-annotated subset, handle GPT-image geometry explicitly, or explicitly re-scope the dataset as detection-only without localization ground truth. As written, the localization results cannot be separated from ground-truth quality.","section":"§3.2.1, §4.0.3, Conclusion"},{"comment":"The perceptual-impact analysis is self-referential and should be disclosed as such. The edits are generated by Gemini-2.0-Flash-Image-Generation, and Prompt 3.2 asks Gemini-2.0-Flash to judge the resulting changes in emotion, identity, narrative, intent, and ethics. The word clouds in Figure 3 and the ethical-impact percentages (81% mild, 14.2% moderate, 0.3% severe) are therefore the editing model family's self-assessments, not independent measurements of human perception. The paper should reframe these as 'VLM-assessed' properties, add an explicit limitation noting the circularity, and ideally validate a subset with human raters before using the word clouds to claim that the manipulations are 'meaningful' or have particular ethical implications.","section":"§3.2.4, Figure 3, Prompt 3.2"},{"comment":"The finetuning results are reported inconsistently. The text states that 'we observe a performance improvement in both CNNSpot and SIDA-13B' and that CNNSpot surpasses SIDA-13B 'by 1.92%' in accuracy and 'by 1.97%' in F1-score. However, Table 4 shows that after finetuning, CNNSpot's overall accuracy drops from 50.02 to 45.88 and SIDA-13B's overall accuracy drops from 55.97 to 38.01; the F1 scores increase, but the accuracy gaps are 7.87 and 33.12 percentage points, respectively, not 1.92 and 1.97. The arrow for CNNSpot overall accuracy is also marked '↑' when the value decreases. The authors should clarify which metric the percentage differences refer to, correct the arrow and text, or revise the table so that the finetuning narrative matches the reported numbers.","section":"§4.0.2, Table 4"}],"minor_comments":[{"comment":"The text refers to 'Prompt 3.2.3' but the prompt is numbered 3.2; the cross-reference should be corrected.","section":"§3.2.4"},{"comment":"The paper is ambiguous about which editing VLMs (Gemini, GPT-Image-1, ICEdit) contribute images to the final dataset of 758,041 fakes. Section 3.1.2 says Gemini emerged as the best after observing 22K generated images, yet Section 3.2.7 reports costs for GPT-Image-1 and ICEdit. Please clarify the per-model composition of the released dataset, because the localization-mask validity (see major comment above) depends on knowing whether GPT-Image-1 images with altered aspect ratios are included.","section":"§3.1.2, §3.2.7"},{"comment":"The user study uses only 18 participants and 50 images; the authors should state that this is a pilot-level evaluation and consider reporting participant demographics or inter-rater agreement, which would help readers calibrate the 61.67% human accuracy figure.","section":"§3.2.6"},{"comment":"The x-axis label for the histogram of edited-area ratios is missing; it should be something like 'ratio of edited area to total image area' to match the text.","section":"Figure 3a"},{"comment":"The method name 'CnnSpot' is spelled inconsistently; elsewhere it is 'CNNSpot'. Please standardize the spelling.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The dataset itself is potentially a strong contribution to the deepfake detection community, and the zero-shot detection results are convincing evidence of difficulty. The main blocker is the localization ground-truth pipeline: as written, the masks are not reproducible, not validated, and likely invalid for at least the GPT-Image-1 subset due to aspect-ratio changes. If the authors can release exact mask-generation parameters and validate on human annotations (or drop the localization claim), the paper could become acceptable. The self-referential perceptual analysis and the finetuning table inconsistencies are fixable but need explicit attention. I do not see grounds for rejection, provided the localization issue is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: this is a genuinely useful dataset for person-centric semantic manipulations, and the detection results are believable. But the localization half of the paper is not ready as presented, and the finetuning narrative overstates what Table 4 actually shows.\n\nWhat's new: MultiFakeVerse targets manipulations that shift perceived traits (naive, proud, remorseful, etc.) rather than face swaps. That fills a real gap in the deepfake benchmark space, and the scale is impressive—758K fakes over 87K real images. The zero-shot numbers (best 66.87% accuracy, humans at 61.67%) support the claim that these edits are hard to detect. The paper also releases code and an EULA-gated dataset, which is reproducible in principle.\n\nWhere it gets soft. The localization ground truth is the biggest problem. Masks are derived by thresholding pixel-wise differences between original and edited images, with connected-component analysis, but the threshold and noise filters are never reported. More importantly, there is no human validation of these masks. The paper calls MultiFakeVerse 'the largest image-based dataset for spatial deepfake localization' based entirely on these auto-masks. If they don't align with perceptually edited regions, the IoU/AUC numbers are not meaningful. And the GPT-Image-1 aspect-ratio issue is a real aggravating factor: if output resolutions differ, pixel-wise differencing without registration produces resampling artifacts that have nothing to do with edits. That subset's masks may be corrupted.\n\nThe finetuning story is also misstated. The paper says both retrained models improve, but Table 4 shows overall accuracy drops for both (CnnSpot 50.02→45.88, SIDA 55.97→38.01); only F1 improves, which is expected under class imbalance. The Limitation acknowledges the imbalance, but the main text should not claim an improvement in accuracy.\n\nThe perceptual-impact analysis uses Gemini-2.0-Flash to judge edits that Gemini suggested and executed. That is self-referential. The word-cloud results are descriptive at best, not evidence about viewer perception. The human study is small (18 participants, 50 images, no error bars), so the 61.67% figure is a pilot result, not a solid estimate.\n\nNumeric inconsistencies (86,952 vs 87,245 real images) and Venn percentages not summing to 100 add noise but are minor.\n\nBottom line: the dataset is a legitimate contribution to the deepfake-detection community, and the detection benchmark stands on its own. The localization claim and the perceptual analysis need major revision. A serious editor should send this to peer review, but the authors need to either validate the masks or narrow the claims, fix the finetuning description, and report the human study with confidence intervals.","headline":"A useful person-centric deepfake dataset whose detection benchmark holds up, but whose localization ground truth is unvalidated and whose finetuning narrative misreads its own table.","tokens_in":13821,"tokens_out":3371,"would_cite":true,"duration_ms":28218,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MultiFakeVerse, a dataset of 845,286 images in which vision-language models make semantic, person-centric edits to photographs, and reports that current deepfake detectors and humans largely miss these subtle changes.","keywords":["MultiFakeVerse","deepfake detection","person-centric manipulation","vision-language models","image manipulation","semantic manipulation","forgery localization","benchmark dataset"],"falsifier":"A human-annotation study in which, for example, 200 edited images are given pixel-level masks by multiple annotators; if the dataset's auto-masks have low intersection over union with the human masks, the reported localization numbers are not measuring what they claim.","tokens_in":12534,"feed_emoji":"🎭","tokens_out":3855,"duration_ms":34330,"temperature":0.7,"pith_summary":"The paper introduces MultiFakeVerse, a dataset of 845,286 images built by asking vision-language models to make minimal edits to real photos of people so that the viewer's perception of the scene changes. The edits target the most important person's apparent emotion, character, status, or the scene's narrative, rather than swapping faces. On this dataset, the best zero-shot detector reaches 66.87% class-wise accuracy and human observers achieve 61.67%, so the manipulations largely evade both. The authors argue this exposes a gap: detectors trained on face swaps and object inpainting are not ready for semantic, person-centric manipulation.","feed_headline":"VLM-made person edits fool deepfake detectors","feed_subtitle":"Large benchmark shows current detectors and humans miss story-changing image edits.","key_machinery":"The key machinery is a two-stage VLM pipeline: first, a vision-language model (Gemini-2.0-Flash or ChatGPT-4o-latest) reads the image, identifies the most important person, and proposes six minimal edits, each expressed as a referring expression plus an edit instruction; second, an image-editing VLM (GPT-Image-1, Gemini-2.0-Flash-Image-Generation, or ICEdit) executes the edit while leaving everything else unchanged. The dataset is built on 86,952 real images from EMOTIC, PISC, PIPA, and PIC 2.0, and each manipulation is subsequently analyzed by a VLM for its perceptual and ethical impact.","core_discovery":"The central claim is that current state-of-the-art deepfake detection methods and human observers cannot reliably detect person-centric, semantically meaningful image manipulations generated by vision-language models. The paper supports this by constructing MultiFakeVerse, in which each real image is paired with several manipulated versions produced from VLM-generated edit instructions targeting perception of the most important person. Benchmarking shows that AntiFakePrompt, the best zero-shot detector, achieves 66.87% class-wise accuracy and 55.55% F1 score, while humans reach only 61.67% accuracy with a 24.96% intersection over union on manipulation-level identification. After finetuning, CNNSpot and SIDA improve but still fall short of reliable detection.","pith_inferences":["If the auto-generated localization masks are unreliable, the reported localization numbers may not reflect true localization quality; a human-annotation validation study could settle this.","The VLM-driven manipulation pipeline could be extended to video or audio-visual scenarios, where narrative-shifting edits may be even harder to detect.","The dataset could serve as a stress test for detector robustness to distribution shift, since the manipulations are semantically coherent rather than artifact-driven.","The low human accuracy on a 50-image sample hints that the broader population may struggle similarly, raising questions about how to communicate the existence of such edits."],"forward_implications":["Detectors trained on existing GAN- and inpainting-based datasets will systematically miss VLM-driven semantic edits, so new training or domain adaptation is needed.","The dataset provides a benchmark for both detection and localization of person-centric manipulations, with localization masks computed from pixel differences.","Human performance near chance suggests that person-centric semantic manipulations are a realistic threat not captured by current deepfake defenses.","Finetuning on MultiFakeVerse improves detection but does not bring it to reliable levels, indicating the task is not solved by simple supervised adaptation."],"supporting_citations":[{"why":"EMOTIC dataset is one of the four real-image sources used to build the person-centric benchmark.","marker":"[17]"},{"why":"PISC dataset is another real-image source used for the benchmark.","marker":"[18]"},{"why":"PIPA dataset supplies real person images for the benchmark.","marker":"[43]"},{"why":"PIC 2.0 dataset supplies the remaining real images and enables human-object relation editing.","marker":"[19]"},{"why":"ChatGPT-4o-latest and GPT-Image-1 are the VLMs used both to propose and to execute edits.","marker":"[25]"},{"why":"Gemini-2.0-Flash and Gemini-2.0-Flash-Image-Generation are the VLMs that produced most of the manipulations and the perceptual analyses.","marker":"[1]"},{"why":"ICEdit is the third image-editing VLM used to generate manipulated images.","marker":"[44]"},{"why":"CNNSpot serves as a zero-shot detector baseline that classify almost all images as real.","marker":"[37]"},{"why":"AntiFakePrompt is the best zero-shot detector baseline in the benchmark.","marker":"[9]"},{"why":"SIDA is both a detection and localization baseline; its localization performance is evaluated on the dataset.","marker":"[15]"}],"fun_headline_variants":["845K deepfakes show AI, humans miss story-changing edits","VLM-generated person edits evade top deepfake detectors","New dataset exposes blind spots in deepfake detection","Semantic deepfakes: benchmark shows detectors and humans fail","MultiFakeVerse: 845K images expose subtle deepfake edits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's localization results rest on masks produced by thresholding pixel differences between each real and edited image, with no human validation that those masks actually cover the manipulated regions.","fun_headline_variants_meta":{"raw":{"variants":["845K deepfakes show AI, humans miss story-changing edits","VLM-generated person edits evade top deepfake detectors","New dataset exposes blind spots in deepfake detection","Semantic deepfakes: benchmark shows detectors and humans fail","MultiFakeVerse: 845K images expose subtle deepfake edits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000723,"raw_usage":{"total_tokens":3212,"prompt_tokens":883,"completion_tokens":2329,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":2245}},"tokens_in":499,"tokens_out":2329,"duration_ms":16385,"temperature":1.0,"reasoning_tokens":2245,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:57:32.901111+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-annotation study in which, for example, 200 edited images are given pixel-level masks by multiple annotators; if the dataset's auto-masks have low intersection over union with the human masks, the reported localization numbers are not measuring what they claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"EMOTIC dataset is one of the four real-image sources used to build the person-centric benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PISC dataset is another real-image source used for the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PIPA dataset supplies real person images for the benchmark."}],"review_version":1}