{"id":"8684fa97-1c7e-4f48-83dc-ede21ec15b63","arxiv_id":"2412.05676","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"Black-box genetic attacks flip 70% of correct fake detections in a retrained patch-based detector, GPT-4o reaches 73% AUC zero-shot on a Celeb-DF subset, and a 6.64% typographic attack degrades it.","lead":"This paper asks whether today's best deepfake detectors can be trusted against attackers, and whether vision-language models like GPT-4o do better. It shows a patch-based detector fails badly under a query-based attack, while GPT-4o is strong but can be confused by faint text added to an image.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's SOTA-robustness and GPT-4o-superiority claims rest on retrained detectors that underperform published benchmarks by a wide margin, so the reported attack ASRs may not transfer to the actual methods.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the attacked models are retrained stand-ins, not the published detectors. This is the linchpin of the paper because all three central claims depend on it. Claim (1) that state-of-the-art detectors are susceptible to black-box attacks would be invalid if the attacked model is a poor reimplementation; claim (2) that GPT-4o outperforms SOTA would be invalid if the baselines are weak reimplementations; claim (3) about the typographic attack is a small proof-of-concept but does not salvage the headline. The evidence in Sections 6.2 and 7.2 shows dramatic underperformance compared to published results, and the paper's own terminology argument admits the retrained LaDeDa is a different model from the WildRF-trained original. A concrete check on the official released weights would settle whether the attack transferability holds. Since the central comparison is unsupported as presented, the reader's REJECT verdict is appropriate, and no adjustment is needed.","tokens_in":10604,"tokens_out":5178,"duration_ms":46105,"concrete_test":"Obtain the official WildRF-pretrained LaDeDa weights and the official CLIPping the Deception checkpoint (if released), and rerun the genetic black-box attack with the same hyperparameters (m=100, n=10, k=5, eps=10/255) on the same FF++ test frames used in Table 1; if the ASR against the official models differs materially from the reported 70% and 30%, the core robustness comparison and the 'GPT-4o better than SOTA' claim are unsupported. Alternatively, retrain LaDeDa using the original WildRF protocol and verify that its benchmark performance matches the published near-perfect results before attacking; if reproduction fails, the claims should be scoped to 'LaDeDa-style' models rather than state-of-the-art detectors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims depend on attacking the authors' retrained versions of LaDeDa and CLIPping the Deception, not the published state-of-the-art models. Section 6.1 trains LaDeDa on FaceForensics++ (the original method was trained on WildRF), and Section 6.2 reports AUC 0.488 on Celeb-DF (chance-level) and 0.7963 on FF++; Section 7.2 reports CoOp AUC 0.6382 on FF++ and 0.5884 on Celeb-DF. These are far below the near-perfect published numbers the authors themselves acknowledge in Section 6.2. The black-box genetic attack (ASR 70% vs 30%, Table 1) is measured on these weak reimplementations; if the retraining is not a faithful reproduction, the comparative robustness conclusion does not transfer to actual SOTA detectors. The 'GPT-4o better than SOTA' claim (Section 8.2) is likewise a comparison against these weak baselines, and GPT-4o's Celeb-DF accuracy (0.6425) is below the 0.6552 majority-class baseline for the sampled 570/300 split, further weakening the claim. The terminology argument in Section 6.2 clarifies that the retrained model differs from the published one, but it does not resolve the mismatch for the attack experiments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates the adversarial robustness of deepfake detectors and argues that high-level semantic representations are more robust than low-level visual forensics. The authors retrain LaDeDa and CLIPping the Deception on FaceForensics++, attack them with a genetic black-box algorithm, evaluate zero-shot GPT-4o on a subset of Celeb-DF, and introduce a typographic text-overlay attack on GPT-4o. They conclude that recently developed state-of-the-art detectors are vulnerable to classical black-box attacks, that GPT-4o outperforms current state-of-the-art deepfake detectors, and that hybridising low-level and high-level detectors is a promising direction.","tokens_in":10883,"tokens_out":4898,"duration_ms":47627,"significance":"If the claims were fully supported, the paper would make a useful contribution by connecting adversarial robustness of deepfake detectors to the distinction between semantic and forensic features, by demonstrating a realistic black-box attack, and by introducing a proof-of-concept typographic semantic attack. The authors are transparent about their hypotheses and report concrete numerical results, which makes their claims falsifiable. However, the central comparisons rest on reimplementations that are far below published benchmark performance, so the headline conclusions about state-of-the-art detectors and about GPT-4o superiority are not supported as stated. The typographic attack and the hybridisation discussion are interesting but preliminary.","major_comments":[{"comment":"The retrained LaDeDa model achieves an AUC of 0.488 on Celeb-DF, which is at chance level, and 0.7963 on FaceForensics++, while the authors themselves note that the original method reports near-perfect benchmark performance. The genetic black-box attack with 70% ASR is run on this retrained model, not on the released LaDeDa model. Because the reproduction is not a faithful stand-in for the published state-of-the-art detector, the attack results do not transfer to the actual method, and the Abstract's claim that \"recently developed state-of-the-art detectors are susceptible\" is unsupported. The training in Section 6.1 uses FaceForensics++ rather than the original WildRF training set, and the terminology argument in Section 6.2 about deepfakes versus AI-generated content does not repair this mismatch.","section":"Section 6.2, Table 1"},{"comment":"The CLIPping the Deception reproduction scores only 63.82% AUC on FaceForensics++ and 58.84% AUC on Celeb-DF, values far below the published results for that method. The comparison between the 30% ASR for this model and the 70% ASR for the LaDeDa reproduction therefore confounds the detector paradigm (semantic versus forensic) with training quality and reproduction fidelity. Without a faithful and comparably trained baseline, the paper cannot attribute the robustness gap to reliance on high-level semantic embeddings.","section":"Section 7.2, Table 1"},{"comment":"The claim that GPT-4o performs zero-shot deepfake detection \"better than current state-of-the-art methods\" is not supported by the evidence. The comparison is made only against the authors' weak reimplementations, not against the published results of LaDeDa or other state-of-the-art detectors. Moreover, on the Celeb-DF subset with 570 fake and 300 real frames, the majority-class baseline accuracy is 570/870 = 0.6552, while GPT-4o achieves 0.6425, which is below the trivial always-fake baseline. The stated superiority claim therefore collapses even relative to the paper's own evaluation set.","section":"Section 8.2, Table 1"},{"comment":"The typographic attack is presented as a proof of concept, and the 6.64% attack success rate is appropriately modest. However, the paper also states that the attack is \"without being obvious to a human observer,\" yet no human evaluation, no ablation over text content or opacity, and no confidence intervals are provided. As a result, the imperceptibility claim is unsupported, and the later discussion in Section 10.1 that overlayed text is \"easily detectable by a low-level forensics model\" is speculative rather than demonstrated.","section":"Section 9"}],"minor_comments":[{"comment":"The genetic algorithm description omits the mutation probability \"p\" and mutation weight \"w\" used in the experiments, and the phrase \"performing on average 1,010 queries per input image batch\" is ambiguous: it should state whether queries are counted per image or per batch and how many images were attacked.","section":"Section 5, Section 6.2"},{"comment":"The zero-shot GPT-4o evaluation averages five outputs per image, but the paper does not report variance across the five samples, the number of times the model refused to respond, or the number of retries; these details matter because the model is stochastic and the evaluation subset is small.","section":"Section 8.1"},{"comment":"In Table 1, the \"NQ\" entry for the typographic attack is listed as 0, which is unclear because the zero-shot evaluation is not an iterative attack; the column should be defined more carefully. In Section 10.2, \"their potentially susceptibility\" should read \"their potential susceptibility.\"","section":"Table 1, Section 10.2"}],"recommendation":"reject","confidential_remarks":"The manuscript reads as a promising workshop-level study, but its central claims reach far beyond what the experiments support: both reimplementations are far below published performance, the GPT-4o result is below a majority-class baseline, and the attack experiments are not performed on the official state-of-the-art models. These are load-bearing issues that would require re-running the main experiments with faithful reproductions or official weights, comparing against published state-of-the-art numbers, and adding proper baselines for the typographic attack. The typographic attack and the hybridisation discussion could be the seed of a future substantial revision or a new paper, but in the current form the headline conclusions cannot stand."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has one genuinely new idea—a typographic attack that overlays text mimicking a file path onto an image and flips GPT-4o's deepfake verdict—but the headline comparisons rest on reimplementations that are far weaker than the published models, so the main claims don't hold as stated.\n\nWhat's actually new: the file-path typographic attack is, as far as I know, not in the literature. The GA-based comparison between a patch-based detector and a CLIP-based detector is also a fresh experimental design, even if the execution is flawed. The authors are transparent about their poor reproduction numbers, which is more than many papers do.\n\nThe problem is that the numbers used for the central claims come from their own retrained models, not the official ones. Their LaDeDa scores 48.8% AUC on Celeb-DF, which is chance-level, and 69.41% accuracy on FF++, far below the near-perfect published results they cite. The GA attack is run on these weak models. A 70% ASR on a detector that can't even beat chance on the target dataset tells us little about whether the real LaDeDa is robust. The authors try to hand-wave the gap with a terminology argument—WildRF is AIGC, not deepfakes—but that doesn't rescue the attack experiments. If the model is broken, the attack numbers are about the broken model.\n\nThe GPT-4o claim is similarly unsupported. They compare against their own weak baselines, not published SOTA. On their 570/300 Celeb-DF sample, GPT-4o's accuracy (64.25%) is actually below the majority-class baseline of 65.52%, so \"better than current state-of-the-art\" doesn't survive contact with the data. The typographic attack itself is a legitimate proof of concept, but 6.64% ASR on a single run, with no error bars and no code, is a small seed rather than a demonstrated vulnerability.\n\nThere are also basic procedural gaps: no code or data, single-run point estimates, and a handful of tunable GA and typographic parameters that are not swept.\n\nWho gets value from this? Someone working on VLM-based deepfake detection might find the typographic attack worth reading about, and the paper is a useful case study in how reproduction failures can derail an otherwise reasonable comparison. But as it stands, the central claims are overreached.\n\nI'd send it to peer review rather than desk-reject—the typographic attack idea is worth referee time and the flaws are identifiable and fixable—but the authors would need to either use the official models or show their retraining is faithful, add published SOTA baselines, and report multiple runs. Without that, the paper is a course project with an interesting seed.","headline":"A novel typographic attack on VLM deepfake detection is buried under unsupported comparisons against weak reimplementations.","tokens_in":11427,"tokens_out":4206,"would_cite":false,"duration_ms":39079,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Recent deepfake detectors that top benchmarks are vulnerable to query-only adversarial attacks, while a zero-shot GPT-4o prompt detects Celeb-DF deepfakes more accurately.","keywords":["deepfake detection","adversarial robustness","black-box attack","genetic algorithm","vision-language models","GPT-4o","typographic attack","CLIP"],"falsifier":"Run the same bounded genetic black-box attack against the official released weights of LaDeDa and of the CLIP-based detector, with identical epsilon and query budget: if the official LaDeDa shows near-perfect benign AUC on Celeb-DF and its attack success rate is not around 70%, the paper's central robustness comparison fails. Separately, test the typographic overlay with the text removed or placed elsewhere on the same Celeb-DF frames to confirm that the verdict flips are caused by the text rather than by cropping or compression.","tokens_in":10373,"feed_emoji":"🎭","tokens_out":6130,"duration_ms":52867,"temperature":0.7,"pith_summary":"The paper argues that recent benchmark-topping deepfake detectors are not safe to deploy, because they rely on local visual artifacts that classical adversarial perturbations can erase. On its re-trained copy of LaDeDa, a patch-based detector, a genetic black-box attack flips about 70% of correct fake detections on FaceForensics++, while the same attack succeeds only about 30% of the time against a CLIP-based detector. The paper then shows that GPT-4o with a simple \"real or fake\" prompt outperforms both dedicated detectors on Celeb-DF, and introduces a faint typographic overlay that flips GPT-4o verdicts with a 6.64% attack success rate. The conclusion is that robust deepfake detection needs both low-level forensics and high-level semantics, and that either family alone is attackable.","feed_headline":"GPT-4o beats deepfake detectors—until faint text flips it","feed_subtitle":"Local-artifact detectors fall to simple black-box attacks, but semantic models resist—and reveal a new text-overlay weakness.","key_machinery":"The load-bearing objects are two detector families and the attacks that expose their limits. LaDeDa is a ResNet50 variant whose 1x1 convolutions restrict the receptive field to 9x9 patches, producing per-patch deepfake scores that are average-pooled; this local scoring is what makes it brittle. CLIPping the Deception is a frozen CLIP image-text model adapted by CoOp prompt tuning, where 16 learnable context tokens are optimized while the encoders stay frozen; this is what gives it semantic, less locally brittle features. Against both, the paper uses a genetic black-box attack with population, elites, crossover and mutation, bounded by epsilon, requiring only query access and averaging 1,010 queries per image. The novel typographic attack overlays faint source-path-looking text to manipulate GPT-4o's semantic reading of the image without altering its low-level artifacts.","core_discovery":"On the paper's own terms, the central discovery is that recent gains in deepfake detection do not transfer to adversarial settings. The local-patch detector LaDeDa, which scores 9x9 patches independently and pools them, reaches 69.41% accuracy and AUC 79.63% on FaceForensics++ but only AUC 48.8% on Celeb-DF, and a black-box genetic attack flips 70% of its true-positive detections. The same attack against CLIPping the Deception, a prompt-tuned CLIP detector, flips only about 30%, evidence that semantic embeddings are harder to perturb. A zero-shot GPT-4o prompt achieves AUC 73.18% on Celeb-DF, above the dedicated detectors, but adding white 7%-opacity text that looks like a file path succeeds on 6.64% of its verdict flips. The paper therefore positions hybrid low-level plus high-level detection as the necessary path.","pith_inferences":["An implicit corollary is that the 70% versus 30% attack gap should be re-measured on the official released checkpoints, since the paper's re-trained LaDeDa falls below chance on Celeb-DF and far below the published benign performance.","A testable extension is to run the typographic attack against a human baseline and against other vision-language models, to see whether the 6.64% success rate reflects a general semantic vulnerability or a GPT-4o-specific quirk.","The complementarity argument predicts that adversarial examples crafted against a low-level detector will transfer poorly to a semantic detector and vice versa; that transfer experiment is the natural next measurement.","If the paper's framing is right, deepfake detection should be treated as an adversarial game rather than a fixed benchmark, and evaluation protocols should include perturbation budgets and semantic-manipulation probes from the start."],"forward_implications":["Benchmark accuracy on standard deepfake datasets does not imply robustness: a detector that posts near-perfect results can be reduced to about 30% accuracy by a query-only adversarial attack.","Detectors built on high-level semantic embeddings, such as CLIP-derived features, are substantially more resistant to bounded pixel perturbations than local patch-based detectors.","Large visuo-lingual models can perform useful zero-shot deepfake detection with no fine-tuning, and here GPT-4o outperforms the dedicated detectors evaluated on Celeb-DF.","Semantic detectors are vulnerable to a new attack surface: text overlaid on the image can flip verdicts even when the model's stated reason does not mention the text.","A hybrid detector combining visual artifacts and high-level semantics should be more robust because the two paradigms fail in complementary ways."],"supporting_citations":[{"why":"LaDeDa patch-based detector; the main target of the black-box attack and the state-of-the-art method being questioned.","marker":"[7]"},{"why":"CLIPping the Deception, the prompt-tuned CLIP detector that provides the semantic-embedding comparison.","marker":"[20]"},{"why":"Genetic algorithm black-box attack method that the paper adapts for its query-based attacks.","marker":"[3]"},{"why":"Non-robust features hypothesis used to explain why local-patch detectors are vulnerable to small perturbations.","marker":"[17]"},{"why":"CoOp prompt tuning procedure used to train the CLIP-based detector with frozen encoders.","marker":"[43]"},{"why":"CLIP model whose frozen encoders supply the semantic visual representations for the visuo-lingual detector.","marker":"[31]"},{"why":"Prior typographic attack results on vision-language models that the paper's novel image-context attack builds on.","marker":"[9]"},{"why":"Prior study using multimodal LLMs for deepfake forensics, providing the baseline that the GPT-4o zero-shot results extend.","marker":"[18]"}],"fun_headline_variants":["Local deepfake detectors crumble under black-box attacks; semantic models hold","Black-box attacks flips most local-patch deepfake flags; semantic models resist","Semantic embeddings make deepfake detectors far harder to fool","Hybrid low-level and semantic cues are the future of deepfake defense","GPT-4o detects deepfakes zero-shot, but text overlays still trick it"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the assumption that the authors' re-trained LaDeDa and CLIP detectors faithfully reproduce the published state-of-the-art models; if the training pipeline did not reproduce them, the 70% versus 30% attack gap may say nothing about the real detectors.","fun_headline_variants_meta":{"raw":{"variants":["Local deepfake detectors crumble under black-box attacks; semantic models hold","Black-box attacks flips most local-patch deepfake flags; semantic models resist","Semantic embeddings make deepfake detectors far harder to fool","Hybrid low-level and semantic cues are the future of deepfake defense","GPT-4o detects deepfakes zero-shot, but text overlays still trick it"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00076,"raw_usage":{"total_tokens":3358,"prompt_tokens":912,"completion_tokens":2446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":2347}},"tokens_in":528,"tokens_out":2446,"duration_ms":17743,"temperature":1.0,"reasoning_tokens":2347,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:28:14.911188+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same bounded genetic black-box attack against the official released weights of LaDeDa and of the CLIP-based detector, with identical epsilon and query budget: if the official LaDeDa shows near-perfect benign AUC on Celeb-DF and its attack success rate is not around 70%, the paper's central robustness comparison fails. Separately, test the typographic overlay with the text removed or placed elsewhere on the same Celeb-DF frames to confirm that the verdict flips are caused by the text rather than by cropping or compression.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CLIPping the Deception, the prompt-tuned CLIP detector that provides the semantic-embedding comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Non-robust features hypothesis used to explain why local-patch detectors are vulnerable to small perturbations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior study using multimodal LLMs for deepfake forensics, providing the baseline that the GPT-4o zero-shot results extend."}],"review_version":1}