Pith. sign in

REVIEW 2 cited by

Adversarial Robustness of AI-Generated Image Detectors in the Real World

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01574 v4 pith:WKT35PYF submitted 2024-10-02 cs.CV cs.LG

classification cs.CVcs.LG
keywords robustnessgenaimethodsadversarialai-generatedattacksdemonstratedetection
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The rapid advancement of Generative Artificial Intelligence (GenAI) capabilities is accompanied by a concerning rise in its misuse. In particular the generation of credible misinformation in the form of images poses a significant threat to the public trust in democratic processes. Consequently, there is an urgent need to develop tools to reliably distinguish between authentic and AI-generated content. The majority of detection methods are based on neural networks that are trained to recognize forensic artifacts. In this work, we demonstrate that current state-of-the-art classifiers are vulnerable to adversarial examples under real-world conditions. Through extensive experiments, comprising four detection methods and five attack algorithms, we show that an attacker can dramatically decrease classification performance, without internal knowledge of the detector's architecture. Notably, most attacks remain effective even when images are degraded during the upload to, e.g., social media platforms. In a case study, we demonstrate that these robustness challenges are also found in commercial tools by conducting black-box attacks on HIVE, a proprietary online GenAI media detector. In addition, we evaluate the robustness of using generated features of a robust pre-trained model and showed that this increases the robustness, while not reaching the performance on benign inputs. These results, along with the increasing potential of GenAI to erode public trust, underscore the need for more research and new perspectives on methods to prevent its misuse.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Introduces the CIFAR Synthetic Evidence Corpus, a multi-family dataset of AI-manipulated documents with source-separated train/test splits for evaluating detectors of AI-generated legal evidence.

  2. Adversarially Robust AI-Generated Image Detection for Free: An Information Theoretic Perspective

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A training-free detector-side defense, TRIM, flips predictions flagged by entropy and KL-divergence thresholds, reporting large robustness gains on ProGAN, GenImage, and SDv1.4.

Pith tools