Pith. sign in

REVIEW 8 cited by

GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08571 v2 pith:C7325OOT submitted 2023-06-14 cs.CV

classification cs.CV
keywords imagesimagedetectorsai-generatedgenimagedatasetadvancedadvantages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The extraordinary ability of generative models to generate photographic images has intensified concerns about the spread of disinformation, thereby leading to the demand for detectors capable of distinguishing between AI-generated fake images and real images. However, the lack of large datasets containing images from the most advanced image generators poses an obstacle to the development of such detectors. In this paper, we introduce the GenImage dataset, which has the following advantages: 1) Plenty of Images, including over one million pairs of AI-generated fake images and collected real images. 2) Rich Image Content, encompassing a broad range of image classes. 3) State-of-the-art Generators, synthesizing images with advanced diffusion models and GANs. The aforementioned advantages allow the detectors trained on GenImage to undergo a thorough evaluation and demonstrate strong applicability to diverse images. We conduct a comprehensive analysis of the dataset and propose two tasks for evaluating the detection method in resembling real-world scenarios. The cross-generator image classification task measures the performance of a detector trained on one generator when tested on the others. The degraded image classification task assesses the capability of the detectors in handling degraded images such as low-resolution, blurred, and compressed images. With the GenImage dataset, researchers can effectively expedite the development and evaluation of superior AI-generated image detectors in comparison to prevailing methodologies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    SynCred-Bench shows that 15 MLLMs reach only 10.5% TPR, open-source detectors under 5%, commercial APIs 57.6%, and humans 63% TPR at 5% FPR when identifying AI-generated images with synthetic credibility.

  2. SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation

    cs.CV 2026-04 conditional novelty 7.0 of 10

    SurFITR is a new collection of 137k+ surveillance-style forged images that causes existing detectors to degrade while enabling substantial gains when used for training in both in-domain and cross-domain settings.

  3. XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A million-scale deepfake benchmark with Edit-Check filtering, dual expert/lay explanations, and EntityScore/EvidenceScore shows fine-tuned detectors collapse under generator shift while surface fluency remains.

  4. AI-generated Images Challenge Visual Trust in High-risk Scenarios

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On SafeIMG, a new safety-focused benchmark of 1,131 GPT Image 2 images, the best VLM detects 49.5% of generated images and the best specialized detector 33.1%, versus 81.7% for humans.

  5. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.

  6. GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GenSyn10 provides 60k CIFAR-10-aligned images from FLUX.2, HunyuanImage-3.0, and Qwen-Image-2512, showing detectors lose 4–18 points of accuracy on an unseen generator.

  7. Deepfakes in the 2025 Canadian Election: Prevalence, Partisanship, and Platform Dynamics

    cs.SI 2025-12 unverdicted novelty 5.0 of 10

    During the 2025 Canadian election, 5.86% of election images on three major platforms were deepfakes, shared more by right-leaning accounts, but harmful ones had minimal reach and most were benign.

  8. DeepFake Forensics AI: A Multi-Modal Detection and Blockchain-Anchored Evidence Management Platform

    cs.CR 2026-05 unverdicted novelty 4.0 of 10

    Describes a multi-modal deepfake detection system with image, video, audio, and GAN fingerprinting models integrated with Ethereum blockchain for evidence anchoring, reporting AUC 0.9868, 0.9628, EER 18.63%, and 99.88...

Pith tools