Pith. sign in

REVIEW 2 cited by

Community Forensics: Using Thousands of Generators to Train Fake Image Detectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.04125 v2 pith:KBPOM3NG submitted 2024-11-06 cs.CV

Community Forensics: Using Thousands of Generators to Train Fake Image Detectors

classification cs.CV
keywords modelsimagesdatasetdetectorsimagearchitecturesbeencommunity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this problem, and we propose a new dataset that is significantly larger and more diverse than prior work. As part of creating this dataset, we systematically download thousands of text-to-image latent diffusion models and sample images from them. We also collect images from dozens of popular open source and commercial models. The resulting dataset contains 2.7M images that have been sampled from 4803 different models. These images collectively capture a wide range of scene content, generator architectures, and image processing settings. Using this dataset, we study the generalization abilities of fake image detectors. Our experiments suggest that detection performance improves as the number of models in the training set increases, even when these models have similar architectures. We also find that detection performance improves as the diversity of the models increases, and that our trained detectors generalize better than those trained on other datasets. The dataset can be found in https://jespark.net/projects/2024/community_forensics

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

    cs.CV 2025-10 unverdicted novelty 6.0

    Locate-Then-Examine improves AI-generated image detection by localizing suspicious regions first then performing region-aware re-examination, while releasing the TRACE dataset of 20k annotated images.

  2. VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

    cs.CV 2026-07 conditional novelty 5.0

    A 100-image cross-paradigm benchmark of 36 deepfake detectors reveals that ROC-AUC and MCC diverge sharply, meaning strong class-separation ranking does not guarantee reliable default-threshold decisions.