A 36-model cross-paradigm benchmark on a hard 100-image corpus shows commercial APIs lead on MCC, open-source detectors trail on average, and a subset of strong rankers are miscalibrated at their default threshold.
Dire for diffusion-generated image detection
7 Pith papers cite this work. Polarity classification is still indexing.
abstract
Diffusion models have shown remarkable success in visual synthesis, but have also raised concerns about potential abuse for malicious purposes. In this paper, we seek to build a detector for telling apart real images from diffusion-generated images. We find that existing detectors struggle to detect images generated by diffusion models, even if we include generated images from a specific diffusion model in their training data. To address this issue, we propose a novel image representation called DIffusion Reconstruction Error (DIRE), which measures the error between an input image and its reconstruction counterpart by a pre-trained diffusion model. We observe that diffusion-generated images can be approximately reconstructed by a diffusion model while real images cannot. It provides a hint that DIRE can serve as a bridge to distinguish generated and real images. DIRE provides an effective way to detect images generated by most diffusion models, and it is general for detecting generated images from unseen diffusion models and robust to various perturbations. Furthermore, we establish a comprehensive diffusion-generated benchmark including images generated by eight diffusion models to evaluate the performance of diffusion-generated image detectors. Extensive experiments on our collected benchmark demonstrate that DIRE exhibits superiority over previous generated-image detectors. The code and dataset are available at https://github.com/ZhendongWang6/DIRE.
citation-role summary
citation-polarity summary
fields
cs.CV 7roles
background 1polarities
background 1representative citing papers
A new six-domain benchmark shows existing AI-image detectors are highly inconsistent on text-rich images and fail badly under JPEG compression, while a vision-language model is stronger but still weak on tables.
GAPL learns a compact set of canonical forgery prototypes and applies two-stage LoRA training to build a low-variance feature space that improves generalization across GAN and diffusion generators.
The ITW-SM dataset and targeted optimization of detector design choices yield a 26.87% average AUC improvement for state-of-the-art AI-generated image detectors under real-world social media conditions.
BLINK benchmark shows multimodal LLMs reach only 45-51 percent accuracy on core visual perception tasks where humans achieve 95 percent, indicating these abilities have not emerged.
Benchmarks Vision Mamba variants for AI-generated image detection against CNN, ViT, and VLM detectors on diverse datasets and synthetic sources, reporting promise alongside limitations.
citing papers explorer
-
VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
A 36-model cross-paradigm benchmark on a hard 100-image corpus shows commercial APIs lead on MCC, open-source detectors trail on average, and a subset of strong rankers are miscalibrated at their default threshold.
-
TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
A new six-domain benchmark shows existing AI-image detectors are highly inconsistent on text-rich images and fail badly under JPEG compression, while a vision-language model is stronger but still weak on tables.
-
Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
GAPL learns a compact set of canonical forgery prototypes and applies two-stage LoRA training to build a low-variance feature space that improves generalization across GAN and diffusion generators.
-
Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters?
The ITW-SM dataset and targeted optimization of detector design choices yield a 26.87% average AUC improvement for state-of-the-art AI-generated image detectors under real-world social media conditions.
-
BLINK: Multimodal Large Language Models Can See but Not Perceive
BLINK benchmark shows multimodal LLMs reach only 45-51 percent accuracy on core visual perception tasks where humans achieve 95 percent, indicating these abilities have not emerged.
-
Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation
Benchmarks Vision Mamba variants for AI-generated image detection against CNN, ViT, and VLM detectors on diverse datasets and synthetic sources, reporting promise alongside limitations.
- DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection