REVIEW 3 cited by
SKDU at De-Factify 4.0: Vision Transformer with Data Augmentation for AI-Generated Image Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
SKDU at De-Factify 4.0: Vision Transformer with Data Augmentation for AI-Generated Image Detection
read the original abstract
The aim of this work is to explore the potential of pre-trained vision-language models, e.g. Vision Transformers (ViT), enhanced with advanced data augmentation strategies for the detection of AI-generated images. Our approach leverages a fine-tuned ViT model trained on the Defactify-4.0 dataset, which includes images generated by state-of-the-art models such as Stable Diffusion 2.1, Stable Diffusion XL, Stable Diffusion 3, DALL-E 3, and MidJourney. We employ perturbation techniques like flipping, rotation, Gaussian noise injection, and JPEG compression during training to improve model robustness and generalisation. The experimental results demonstrate that our ViT-based pipeline achieves state-of-the-art performance, significantly outperforming competing methods on both validation and test datasets.
Forward citations
Cited by 3 Pith papers
-
Findings of the Counter Turing Test: AI-Generated Image Detection
The Counter Turing Test competition finds F1-scores above 0.83 for binary real-vs-AI classification but only 0.4986 at best for identifying the specific generative model.
-
Findings of the Counter Turing Test: AI-Generated Image Detection
A competition using a new 50k-image dataset found high accuracy in binary real-vs-AI detection but only modest success in identifying the exact generative model.
-
Findings of the Counter Turing Test: AI-Generated Image Detection
Binary AI vs. real image classification reaches F1 > 0.83 while identifying the exact generative model achieves a highest F1 of 0.4986 on the MS COCOAI dataset.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.