Pith. sign in

REVIEW 2 cited by

ImagiNet: A Multi-Content Benchmark for Synthetic Image Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20020 v3 pith:VD4YRCCP submitted 2024-07-29 cs.CV cs.LG

classification cs.CVcs.LG
keywords syntheticcontentimagesimaginetdetectorsmodelrealbalanced
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent generative models produce images with a level of authenticity that makes them nearly indistinguishable from real photos and artwork. Potential harmful use cases of these models, necessitate the creation of robust synthetic image detectors. However, current datasets in the field contain generated images with questionable quality or have examples from one predominant content type which leads to poor generalizability of the underlying detectors. We find that the curation of a balanced amount of high-resolution generated images across various content types is crucial for the generalizability of detectors, and introduce ImagiNet, a dataset of 200K examples, spanning four categories: photos, paintings, faces, and miscellaneous. Synthetic images in ImagiNet are produced with both open-source and proprietary generators, whereas real counterparts for each content type are collected from public datasets. The structure of ImagiNet allows for a two-track evaluation system: i) classification as real or synthetic and ii) identification of the generative model. To establish a strong baseline, we train a ResNet-50 model using a self-supervised contrastive objective (SelfCon) for each track which achieves evaluation AUC of up to 0.99 and balanced accuracy ranging from 86% to 95%, even under conditions that involve compression and resizing. The provided model is generalizable enough to achieve zero-shot state-of-the-art performance on previous synthetic detection benchmarks. We provide ablations to demonstrate the importance of content types and publish code and data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples

    cs.CV 2025-09 conditional novelty 5.0 of 10

    OmniDFA performs few-shot, open-set attribution of AI-generated images, identifying the source generator from just five support samples across 45 known and unseen generators.

  2. VisionTrap: Unanswerable Questions On Visual Data

    cs.CV 2025-07 conditional novelty 5.0 of 10

    VisionTrap shows that GPT-4o, GPT-4.1, Gemini Flash 2.5, and LLaVA tend to answer unanswerable visual questions rather than abstain, especially when given multiple-choice options.

Pith tools