A LoRA-tuned CLIP-ViT with deep-shallow feature fusion and triplet loss detects images from unseen GANs and diffusion models at 89% average accuracy when trained only on four ProGAN classes.
In: Proceedings of CVPR
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images
A LoRA-tuned CLIP-ViT with deep-shallow feature fusion and triplet loss detects images from unseen GANs and diffusion models at 89% average accuracy when trained only on four ProGAN classes.