Pith. sign in

REVIEW 14 cited by

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.02664 v2 pith:XFTJAAUU submitted 2025-07-03 cs.CV

classification cs.CV
keywords ai-generatedexplanationsaigiaigi-holmesdatasetdetectionexpertpreference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development of AI-generated content (AIGC) technology has led to the misuse of highly realistic AI-generated images (AIGI) in spreading misinformation, posing a threat to public information security. Although existing AIGI detection techniques are generally effective, they face two issues: 1) a lack of human-verifiable explanations, and 2) a lack of generalization in the latest generation technology. To address these issues, we introduce a large-scale and comprehensive dataset, Holmes-Set, which includes the Holmes-SFTSet, an instruction-tuning dataset with explanations on whether images are AI-generated, and the Holmes-DPOSet, a human-aligned preference dataset. Our work introduces an efficient data annotation method called the Multi-Expert Jury, enhancing data generation through structured MLLM explanations and quality control via cross-model evaluation, expert defect filtering, and human preference modification. In addition, we propose Holmes Pipeline, a meticulously designed three-stage training framework comprising visual expert pre-training, supervised fine-tuning, and direct preference optimization. Holmes Pipeline adapts multimodal large language models (MLLMs) for AIGI detection while generating human-verifiable and human-aligned explanations, ultimately yielding our model AIGI-Holmes. During the inference stage, we introduce a collaborative decoding strategy that integrates the model perception of the visual expert with the semantic reasoning of MLLMs, further enhancing the generalization capabilities. Extensive experiments on three benchmarks validate the effectiveness of our AIGI-Holmes.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    ForeAgent combines a Perception-Verdict MLLM architecture with hindsight-driven self-refining via sampling-reflection-evolution to reach 82.18% accuracy on Chameleon and 93.3% mean accuracy across 16 generators on AIG...

  2. ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    ReAlign distills LLM-generated reasoning texts into a lightweight AIGI forgery detector via contrastive image-text alignment to improve generalization on complex forgeries.

  3. XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A million-scale deepfake benchmark with Edit-Check filtering, dual expert/lay explanations, and EntityScore/EvidenceScore shows fine-tuned detectors collapse under generator shift while surface fluency remains.

  4. HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    HydraPrompt uses an Asymmetric Prompt Adapter with fixed real prompts and adaptive fake prompts plus a Conditional Supervised Contrastive loss to achieve SOTA synthetic image detection on benchmarks.

  5. DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    DVAR turns video authenticity detection into an iterative debate between a generative hypothesis agent and a natural mechanism agent, resolved via minimum description length and a knowledge base for better generalizat...

  6. Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    MAFL uses adversarial training to suppress pattern and content biases, guiding models to learn shared generative features for better cross-model generalization in detecting AI images.

  7. AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection

    cs.CV 2026-03 conditional novelty 6.0 of 10

    An LLM agent guided by Expert and Clustering Profiles fuses heterogeneous AIGI detectors, resolves conflicts, and outputs explainable forensic reports that beat single experts and standard ensembles on high-conflict a...

  8. Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Frozen features from vision foundation models enable a linear probe to outperform specialized AIGI detectors by over 30% on in-the-wild data due to emergent forgery knowledge from pre-training.

  9. From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A multi-agent forensic system integrates multiple evidence sources and debate to detect AI-generated images, reporting 97.05% accuracy on a 6,000-image benchmark while outperforming traditional classifiers.

  10. HunyuanImage 3.0 Technical Report

    cs.CV 2025-09 accept novelty 6.0 of 10

    HunyuanImage 3.0 delivers an 80B-parameter MoE model unifying multimodal understanding and generation that matches prior state-of-the-art results while being fully open-sourced.

  11. HunyuanImage 3.0 Technical Report

    cs.CV 2025-09 conditional novelty 6.0 of 10

    HunyuanImage 3.0 is an open 80B-parameter multimodal autoregressive image generator that reportedly matches leading closed models on in-house benchmarks.

  12. SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Frozen multimodal encoders enable robust AI-generated image detection via linear classification on a 10K-image curated training set that improves generalization over larger datasets.

  13. HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    HiMix combines mixup augmentation to create transitional real-fake samples with hierarchical global-local artifact feature fusion to achieve better generalization in detecting AI-generated images from unseen generators.

  14. UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    UniGenDet unifies generative and discriminative models through symbiotic self-attention and detector-guided alignment to co-evolve image generation and authenticity detection.

Pith tools