Pith. sign in

REVIEW 3 major objections 5 minor 58 references

This paper argues that adding a DDIM reconstruction-residual stream to a large multimodal model improves explainable deepfake detection and artifact localization beyond RGB evidence alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 00:58 UTC pith:37ZIOWEN

load-bearing objection A credible, novel dual-stream deepfake detector whose central residual-stream claim is confounded by its own reward function—worth refereeing, but the causal evidence needs a cleaner experiment. the 3 major comments →

arxiv 2607.25962 v1 pith:37ZIOWEN submitted 2026-07-28 cs.CV

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

classification cs.CV
keywords deepfake detectionexplainable forensicsDDIM inversion-reconstructionreconstruction residualmultimodal reasoningchain-of-thoughtGRPOartifact localization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that a reconstruction-based compatibility signal — how well a frozen Stable Diffusion reference can invert and rebuild an image — meaningfully augments RGB appearance cues for explainable deepfake detection. The claim is that the residual map R=|x−x̂|, encoded as a second visual stream and fused with a CLIP semantic stream, improves both artifact localization and cross-generator image-level detection, giving the highest reported accuracy on four of seven generator families in the UniversalFakeDetect benchmark. A cascade of ablations supports the claim: removing the residual stream costs 10.4 mIoU and 21.9 F1 points; a multi-step DDIM residual beats a single-step VAE residual; and counterfactual map replacement changes predicted masks, indicating the model uses the cue rather than ignoring it. The paper also argues that a Group Relative Policy Optimization stage with format and evidence-reference rewards aligns text output and masks, while explicitly limiting the claim that free-form text is semantically faithful. A sympathetic reader would care because it points toward detectors that can say where and why an image is synthetic, not just whether.

Core claim

The central discovery is that a pixel-level compatibility residual, computed by deterministic DDIM inversion-reconstruction against a fixed Stable Diffusion v1.5 reference at T=50, acts as a transferable forensic cue when fused with RGB semantics. The paper reports that this dual-stream representation, trained only on ProGAN and evaluated on UniversalFakeDetect, reaches 97.23% accuracy on GANs, 92.18% on Diffusion, 98.98% on CRN, and 98.62% on IMLE — the highest reported numbers in those columns — while a separate reasoning model using the same cue with a LLaMA-2-7B backbone and SAM decode achieves 72.19 mIoU and 63.62 F1 on a curated SynthScars-CoT split, versus 61.83 and 41.71 when the res

What carries the argument

The load-bearing object is the DDIM inversion-reconstruction residual R=|x−x̂|, computed with a frozen Stable Diffusion v1.5 and a fixed T=50 deterministic schedule. This map is fed through the same frozen CLIP vision encoder as the RGB image but with an independent projector, so the language model can attend to both streams. The reasoning side is a structured Where-What-Why chain-of-thought with a [SEG] token decoded by SAM, optimized first by supervised fine-tuning and then by GRPO with a reward that mixes mask IoU with format and evidence-reference terms. The image-level detector is a simple MLP over concatenated [CLS] tokens from both streams. The residual stream is the mechanism that ca

Load-bearing premise

The entire argument hinges on the premise that the DDIM residual map reflects manipulation-induced incompatibility rather than merely benign textures, compression, resampling, or domain shift — a distinction the paper itself flags as open in Section 3.2 and the Limitations.

What would settle it

Run the same architecture on a corpus of clean, unmanipulated, high-texture photographs (hair, woven fabric, foliage, specular highlights); if the residual-driven detector's false-positive rate rises sharply or the localization model produces masks on these pristine images with nontrivial IoU against empty ground truth, then the cue is tracking benign texture rather than manipulation.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A single fixed reconstruction reference can serve as a shared forensic cue across generator families, including GANs and CRN that do not follow diffusion trajectories — a point the paper explicitly limits to the evaluated protocol.
  • Reconstruction residuals and RGB semantics carry complementary information; the multi-step DDIM residual outperforms both RGB-only and single-step VAE residuals in the paper's ablations.
  • Combining mask IoU with format and evidence-reference rewards in GRPO improves localization over mask-only reinforcement, indicating that structural constraints stabilize policy optimization for segmentation.
  • Counterfactual map interventions provide a way to test whether a multimodal forensic model actually uses its evidence stream rather than merely being affected by its removal.
  • The documented sensitivity to benign high-frequency textures and the lack of a verifier for free-form textual truth bound the approach's scope to settings without strong post-processing.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be to evaluate the same residual construction as an explicit fusion input for non-diffusion detectors, such as GAN or editing-attribution models; if it transfers, the residual would act as a general 'compatibility fingerprint' rather than a diffusion-specific artifact.
  • The zero-map and donor-map intervention procedure could be adopted as a standard sanity check for any explainable forensic model that claims to ground its explanations in a given evidence map.
  • Because the residual is computed against one fixed reconstruction reference, the framework inherits that reference's biases; pooling residuals from a family of frozen references could improve robustness, though the paper does not explore this.
  • The paper's distinction between 'evidence reference' and 'semantic faithfulness' suggests that a future entailment-based verification reward could complement the current keyword and structure rewards in GRPO.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes LaP-Forensics, a two-stream framework that augments RGB image semantics with a DDIM inversion-reconstruction residual computed against a frozen Stable Diffusion v1.5 reference. The residual map is encoded as a second visual stream, and a LLaMA-2-7B MLLM with a SAM decoder performs structured Where–What–Why reasoning and artifact localization. A separately trained image-level head concatenates RGB and residual CLS features for detection. Training uses supervised fine-tuning followed by GRPO with mask, format, and evidence-reference rewards. Experiments are reported on UniversalFakeDetect for detection and on SynthScars, LOKI, and RichHF for localization, together with component, cue-construction, horizon, reward, and counterfactual ablations. The paper carefully separates the official SynthScars protocol from an internally curated CoT split, and it repeatedly and explicitly disclaims that the residual is a calibrated manipulation probability or a source-identification cue. The central claim is that the residual stream is the most critical component, but, as detailed below, the evidence for this claim is weakened by a confounded ablation and by the absence of a same-protocol RGB-only detector baseline.

Significance. If the residual-stream benefit were convincingly established, the paper would be a useful contribution to explainable deepfake detection: it combines a fixed reconstruction reference with MLLM-based reasoning and pixel-level localization, and it evaluates the model's dependence on the residual through counterfactual interventions. The paper deserves credit for explicitly bounding the interpretation of the residual map, for separating official-benchmark and curated-protocol results, and for acknowledging that the text-side rewards do not verify semantic faithfulness. However, the current evidence is not yet sufficient to support the headline contribution. The main residual-stream ablation is confounded by the GRPO evidence-reference reward and the CoT targets, and the detection experiments do not include an RGB-only control trained under the same protocol. These are fixable with additional experiments and clarifications, but they are load-bearing for the paper's central claim.

major comments (3)
  1. [§4.6, Table 4; §3.4, Eq. (8)] The residual-stream ablation is not isolated. The full GRPO reward (Eq. 8) includes a format/evidence-reference term that explicitly rewards references to the Latent-Pixel Consistency Map, and the SFT CoT targets (Section 3.3) include a step that relates observations to that map. For the 'w/o residual stream' and 'RGB input only' conditions, the manuscript does not state that these targets/rewards were modified. If they were kept unchanged, the ablated model is penalized for failing to reference an input that no longer exists, so the 10.36 mIoU / 21.91 F1 drop conflates loss of forensic information with an objective that is partially unsatisfiable. This concern is material: Table 5 shows that reward-configuration changes alone move mIoU by roughly 17 points (mask-only 55.46 vs full 72.19), comparable to the residual ablation. Please report the exact training setup for the ablated variant
  2. [§4.4, Table 2; §4.2] The standalone detector's residual contribution is unisolated. Table 2 compares the RGB+DDIM-residual detector against published baselines with different training objectives, data sources, and checkpoints. There is no same-protocol RGB-only detector row. Consequently the results in Table 2 and the sentence in §4.4 that the 'RGB-plus-reconstruction representation transfers numerically' support a system-level benchmark claim, not the utility of the residual stream for detection. Please add an RGB-only detector trained under the identical protocol (and ideally a residual-only detector) to Table 2, or explicitly restrict the claim to the full system.
  3. [§4.6, Tables 4–5; §4.2] Hyperparameters appear to be selected on the same curated test split used for reporting. The manuscript states that T=50 'provides the best balance among the tested inversion horizons' and Tables 4–5 highlight best configurations on the 246-entry curated SynthScars-CoT test split, but no separate validation split is mentioned. This raises the risk that the reported gains for T=50, alpha_mask=0.7, alpha_format=0.6, and the IoU bonus threshold are selection results rather than unbiased estimates. Use a validation split for model selection and report performance on a held-out test split, or demonstrate that the conclusions are stable across multiple random splits.
minor comments (5)
  1. [Table 4] The 'w/o residual stream' row and the 'RGB input only' row report identical numbers (61.83/41.71). If they are the same configuration, this should be stated explicitly; if they are different, the identical values should be explained.
  2. [§4.2, Standalone Detection Training] The detection head training description gives the learning rate, optimizer, and precision but omits the number of epochs and effective batch size. These details are needed for reproducibility.
  3. [§3.4, Eq. (6)] The GRPO objective is written as the group-relative policy term only, while the text says the implementation uses clipping and KL regularization. Consider presenting the full objective or explicitly labeling Eq. (6) as a simplified version.
  4. [Tables 1 and 2] The reported benchmark numbers are point estimates without confidence intervals or significance tests. For claims such as 'highest reported accuracy' and for the family-level differences, at least a brief statement of available variance or the absence of repeated trials would be helpful.
  5. [§3.4, Eq. (8)] The 'evidence-reference' reward is described only abstractly via 'evidence keywords.' A few concrete examples of the keywords and the exact structural markers used in R_pos would improve transparency, especially because CoT-Fmt in Table 5 is a format-only measure.

Circularity Check

0 steps flagged

No significant circularity: core claims are empirical against external benchmarks; only minor self-citations and a disclosed reward-induced dependence on the consistency map.

full rationale

The paper does not derive a target from its inputs; it reports empirical evaluations. The central detection claim is benchmarked on UniversalFakeDetect with a ProGAN-trained detector, and localization is compared on official SynthScars, LOKI, and RichHF splits. The DDIM residual R=|x−x̂| is a fixed, externally defined cue from a frozen reference model, and the ground-truth masks are expert annotations, so mask supervision is not derived from R. The closest candidate for circularity is the 'w/o residual stream' ablation (Table 4): the GRPO reward (Eq. 8) explicitly rewards 'explicit references to the consistency map,' and SFT CoT targets require a second step referencing the map. If those objectives are left unchanged, the drop partly reflects the objective's dependence on a now-absent input rather than the residual's forensic content. However, this is a validity/confound concern, not a definitional equivalence: the mask reward is supervised by external ground truth, counterfactual map interventions test spatial dependence independently, and the paper repeatedly disclaims that text-side rewards verify semantic faithfulness (Abstract, Sec. 3.4, Limitations). Self-citations [5,6,31-35] are contextual related-work citations and are not load-bearing for the main results. Overall, the paper is self-contained against external benchmarks; score 2 reflects minor self-citation and the disclosed reward-induced coupling, not circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The paper introduces no new physical or conceptual entities; the residual map and Where-What-Why protocol are constructed from existing components. The central free parameters are the DDIM horizon, GRPO reward weights, and related hyperparameters, all chosen empirically. The key domain assumptions concern the forensic interpretability of the residual and the reliability of the pseudo-labeled CoT data.

free parameters (4)
  • DDIM inversion horizon T = 50
    Selected as best among T=20,50,100 on the curated SynthScars-CoT test split (Table 4); used for all main residual maps.
  • GRPO reward weights alpha_mask, alpha_format = 0.7 / 0.6
    Set by hand without sensitivity analysis; Table 5 shows reward composition materially changes localization and format compliance.
  • Mask IoU bonus threshold and bonus value = IoU>0.7 -> +0.5
    Hand-selected constants in Eq. 7 that shape the GRPO mask reward and affect optimization behavior.
  • GRPO group size and sampling temperature = G=4, temperature=0.8
    Hyperparameters for the RL stage; no sensitivity study is provided.
axioms (4)
  • domain assumption The frozen Stable Diffusion v1.5 DDIM inversion-reconstruction residual R=|x-xhat| is a meaningful local compatibility cue for AI-generated and manipulated imagery across unseen generators.
    Section 3.2 defines the cue; Section 4.6 and the Limitations concede it also responds to benign high-frequency textures, compression, and domain shift.
  • domain assumption Feeding R through the same frozen CLIP-ViT-L/14 encoder as RGB produces features from which lightweight projectors and the MLLM can extract forensic signals.
    Eq. 3 uses shared-weight CLIP for R; the paper provides no dedicated validation that CLIP features of residual maps are informative, only stream-removal ablations.
  • domain assumption Qwen3-VL-Plus-generated CoT annotations, after human screening, are reliable enough to train structured forensic reasoning and localization.
    Section 4.1 describes the semi-automatic pipeline; no inter-annotator agreement or quality metrics are reported, and free-form faithfulness is conceded as unverified.
  • domain assumption GRPO with the specified rewards improves desired behavior rather than merely gaming format checks.
    Section 3.4/Table 5; the format reward measures surface markers and the authors state it is not a semantic verifier.

pith-pipeline@v1.3.0-alltime-deepseek · 15981 in / 12925 out tokens · 118113 ms · 2026-08-01T00:58:43.661780+00:00 · methodology

0 comments
read the original abstract

Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence. A frozen Stable Diffusion DDIM inversion-reconstruction model provides a fixed reconstruction reference, and its residual map measures local compatibility with that reference. Independent projectors encode the RGB image and residual map before a structured Where-What-Why model predicts a textual analysis and an artifact mask.Supervised fine-tuning is followed by Group Relative Policy Optimization (GRPO), whose reward combines mask overlap with output-structure and evidence-reference terms. These text-side terms encourage the model to refer to the consistency map but do not constitute a verifier of free-form textual truth. A separate image-level head fuses RGB and DDIM-residual class features. Experiments show cross-generator detection on UniversalFakeDetect and competitive artifact localization on the official SynthScars benchmark. Controlled cue-construction, inversion-horizon, component, reward-term, and counterfactual analyses support the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.

Figures

Figures reproduced from arXiv: 2607.25962 by Canran Xiao, Can Wang, Fei Shen, Yuhao Wang, Yushe Cao.

Figure 1
Figure 1. Figure 1: Motivation for LaP-Forensics. RGB-only forensic [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the dual-stream interface. RGB and DDIM-residual inputs share a frozen CLIP encoder but use independent [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Two-stage optimization of the localiza￾tion/reasoning model. Stage I jointly learns structured text generation and artifact segmentation. Stage II uses mask quality, output-format checks, residual-map references, and repetition control, with a clipped policy update and KL regularization to the frozen SFT reference policy. These observable reward terms encourage evidence-referencing outputs but do not direc… view at source ↗
Figure 4
Figure 4. Figure 4: Family-level detection accuracy on Universal [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison against LEGION on six se [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 4 canonical work pages · 2 internal anchors

  1. [1]

    George Cazenavette, Avneesh Sud, Thomas Leung, and Ben Usman. 2024. FakeIn- version: Learning to Detect Images from Unseen Text-to-Image Models by In- verting Stable Diffusion. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Los Alamitos, CA, USA, 10759–10769. https://doi.org/10.1109/CVPR52733.2024.01023

  2. [2]

    Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. 2020. What Makes Fake Images Detectable? Understanding Properties that Generalize. https: //doi.org/10.48550/arXiv.2008.10588

  3. [3]

    Baoying Chen, Jishen Zeng, Jianquan Yang, and Rui Yang. 2024. DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated Images. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, Vienna, Austria, 7621–7639. https://proceedings.mlr.p...

  4. [4]

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al . 2024. InternVL: Scal- ing up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 24185–24198

  5. [5]

    Jingtong Dou, Chuancheng Shi, Jian Wang, Fei Shen, Zhiyong Wang, and Tat- Seng Chua. 2026. Beyond surface artifacts: Capturing shared latent forgery knowledge across modalities.arXiv preprint arXiv:2604.07763(2026)

  6. [6]

    Jingtong Dou, Chuancheng Shi, Yemin Wang, Shiming Guo, Anqi Yi, Wenhua Wu, Li Zhang, Fei Shen, and Tat-Seng Chua. 2026. DNA: Uncovering Universal Latent Forgery Knowledge.arXiv preprint arXiv:2601.22515(2026)

  7. [7]

    Fabrizio Guillaro, Davide Cozzolino, Avneesh Sud, Nicholas Dufour, and Luisa Verdoliva. 2023. TruFor: Leveraging All-Round Clues for Trustworthy Image Forgery Detection and Localization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 20606–20615

  8. [8]

    Xiao Guo, Xiaohong Liu, Zhiyuan Ren, Steven Grosz, Iacopo Masi, and Xiaoming Liu. 2023. Hierarchical Fine-Grained Image Forgery Detection and Localiza- tion. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 3155–3165

  9. [9]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. https://doi.org/10.48550/arXiv.2106.09685

  10. [10]

    Zhenglin Huang, Jinwei Hu, Xiangtai Li, Yiwei He, Xingyu Zhao, Bei Peng, Baoyuan Wu, Xiaowei Huang, and Guangliang Cheng. 2025. SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model. https://doi.org/10.48550/arXiv.2412.04292

  11. [11]

    Yikun Ji, Yan Hong, Bowen Deng, Jun Lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, and Jianfu Zhang. 2025. Locate-Then-Examine: Grounded Region Rea- soning Improves Detection of AI-Generated Images. https://doi.org/10.48550/ arXiv.2510.04225

  12. [12]

    Yikun Ji, Hong Yan, Jun Lan, Huijia Zhu, Weiqiang Wang, Qi Fan, Liqing Zhang, and Jianfu Zhang. 2025. Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs. https://doi.org/10.48550/arXiv.2506. 07045

  13. [13]

    Zhengyuan Jiang, Yuyang Zhang, Moyang Guo, and Neil Zhenqiang Gong. 2025. EditTrack: Detecting and Attributing AI-assisted Image Editing. https://doi.org/ 10.48550/arXiv.2510.01173

  14. [14]

    Hengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye, Weijia Li, Peilin Feng, Baichuan Zhou, Bin Wang, Dahua Lin, Linfeng Zhang, and Conghui He. 2025. LEGION: Learning to Ground and Explain for Synthetic Image Detection. https://doi.org/10.48550/arXiv.2503.15264

  15. [15]

    Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. 2023. Segment Anything. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Paris, France, 4015– 4026

  16. [16]

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. 2024. LISA: Reasoning Segmentation via Large Language Model. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 9579–9589

  17. [17]

    Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. 2020. Face X-Ray for More General Face Forgery Detection. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 5001–5010

  18. [18]

    Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Junyan Ye, Ke-Yue Zhang, Yue Zhou, Peng Jin, Bin Li, Taiping Yao, and Shouhong Ding. 2025. Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection. https://doi.org/10.48550/arXiv.2509.25502

  19. [19]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023. Improved Baselines with Visual Instruction Tuning. https://doi.org/10.48550/arXiv.2310.03744

  20. [20]

    Huan Liu, Zichang Tan, Chuangchuang Tan, Yunchao Wei, Jingdong Wang, and Yao Zhao. 2024. Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 10770–10780

  21. [21]

    Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. 2021. Generalizing Face Forgery Detection with High-Frequency Features. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Nashville, TN, USA, 16317–16326

  22. [22]

    Lianrui Mu, Xingze Zou, Jianhong Bai, Jiaqi Hu, Wenjie Zheng, Jiangnan Ye, Jiedong Zhuang, Mudassar Ali, Jing Wang, and Haoji Hu. 2025. No Pixel Left Be- hind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection. https://doi.org/10.48550/arXiv.2508.17346

  23. [23]

    Bappy, Amit K

    Lakshmanan Nataraj, Tajuddin Manhar Mohammed, Shivkumar Chandrasekaran, Arjuna Flenner, Jawadul H. Bappy, Amit K. Roy-Chowdhury, and B. S. Manjunath

  24. [24]

    Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. 2023. Towards Universal Fake Image Detectors that Generalize Across Generative Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 24480–24489

  25. [25]

    Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. 2020. Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues. In Computer Vision – ECCV 2020. Springer, Glasgow, UK, 86–103

  26. [26]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models from Natural Language Supervision. https://doi.org/10.48550/arXiv.2103.00020

  27. [27]

    Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. FaceForensics++: Learning to Detect Manipulated Fa- cial Images. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Seoul, Korea, 1–11

  28. [28]

    Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Gurpreet Singh, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Amitava Das, Amit Sheth, Gaytri Jena, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, and Aman Chadha. 2026. A Comprehensive Dataset for Human v...

  29. [29]

    Zeyang Sha, Yicong Tan, Mingjie Li, Michael Backes, and Yang Zhang. 2024. ZeroFake: Zero-Shot Detection of Fake Images Generated and Edited by Text-to- Image Generation Models. InProceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security. Association for Computing Machinery, New York, NY, USA, 4852–4866. https://doi.org/10.1145/...

  30. [30]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, Daya Guo, et al. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. https: //doi.org/10.48550/arXiv.2402.03300

  31. [31]

    Fei Shen, Xin Jiang, Xin He, Hu Ye, Cong Wang, Xiaoyu Du, Zechao Li, and Jinhui Tang. 2025. Imagdressing-v1: Customizable virtual dressing. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 6795–6804

  32. [32]

    Fei Shen and Jinhui Tang. 2024. Imagpose: A unified conditional framework for pose-guided person generation.Advances in neural information processing systems37 (2024), 6246–6266

  33. [33]

    Fei Shen, Hu Ye, Sibo Liu, Jun Zhang, Cong Wang, Xiao Han, and Yang Wei. 2025. Boosting consistency in story visualization with rich-contextual conditional diffusion models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 6785–6794

  34. [34]

    Fei Shen, Hu Ye, Jun Zhang, Cong Wang, Xiao Han, and Yang Wei. 2024. Ad- vancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=rHzapPnCgT

  35. [35]

    Fei Shen, Jian Yu, Cong Wang, Xin Jiang, Xiaoyu Du, and Jinhui Tang. 2025. IMAGGarment-1: Fine-Grained Garment Generation for Controllable Fashion Design.arXiv preprint arXiv:2504.13176(2025)

  36. [36]

    Zenan Shi, Wenyu Liu, and Haipeng Chen. 2025. Face Reconstruction-Based Generalized Deepfake Detection Model with Residual Outlook Attention.ACM Transactions on Multimedia Computing, Communications, and Applications21, 4 (2025), 1–19. https://doi.org/10.1145/3686162

  37. [37]

    Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising Diffusion Implicit Models. https://doi.org/10.48550/arXiv.2010.02502

  38. [38]

    Ke Sun, Shen Chen, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, Chia-Wen Lin, and Rongrong Ji. 2025. Towards General Visual-Linguistic Face Forgery Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Nashville, TN, USA, 19576–19586

  39. [39]

    Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. 2024. Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning.Proceedings of the AAAI Conference on Artificial Intelligence38, 5 (2024), 5052–5060. https://doi.org/10.1609/aaai. v38i5.28310

  40. [40]

    Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. 2024. Rethinking the Up-Sampling Operations in CNN-Based Generative 9 Wang et al., Can Wang, Yuhao Wang, Yushe Cao, Canran Xiao, and Fei Shen Network for Generalizable Deepfake Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...

  41. [41]

    Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei. 2023. Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 12105–12114

  42. [42]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. https://doi.org/10.48550/arXiv.2307.09288

  43. [43]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. https: //doi.org/10.48550/arXiv.2409.12191

  44. [44]

    Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. 2020. CNN-Generated Images Are Surprisingly Easy to Spot...for Now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR). IEEE, Seattle, WA, USA, 8695–8704

  45. [45]

    Yabin Wang, Zhiwu Huang, and Xiaopeng Hong. 2025. OpenSDI: Spotting Diffusion-Generated Images in the Open World. https://doi.org/10.48550/arXiv. 2503.19653

  46. [46]

    Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, and Houqiang Li. 2023. DIRE for Diffusion-Generated Image Detection. https://doi.org/10.48550/arXiv.2303.09295

  47. [47]

    Shiyu Wu, Shuyan Li, Jing Li, Jing Liu, and Yequan Wang. 2025. Few-Shot Syn- thetic Image Attribution: Identifying Unseen Generators with Limited Samples. https://doi.org/10.48550/arXiv.2509.25682

  48. [48]

    Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai, Kaiwei Zhang, Qiang Hu, Jing Liu, Xiongkuo Min, and Guangtao Zhai. 2025. ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation. https://doi.org/10.48550/arXiv.2511.14259

  49. [49]

    Bosheng Yan, Chang-Tsun Li, and Xuequan Lu. 2024. JRC: Deepfake detection via joint reconstruction and classification.Neurocomputing598 (2024), 127862. https://doi.org/10.1016/j.neucom.2024.127862

  50. [50]

    Shilin Yan, Ouxiang Li, Jiayin Cai, Yanbin Hao, Xiaolong Jiang, Yao Hu, and Weidi Xie. 2025. A Sanity Check for AI-generated Image Detection. https: //doi.org/10.48550/arXiv.2406.19435

  51. [51]

    Yongqi Yang, Zhihao Qian, Ye Zhu, Olga Russakovsky, and Yu Wu. 2025. D 3: Scaling Up Deepfake Detection by Learning from Discrepancy. https://doi.org/ 10.48550/arXiv.2404.04584

  52. [52]

    Zheng Yang, Ruoxin Chen, Zhiyuan Yan, Ke-Yue Zhang, Xinghe Fu, Shuang Wu, Xiujun Shu, Taiping Yao, Shouhong Ding, Zequn Qin, and Xi Li. 2025. All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning. https://doi.org/10.48550/arXiv.2504.01396

  53. [53]

    Lingzhi Zhang, Zhengjie Xu, Connelly Barnes, Yuqian Zhou, Qing Liu, He Zhang, Sohrab Amirghodsi, Zhe Lin, Eli Shechtman, and Jianbo Shi. 2023. Perceptual Artifacts Localization for Image Synthesis Tasks. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Paris, France, 7579– 7590

  54. [54]

    Xu Zhang, Svebor Karaman, and Shih-Fu Chang. 2019. Detecting and Simulating Artifacts in GAN Fake Images. In2019 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, Delft, Netherlands, 1–6

  55. [55]

    Morariu, and Larry S

    Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. 2018. Learning Rich Features for Image Manipulation Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Salt Lake City, UT, USA, 1053–1061

  56. [56]

    Yuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin, Linkai Liu, Zipeng Guo, Hao Fei, Xiaobo Xia, and Chao Gou. 2025. Where, What, Why: Towards Explainable Driver Attention Prediction. https://doi.org/10.48550/arXiv.2506.23088

  57. [57]

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. https://doi.org/10.48550/arXiv.2304.10592 10

  58. [2019]

    https://doi.org/10.48550/arXiv.1903.06836

    Detecting GAN Generated Fake Images Using Co-Occurrence Matrices. https://doi.org/10.48550/arXiv.1903.06836