REVIEW 3 major objections 5 minor 58 references
This paper argues that adding a DDIM reconstruction-residual stream to a large multimodal model improves explainable deepfake detection and artifact localization beyond RGB evidence alone.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 00:58 UTC pith:37ZIOWEN
load-bearing objection A credible, novel dual-stream deepfake detector whose central residual-stream claim is confounded by its own reward function—worth refereeing, but the causal evidence needs a cleaner experiment. the 3 major comments →
LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that a pixel-level compatibility residual, computed by deterministic DDIM inversion-reconstruction against a fixed Stable Diffusion v1.5 reference at T=50, acts as a transferable forensic cue when fused with RGB semantics. The paper reports that this dual-stream representation, trained only on ProGAN and evaluated on UniversalFakeDetect, reaches 97.23% accuracy on GANs, 92.18% on Diffusion, 98.98% on CRN, and 98.62% on IMLE — the highest reported numbers in those columns — while a separate reasoning model using the same cue with a LLaMA-2-7B backbone and SAM decode achieves 72.19 mIoU and 63.62 F1 on a curated SynthScars-CoT split, versus 61.83 and 41.71 when the res
What carries the argument
The load-bearing object is the DDIM inversion-reconstruction residual R=|x−x̂|, computed with a frozen Stable Diffusion v1.5 and a fixed T=50 deterministic schedule. This map is fed through the same frozen CLIP vision encoder as the RGB image but with an independent projector, so the language model can attend to both streams. The reasoning side is a structured Where-What-Why chain-of-thought with a [SEG] token decoded by SAM, optimized first by supervised fine-tuning and then by GRPO with a reward that mixes mask IoU with format and evidence-reference terms. The image-level detector is a simple MLP over concatenated [CLS] tokens from both streams. The residual stream is the mechanism that ca
Load-bearing premise
The entire argument hinges on the premise that the DDIM residual map reflects manipulation-induced incompatibility rather than merely benign textures, compression, resampling, or domain shift — a distinction the paper itself flags as open in Section 3.2 and the Limitations.
What would settle it
Run the same architecture on a corpus of clean, unmanipulated, high-texture photographs (hair, woven fabric, foliage, specular highlights); if the residual-driven detector's false-positive rate rises sharply or the localization model produces masks on these pristine images with nontrivial IoU against empty ground truth, then the cue is tracking benign texture rather than manipulation.
If this is right
- A single fixed reconstruction reference can serve as a shared forensic cue across generator families, including GANs and CRN that do not follow diffusion trajectories — a point the paper explicitly limits to the evaluated protocol.
- Reconstruction residuals and RGB semantics carry complementary information; the multi-step DDIM residual outperforms both RGB-only and single-step VAE residuals in the paper's ablations.
- Combining mask IoU with format and evidence-reference rewards in GRPO improves localization over mask-only reinforcement, indicating that structural constraints stabilize policy optimization for segmentation.
- Counterfactual map interventions provide a way to test whether a multimodal forensic model actually uses its evidence stream rather than merely being affected by its removal.
- The documented sensitivity to benign high-frequency textures and the lack of a verifier for free-form textual truth bound the approach's scope to settings without strong post-processing.
Where Pith is reading between the lines
- A natural extension would be to evaluate the same residual construction as an explicit fusion input for non-diffusion detectors, such as GAN or editing-attribution models; if it transfers, the residual would act as a general 'compatibility fingerprint' rather than a diffusion-specific artifact.
- The zero-map and donor-map intervention procedure could be adopted as a standard sanity check for any explainable forensic model that claims to ground its explanations in a given evidence map.
- Because the residual is computed against one fixed reconstruction reference, the framework inherits that reference's biases; pooling residuals from a family of frozen references could improve robustness, though the paper does not explore this.
- The paper's distinction between 'evidence reference' and 'semantic faithfulness' suggests that a future entailment-based verification reward could complement the current keyword and structure rewards in GRPO.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LaP-Forensics, a two-stream framework that augments RGB image semantics with a DDIM inversion-reconstruction residual computed against a frozen Stable Diffusion v1.5 reference. The residual map is encoded as a second visual stream, and a LLaMA-2-7B MLLM with a SAM decoder performs structured Where–What–Why reasoning and artifact localization. A separately trained image-level head concatenates RGB and residual CLS features for detection. Training uses supervised fine-tuning followed by GRPO with mask, format, and evidence-reference rewards. Experiments are reported on UniversalFakeDetect for detection and on SynthScars, LOKI, and RichHF for localization, together with component, cue-construction, horizon, reward, and counterfactual ablations. The paper carefully separates the official SynthScars protocol from an internally curated CoT split, and it repeatedly and explicitly disclaims that the residual is a calibrated manipulation probability or a source-identification cue. The central claim is that the residual stream is the most critical component, but, as detailed below, the evidence for this claim is weakened by a confounded ablation and by the absence of a same-protocol RGB-only detector baseline.
Significance. If the residual-stream benefit were convincingly established, the paper would be a useful contribution to explainable deepfake detection: it combines a fixed reconstruction reference with MLLM-based reasoning and pixel-level localization, and it evaluates the model's dependence on the residual through counterfactual interventions. The paper deserves credit for explicitly bounding the interpretation of the residual map, for separating official-benchmark and curated-protocol results, and for acknowledging that the text-side rewards do not verify semantic faithfulness. However, the current evidence is not yet sufficient to support the headline contribution. The main residual-stream ablation is confounded by the GRPO evidence-reference reward and the CoT targets, and the detection experiments do not include an RGB-only control trained under the same protocol. These are fixable with additional experiments and clarifications, but they are load-bearing for the paper's central claim.
major comments (3)
- [§4.6, Table 4; §3.4, Eq. (8)] The residual-stream ablation is not isolated. The full GRPO reward (Eq. 8) includes a format/evidence-reference term that explicitly rewards references to the Latent-Pixel Consistency Map, and the SFT CoT targets (Section 3.3) include a step that relates observations to that map. For the 'w/o residual stream' and 'RGB input only' conditions, the manuscript does not state that these targets/rewards were modified. If they were kept unchanged, the ablated model is penalized for failing to reference an input that no longer exists, so the 10.36 mIoU / 21.91 F1 drop conflates loss of forensic information with an objective that is partially unsatisfiable. This concern is material: Table 5 shows that reward-configuration changes alone move mIoU by roughly 17 points (mask-only 55.46 vs full 72.19), comparable to the residual ablation. Please report the exact training setup for the ablated variant
- [§4.4, Table 2; §4.2] The standalone detector's residual contribution is unisolated. Table 2 compares the RGB+DDIM-residual detector against published baselines with different training objectives, data sources, and checkpoints. There is no same-protocol RGB-only detector row. Consequently the results in Table 2 and the sentence in §4.4 that the 'RGB-plus-reconstruction representation transfers numerically' support a system-level benchmark claim, not the utility of the residual stream for detection. Please add an RGB-only detector trained under the identical protocol (and ideally a residual-only detector) to Table 2, or explicitly restrict the claim to the full system.
- [§4.6, Tables 4–5; §4.2] Hyperparameters appear to be selected on the same curated test split used for reporting. The manuscript states that T=50 'provides the best balance among the tested inversion horizons' and Tables 4–5 highlight best configurations on the 246-entry curated SynthScars-CoT test split, but no separate validation split is mentioned. This raises the risk that the reported gains for T=50, alpha_mask=0.7, alpha_format=0.6, and the IoU bonus threshold are selection results rather than unbiased estimates. Use a validation split for model selection and report performance on a held-out test split, or demonstrate that the conclusions are stable across multiple random splits.
minor comments (5)
- [Table 4] The 'w/o residual stream' row and the 'RGB input only' row report identical numbers (61.83/41.71). If they are the same configuration, this should be stated explicitly; if they are different, the identical values should be explained.
- [§4.2, Standalone Detection Training] The detection head training description gives the learning rate, optimizer, and precision but omits the number of epochs and effective batch size. These details are needed for reproducibility.
- [§3.4, Eq. (6)] The GRPO objective is written as the group-relative policy term only, while the text says the implementation uses clipping and KL regularization. Consider presenting the full objective or explicitly labeling Eq. (6) as a simplified version.
- [Tables 1 and 2] The reported benchmark numbers are point estimates without confidence intervals or significance tests. For claims such as 'highest reported accuracy' and for the family-level differences, at least a brief statement of available variance or the absence of repeated trials would be helpful.
- [§3.4, Eq. (8)] The 'evidence-reference' reward is described only abstractly via 'evidence keywords.' A few concrete examples of the keywords and the exact structural markers used in R_pos would improve transparency, especially because CoT-Fmt in Table 5 is a format-only measure.
Circularity Check
No significant circularity: core claims are empirical against external benchmarks; only minor self-citations and a disclosed reward-induced dependence on the consistency map.
full rationale
The paper does not derive a target from its inputs; it reports empirical evaluations. The central detection claim is benchmarked on UniversalFakeDetect with a ProGAN-trained detector, and localization is compared on official SynthScars, LOKI, and RichHF splits. The DDIM residual R=|x−x̂| is a fixed, externally defined cue from a frozen reference model, and the ground-truth masks are expert annotations, so mask supervision is not derived from R. The closest candidate for circularity is the 'w/o residual stream' ablation (Table 4): the GRPO reward (Eq. 8) explicitly rewards 'explicit references to the consistency map,' and SFT CoT targets require a second step referencing the map. If those objectives are left unchanged, the drop partly reflects the objective's dependence on a now-absent input rather than the residual's forensic content. However, this is a validity/confound concern, not a definitional equivalence: the mask reward is supervised by external ground truth, counterfactual map interventions test spatial dependence independently, and the paper repeatedly disclaims that text-side rewards verify semantic faithfulness (Abstract, Sec. 3.4, Limitations). Self-citations [5,6,31-35] are contextual related-work citations and are not load-bearing for the main results. Overall, the paper is self-contained against external benchmarks; score 2 reflects minor self-citation and the disclosed reward-induced coupling, not circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- DDIM inversion horizon T =
50
- GRPO reward weights alpha_mask, alpha_format =
0.7 / 0.6
- Mask IoU bonus threshold and bonus value =
IoU>0.7 -> +0.5
- GRPO group size and sampling temperature =
G=4, temperature=0.8
axioms (4)
- domain assumption The frozen Stable Diffusion v1.5 DDIM inversion-reconstruction residual R=|x-xhat| is a meaningful local compatibility cue for AI-generated and manipulated imagery across unseen generators.
- domain assumption Feeding R through the same frozen CLIP-ViT-L/14 encoder as RGB produces features from which lightweight projectors and the MLLM can extract forensic signals.
- domain assumption Qwen3-VL-Plus-generated CoT annotations, after human screening, are reliable enough to train structured forensic reasoning and localization.
- domain assumption GRPO with the specified rewards improves desired behavior rather than merely gaming format checks.
read the original abstract
Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB semantics with reconstruction-based forensic evidence. A frozen Stable Diffusion DDIM inversion-reconstruction model provides a fixed reconstruction reference, and its residual map measures local compatibility with that reference. Independent projectors encode the RGB image and residual map before a structured Where-What-Why model predicts a textual analysis and an artifact mask.Supervised fine-tuning is followed by Group Relative Policy Optimization (GRPO), whose reward combines mask overlap with output-structure and evidence-reference terms. These text-side terms encourage the model to refer to the consistency map but do not constitute a verifier of free-form textual truth. A separate image-level head fuses RGB and DDIM-residual class features. Experiments show cross-generator detection on UniversalFakeDetect and competitive artifact localization on the official SynthScars benchmark. Controlled cue-construction, inversion-horizon, component, reward-term, and counterfactual analyses support the utility of the residual stream under the evaluated settings, while free-form textual faithfulness and reliability under post-processing remain open limitations.
Figures
Reference graph
Works this paper leans on
-
[1]
George Cazenavette, Avneesh Sud, Thomas Leung, and Ben Usman. 2024. FakeIn- version: Learning to Detect Images from Unseen Text-to-Image Models by In- verting Stable Diffusion. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Los Alamitos, CA, USA, 10759–10769. https://doi.org/10.1109/CVPR52733.2024.01023
arXiv 2024
-
[2]
Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. 2020. What Makes Fake Images Detectable? Understanding Properties that Generalize. https: //doi.org/10.48550/arXiv.2008.10588
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2008.10588 2020
-
[3]
Baoying Chen, Jishen Zeng, Jianquan Yang, and Rui Yang. 2024. DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated Images. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235). PMLR, Vienna, Austria, 7621–7639. https://proceedings.mlr.p...
2024
-
[4]
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al . 2024. InternVL: Scal- ing up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 24185–24198
2024
-
[5]
Jingtong Dou, Chuancheng Shi, Jian Wang, Fei Shen, Zhiyong Wang, and Tat- Seng Chua. 2026. Beyond surface artifacts: Capturing shared latent forgery knowledge across modalities.arXiv preprint arXiv:2604.07763(2026)
Pith/arXiv arXiv 2026
-
[6]
Jingtong Dou, Chuancheng Shi, Yemin Wang, Shiming Guo, Anqi Yi, Wenhua Wu, Li Zhang, Fei Shen, and Tat-Seng Chua. 2026. DNA: Uncovering Universal Latent Forgery Knowledge.arXiv preprint arXiv:2601.22515(2026)
arXiv 2026
-
[7]
Fabrizio Guillaro, Davide Cozzolino, Avneesh Sud, Nicholas Dufour, and Luisa Verdoliva. 2023. TruFor: Leveraging All-Round Clues for Trustworthy Image Forgery Detection and Localization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 20606–20615
2023
-
[8]
Xiao Guo, Xiaohong Liu, Zhiyuan Ren, Steven Grosz, Iacopo Masi, and Xiaoming Liu. 2023. Hierarchical Fine-Grained Image Forgery Detection and Localiza- tion. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 3155–3165
2023
-
[9]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. https://doi.org/10.48550/arXiv.2106.09685
-
[10]
Zhenglin Huang, Jinwei Hu, Xiangtai Li, Yiwei He, Xingyu Zhao, Bei Peng, Baoyuan Wu, Xiaowei Huang, and Guangliang Cheng. 2025. SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model. https://doi.org/10.48550/arXiv.2412.04292
-
[11]
Yikun Ji, Yan Hong, Bowen Deng, Jun Lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, and Jianfu Zhang. 2025. Locate-Then-Examine: Grounded Region Rea- soning Improves Detection of AI-Generated Images. https://doi.org/10.48550/ arXiv.2510.04225
-
[12]
Yikun Ji, Hong Yan, Jun Lan, Huijia Zhu, Weiqiang Wang, Qi Fan, Liqing Zhang, and Jianfu Zhang. 2025. Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs. https://doi.org/10.48550/arXiv.2506. 07045
-
[13]
Zhengyuan Jiang, Yuyang Zhang, Moyang Guo, and Neil Zhenqiang Gong. 2025. EditTrack: Detecting and Attributing AI-assisted Image Editing. https://doi.org/ 10.48550/arXiv.2510.01173
-
[14]
Hengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye, Weijia Li, Peilin Feng, Baichuan Zhou, Bin Wang, Dahua Lin, Linfeng Zhang, and Conghui He. 2025. LEGION: Learning to Ground and Explain for Synthetic Image Detection. https://doi.org/10.48550/arXiv.2503.15264
-
[15]
Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. 2023. Segment Anything. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Paris, France, 4015– 4026
2023
-
[16]
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. 2024. LISA: Reasoning Segmentation via Large Language Model. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 9579–9589
2024
-
[17]
Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. 2020. Face X-Ray for More General Face Forgery Detection. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 5001–5010
2020
-
[18]
Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Junyan Ye, Ke-Yue Zhang, Yue Zhou, Peng Jin, Bin Li, Taiping Yao, and Shouhong Ding. 2025. Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection. https://doi.org/10.48550/arXiv.2509.25502
-
[19]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023. Improved Baselines with Visual Instruction Tuning. https://doi.org/10.48550/arXiv.2310.03744
-
[20]
Huan Liu, Zichang Tan, Chuangchuang Tan, Yunchao Wei, Jingdong Wang, and Yao Zhao. 2024. Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, WA, USA, 10770–10780
2024
-
[21]
Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. 2021. Generalizing Face Forgery Detection with High-Frequency Features. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Nashville, TN, USA, 16317–16326
2021
-
[22]
Lianrui Mu, Xingze Zou, Jianhong Bai, Jiaqi Hu, Wenjie Zheng, Jiangnan Ye, Jiedong Zhuang, Mudassar Ali, Jing Wang, and Haoji Hu. 2025. No Pixel Left Be- hind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection. https://doi.org/10.48550/arXiv.2508.17346
-
[23]
Bappy, Amit K
Lakshmanan Nataraj, Tajuddin Manhar Mohammed, Shivkumar Chandrasekaran, Arjuna Flenner, Jawadul H. Bappy, Amit K. Roy-Chowdhury, and B. S. Manjunath
-
[24]
Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. 2023. Towards Universal Fake Image Detectors that Generalize Across Generative Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 24480–24489
2023
-
[25]
Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. 2020. Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues. In Computer Vision – ECCV 2020. Springer, Glasgow, UK, 86–103
2020
-
[26]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models from Natural Language Supervision. https://doi.org/10.48550/arXiv.2103.00020
-
[27]
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. FaceForensics++: Learning to Detect Manipulated Fa- cial Images. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Seoul, Korea, 1–11
2019
-
[28]
Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Gurpreet Singh, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Amitava Das, Amit Sheth, Gaytri Jena, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, and Aman Chadha. 2026. A Comprehensive Dataset for Human v...
-
[29]
Zeyang Sha, Yicong Tan, Mingjie Li, Michael Backes, and Yang Zhang. 2024. ZeroFake: Zero-Shot Detection of Fake Images Generated and Edited by Text-to- Image Generation Models. InProceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security. Association for Computing Machinery, New York, NY, USA, 4852–4866. https://doi.org/10.1145/...
arXiv 2024
-
[30]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, Daya Guo, et al. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. https: //doi.org/10.48550/arXiv.2402.03300
-
[31]
Fei Shen, Xin Jiang, Xin He, Hu Ye, Cong Wang, Xiaoyu Du, Zechao Li, and Jinhui Tang. 2025. Imagdressing-v1: Customizable virtual dressing. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 6795–6804
2025
-
[32]
Fei Shen and Jinhui Tang. 2024. Imagpose: A unified conditional framework for pose-guided person generation.Advances in neural information processing systems37 (2024), 6246–6266
2024
-
[33]
Fei Shen, Hu Ye, Sibo Liu, Jun Zhang, Cong Wang, Xiao Han, and Yang Wei. 2025. Boosting consistency in story visualization with rich-contextual conditional diffusion models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 6785–6794
2025
-
[34]
Fei Shen, Hu Ye, Jun Zhang, Cong Wang, Xiao Han, and Yang Wei. 2024. Ad- vancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=rHzapPnCgT
2024
-
[35]
Fei Shen, Jian Yu, Cong Wang, Xin Jiang, Xiaoyu Du, and Jinhui Tang. 2025. IMAGGarment-1: Fine-Grained Garment Generation for Controllable Fashion Design.arXiv preprint arXiv:2504.13176(2025)
Pith/arXiv arXiv 2025
-
[36]
Zenan Shi, Wenyu Liu, and Haipeng Chen. 2025. Face Reconstruction-Based Generalized Deepfake Detection Model with Residual Outlook Attention.ACM Transactions on Multimedia Computing, Communications, and Applications21, 4 (2025), 1–19. https://doi.org/10.1145/3686162
-
[37]
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising Diffusion Implicit Models. https://doi.org/10.48550/arXiv.2010.02502
-
[38]
Ke Sun, Shen Chen, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, Chia-Wen Lin, and Rongrong Ji. 2025. Towards General Visual-Linguistic Face Forgery Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Nashville, TN, USA, 19576–19586
2025
-
[39]
Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. 2024. Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning.Proceedings of the AAAI Conference on Artificial Intelligence38, 5 (2024), 5052–5060. https://doi.org/10.1609/aaai. v38i5.28310
doi:10.1609/aaai 2024
-
[40]
Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. 2024. Rethinking the Up-Sampling Operations in CNN-Based Generative 9 Wang et al., Can Wang, Yuhao Wang, Yushe Cao, Canran Xiao, and Fei Shen Network for Generalizable Deepfake Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR...
2024
-
[41]
Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei. 2023. Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Vancouver, BC, Canada, 12105–12114
2023
-
[42]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. https://doi.org/10.48550/arXiv.2307.09288
-
[43]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. https: //doi.org/10.48550/arXiv.2409.12191
-
[44]
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. 2020. CNN-Generated Images Are Surprisingly Easy to Spot...for Now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR). IEEE, Seattle, WA, USA, 8695–8704
2020
-
[45]
Yabin Wang, Zhiwu Huang, and Xiaopeng Hong. 2025. OpenSDI: Spotting Diffusion-Generated Images in the Open World. https://doi.org/10.48550/arXiv. 2503.19653
-
[46]
Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, and Houqiang Li. 2023. DIRE for Diffusion-Generated Image Detection. https://doi.org/10.48550/arXiv.2303.09295
-
[47]
Shiyu Wu, Shuyan Li, Jing Li, Jing Liu, and Yequan Wang. 2025. Few-Shot Syn- thetic Image Attribution: Identifying Unseen Generators with Limited Samples. https://doi.org/10.48550/arXiv.2509.25682
-
[48]
Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai, Kaiwei Zhang, Qiang Hu, Jing Liu, Xiongkuo Min, and Guangtao Zhai. 2025. ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation. https://doi.org/10.48550/arXiv.2511.14259
-
[49]
Bosheng Yan, Chang-Tsun Li, and Xuequan Lu. 2024. JRC: Deepfake detection via joint reconstruction and classification.Neurocomputing598 (2024), 127862. https://doi.org/10.1016/j.neucom.2024.127862
arXiv 2024
-
[50]
Shilin Yan, Ouxiang Li, Jiayin Cai, Yanbin Hao, Xiaolong Jiang, Yao Hu, and Weidi Xie. 2025. A Sanity Check for AI-generated Image Detection. https: //doi.org/10.48550/arXiv.2406.19435
-
[51]
Yongqi Yang, Zhihao Qian, Ye Zhu, Olga Russakovsky, and Yu Wu. 2025. D 3: Scaling Up Deepfake Detection by Learning from Discrepancy. https://doi.org/ 10.48550/arXiv.2404.04584
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2404.04584 2025
-
[52]
Zheng Yang, Ruoxin Chen, Zhiyuan Yan, Ke-Yue Zhang, Xinghe Fu, Shuang Wu, Xiujun Shu, Taiping Yao, Shouhong Ding, Zequn Qin, and Xi Li. 2025. All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning. https://doi.org/10.48550/arXiv.2504.01396
-
[53]
Lingzhi Zhang, Zhengjie Xu, Connelly Barnes, Yuqian Zhou, Qing Liu, He Zhang, Sohrab Amirghodsi, Zhe Lin, Eli Shechtman, and Jianbo Shi. 2023. Perceptual Artifacts Localization for Image Synthesis Tasks. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Paris, France, 7579– 7590
2023
-
[54]
Xu Zhang, Svebor Karaman, and Shih-Fu Chang. 2019. Detecting and Simulating Artifacts in GAN Fake Images. In2019 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, Delft, Netherlands, 1–6
2019
-
[55]
Morariu, and Larry S
Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. 2018. Learning Rich Features for Image Manipulation Detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Salt Lake City, UT, USA, 1053–1061
2018
-
[56]
Yuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin, Linkai Liu, Zipeng Guo, Hao Fei, Xiaobo Xia, and Chao Gou. 2025. Where, What, Why: Towards Explainable Driver Attention Prediction. https://doi.org/10.48550/arXiv.2506.23088
-
[57]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. https://doi.org/10.48550/arXiv.2304.10592 10
-
[2019]
https://doi.org/10.48550/arXiv.1903.06836
Detecting GAN Generated Fake Images Using Co-Occurrence Matrices. https://doi.org/10.48550/arXiv.1903.06836
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.