REVIEW 4 major objections 5 minor 58 references
Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read By splitting denoising into three stages and giving each its own reward, SGPO raises diffusion-model generative quality by 26.7% on average and speeds convergence by 36.7% while reducing reward hacking.
desk verdict A stage-aware reward scheme with genuinely careful anti-reward-hacking evaluation, but the dense reward's relation to the true objective is unproven and must be fixed before acceptance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the adaptive stage indicator $\pi(t)$, which partitions the denoising trajectory into three intervals using two switch points: $t_c^{(1)}$, where the second derivative of the signal-to-noise ratio $\gamma(t)$ tends to zero, and $t_c^{(2)}$, where the second derivative of the semantic-evolution rate $\Delta E(t)=\phi(\hat{x}_0(x_t),z)-\phi(\hat{x}_0(x_{t-1}),z)$ tends to zero. Each interval gets a distinct reward: $r_t^{(I)}=\lambda_t\|x_t-x_T\|^2$ to flee chaos; $r_t^{(II)}=r(x_0,z)+\Delta E(t)\|\hat{x}_0(x_t)-\hat{x}_0(x_{t-1})\|^2$ to optimize preference while exploring under semantic constraint; and $r_t^{(III)}=-\frac{1}{2}\|\hat{x}_0(x_t)-x_0^{\mathrm{pre}}\|^2$ to converge stably to the pretrained output. Theorem 3.1, which states that the posterior variance of $\hat{x}_0$ decreases monotonically during generation, supplies the theoretical justification for why uncertainty is high early, moderate in the middle, and low late, matching the three reward priorities.
What would settle it
Run SGPO with stage boundaries randomized—same three reward formulas but with the chaotic, stable, and convergent segments assigned to random portions of each trajectory. If randomized boundaries match SGPO's 26.7% quality gain and 36.7% speedup, then the specific online stage detection is not the causal source of the improvement and the temporal-alignment claim fails.
Extended reading notes
Core claim
The paper's central claim is that the temporal homogenization of rewards—giving every denoising step the same final-reward signal—is a root cause of reward hacking in RL fine-tuning of diffusion models. SGPO replaces that with a stage-aware dense reward $R_{\mathrm{total}}$: during the chaotic Stage I it maximizes distance to the initial noise to exit noise-dominated states; during the stable Stage II it optimizes the preference reward $r(x_0,z)$ plus a semantic-modulated exploration term; during the convergent Stage III it drives the predicted clean image toward the pretrained model's output to avoid overfitting. Stage boundaries are detected online from the second derivative of SNR and of the CLIP-based semantic-change signal $\Delta E(t)$. The empirical claim is that this stagewise schedule improves reward learning, preserves diversity, and mitigates reward hacking across SDE-based and flow-matching backbones, with 26.7% average quality gains and 36.7% higher convergence speed.
Load-bearing premise
The load-bearing premise is that the dense per-stage rewards genuinely decompose the final preference score, so optimizing them is the same as maximizing the true preference; the paper does not prove that this reward replacement preserves the optimal policy, and if it does not, the reported gains may come from a different objective rather than from stage alignment.
Editorial extensions
If this is right
- Existing RL fine-tuning methods that propagate a single final reward to all denoising steps can be replaced by stage-aware reward schedules without changing the underlying policy-gradient algorithm.
- Reward hacking, measured by quality and diversity degradation at matched reward scores, should be reduced because no single reward is optimized uniformly across all steps.
- The same stage-switching mechanism should transfer to flow-matching diffusion backbones such as SD3.5 and FLUX, since the rewards and stage boundaries are defined on latents and CLIP semantic change rather than architecture-specific components.
- Convergence speed improves, so fewer queries are needed to reach a target reward score, lowering compute cost in wall-clock terms.
- Diversity is preserved during preference optimization because the Stage II exploration term is modulated by semantic improvement, discouraging mode collapse.
Reading between the lines
- The paper never proves that its dense stage rewards constitute reward shaping of the sparse final reward; a potential-based shaping correction could isolate whether the gains come from stage alignment or from an altered objective.
- Stage boundaries from second derivatives of SNR and semantic change could be predicted analytically for a given sampler schedule rather than measured online, making the method cheaper and more portable.
- The semantic scorer used for stage detection, CLIP, is itself a learned model that could drift or be gamed during fine-tuning; the paper does not analyze this second reward-model-like object.
- A control that shuffles stage assignments across trajectories while keeping the three reward formulas fixed would clarify whether the specific chaotic-stable-convergent ordering is the causal driver of the reported gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Stage-Guided Per-Step Optimization (SGPO), an RL fine-tuning method for text-to-image diffusion models that divides the denoising trajectory into chaotic, stable, and convergent stages and assigns a different reward to each stage (Eqs. 6–8). Stage boundaries are detected from second derivatives of the SNR and of a CLIP-based semantic-change signal, and the composite reward R_total in Eq. (9) is used for policy-gradient updates. The authors report 26.7% average gains in generative quality, 36.7% higher convergence speed, and improved diversity and detail fidelity over DDPO, DPOK, D3PO, TDPO, DenseReward, and B2DiffusionRL, with additional experiments on flow-matching models and a multi-metric reward-hacking evaluation. The paper also states Theorem 3.1, which shows that posterior variance decreases during generation, as motivation for stage-dependent optimization priorities.
Significance. If the central claim holds, SGPO would be a useful contribution to diffusion-model alignment, providing a concrete way to avoid the temporal homogenization of rewards in RL fine-tuning and to mitigate reward hacking while preserving diversity. The experimental effort is substantial: 16 comparative experiments, multiple backbones spanning SDE-based and flow-matching models, four datasets, an ablation study, a user study, and a multi-metric reward-hacking evaluation. The paper also explicitly attempts to evaluate reward hacking with cross-metrics rather than only reporting the trained reward. However, the central method currently rests on an unproven identification between the dense composite reward R_total and the sparse preference reward r(x0,z), and the stage-detection thresholds and reward scales are underspecified. These issues make the main quantitative claims not yet fully reproducible or fully interpretable as preference-reward optimization.
major comments (4)
- [§3.4, Eq. (9)] The training reward R_total replaces the sparse preference reward r(x0,z) from Eq. (2) with stage-dependent auxiliary terms: λ_t||x_t-x_T||^2 in Stage I, r(x0,z) plus an exploration term in Stage II, and -1/2||x̂0(x_t)-x0_pre||^2 in Stage III. The paper does not prove that maximizing E[R_total] is equivalent to, or a faithful proxy for, maximizing E[r(x0,z)]. Policy-gradient updates therefore optimize a different objective, and the reported 26.7% reward gain is, strictly speaking, a gain on R_total rather than on the preference reward. Please provide a reward-shaping argument (for example, potential-based shaping or a zero-mean decomposition of the auxiliary terms) or explicitly reframe the contribution as optimizing a stage-wise composite objective and evaluate that objective directly. Theorem 3.1 is not sufficient for this purpose, because monotone decreasing posterior variance does not establish that uncertainty tracks action–reward attribution.
- [§3.3, Eq. (5)] The stage boundaries t_c^(1) and t_c^(2) are defined only by the informal conditions d^2γ/dt^2 → 0 and d^2ΔE/dt^2 → 0. No threshold, detection algorithm, or sensitivity analysis is given, even though the composite reward in Eq. (9) switches at these trajectory-dependent boundaries. The method is therefore not fully specified and the main experiments cannot be reproduced from the paper as written. Please describe the exact detection procedure, report the threshold values used in the experiments, and include an ablation over reasonable threshold choices.
- [§3.4, Eq. (8)] The Stage III reward anchors to x0_pre, described as the output of the pretrained diffusion model under the same prompt. Because the policy being fine-tuned is initialized from that same pretrained model, this term penalizes deviation from a self-referential reference and may cap the achievable preference improvement or, conversely, stabilize the model depending on the anchor's definition. It is also unclear whether x0_pre is a single fixed sample per prompt or whether it is recomputed during training. Please clarify the anchor's definition and justify why constraining late-stage latents toward the pretrained output is compatible with the paper's stated goal of amplifying local details without overfitting.
- [§3.4, Eqs. (6)–(7)] The scale of the Stage I reward and the Stage II exploration term is not specified: λ_t is only described as 'obtained by scaling SNR(t) with a constant factor', and the exploration term ΔE(t)||x̂0(x_t)-x̂0(x_{t-1})|| has no reported coefficient. Since these terms are added to or substituted for the preference reward, their relative magnitudes determine the effective objective, and the reported improvements may be sensitive to these undocumented hyperparameters. Please report all scaling constants and include sensitivity ablations over them.
minor comments (5)
- [Abstract and Section 4] The abstract and Section 4 claim '36.7% higher convergence speed', but no table or figure in the manuscript directly reports a convergence-speed metric; please define the metric and give the underlying numbers.
- [§3.3, Eq. (4)] The notation t is used both for the forward noising step and for generation progress, and ΔE(t) is defined as E(t)-E(t-1) without an explicit convention for which direction corresponds to the denoising process; please clarify the indexing.
- [Table 2] The 'Fixed' ablation row is not defined in the text; please specify what 'fixed exploration reward at every stage' means and whether it replaces all three stage rewards.
- [Figure 15] Figure 15 in the Appendix compares backpropagating the final reward to the last step, the full trajectory, and the stable middle stage, but the main text does not describe how this comparison was performed or what the quantitative conclusion is; please add a description and results.
- [Section 4.2] The statement that TDPO 'leads to reward hacking manifested as text–image misalignment' is supported only by an Appendix figure; please provide the corresponding quantitative evidence in the text.
Circularity Check
No significant circularity: SGPO's stage-specific rewards are hand-designed objectives evaluated on external benchmarks, not fitted re-derivations of the final preference reward.
full rationale
The paper's central claim is that assigning different reward signals to the chaotic, stable, and convergent denoising stages improves alignment, diversity, and convergence speed. This claim is not obtained by fitting the final reward and re-deriving it. The preference reward r(x0,z) is defined independently in Eq. (2) and is explicitly kept as the Stage II reward component in Eq. (7); the auxiliary exploration and convergence terms in Eqs. (6) and (8) are hand-designed training signals, not parameters fitted to the evaluation metrics. Stage boundaries are detected from external signals (SNR and semantic change, Eqs. (3)-(5)), and the headline evaluations use external benchmarks (AES, PickScore, FID, LPIPS, etc.) against external baselines rather than against the SGPO objective itself. The Stage III anchor to the frozen pretrained model's own output, x0_pre in Eq. (8), is self-referential as a regularizer (it pulls late-stage latents back toward the backbone), but it is a deliberate design choice and not a claim that this anchor independently predicts quality; the main preference-optimization result rests on Stage II and on external metrics. The monotone posterior variance result (Theorem 3.1) is a standard property of the forward diffusion posterior, proven in the appendix, and is used heuristically to motivate stage priorities rather than to define the rewards. The paper does cite prior work by the same authors (e.g., [42] and [43]), but these citations are contextual and not load-bearing: removing them would not change any equation or evaluation. The unproven assumption that R_total is a faithful proxy for the final preference reward is a correctness risk (a reward-shaping gap), but a missing proof is not circularity. Hence no circular step can be exhibited by quoting an equation that reduces to its own input.
Assumptions & free parameters
free parameters (3)
- lambda_t reward scale for Stage I =
not reported
- Stage transition thresholds t_c^(1) and t_c^(2) =
not reported
- Stage II exploration weight =
implicitly 1
assumptions (4)
- domain assumption Gaussian prior p(x0)=N(0, sigma_0^2 I) for the posterior-variance theorem
- ad hoc to paper Stage boundaries are detectable from second derivatives of SNR and CLIP semantic change, and these boundaries align with the claimed chaotic, stable, and convergent regimes
- domain assumption Dense per-step rewards in Eq. (9) can replace the true sparse reward in Eq. (2) without biasing the RL objective
- ad hoc to paper CLIP semantic score is a valid signal for detecting generation stages
Cite this review
Pith. "Pith review of Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models." pith.science (2026). https://pith.science/paper/CPL3YYIW
@misc{pith2026260806768,
author = {Pith},
title = {Pith review of: Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/CPL3YYIW}},
note = {Machine review of arXiv:2608.06768}
}
read the original abstract
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences. Reinforcement learning (RL) for preference alignment in diffusion models is promising but limited by reward sparsity. Since a single reward cannot support optimization, existing RL methods usually backpropagate the final reward to all previous steps. However, denoising is stage-wise, with distinct semantics and controllability. Repeating the final reward across all steps creates a temporal objective mismatch, encouraging reward shortcuts that lead to reward hacking. At the same time, due to reward backfilling, each time step receives the same reward, making it impossible to distinguish between actions, thereby weakening the optimization process. To resolve this issue, we propose Stage-Guided Per-Step Optimization (SGPO) for diffusion models, which jointly leverages signal-to-noise ratio and semantic changes to identify generation stages and adaptively assign stage-specific objectives. Early denoising is chaotic and far from the final reward, resulting in weak reward-behavior correlation. This stage should prioritize exiting the chaotic state. In the mid stage, the latent transitions to a stable structure, where the final reward better corresponds to generative behavior. Therefore, this stage optimizes the final reward while exploring diversity to avoid early convergence to a single mode. In the late stage, the latent's core structure is largely fixed, and preference optimization mainly amplifies local details, risking overfitting. Therefore, stable convergence is preferred to avoid quality degradation. Results from 16 comparative experiments validate SGPO. Our method achieves 26.7% average gains in generative quality and 36.7% higher convergence speed.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Surjeet Balhara, Nishu Gupta, Ahmed Alkhayyat, Isha Bharti, Rami Q Malik, Sarmad Nozad Mahmood, and Firas Abedi. 2025. A survey on deep reinforcement learning architectures, applications and emerging trends.IET Communications 19, 1 (2025), e12447
work page 2025
-
[3]
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine
-
[4]
Alex James Chan, Hao Sun, Samuel Holt, and Mihaela van der Schaar. [n. d.]. Dense Reward for Free in Reinforcement Learning from Human Feedback. In International Conference on Machine Learning
-
[5]
Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, and Bruno Castro da Silva. 2025. Rlhf deciphered: A critical analysis of reinforcement learning from human feedback for llms.Comput. Surveys58, 2 (2025), 1–37
work page 2025
-
[6]
Jooyoung Choi, Jungbeom Lee, Chaehun Shin, Sungwon Kim, Hyunwoo Kim, and Sungroh Yoon. 2022. Perception prioritized training of diffusion models. In Computer Vision and Pattern Recognition. 11472–11481
work page 2022
-
[7]
Kevin Clark, Paul Vicol, Kevin Swersky, and David J Fleet. 2023. Directly fine- tuning diffusion models on differentiable rewards.arXiv preprint arXiv:2309.17400 (2023)
arXiv 2023
-
[8]
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D’Amour, DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachan- dran, et al. 2023. Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking.arXiv preprint arXiv:2312.09244(2023)
arXiv 2023
Show all 58 references
-
[9]
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, and Kimin Lee. 2024. Reinforcement learning for fine-tuning text-to-image diffusion models.Advances in Neural Information Processing Systems36 (2024)
2024
-
[10]
Yaozhong Gan, Renye Yan, Xiaoyang Tan, Zhe Wu, and Junliang Xing. 2024. Trans- ductive off-policy proximal policy optimization.arXiv preprint arXiv:2406.03894 (2024)
2024 arXiv
-
[11]
Yaozhong Gan, Renye Yan, Zhe Wu, and Junliang Xing. 2024. Reflective policy optimization.arXiv preprint arXiv:2406.03678(2024)
2024 arXiv
-
[12]
Rohit Gandikota, Zongze Wu, Richard Zhang, David Bau, Eli Shechtman, and Nick Kolkin. 2025. Sliderspace: Decomposing the visual capabilities of diffusion models. InProceedings of the IEEE/CVF International Conference on Computer Vision. 15994–16003
2025
-
[13]
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. 2023. Geneval: An object-focused framework for evaluating text-to-image alignment.NeurIPS36 (2023), 52132–52152
2023
-
[14]
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan. 2017. Inverse reward design.Advances in neural information processing systems30 (2017)
2017
-
[15]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in Neural Information Processing Systems30 (2017)
2017
-
[16]
Zijing Hu, Fengda Zhang, Long Chen, Kun Kuang, Jiahui Li, Kaifeng Gao, Jun Xiao, Xin Wang, and Wenwu Zhu. 2025. Towards better alignment: Training diffusion models with reinforcement learning against sparse rewards. InProceedings of the Computer Vision and Pattern Recognition ...
2025
-
[17]
Francisco Ibarrola and Kazjon Grace. 2024. Measuring diversity in co-creative image generation.arXiv preprint arXiv:2403.13826(2024)
2024 arXiv
-
[18]
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore. 1996. Rein- forcement learning: A survey.Journal of artificial intelligence research4 (1996), 237–285
1996
-
[19]
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023. Pick-a-pic: An open dataset of user preferences for text-to- image generation.Advances in Neural Information Processing Systems(2023)
2023
-
[20]
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. 2019. Improved precision and recall metric for assessing generative models. Advances in Neural Information Processing Systems32 (2019)
2019
-
[21]
Cassidy Laidlaw, Shivam Singhal, and Anca Dragan. 2024. Correlated proxies: A new definition and improved mitigation for reward hacking.arXiv preprint arXiv:2403.03185(2024)
2024 arXiv
-
[22]
Preeti Lamba, Kiran Ravish, Ankita Kushwaha, and Pawan Kumar. 2025. Align- ment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey.arXiv preprint arXiv:2505.17352(2025)
2025 arXiv
-
[23]
Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu, Xuefeng Xiao, Rui Wang, Min Zheng, Xin Pan, Fei Chao, and Rongrong Ji. 2023. Autodiffusion: Training-free optimiza- tion of time steps and architectures for automated diffusion model acceleration. InInternational Conference on Comput...
2023
-
[24]
Lixiong Liu, Bao Liu, Hua Huang, and Alan Conrad Bovik. 2014. No-reference im- age quality assessment based on spatial and spectral entropies.Signal Processing: Image communication29, 8 (2014), 856–863
2014
-
[25]
Panagiotis Michailidis, Iakovos Michailidis, and Elias Kosmatopoulos. 2025. Re- inforcement learning for optimizing renewable energy utilization in buildings: A review on applications and innovations.Energies18, 7 (2025), 1724
2025
-
[26]
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. 2012. No- reference image quality assessment in the spatial domain.IEEE Transactions on Image Processing21, 12 (2012), 4695–4708
2012
-
[27]
completely blind
Anish Mittal, Rajiv Soundararajan, and Alan C Bovik. 2012. Making a “completely blind” image quality analyzer.IEEE Signal Processing Letters20, 3 (2012), 209–212
2012
-
[28]
Alexander Pan, Erik Jones, Meena Jagadeesan, and Jacob Steinhardt. 2024. Feed- back loops with language models drive in-context reward hacking.arXiv preprint arXiv:2402.06627(2024)
2024 arXiv
-
[29]
Nikolaos Pippas, Elliot A Ludvig, and Cagatay Turkay. 2025. The Evolution of Reinforcement Learning in Quantitative Finance: A Survey.Comput. Surveys57, 11 (2025), 1–51
2025
-
[30]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. [n. d.]. Learning transferable visual models from natural language supervision. InInternational conference on machine le...
-
[31]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...
2021
-
[32]
Rafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sushil Sikchi, Joey Hejna, Brad Knox, Chelsea Finn, and Scott Niekum. 2024. Scaling laws for reward model overoptimization in direct alignment algorithms.Advances in Neural Information Processing Systems37 (2024), 126207–126242
2024
-
[33]
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training gans.Advances in neural information processing systems29 (2016)
2016
-
[34]
Aziz, Rosli Besar, and Anith Khairunnisa Ghazali
Masuda Begum Sampa, Nor Hidayati Abdul Aziz, Md Siddikur Rahman, Nor Azlina Ab. Aziz, Rosli Besar, and Anith Khairunnisa Ghazali. 2026. Rein- forcement learning for medical image analysis: a systematic review of algorithms, engineering challenges, and clinical deployment.Compu...
2026
-
[35]
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in...
2022
-
[36]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
-
[37]
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. Defining and characterizing reward gaming.Advances in Neural Information Processing Systems35 (2022), 9460–9471
2022
-
[38]
Chen Tang, Ben Abbatematteo, Jiaheng Hu, Rohan Chandra, Roberto Martín- Martín, and Peter Stone. 2025. Deep reinforcement learning for robotics: A survey of real-world successes.Annual Review of Control, Robotics, and Autonomous Systems8, 1 (2025), 153–188
2025
-
[39]
Andrew Wagenmaker, Mitsuhiko Nakamoto, Yunchu Zhang, Seohong Park, Waleed Yagoub, Anusha Nagabandi, Abhishek Gupta, and Sergey Levine. 2025. Steering Your Diffusion Policy with Latent Space Reinforcement Learning.arXiv preprint arXiv:2506.15799(2025)
2025 arXiv
-
[40]
Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023. Human preference score v2: A solid benchmark for evaluat- ing human preferences of text-to-image synthesis.arXiv preprint arXiv:2306.09341 (2023)
2023 arXiv
-
[41]
Xin Xie and Dong Gong. 2025. DyMO: Training-Free Diffusion Model Align- ment with Dynamic Multi-Objective Scheduling. InComputer Vision and Pattern Recognition Conference. 13220–13230
2025
-
[42]
RenYe Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, You Wu, Yunfan Yang, Liang Ling, Jinlong Lin, Yeshuang Zhu, Jie Zhou, et al. 2025. Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment. InProceedings of the IEEE/CVF International Conference on Computer ...
2025
-
[43]
Renye Yan, Jikang Cheng, Shikun Sun, Yi Sun, You Wu, Wei Peng, Zongwei Wang, Ling Liang, Junliang Xing, and Yimao Cai. 2026. Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?arXiv preprint arXiv:2605.15855(2026)
2026 arXiv
-
[44]
Renye Yan, Jikang Cheng, You Wu, Ling Liang, Wei Peng, Athanasios V Vasilakos, Qingyu Zhao, Yu Zhang, Ehsan Adeli, Kilian M Pohl, et al . 2026. Pixel-Space Diffusion Transformers.arXiv preprint arXiv:2607.17585(2026)
2026 arXiv
-
[45]
Renye Yan, Yaozhong Gan, You Wu, Ling Liang, Junliang Xing, Yimao Cai, and Ru Huang. 2024. The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective.arXiv preprint arXiv:2408.09974(2024). Conference’17, July 2017, Washington, DC, USA Yan et al
2024 arXiv
-
[46]
Kai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge, Jiaxin Chen, Weihan Shen, Xiaolong Zhu, and Xiu Li. 2024. Using human feedback to fine-tune diffusion models without any reward model. InComputer Vision and Pattern Recognition. 8941– 8951
2024
-
[47]
Mingyang Yi, Aoxue Li, Yi Xin, and Zhenguo Li. 2024. Towards understand- ing the working mechanism of text-to-image diffusion model.arXiv preprint arXiv:2405.15330(2024)
2024 arXiv
-
[48]
Kevin Zhai, Utsav Singh, Anirudh Thatipelli, Souradip Chakraborty, Anit Kumar Sahu, Furong Huang, Amrit Singh Bedi, and Mubarak Shah. 2025. Mira: Towards mitigating reward hacking in inference-time alignment of t2i diffusion models. arXiv preprint arXiv:2510.01549(2025)
2025
-
[49]
Kaiyan Zhang, Yuxin Zuo, Bingxiang He, Youbang Sun, Runze Liu, Che Jiang, Yuchen Fan, Kai Tian, Guoli Jia, Pengfei Li, et al. 2025. A survey of reinforcement learning for large reasoning models.arXiv preprint arXiv:2509.08827(2025)
2025 arXiv
-
[50]
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang
-
[51]
Xueyi Zhang, Peiyin Zhu, Yuan Liao, Xiyu Wang, Mingrui Lao, Siqi Cai, Yanming Guo, and Haizhou Li. 2025. Trustclip: Learning from noisy labels via semantic label verification and trust-aligned gradient projection. InProceedings of the 33rd ACM International Conference on Multi...
2025
-
[52]
Xueyi Zhang, Peiyin Zhu, Chengwei Zhang, Zhiyuan Yan, Jikang Cheng, Mingrui Lao, Siqi Cai, and Yanming Guo. 2025. Generalization-preserved learning: Closing the backdoor to catastrophic forgetting in continual deepfake detection. In2025 IEEE/CVF International Conference on Com...
2025
-
[53]
Ziyi Zhang, Sen Zhang, Yibing Zhan, Yong Luo, Yonggang Wen, and Dacheng Tao
-
[54]
Enmin Zhao, Renye Yan, Jinqiu Li, Kai Li, and Junliang Xing. 2022. Alphaholdem: High-performance artificial intelligence for heads-up no-limit poker via end-to- end reinforcement learning. InAAAI Conference on Artificial Intelligence, Vol. 36. 4689–4697. Explore or Converge? S...
2022
-
[2017]
Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347 (2017)
2017 arXiv
-
[2018]
In Computer Vision and Pattern Recognition
The unreasonable effectiveness of deep features as a perceptual metric. In Computer Vision and Pattern Recognition. 586–595
-
[2023]
Training diffusion models with reinforcement learning.arXiv preprint arXiv:2305.13301(2023)
2023 arXiv
-
[2024]
InInternational Conference on Machine Learning
Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases. InInternational Conference on Machine Learning. PMLR, 60396–60413
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.