REVIEW 4 major objections 5 minor 64 references
Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read By estimating the terminal denoised distribution from any intermediate step, DDE turns terminal-only preference labels into per-step training signal and concentrates optimization on the middle of the denoising trajectory.
desk verdict A usable heuristic DPO variant with a broken derivation; the paper should be revised to drop the 'naturally derives' claim unless the math is fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is Denoised Distribution Estimation itself: the two-segment estimate of the terminal distribution $p_\theta(x_0)$ from an intermediate step $t$. Its workhorse identity is $p_\theta(x_0) \approx \exp\{\sum_{k=t}^{T-1} r_k\} \, \mathbb{E}_{x_t \sim q(x_t|x_0)}[p_\theta(x_0|x_t)]$, which lets the terminal preference label be differentiated through every step of the trajectory. The stepwise estimate supplies the correction terms $r_k$ (EMA-calibrated log-ratios between model and conditional denoising distributions), and the single-shot DDIM estimate supplies the one-pass projection $p_\theta(\hat{x}_0|x_t)$; together they determine how much gradient credit each step receives, and both components are needed for the reported behavior.
What would settle it
Train an off-the-shelf diffusion DPO loss while masking out the first and last 15% of denoising steps in the objective, using the same Pick-a-Pic-V2 data and SD15/SDXL backbones. If this masking reproduces DDE's reported CLIP, HPS, and PS gains, the middle-step credit assignment is the operative mechanism; if the gains vanish, the specific estimation terms in Eq. 10 are doing essential work beyond a rough weighting scheme.
Extended reading notes
Core claim
The paper's central claim is that DDE solves the terminal-only preference problem by replacing the intractable marginal $p_\theta(x_0)$ with an explicit estimate built from two segments around a sampled training step $t$. For the segment $T \to t$, each model denoising distribution $p_\theta(x_k|x_{k+1})$ is replaced by $\exp\{r_k\}q(x_k|x_{k+1},x_0)$, with $q$ the known forward conditional and $r_k$ a non-gradient calibration coefficient maintained by exponential moving average; this collapses the segment into the single factor $q(x_t|x_0)$ times a correction term. For the segment $t \to 0$, a single DDIM pass maps the intermediate latent to a predicted $\hat{x}_0$. Substituting this estimate into the DPO log-ratio yields a loss in which the calibration terms and DDIM coefficients enter inside the logistic function, pushing the early and late steps toward gradient saturation and thereby assigning more optimization credit to middle denoising steps. The authors argue, and support with ablations, that this naturally derived middle-step credit assignment is what makes the method outperform hand-crafted uniform or discounted schemes without any auxiliary model.
Load-bearing premise
The load-bearing premise is that the running calibration coefficients keep the estimated per-step distributions close to the model's own denoising distributions throughout training, even though the coefficients are computed from the very model being updated and then treated as constants.
Editorial extensions
If this is right
- Terminal preference labels become usable at every denoising step without training an auxiliary reward model; the optimization signal is derived directly from estimating $p_\theta(x_0)$.
- The gradient signal concentrates on the middle of the denoising trajectory, so early near-noise steps and late near-clean steps are not over-weighted by a label that only observes the final image.
- On SD15, DDE reports consistent improvements of roughly 3.3% to 6.7% over the base model on CLIP, HPS, and PS scores; on SDXL the reported gains are 1.0% to 3.1%.
- Each estimation strategy is necessary: ablations that drop the calibration coefficients or perform optimization directly on the intermediate noisy sample both degrade performance.
Reading between the lines
- A direct testable consequence the paper does not claim: if the middle-step emphasis is the active ingredient, then masking out the first and last 15% of denoising steps in an off-the-shelf diffusion DPO loss should reproduce most of DDE's reported gain; the paper's own Fig. 5(c) points in that direction.
- The converged calibration coefficients could be read as a diagnostic of where the pretrained model's denoising distribution drifts from the ideal forward conditional, which could inform targeted preference-data collection or per-step learning-rate schedules.
- The same two-segment estimation recipe should transfer to other generative models with a known forward noising process and a deterministic one-step inverse map, such as flow-matching or consistency models; if transfer fails, it would localize the mechanism to DDPM/DDIM-specific structure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Denoised Distribution Estimation (DDE), a DPO-style preference alignment method for text-to-image diffusion models. DDE splits the denoising trajectory at a sampled step t into a T→t segment estimated by stepwise Gaussian transitions with scalar calibration coefficients and a t→0 segment estimated by a single DDIM step. The resulting loss is a DPO-like objective with correction terms that the authors claim naturally derive a credit assignment scheme emphasizing intermediate denoising steps. Experiments on SD15 and SDXL with CLIP, HPS, and PS annotators show consistent improvements over uniform and discounted baselines without auxiliary reward models.
Significance. If the derivation were sound, DDE would be a valuable contribution: it addresses the terminal-only preference-label problem without auxiliary models, provides a simple training objective, and its empirical results on two base models are encouraging. The paper includes useful ablations (DDE-Single, DDE-Step), an analysis of step prioritization, and qualitative comparisons. However, the central mathematical derivation has load-bearing gaps, and the claim that the credit assignment scheme is 'naturally derived' is not supported by the current equations. The empirical results may still stand as a heuristic, but the paper's framing and theoretical claims require substantial revision.
major comments (4)
- [Section 3.1, Eq. (7) and Suppl. Eq. (11)] The factorization of exp(Σ r_k) out of the path integral is valid only if r_k is independent of the integration variable x_k. But r_k is defined as log[p_θ(x_k|x_{k+1})/q(x_k|x_{k+1},x0)], which depends on x_k. For the Gaussian DDPM parameterization in this paper, the two densities have the same covariance, so the log-ratio is affine in x_k, not constant. A single scalar coefficient cannot make exp{r_k}q(x_k|x_{k+1},x0) equal to p_θ(x_k|x_{k+1}) for all x_k. Therefore Eq. (7) is not a valid estimate of the DPO objective; the correction term is a heuristic weighting rather than a derived quantity. This undermines the central claim that DDE naturally derives a credit assignment scheme from terminal distribution estimation.
- [Section 3.2 / Suppl. Eq. (15)] The replacement of log(E[p_θ(x0|xt)]/E[p_ref(x0|xt)]) by -||x0-μ_θ||^2 + ||x0-μ_ref||^2 drops the variance and normalization terms of the Gaussian densities and treats a ratio of expectations as if it were the ratio of single-point evaluations. Even when the same xt is used for both models, the equality does not hold in expectation; it is at best a biased approximation. The paper does not analyze this bias or its effect on preference optimization. Since this equation is the basis of the actual loss in Eq. (10), the claimed equivalence to DPO is not established.
- [Section 3.1 / Algorithm 1] The calibration coefficients r_k are updated by EMA using log-ratios evaluated on the current target model and then treated as non-gradient constants. This is a moving-target objective: the loss depends on statistics of the same model being trained, and the paper provides only empirical convergence of the coefficients (Fig. 5(b)), not a proof or analysis of training stability. If the EMA values drift or the approximation degrades at certain steps, the derived credit assignment could misweight the preference loss. The assertion that the EMA 'does not adversely affect the training process' is not established by the current evidence.
- [Eq. (10) vs Suppl. Eq. (16) and Algorithm 1] There is an index inconsistency: Eq. (10) in the main text sums k=t to T-1, while Suppl. Eq. (16) sums k=t to T; the coefficient array has length T with indices 0..T-1, so the upper limit T in the supplement is out of range. Algorithm 1 updates r[t-1] on line 9 while the loss uses r[t..T-1]; the relationship between the updated index and the summed range is not explained. These inconsistencies reinforce that the derivation is not pinned down.
minor comments (5)
- [Abstract / Introduction] There are minor typos such as 'etimating' in Section 1 and inconsistent use of 'DDE' vs. 'our DDE' in the same paragraph.
- [Table 2] In the DDE-Step row, the HPS value is reported as 2.600±0.671, but the standard deviation in other rows is around 0.2; please verify that this is not a typo.
- [Figure 5] In Fig. 5(c), the caption says the red dashed line denotes the original SD15 score of 0.320, but Table 2 reports CLIP 3.200; please clarify whether the figure uses a normalized scale or a different metric.
- [Section 4.5] The sentence 'The red dashed line denotes the performance of the original SD15 with a score of 0.320' is confusing because Table 2 lists the SD15 CLIP score as 3.200; please reconcile the numbers.
- [References] Reference [39] has 'abs/2404.3715' which appears to be a malformed arXiv identifier; please correct it.
Circularity Check
The calibration-coefficient correction term in Eq. 10 is populated by EMA fits of the trained model's own log-density ratio, so the 'naturally derived' credit assignment scheme partially reduces to a self-referential fit.
-
self definitional
[Sec. 3.1, 'Calculation of calibration coefficients rk'; Eq. 7; Eq. 10]
"Since we want exp{rk}q(xk|xk+1, x0) ≈ pθ(xk|xk+1), the rk should equals to log pθ(xk|xk+1) q(xk|xk+1,x0). ... we maintain an array of length T for recording rk and employ exponential moving average (EMA) to update this value throughout the training process."
The stepwise 'estimate' is defined to be exact at the sampled point: rk is set to the log-ratio of the very model being trained to q(xk|xk+1,x0), so exp{rk}q equals pθ(xk|xk+1) by construction there. These same EMA-fitted values are inserted into Eq. 10 as Sigma(r_w_theta,k - r_w_ref,k - r_l_theta,k + r_l_ref,k), the correction terms that produce the claimed middle-step credit assignment. The correction weighting is therefore not derived from terminal preference labels or from an independent distribution estimate; it is a moving average of the target model's own current log-density ratios, i.e., a fitted input renamed as a 'naturally derived' calibration.
full rationale
No problematic self-citation chain exists: the authors do not lean on their own prior uniqueness theorems or ansatz-by-citation. The empirical performance claims (Tables 2-4, qualitative comparisons) are benchmarked against external annotators and baselines, so the main quantitative contribution is not circular. The partial circularity is in the derivation narrative. The correction coefficients rk are explicitly defined as log[pθ(xk|xk+1)/q(xk|xk+1,x0)] for the model under training, and then used in the loss as constants that determine the credit-assignment weighting. Thus the 'DDE naturally derives a novel automatic credit assignment scheme' claim reduces, for the correction-term component, to fitting the model's own current density ratio via EMA. The MSE terms anchored to the reference model are independent and keep the method from being wholly tautological, and the middle-step emphasis is partially corroborated by the analytic DDIM coefficients and by the Fig. 5(c) experiment. However, the specific calibration-factor part of the derivation is self-referential by construction, meriting a 6 rather than a lower score. The additional mathematical concern that Eq. 7 pulls a scalar rk out of an integral when the density ratio is x_k-dependent is a correctness issue, not a circularity issue, and is not scored here.
Assumptions & free parameters
free parameters (3)
- DPO temperature beta =
5000
- EMA decay mu =
0.1
- Calibration coefficient arrays r_k =
Per-step EMA values, one array per sample type (w/l) and model (theta/ref)
assumptions (4)
- standard math The forward process posterior q(x_{t-1}|x_t,x_0) is Gaussian and computable in closed form (Eq. 3).
- ad hoc to paper For all k in the segment T to t, p_theta(x_k|x_{k+1}) is well approximated by exp{r_k} q(x_k|x_{k+1},x_0) with step-specific constants r_k.
- domain assumption A single DDIM step from x_t to x_0 (Eq. 9) provides a sufficiently accurate estimate of the terminal distribution for preference optimization, because relative differences between pairs matter more than absolute accuracy.
- ad hoc to paper Treating the EMA calibration coefficients r_k as non-gradient constants while updating them from the same trained model yields stable convergence.
invented entities (1)
-
Calibration coefficient arrays r_k (for winning/losing and target/reference)
Cite this review
Pith. "Pith review of Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation." pith.science (2026). https://pith.science/paper/J7TSNTAG
@misc{pith2026241114871,
author = {Pith},
title = {Pith review of: Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/J7TSNTAG}},
note = {Machine review of arXiv:2411.14871}
}
read the original abstract
Diffusion models have shown remarkable success in text-to-image generation, making preference alignment for these models increasingly important. The preference labels are typically available only at the terminal of denoising trajectories, which poses challenges in optimizing the intermediate denoising steps. In this paper, we propose to conduct Denoised Distribution Estimation (DDE) that explicitly connects intermediate steps to the terminal denoised distribution. Therefore, preference labels can be used for the entire trajectory optimization. To this end, we design two estimation strategies for our DDE. The first is stepwise estimation, which utilizes the conditional denoised distribution to estimate the model denoised distribution. The second is single-shot estimation, which converts the model output into the terminal denoised distribution via DDIM modeling. Analytically and empirically, we reveal that DDE equipped with two estimation strategies naturally derives a novel credit assignment scheme that prioritizes optimizing the middle part of the denoising trajectory. Extensive experiments demonstrate that our approach achieves superior performance, both quantitatively and qualitatively.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Direct preference optimization with an offset
Afra Amini, Tim Vieira, and Ryan Cotterell. Direct preference optimization with an offset. In Findings of the Association for Computational Linguistics, 2024. 1
work page 2024
-
[2]
A general theoretical paradigm to understand learning from human prefer- ences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rémi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello. A general theoretical paradigm to understand learning from human prefer- ences. In International Conference on Artificial Intelli- gence and Statistics, 2024. 1
work page 2024
-
[3]
Anirudhan Badrinath, Prabhat Agarwal, and Jiajing Xu. Hybrid preference optimization: Augmenting di- rect preference optimization with auxiliary objectives. arXiv preprint arXiv:2405.17956, 2024. 8
arXiv 2024
-
[4]
Training diffusion models with reinforcement learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine. Training diffusion models with reinforcement learning. In The Twelfth International Conference on Learning Representations, 2024. 1, 8
work page 2024
-
[5]
Align your latents: High-resolution video syn- thesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video syn- thesis with latent diffusion models. In IEEE/CVF Con- ference on Computer Vision and Pattern Recognition,
-
[6]
Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James T. Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to- image synthesis. In The Twelfth International Confer- ence on Learning Representations, 2024. 8
work page 2024
-
[7]
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. Self-play fine-tuning converts weak language models to strong language models. In Forty- first International Conference on Machine Learning ,
-
[8]
Paul F. Christiano, Jan Leike, Tom B. Brown, Mil- jan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems,
Show all 64 references
-
[9]
Diffu- sion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Quinn Nichol. Diffu- sion models beat gans on image synthesis. In Advances in Neural Information Processing Systems, 2021. 8
2021
-
[10]
Scaling rectified flow transformers for high- resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rom- bach. Scaling rectified flow transformers for high- resolution image synt...
2024
-
[11]
KTO: model align- ment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. KTO: model align- ment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306, 2024. 8
2024 arXiv
-
[12]
DPOK: reinforcement learning for fine-tuning text-to-image diffusion models
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mo- hammad Ghavamzadeh, Kangwook Lee, and Kimin Lee. DPOK: reinforcement learning for fine-tuning text-to-image diffusion models. Advances in Neural Information Processing Systems, 2023. 1, 8
2023
-
[13]
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Interna- tional Conference on Machine Learning, 2023. 8
2023
-
[14]
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Confer- ence on Machine Learning, 2018. 8
2018
-
[15]
Brandt, and Tomer Michaeli
René Haas, Inbar Huberman-Spiegelglas, Rotem Mu- layoff, Stella Graßhof, Sami S. Brandt, and Tomer Michaeli. Discovering interpretable directions in the semantic latent space of diffusion models. In 18th IEEE International Conference on Automatic Face and Gesture Recognition, 2024. 8
2024
-
[16]
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020. 1, 2, 4
2020
-
[17]
Estimation of non-normalized statis- tical models by score matching
Aapo Hyvärinen. Estimation of non-normalized statis- tical models by score matching. J. Mach. Learn. Res.,
-
[18]
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. Reward learning from human preferences and demonstrations in atari. In Advances in Neural Information Processing Systems,
-
[19]
Ryzhakov, Andrei Chertkov, and Ivan V
Valentin Khrulkov, Gleb V . Ryzhakov, Andrei Chertkov, and Ivan V . Oseledets. Understanding DDPM latent codes through optimal transport. In The Eleventh International Conference on Learning Repre- sentations, 2023. 8
2023
-
[20]
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbu- land Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. In Advances in Neural Information Pro- cessing Systems, 2023. 6
2023
-
[21]
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbu- land Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. 2023. 5
2023
-
[22]
Bradley Knox and Peter Stone
W. Bradley Knox and Peter Stone. Interactively shap- ing agents via human reinforcement: the TAMER framework. In Proceedings of the 5th International Conference on Knowledge Capture, 2009. 8
2009
-
[23]
Dif- fusion models already have A semantic latent space
Mingi Kwon, Jaeseok Jeong, and Youngjung Uh. Dif- fusion models already have A semantic latent space. 9 In The Eleventh International Conference on Learning Representations, 2023. 8
2023
-
[24]
Aligning diffusion models by optimizing human utility
Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, and Kazuki Kozuka. Aligning diffusion models by optimizing human utility. arXiv preprint arXiv:2404.04465, 2024. 8
2024 arXiv
-
[25]
Step-aware preference optimization: Aligning preference with denoising performance at each step
Zhanhao Liang, Yuhui Yuan, Shuyang Gu, Bohan Chen, Tiankai Hang, Ji Li, and Liang Zheng. Step-aware preference optimization: Aligning preference with denoising performance at each step. arXiv preprint arXiv:2406.04314, 2024. 1, 8
2024 arXiv
-
[26]
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Har- rison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. In The Twelfth International Con- ference on Learning Representations, 2024. 1
2024
-
[27]
Lillicrap, Jonathan J
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. In 4th International Con- ference on Learning Representations, 2016. 1
2016
-
[28]
Alignment of diffusion models: Fun- damentals, challenges, and future
Buhua Liu, Shitong Shao, Bao Li, Lichen Bai, Zhiqiang Xu, Haoyi Xiong, James Kwok, Sumi Helal, and Zeke Xie. Alignment of diffusion models: Fun- damentals, challenges, and future. arXiv preprint arXiv:2409.07253, 2024. 1
2024
-
[29]
Interpretation and generalization of score matching
Siwei Lyu. Interpretation and generalization of score matching. In Proceedings of the Twenty-Fifth Confer- ence on Uncertainty in Artificial Intelligence , 2009. 8
2009
-
[30]
Ho, Robert Tyler Loftin, Bei Peng, Guan Wang, David L
James MacGlashan, Mark K. Ho, Robert Tyler Loftin, Bei Peng, Guan Wang, David L. Roberts, Matthew E. Taylor, and Michael L. Littman. Interactive learning from policy-dependent human feedback. In Proceed- ings of the 34th International Conference on Machine Learning, 2017. 8
2017
-
[31]
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen. Simpo: Simple preference optimization with a reference-free reward. In Advances in Neural Information Processing Systems, 2024. 1
2024
-
[32]
Riedmiller
V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. Playing atari with deep rein- forcement learning. arXiv preprint arXiv:1312.5602,
-
[33]
Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu
V olodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In Proceed- ings of the 33nd International Conference on Machine Learning, 2016. 8
2016
-
[34]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...
2022
-
[35]
Sdxl: Improving latent diffu- sion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffu- sion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 6
2023 arXiv
-
[36]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...
2021
-
[37]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christo- pher D. Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, 2023. 1, 3, 8
2023
-
[38]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 1, 6
2022
-
[39]
Direct nash optimization: Teaching language models to self-improve with general preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie. Direct nash optimization: Teaching language models to self-improve with general preferences. arXiv preprint arXiv:2404.03715, abs/2404.3715, 2024. 8
2024 arXiv
-
[40]
Jordan, and Philipp Moritz
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz. Trust region policy optimization. In Proceedings of the 32nd In- ternational Conference on Machine Learning , 2015. 8
2015
-
[41]
Proximal policy optimiza- tion algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimiza- tion algorithms. arXiv preprint arXiv:1707.06347 ,
-
[42]
Denoising diffusion models on model-based latent space
Carmelo Scribano, Danilo Pezzi, Giorgia Franchini, and Marco Prato. Denoising diffusion models on model-based latent space. Algorithms, 2023. 2
2023
-
[43]
Riedmiller
David Silver, Guy Lever, Nicolas Heess, Thomas De- gris, Daan Wierstra, and Martin A. Riedmiller. Deter- ministic policy gradient algorithms. In Proceedings of the 31th International Conference on Machine Learn- ing, 2014. 8
2014
-
[44]
Weiss, Niru Mah- eswaranathan, and Surya Ganguli
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Mah- eswaranathan, and Surya Ganguli. Deep unsupervised 10 learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, 2015. 8
2015
-
[45]
Prefer- ence ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. Prefer- ence ranking optimization for human alignment. In Thirty-Eighth AAAI Conference on Artificial Intelli- gence, 2024. 8
2024
-
[46]
De- noising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. De- noising diffusion implicit models. In 9th International Conference on Learning Representations, 2021. 1, 2, 8
2021
-
[47]
Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In 9th International Conference on Learning Representations, 2021. 2
2021
-
[48]
Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F. Christiano. Learning to summarize from human feedback. Advances in Neural Information Processing Systems, 2020. 1, 8
2020
-
[49]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto.Reinforcement learning - an introduction. MIT Press, 1998. 8
1998
-
[50]
A connection between score matching and denoising autoencoders
Pascal Vincent. A connection between score matching and denoising autoencoders. Neural Comput., 2011. 8
2011
-
[51]
Diffusion model alignment using direct preference op- timization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Er- mon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model alignment using direct preference op- timization. In IEEE/CVF Conference on Computer Vision and Pattern Recogn...
2024
-
[52]
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. Aligning large language models with human: A survey. arXiv preprint arXiv:2307.12966, 2023. 1
2023 arXiv
-
[53]
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas. Dueling network architectures for deep reinforcement learning. In Proceedings of the 33nd International Conference on Machine Learning, 2016. 8
2016
-
[54]
Waytowich, Vernon Lawh- ern, and Peter Stone
Garrett Warnell, Nicholas R. Waytowich, Vernon Lawh- ern, and Peter Stone. Deep TAMER: interactive agent shaping in high-dimensional state spaces. In Proceed- ings of the Thirty-Second AAAI Conference on Artificial Intelligence, 2018. 8
2018
-
[55]
β-dpo: Direct preference optimization with dynamic β
Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jinyang Gao, Bolin Ding, Xiang Wang, and Xi- angnan He. β-dpo: Direct preference optimization with dynamic β. In Advances in Neural Information Processing Systems, 2024. 1
2024
-
[56]
Human pref- erence score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis
Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human pref- erence score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv preprint arXiv:2306.09341, 2023. 6
2023 arXiv
-
[57]
Using human feedback to fine-tune diffusion models without any reward model
Kai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge, Jiaxin Chen, Weihan Shen, Xiaolong Zhu, and Xiu Li. Using human feedback to fine-tune diffusion models without any reward model. In IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2024. 1, 6, 7, 8
2024
-
[58]
Diffusion models: A comprehen- sive survey of methods and applications
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. Diffusion models: A comprehen- sive survey of methods and applications. ACM Comput. Surv., 2024. 8
2024
-
[59]
A dense reward view on aligning text-to-image diffusion with preference
Shentao Yang, Tianqi Chen, and Mingyuan Zhou. A dense reward view on aligning text-to-image diffusion with preference. In Forty-first International Conference on Machine Learning, 2024. 1, 6, 7, 8
2024
-
[60]
RRHF: rank re- sponses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. RRHF: rank re- sponses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302, 2023. 1
2023 arXiv
-
[61]
Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Chris- tiano, and Geoffrey Irving. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019. 1, 8 11 Preference Alignment for Diffusion Model via E...
1909 arXiv
-
[62]
Derivation of the Loss Function Defined in Eq
Derivation 7.1. Derivation of the Loss Function Defined in Eq. 7 By substituting pθ(xk|xk+1) with erk q(xk|xk+1, x0) for all k ∈ {t, ..., T− 1}in Eq. 6, we obtain: pθ(x0) = Z x1:T q(xT )pθ(xT −1|xT )...pθ(x0|x1)dx1:T = exp{ T −1X k=t rk} Z x1:T q(xT ) tY k=T −1 q(xk|xk+1, x0) ...
-
[63]
follow the same procedure. Consequently, the total loss function is given by: LDDE = Exw 0 ,xl 0,xw t ∼q(xw t |xw 0 ),xl t∼q(xl t|xl 0)[− log σ(β( − ||xw 0 − ˆµθ,t′=0(xw t , t)||2 2 + ||xw 0 − ˆµref,t ′=0(xw t , t)||2 2 + ||xl 0 − ˆµθ,t′=0(xl t, t)||2 2 − ||xl 0 − ˆµref,t ′=0(...
-
[64]
Implementation Details We employ a constant learning rate with a warm-up sched- ule, finalizing at 2.05 × 10−5
Extended Experiments 8.1. Implementation Details We employ a constant learning rate with a warm-up sched- ule, finalizing at 2.05 × 10−5. The hyper-parameters β and µ are set to 5000 and 0.1 respectively. To optimize compu- tational efficiency, both gradient accumulation and g...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.