Pith. sign in

REVIEW 28 cited by

Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13734 v1 pith:RRY34ZER submitted 2024-07-18 cs.LG cs.AIq-bio.QMstat.ML

Understanding Reinforcement Learning-Based Fine-Tuning of Diffusion Models: A Tutorial and Review

classification cs.LG cs.AIq-bio.QMstat.ML
keywords diffusionfine-tuningmodelsrl-basedtutorialalgorithmsdistributionslearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This tutorial provides a comprehensive survey of methods for fine-tuning diffusion models to optimize downstream reward functions. While diffusion models are widely known to provide excellent generative modeling capability, practical applications in domains such as biology require generating samples that maximize some desired metric (e.g., translation efficiency in RNA, docking score in molecules, stability in protein). In these cases, the diffusion model can be optimized not only to generate realistic samples but also to explicitly maximize the measure of interest. Such methods are based on concepts from reinforcement learning (RL). We explain the application of various RL algorithms, including PPO, differentiable optimization, reward-weighted MLE, value-weighted sampling, and path consistency learning, tailored specifically for fine-tuning diffusion models. We aim to explore fundamental aspects such as the strengths and limitations of different RL-based fine-tuning algorithms across various scenarios, the benefits of RL-based fine-tuning compared to non-RL-based approaches, and the formal objectives of RL-based fine-tuning (target distributions). Additionally, we aim to examine their connections with related topics such as classifier guidance, Gflownets, flow-based diffusion models, path integral control theory, and sampling from unnormalized distributions such as MCMC. The code of this tutorial is available at https://github.com/masa-ue/RLfinetuning_Diffusion_Bioseq

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 8.0

    FMRG reformulates guidance as deterministic optimal control, deriving a single-trajectory method using the flow map that matches or exceeds baselines on reward-guided generation and inverse problems with 3 NFEs at tex...

  2. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  3. Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 7.0

    A Gaussian process surrogate gate inserted between generative crystal models and property oracles matches or exceeds ungated fine-tuning while using roughly one-fifth the oracle calls for heat capacity and bulk modulus.

  4. A Markov Chain Approach to Preference Alignment

    cs.LG 2026-06 unverdicted novelty 7.0

    MCHF defines a Markov kernel from pairwise utilities U and proves geometric convergence to its stationary distribution at a rate set by the seminorm measuring non-transitivity of U, with first-order equivalence to RLH...

  5. Active Flow Expansion for Out-of-Distribution Discovery: from Theory to Molecules

    cs.LG 2026-06 unverdicted novelty 7.0

    ActFlow expands the generable set of pre-trained flow models for out-of-distribution molecular and sequence design via active synthetic data generation and verifier feedback, with new statistical guarantees.

  6. Your GFlowNet Secretly Learns an Optimal Transport Plan

    cs.LG 2026-06 unverdicted novelty 7.0

    Minimum-flow GFlowNets on graphs encode optimal transport plans, with the learned policy recovering the optimal coupling between source and target distributions.

  7. Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

    cs.LG 2026-05 unverdicted novelty 7.0

    FAV aligns few-step generative models by amortizing SVGD updates from reward-tilted sampling into generator parameters via fixed-point regression, requiring only sample access, and shows outperformance on robotics tas...

  8. Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion

    cs.LG 2026-05 unverdicted novelty 7.0

    CDM amortizes SMC inference for reward-tilted discrete diffusion by training a parameterized twist function on contrastive samples with closed-form kernels.

  9. Improved techniques for fine-tuning flow models via adjoint matching: a deterministic control pipeline

    cs.AI 2026-05 unverdicted novelty 7.0

    A new adjoint matching framework formulates flow model alignment as optimal control, enabling direct regression training and terminal-trajectory truncation for efficiency gains on models like SiT-XL and FLUX.

  10. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 7.0

    FMRG is a training-free, single-trajectory guidance method for flow models derived from optimal control that achieves strong reward alignment with only 3 NFEs.

  11. Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

    cs.LG 2026-04 conditional novelty 7.0

    Many reward-based fine-tuning methods for diffusion and flow models are special cases of a common score-matching objective whose differences reduce to value-guidance estimator design, temporal weighting, and trust-reg...

  12. Step-level Denoising-time Diffusion Alignment with Multiple Objectives

    cs.LG 2026-04 unverdicted novelty 7.0

    MSDDA derives a closed-form optimal reverse denoising distribution for multi-objective diffusion alignment that is exactly equivalent to step-level RL fine-tuning with no approximation error.

  13. Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models

    cs.CV 2026-03 accept novelty 6.5

    Finite-difference flow optimization (FDFO) post-trains flow-matching image models by paired-trajectory updates that reduce variance and outperform Flow-GRPO on reward, quality, and alignment.

  14. Diffusion Fine-tuning with Rewarded Moment Matching Distillation

    cs.LG 2026-06 unverdicted novelty 6.0

    RMMD simultaneously distills diffusion models and optimizes rewards, yielding better FID-reward trade-offs on ImageNet than DI++, DRaFT and HyperNoise, and a 7.5x faster GenCast model that beats its teacher on 93% of ...

  15. BayesFP: Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling

    cs.RO 2026-06 unverdicted novelty 6.0

    BayesFP provides a unified retraining-free sampler for diffusion and flow policies by casting constrained trajectory generation as posterior sampling via an extended Feynman-Kac corrector.

  16. A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding

    cs.LG 2026-06 unverdicted novelty 6.0

    A2D2 derives the Radon-Nikodym derivative for joint insertion-unmasking paths in discrete diffusion to enable reward-tilted fine-tuning and introduces the Adaptive Joint Decoding loss.

  17. Plug-and-Play Guidance for Discrete Diffusion Models via Gradient-Informed Logit Correction

    cs.LG 2026-06 unverdicted novelty 6.0

    Introduces GILC, a training-free plug-and-play guidance framework for discrete diffusion models that uses Jacobian-free logit correction to achieve SOTA results on DNA, protein, and molecular generation tasks.

  18. Masked Diffusion Modeling for Anomaly Detection

    cs.LG 2026-05 unverdicted novelty 6.0

    MaskDiff-AD uses reconstruction difficulty of masked coordinates in a diffusion model trained only on nominal data to detect anomalies, with a non-parametric variant and theoretical error guarantees, achieving the bes...

  19. AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models

    cs.LG 2026-05 unverdicted novelty 6.0

    AdvantageFlow proposes an advantage-weighted forward-process least-squares loss for RL in rectified flow models, stabilized by rollout policy regularization, and reports better image generation performance than Flow-G...

  20. Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

    cs.CV 2026-05 unverdicted novelty 6.0

    AdaScope adaptively selects optimal RL intervention points during diffusion denoising by monitoring structural and semantic changes, delivering 66% higher performance at 59% lower cost than full-trajectory RL baselines.

  21. When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy

    cs.CV 2026-05 unverdicted novelty 6.0

    Policy entropy remains constant in flow-matching models during RLHF due to fixed noise schedules while perceptual diversity collapses from mode-seeking policy gradients, so perceptual entropy constraints are introduce...

  22. Gradient-Free Noise Optimization for Reward Alignment in Generative Models

    cs.LG 2026-05 unverdicted novelty 6.0

    ZeNO frames noise optimization as a path-integral control problem solvable from zeroth-order reward evaluations, connecting to implicit Langevin dynamics for reward-tilted distributions.

  23. Gradient-Free Noise Optimization for Reward Alignment in Generative Models

    cs.LG 2026-05 unverdicted novelty 6.0

    ZeNO formulates noise optimization for reward alignment as a path-integral control problem solvable via zeroth-order reward evaluations alone, connecting to Langevin dynamics under an Ornstein-Uhlenbeck process.

  24. Conditional Diffusion Under Linear Constraints: Langevin Mixing and Information-Theoretic Guarantees

    cs.LG 2026-05 unverdicted novelty 6.0

    Error in approximating the tangent conditional score by the unconditional score in diffusion models is bounded by dimension-free conditional mutual information, with a projected-Langevin method outperforming baselines...

  25. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 6.0

    FMRG is a training-free single-trajectory guidance framework for flow-based models that matches or exceeds baselines on reward-guided tasks and inverse problems using as few as 3 NFEs.

  26. Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

    cs.LG 2026-04 unverdicted novelty 6.0

    Reward Score Matching unifies reward-based fine-tuning for flow and diffusion models by recasting alignment as score matching to a value-guided target.

  27. How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?

    cs.LG 2026-02 unverdicted novelty 6.0

    ALGD augments the Lagrangian to locally convexify the energy landscape in diffusion models, stabilizing safe RL training and generation without changing optimal policies.

  28. Control-Augmented Autoregressive Diffusion for Data Assimilation

    cs.LG 2025-10 unverdicted novelty 6.0

    An offline-trained controller augments autoregressive diffusion models to perform fast, feed-forward data assimilation in chaotic spatiotemporal PDEs with order-of-magnitude speedups and improved accuracy over baselines.