Pith. sign in

REVIEW 3 cited by

Reward-Directed Score-Based Diffusion Models via q-Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.04832 v2 pith:WZ33RSHD submitted 2024-09-07 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords modelsdiffusiondistributionsformulationfunctionspretrainedscoreunknown
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new reinforcement learning (RL) formulation for training continuous-time score-based diffusion models for generative AI to generate samples that maximize reward functions while keeping the generated distributions close to the unknown target data distributions. Different from most existing studies, ours does not involve any pretrained model for the unknown score functions of the noise-perturbed data distributions, nor does it attempt to learn the score functions. Instead, we formulate the problem as entropy-regularized continuous-time RL and show that the optimal stochastic policy has a Gaussian distribution with a known covariance matrix. Based on this result, we parameterize the mean of Gaussian policies and develop an actor--critic type (little) q-learning algorithm to solve the RL problem. A key ingredient in our algorithm design is to obtain noisy observations from the unknown score function via a ratio estimator. Our formulation can also be adapted to solve pure score-matching and fine-tuning pretrained models. Numerically, we show the effectiveness of our approach by comparing its performance with two state-of-the-art RL methods that fine-tune pretrained models on several generative tasks including high-dimensional image generations. Finally, we discuss extensions of our RL formulation to probability flow ODE implementation of diffusion models and to conditional diffusion models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning

    cs.LG 2026-07 unverdicted novelty 7.0 of 10

    ART-RL learns adaptive diffusion sampling timesteps via continuous-time control and Gaussian actor–critic RL, improving and transferring over hand-designed grids at matched budgets.

  2. Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

    cs.AI 2026-02 conditional novelty 6.0 of 10

    By adding drift g(t)^2 ∇log h(t,y) with h estimated via martingale and covariation losses, diffusion samples can be hard-conditioned on an event.

  3. Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology

    stat.ML 2025-09 conditional novelty 4.0 of 10

    RLHF, RLIF, and soft best-of-N sampling reduce to the same exponential-tilting objective under parameter matching, and test-time scaling can asymptotically implement classifier-free diffusion guidance.

Pith tools