Pith. sign in

hub Canonical reference

Fine- tuning of continuous-time diffusion models as entropy-regularized control

Canonical reference. 80% of citing Pith papers cite this work as background.

25 Pith papers citing it
Background 80% of classified citations
abstract

Diffusion models excel at capturing complex data distributions, such as those of natural images and proteins. While diffusion models are trained to represent the distribution in the training dataset, we often are more concerned with other properties, such as the aesthetic quality of the generated images or the functional properties of generated proteins. Diffusion models can be finetuned in a goal-directed way by maximizing the value of some reward function (e.g., the aesthetic quality of an image). However, these approaches may lead to reduced sample diversity, significant deviations from the training data distribution, and even poor sample quality due to the exploitation of an imperfect reward function. The last issue often occurs when the reward function is a learned model meant to approximate a ground-truth "genuine" reward, as is the case in many practical applications. These challenges, collectively termed "reward collapse," pose a substantial obstacle. To address this reward collapse, we frame the finetuning problem as entropy-regularized control against the pretrained diffusion model, i.e., directly optimizing entropy-enhanced rewards with neural SDEs. We present theoretical and empirical evidence that demonstrates our framework is capable of efficiently generating diverse samples with high genuine rewards, mitigating the overoptimization of imperfect reward models.

hub tools

citation-role summary

background 4 method 1

citation-polarity summary

years

2026 21 2025 4

representative citing papers

How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

cs.LG · 2026-04-29 · unverdicted · novelty 8.0 · 3 refs

FMRG reformulates guidance as deterministic optimal control, deriving a single-trajectory method using the flow map that matches or exceeds baselines on reward-guided generation and inverse problems with 3 NFEs at text-to-image scale.

A Markov Chain Approach to Preference Alignment

cs.LG · 2026-06-21 · unverdicted · novelty 7.0

MCHF defines a Markov kernel from pairwise utilities U and proves geometric convergence to its stationary distribution at a rate set by the seminorm measuring non-transitivity of U, with first-order equivalence to RLHF and NLHF solutions.

citing papers explorer

Showing 25 of 25 citing papers.