Pith. sign in

REVIEW 29 cited by

Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15194 v2 pith:JPSETJC4 submitted 2024-02-23 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords rewarddiffusionmodelsfunctionqualityaestheticcollapsecontrol
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models excel at capturing complex data distributions, such as those of natural images and proteins. While diffusion models are trained to represent the distribution in the training dataset, we often are more concerned with other properties, such as the aesthetic quality of the generated images or the functional properties of generated proteins. Diffusion models can be finetuned in a goal-directed way by maximizing the value of some reward function (e.g., the aesthetic quality of an image). However, these approaches may lead to reduced sample diversity, significant deviations from the training data distribution, and even poor sample quality due to the exploitation of an imperfect reward function. The last issue often occurs when the reward function is a learned model meant to approximate a ground-truth "genuine" reward, as is the case in many practical applications. These challenges, collectively termed "reward collapse," pose a substantial obstacle. To address this reward collapse, we frame the finetuning problem as entropy-regularized control against the pretrained diffusion model, i.e., directly optimizing entropy-enhanced rewards with neural SDEs. We present theoretical and empirical evidence that demonstrates our framework is capable of efficiently generating diverse samples with high genuine rewards, mitigating the overoptimization of imperfect reward models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0 of 10

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  2. TILDE: TILt-based Distributional Erasure for Concept Unlearning

    cs.LG 2026-07 conditional novelty 7.0 of 10

    TILDE derives a minimum-deviation, energy-tilted target distribution for concept unlearning in diffusion models and realizes it via residual ∇-GFlowNet training.

  3. Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    Q-VGM offline RL fine-tuning converts critic Q-gradients into residual velocity targets for flow-matching VLAs, raising LIBERO success from 75.0% to 92.5%.

  4. Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Reward Score Matching unifies reward-based fine-tuning for flow and diffusion models by recasting alignment as score matching to a value-guided target.

  5. Is RL fine-tuning harder than regression? A PDE learning approach for diffusion models

    cs.LG 2025-09 conditional novelty 7.0 of 10

    For diffusion model fine-tuning, RL value estimation reduces to a variational inequality whose solution satisfies a supervised-learning oracle inequality with self-mitigating statistical error.

  6. Provable Maximum Entropy Manifold Exploration via Diffusion Models

    cs.LG 2025-06 conditional novelty 7.0 of 10

    S-MEME iteratively fine-tunes a diffusion model using its own score as the exploration reward, provably converging to the maximum-entropy distribution on the learned manifold.

  7. Efficient Diversity-Preserving Diffusion Alignment via Gradient-Informed GFlowNets

    cs.LG 2024-12 conditional novelty 7.0 of 10

    A gradient-informed GFlowNet objective, residual nabla-DB, finetunes diffusion models to sample according to a reward while preserving diversity and prior knowledge.

  8. CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction

    cs.LG 2026-08 conditional novelty 6.0 of 10

    CrystalGRPO post-trains flow-based crystal generators with joint coordinate-lattice stochastic policies and a hybrid MACE-energy plus structure-matching reward, improving Top-1 recovery in one mode and Top-20 coverage...

  9. Generalized Fine-Tuning of Diffusion Models via Stochastic Control and FBSDEs

    math.OC 2026-06 reject novelty 6.0 of 10

    Diffusion fine-tuning is generalized to arbitrary running costs and solved through HJB/FBSDE, but the main theorems are not proven as stated.

  10. Calibrated Test-Time Guidance for Bayesian Inference

    cs.LG 2026-02 conditional novelty 6.0 of 10

    CBG replaces biased point estimates of the diffused likelihood with consistent Monte Carlo score estimates, and corrects how guidance scales temper the likelihood.

  11. Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach

    cs.AI 2026-02 conditional novelty 6.0 of 10

    By adding drift g(t)^2 ∇log h(t,y) with h estimated via martingale and covariation losses, diffusion samples can be hard-conditioned on an event.

  12. Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Optimizing the null-text embedding in classifier-free guidance aligns diffusion outputs to a target reward while preserving cross-reward quality.

  13. Calibrating Generative Models to Distributional Constraints

    stat.ML 2025-10 conditional novelty 6.0 of 10

    CGM-relax and CGM-reward fine-tune generative models to meet distributional constraints by minimizing a miscalibration penalty or a KL divergence to an estimated maximum-entropy tilt.

  14. Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A noise hypernetwork learns a reward-tilted initial noise distribution for frozen distilled diffusion generators, recovering roughly half of test-time noise-optimization gains at a fraction of the compute.

  15. Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Diffusion Tree Sampling is a Monte Carlo tree search over denoising trajectories that propagates terminal rewards backward to sample from reward-aligned distributions, showing up to 10x compute savings on tested benchmarks.

  16. Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Nabla-R2D3 aligns 3D-native diffusion models with human preferences by backpropagating multi-view 2D reward gradients through the denoising process, improving reward without destroying the pretrained 3D prior.

  17. ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

    cs.RO 2025-05 conditional novelty 6.0 of 10

    ChatVLA-2 uses dynamic mixture-of-experts and a two-stage training recipe to let a vision-language-action model retain pretrained reasoning while following robot instructions.

  18. Scaling Image and Video Generation via Test-Time Evolutionary Search

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Evolutionary search over denoising trajectories improves image and video generation quality and diversity as test-time compute increases, without retraining the generative model.

  19. A First-order Generative Bilevel Optimization Framework for Diffusion Models

    cs.LG 2025-02 reject novelty 6.0 of 10

    A bilevel first-order method tunes entropy-regularization strength and noise schedules in diffusion models without backpropagating through sampling, improving FID and CLIP over hyperparameter search baselines.

  20. Outsourced diffusion sampling: Efficient posterior inference in latent spaces of generative models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Posterior sampling under a generative prior can be performed by training a diffusion model in the generator's noise space and pushing its samples through the generator.

  21. Statistical guarantees for continuous-time policy evaluation: blessing of ellipticity and new tradeoffs

    cs.LG 2025-02 conditional novelty 6.0 of 10

    For continuous-time policy evaluation, the LSTD estimator's H1 error scales as the square root of (approximation error plus m/T), with a trajectory length that can be nearly linear in the number of basis functions whe...

  22. Direct Distributional Optimization for Provable Alignment of Diffusion Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A distribution-level optimization framework, dual averaging plus Doob's h-transform, aligns diffusion models with provable convergence and isoperimetry-free sampling.

  23. Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A robot policy generates its own language reasoning before acting and injects it into a diffusion action decoder, outperforming several VLA baselines on real-robot manipulation.

  24. Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation

    cs.RO 2024-11 conditional novelty 6.0 of 10

    HyDo combines diffusion-model policies with maximum entropy RL in a hybrid discrete/continuous action space, improving success rates on non-prehensile manipulation tasks.

  25. On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...

  26. Time-Reversed BSDEs for Accurate Gradient Estimation in Diffusion Models

    math.OC 2026-03 conditional novelty 5.5 of 10

    Time-reversed BSDEs produce adapted adjoints that yield more stable, lower-variance gradients for SOC fine-tuning of diffusion models than non-adapted adjoint matching.

  27. Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework

    cs.HC 2025-08 unverdicted novelty 5.0 of 10

    A three-layer framework (input, processing, output) for adaptive external human-machine interfaces in autonomous vehicles is introduced to systematize design and analysis.

  28. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...

  29. Score as Action: Fine-Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A continuous-time RL algorithm that treats diffusion scores as actions fine-tunes text-to-image models with a Girsanov-based KL regularizer, showing stability across different denoising step counts.

Pith tools