Pith. sign in

REVIEW 26 cited by

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09685 v2 pith:YL4X2YJM submitted 2025-01-16 cs.AI cs.LGq-bio.QMstat.ML

Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review

classification cs.AI cs.LGq-bio.QMstat.ML
keywords inference-timemodelsalgorithmsdiffusiontutorialfunctionsguidancemethods
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This tutorial provides an in-depth guide on inference-time guidance and alignment methods for optimizing downstream reward functions in diffusion models. While diffusion models are renowned for their generative modeling capabilities, practical applications in fields such as biology often require sample generation that maximizes specific metrics (e.g., stability, affinity in proteins, closeness to target structures). In these scenarios, diffusion models can be adapted not only to generate realistic samples but also to explicitly maximize desired measures at inference time without fine-tuning. This tutorial explores the foundational aspects of such inference-time algorithms. We review these methods from a unified perspective, demonstrating that current techniques -- such as Sequential Monte Carlo (SMC)-based guidance, value-based sampling, and classifier guidance -- aim to approximate soft optimal denoising processes (a.k.a. policies in RL) that combine pre-trained denoising processes with value functions serving as look-ahead functions that predict from intermediate states to terminal rewards. Within this framework, we present several novel algorithms not yet covered in the literature. Furthermore, we discuss (1) fine-tuning methods combined with inference-time techniques, (2) inference-time algorithms based on search algorithms such as Monte Carlo tree search, which have received limited attention in current research, and (3) connections between inference-time algorithms in language models and diffusion models. The code of this tutorial on protein design is available at https://github.com/masa-ue/AlignInversePro

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 8.0

    FMRG reformulates guidance as deterministic optimal control, deriving a single-trajectory method using the flow map that matches or exceeds baselines on reward-guided generation and inverse problems with 3 NFEs at tex...

  2. Active Flow Expansion for Out-of-Distribution Discovery: from Theory to Molecules

    cs.LG 2026-06 unverdicted novelty 7.0

    ActFlow expands the generable set of pre-trained flow models for out-of-distribution molecular and sequence design via active synthetic data generation and verifier feedback, with new statistical guarantees.

  3. Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

    cs.LG 2026-05 unverdicted novelty 7.0

    FAV aligns few-step generative models by amortizing SVGD updates from reward-tilted sampling into generator parameters via fixed-point regression, requiring only sample access, and shows outperformance on robotics tas...

  4. Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

    cs.LG 2026-05 conditional novelty 7.0

    TRI-TSMC is a trust-region framework for learning twisting functions in SMC-based inference-time alignment of diffusion models that yields zero-variance samplers in theory and better alignment on text and image tasks ...

  5. SURGE: Approximation and Training Free Particle Filter for Diffusion Surrogate

    stat.ML 2026-05 unverdicted novelty 7.0

    URGE performs unbiased path-wise importance reweighting via Girsanov estimation for derivative-free inference-time scaling in diffusion models, proving equivalence to particle-wise SMC and outperforming baselines empirically.

  6. LENS: Low-Frequency Eigen Noise Shaping for Efficient Diffusion Sampling

    cs.CV 2026-05 unverdicted novelty 7.0

    LENS shapes low-frequency eigen noise with a lightweight network to enable efficient, high-quality sampling in distilled diffusion models.

  7. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 7.0

    FMRG is a training-free, single-trajectory guidance method for flow models derived from optimal control that achieves strong reward alignment with only 3 NFEs.

  8. Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

    cs.LG 2026-04 conditional novelty 7.0

    Many reward-based fine-tuning methods for diffusion and flow models are special cases of a common score-matching objective whose differences reduce to value-guidance estimator design, temporal weighting, and trust-reg...

  9. Iterative Inference-time Scaling with Adaptive Frequency Steering for Image Super-Resolution

    cs.CV 2025-12 unverdicted novelty 7.0

    IAFS is a training-free iterative inference-time scaling framework that uses adaptive frequency-aware particle fusion to resolve the perception-fidelity conflict in diffusion super-resolution models, outperforming pri...

  10. Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs

    cs.AI 2025-05 unverdicted novelty 7.0

    UniR is a composable reasoning module trained with verifiable rewards and added to frozen LLMs via logit summation, enabling modular composition and weak-to-strong generalization across tasks and model sizes.

  11. Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0

    Under wall-clock budgets, cheap multi-knob drafts plus multi-stage verification outperform guided intermediate search for diffusion T2I inference-time scaling.

  12. Few-step Cofolding with All-Atom Flow Maps

    cs.LG 2026-06 unverdicted novelty 6.0

    DeCAF distills all-atom cofolding diffusion models into few-step flow maps, showing improved or matched accuracy on protein-ligand tasks with 5x fewer inference steps.

  13. Are we really tilting? The mechanics of reward guidance in flow and diffusion models

    cs.LG 2026-06 unverdicted novelty 6.0

    Finite-particle approximation of the Doob h-function causes reward hacking via two failure modes in reward-guided diffusion; a damping schedule corrects within-mode bias in Gaussian settings.

  14. StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

    cs.CV 2026-05 unverdicted novelty 6.0

    StressDream optimizes initial noise in diffusion video world models using VLM semantic and plausibility objectives to steer generations toward specified high-impact outcomes for improved policy evaluation.

  15. Parallel Tempering Initial Sampling in Inference-Time Reward Alignment

    cs.LG 2026-05 unverdicted novelty 6.0

    PATHS applies parallel tempering to improve initial particle sampling for SMC reward alignment, yielding better results on layout-to-image and quantity-aware generation tasks.

  16. SURGE: Approximation and Training Free Particle Filter for Diffusion Surrogate

    stat.ML 2026-05 unverdicted novelty 6.0

    SURGE is an unbiased particle filter that fuses diffusion-model simulations with noisy observations via sequential Monte Carlo reweighting over diffusion trajectories.

  17. Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures

    stat.ML 2026-05 unverdicted novelty 6.0

    URGE performs unbiased inference-time scaling for diffusion models by attaching multiplicative path weights from Girsanov estimation and resampling trajectories, with a proven equivalence to prior particle-wise SMC schemes.

  18. Inference-Time Attribute Distribution Alignment for Unconditional Diffusion

    cs.LG 2026-05 unverdicted novelty 6.0

    An optimal control formulation adds time-dependent perturbations to the reverse diffusion process to match target attribute distributions while preserving sample fidelity.

  19. On the Robustness of Distribution Support under Diffusion Guidance

    cs.LG 2026-05 unverdicted novelty 6.0

    Guided diffusion generates samples near the target distribution support under exact score access, explaining its empirical success in producing plausible outputs.

  20. How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

    cs.LG 2026-04 unverdicted novelty 6.0

    FMRG is a training-free single-trajectory guidance framework for flow-based models that matches or exceeds baselines on reward-guided tasks and inverse problems using as few as 3 NFEs.

  21. Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

    cs.LG 2026-04 unverdicted novelty 6.0

    Reward Score Matching unifies reward-based fine-tuning for flow and diffusion models by recasting alignment as score matching to a value-guided target.

  22. RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling

    cs.CV 2025-10 unverdicted novelty 6.0

    RAPO++ is a three-stage prompt optimization framework combining retrieval-augmented refinement, closed-loop test-time scaling, and LLM fine-tuning to enhance text-to-video generation quality.

  23. Efficient Inference for Coupled Hidden Markov Models in Continuous Time and Discrete Space

    stat.ML 2025-10 unverdicted novelty 6.0

    Proposes Latent Interacting Particle Systems with an efficient parameterization of twist potentials to enable approximate posterior inference for coupled continuous-time hidden Markov models via twisted sequential Mon...

  24. Control-Augmented Autoregressive Diffusion for Data Assimilation

    cs.LG 2025-10 unverdicted novelty 6.0

    An offline-trained controller augments autoregressive diffusion models to perform fast, feed-forward data assimilation in chaotic spatiotemporal PDEs with order-of-magnitude speedups and improved accuracy over baselines.

  25. On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

    cs.LG 2026-07 conditional novelty 5.5

    Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...

  26. On the Robustness of Distribution Support under Diffusion Guidance

    cs.LG 2026-05 unverdicted novelty 4.0

    Establishes robustness of distribution support for guided diffusion processes under exact score access across DDIM, DDPM, and exponential integrator discretizations.