Pith. sign in

REVIEW 23 cited by

Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08252 v5 pith:7FUETTWG submitted 2024-08-15 cs.LG cs.AIq-bio.GNstat.ML

classification cs.LGcs.AIq-bio.GNstat.ML
keywords modelsdiffusionfine-tuninggenerationguidancemethodalgorithmdesign
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models excel at capturing the natural design spaces of images, molecules, DNA, RNA, and protein sequences. However, rather than merely generating designs that are natural, we often aim to optimize downstream reward functions while preserving the naturalness of these design spaces. Existing methods for achieving this goal often require ``differentiable'' proxy models (\textit{e.g.}, classifier guidance or DPS) or involve computationally expensive fine-tuning of diffusion models (\textit{e.g.}, classifier-free guidance, RL-based fine-tuning). In our work, we propose a new method to address these challenges. Our algorithm is an iterative sampling method that integrates soft value functions, which looks ahead to how intermediate noisy states lead to high rewards in the future, into the standard inference procedure of pre-trained diffusion models. Notably, our approach avoids fine-tuning generative models and eliminates the need to construct differentiable models. This enables us to (1) directly utilize non-differentiable features/reward feedback, commonly used in many scientific domains, and (2) apply our method to recent discrete diffusion models in a principled way. Finally, we demonstrate the effectiveness of our algorithm across several domains, including image generation, molecule generation, and DNA/RNA sequence generation. The code is available at \href{https://github.com/masa-ue/SVDD}{https://github.com/masa-ue/SVDD}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Inference-Time Scaling in Diffusion Models through Iterative Partial Refinement

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    IPR improves valid solution rates on MNIST Sudoku from 55.8% to 75.0% by iteratively refining partial regions in sequential diffusion models without external verifiers or reward models.

  2. Step-level Denoising-time Diffusion Alignment with Multiple Objectives

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    MSDDA derives a closed-form optimal reverse denoising distribution for multi-objective diffusion alignment that is exactly equivalent to step-level RL fine-tuning with no approximation error.

  3. VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion

    cs.AI 2026-04 unverdicted novelty 7.0 of 10

    FVD applies Fleming-Viot population dynamics to diffusion model sampling at inference time to reduce diversity collapse while improving reward alignment and FID scores.

  4. Offline Materials Optimization with CliqueFlowmer

    cs.AI 2026-03 unverdicted novelty 7.0 of 10

    CliqueFlowmer combines clique-based model-based optimization with transformer and flow models to generate materials that optimize target properties better than generative baselines.

  5. Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching

    cs.LG 2025-09 conditional novelty 7.0 of 10

    Derives exact guidance transition rates for discrete flow matching models that require only one model evaluation per sampling step and unify prior approximation-based methods.

  6. Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement

    cs.LG 2025-07 conditional novelty 7.0 of 10

    PG-DLM applies particle Gibbs sampling over full trajectories in diffusion language models to enable iterative refinement, yielding higher accuracy on reward-guided generation with theoretical convergence guarantees.

  7. Histogram-constrained Image Generation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    HIG enforces exact histogram constraints on diffusion-generated images by modeling the control task as an optimal transport problem and applying guidance transformations during sampling.

  8. Plug-and-Play Guidance for Discrete Diffusion Models via Gradient-Informed Logit Correction

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Introduces GILC, a training-free plug-and-play guidance framework for discrete diffusion models that uses Jacobian-free logit correction to achieve SOTA results on DNA, protein, and molecular generation tasks.

  9. Plug-and-Play Guidance for Discrete Diffusion Models via Gradient-Informed Logit Correction

    cs.LG 2026-06 conditional novelty 6.0 of 10

    GILC guides discrete diffusion at inference time by adding reward gradients to the prediction logits, matching or beating fine-tuned baselines on DNA, protein, and molecule generation tasks without retraining.

  10. Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Proprio uses flow residuals from latent perturbations in frozen video generators as a self-scoring signal for physical plausibility, yielding reported gains of 16.5% on Physics-IQ and 20.6% on VideoPhy2-hard.

  11. Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures

    stat.ML 2026-05 unverdicted novelty 6.0 of 10

    URGE performs unbiased inference-time scaling for diffusion models by attaching multiplicative path weights from Girsanov estimation and resampling trajectories, with a proven equivalence to prior particle-wise SMC schemes.

  12. LPDP: Inference-Time Reward Control for Variable-Length DNA Generation with Edit Flows

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    LPDP adds a local re-solving operator to edit-flow DNA generators so that reward signals can guide insertions, deletions, and substitutions without retraining.

  13. dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    dFlowGRPO is a new rate-aware RL method for discrete flow models that outperforms prior GRPO approaches on image generation and matches continuous flow models while supporting broad probability paths.

  14. DAG-STL: A Hierarchical Framework for Zero-Shot Trajectory Planning under Signal Temporal Logic Specifications

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    DAG-STL decomposes long-horizon STL planning into decomposition, timed waypoint allocation, and diffusion-based trajectory generation to enable zero-shot planning under unknown dynamics.

  15. VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    VASR separates continuation and residual variance in reward-guided diffusion SMC, using optimal mass allocation and systematic resampling to achieve up to 26% better FID scores and faster runtimes than prior SMC and M...

  16. Superbunched random fiber laser

    physics.optics 2026-03 unverdicted novelty 6.0 of 10

    A fiber-integrated random laser uses Rayleigh scattering, cascaded Brillouin scattering, and four-wave mixing to generate multi-wavelength superbunched light with g(2)(0) up to ~26 and improved temporal ghost imaging.

  17. Calibrated Test-Time Guidance for Bayesian Inference

    cs.LG 2026-02 conditional novelty 6.0 of 10

    CBG replaces biased point estimates of the diffused likelihood with consistent Monte Carlo score estimates, and corrects how guidance scales temper the likelihood.

  18. Solving Inverse Problems with Flow-based Models via Model Predictive Control

    eess.IV 2026-01 conditional novelty 6.0 of 10

    MPC-Flow applies model predictive control to guide pretrained flow models through inverse problems, with a single-step variant that avoids backpropagation and scales to 32B-parameter models on consumer hardware.

  19. Efficient Inference for Coupled Hidden Markov Models in Continuous Time and Discrete Space

    stat.ML 2025-10 unverdicted novelty 6.0 of 10

    Proposes Latent Interacting Particle Systems with an efficient parameterization of twist potentials to enable approximate posterior inference for coupled continuous-time hidden Markov models via twisted sequential Mon...

  20. Feynman-Kac-Flow: Inference Steering of Conditional Flow Matching to an Energy-Tilted Posterior

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Feynman-Kac particle steering, previously diffusion-only, is derived for conditional flow matching and used to generate chirality-correct chemical transition states.

  21. On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...

  22. Multi-Cycle Spatio-Temporal Adaptation in Human-Robot Teaming

    cs.RO 2026-04 unverdicted novelty 5.0 of 10

    RAPIDDS unifies task-level and motion-level adaptation in human-robot teaming by modeling individualized spatial and temporal behaviors across multiple cycles and jointly optimizing schedules and diffusion-based motions.

  23. Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

    cs.LG 2026-07 unverdicted novelty 4.0 of 10

    Discrete diffusion models are re-framed as instances of a tokenization-centric, four-component design space (corruption, denoiser, objective, sampler) in a broad survey with no new experimental or theoretical results.

Pith tools