Pith. sign in

REVIEW 2 major objections 6 minor 67 references

Online fine-tuning of discrete diffusion models finds better molecules when acquisition, reward shaping, and debiasing work together.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 06:43 UTC pith:4EAAVY6V

load-bearing objection Solid empirical design-space study with a usable recipe; the cheap density estimator is a real but contained soft spot, not a collapse of the claim. the 2 major comments →

arxiv 2607.02834 v1 pith:4EAAVY6V submitted 2026-07-03 cs.LG

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

classification cs.LG
keywords discrete diffusionmolecular optimizationonline adaptationThompson samplingCVaR reward shapingdensity entropy regularizationreplay buffertest-time fine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Molecular design usually starts from a pretrained generator that knows what valid molecules look like, but the real goal is to spend a small evaluation budget to push that generator toward high-reward candidates for a specific task. This paper treats the online adaptation loop for discrete diffusion models as a design space: each round must choose which candidates to score, how scores become training rewards, which past feedback to reuse, and how far to leave the pretrained prior. Controlled ablations on six small-molecule binding tasks and three protein-fitness tasks show that Thompson sampling for acquisition, CVaR-style reward shaping, and density-entropy debiasing give complementary gains, especially when good solutions sit far from the prior. Replay and invalid-molecule penalties stabilize the loop rather than drive the top reward. The combined recipe beats offline fine-tuning and pure inference-time search under matched oracle-call and GPU-hour budgets, with the largest gains precisely when the target molecules require a large shift from the starting prior.

Core claim

Inside a full online adaptation loop for discrete diffusion molecular generators, acquisition (Thompson sampling), CVaR reward shaping, and Density Entropy Regularization provide complementary routes to higher reward, while replay and invalid-output penalties act mainly as stabilizers; the resulting combined recipe outperforms offline fine-tuning and inference-time search under matched oracle and compute budgets, and the advantage is largest when high-reward candidates lie far from the pretrained prior.

What carries the argument

The online active-loop harness: each round samples candidates from the current discrete diffusion model, selects a batch via Thompson sampling over a reward-model ensemble, converts oracle scores into CVaR-shaped and density-debiased log-rewards (with an optional invalid-SMILES penalty), updates from a stratified replay buffer, and fine-tunes with a plug-in objective such as DDPP-LB or VIDD.

Load-bearing premise

The cheap single-pass likelihood estimate used for density-entropy debiasing is treated as a faithful enough proxy for the intended model-density penalty that the observed off-prior shifts and reward gains can be attributed to true debiasing.

What would settle it

On a held-out small-molecule target where high-affinity ligands are known to lie far from the pretrained prior, replace the cheap single-pass density estimator with a multi-sample unbiased estimator throughout training and check whether the top-1 and top-10% reward curves and the final anchor-NLL shift still match the paper's reported gains; a collapse of either would undermine the claim that the cheap estimator is doing real debiasing.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper studies online test-time adaptation of pretrained discrete diffusion generators for molecular optimization under fixed oracle and wall-clock budgets. It frames the loop as a finetuner-agnostic harness with five plug-in components—Thompson sampling acquisition, CVaR reward shaping, Density Entropy Regularization (model debiasing), replay, and invalid-output penalties—and runs controlled leave-one-out and multi-knob ablations with DDPP-LB and VIDD on six small-molecule binding-affinity tasks and three protein-fitness tasks. The main claim is that acquisition, CVaR, and debiasing provide complementary reward gains (especially on small molecules that require larger shifts from the prior), while replay and validity penalties stabilize exploration; the combined recipe outperforms offline fine-tuning and inference-time search under matched oracle-call and GPU-hour accounting. Supporting analyses include secondary validity/diversity/QED/SA metrics, anchor-NLL distribution shifts, compute breakdowns, and selected Boltz-2 re-runs.

Significance. If the empirical conclusions hold, this is a useful design-space map for a setting that is increasingly common: expensive oracles, pretrained discrete diffusion priors, and limited online feedback. The work’s strengths include dual finetuners, dual evaluation axes (oracle calls and GPU hours), leave-one-out plus multi-knob ablations, secondary manifold metrics, and public code/results. The domain contrast (broad small-molecule prior vs family-specific protein priors) is a clear, falsifiable organizing principle. The practical recipe is actionable even if some mechanistic attributions remain approximate. The contribution is primarily empirical and systems-level rather than a new theoretical guarantee, but that is appropriate for the stated question.

major comments (2)
  1. [§2.2, Eq. (2); Figs. 5, 7, 35–36] §2.2 and Eq. (2): Density Entropy Regularization is implemented with a detached single-pass masked-token estimator log p_θ(x0|xt), which the paper correctly states is not equivalent to the continuous negative-score regularizer or the LLaDA-Alg. 3 ELBO. The central complementarity narrative (especially Fig. 5 and Fig. 7, where removing DER causes the largest small-molecule drop and the strongest prior collapse) attributes gains to mode debiasing. Supporting evidence is mainly post-hoc NLL histograms on FA/2VT4 (Figs. 35–36) and oracle-call curves vs LLaDA n_mc variants. That shows similar off-prior mass and efficiency, not that the online gradient direction matches the intended density penalty. Please either (i) run the leave-one-out / full-loop comparison with an unbiased (or higher-n_mc) estimator inside training on at least one small-molecule and one protein task, or (ii) substantially
  2. [Table 2; §B.7; Fig. 5] Table 2 and §B.7: free parameters (τ/q, γ, r_inv, M/K/G, ensemble size) are fixed at reference defaults without a systematic sensitivity study except for r_inv (Fig. 40). The claim that components “complement one another” rather than “work at these defaults” is load-bearing for the design-space conclusion. At minimum, report one-dimensional sweeps or a small grid for γ and the CVaR quantile on a representative small-molecule target (and note whether ordering of leave-one-out drops is stable). Without this, readers cannot separate knob identity from a lucky operating point, especially for DER where γ=1.0 is strong.
minor comments (6)
  1. [Eq. (3); Figs. 5–6] Eq. (3) / §B.1: normalized “lift” rescales by per-target min/max over the methods shown in each figure. This is fine for within-figure comparison but can inflate apparent gaps when the method set changes. State this limitation next to Fig. 5–6 and prefer raw reward panels (or a fixed reference set) in the appendix for absolute scale.
  2. [§4.1; Figs. 24–26] §4.1: inference-time search hybrids are argued to fail mainly because generation consumes the wall budget (Figs. 25–26). Consider one matched-gradient-step or matched-candidate-pool control so the conclusion is not solely compute-allocation dependent.
  3. [Abstract; Fig. 5; §5] Protein results are weaker and more acquisition-driven; the abstract and conclusion already note this, but Fig. 5’s pooled bars can still over-sell “complementary routes” as domain-general. A one-sentence domain-conditional summary in the abstract would help.
  4. [§2.2] Notation: logp vs log p, and the cheap estimator \logp, are easy to miss. Define the estimator once in a displayed equation and reuse a single symbol.
  5. [Throughout] Typos / polish: “intop θt+1”, “2VT4R” vs 2VT4, mixed “Density Entropy” / “model debiasing” naming, and occasional doubled words (“we find empirically that it… we find empirically”).
  6. [Appendices G, M] Appendix is very long relative to the main text; consider moving multi-knob and Boltz-2 highlights into the main body if space allows, since they support the transfer and complementarity claims.

Circularity Check

0 steps flagged

No significant circularity: empirical ablation study against external oracles and baselines; no derivation reduces to its inputs by construction.

full rationale

This paper is a controlled empirical design-space study of online fine-tuning loops for discrete diffusion molecular optimization. Its central claims are comparative performance results under matched oracle-call and GPU-hour budgets (full recipe vs. leave-one-out ablations, offline fine-tuning, and inference-time search), measured against external oracles (FlashAffinity, Boltz-2, family fitness surrogates) and pretrained generators. There is no first-principles derivation chain whose conclusion is forced by its premises. The composable knobs (Thompson acquisition, CVaR shaping, Density Entropy Regularization, replay, invalid penalty) are taken from prior literature and ablated; self-citations to coauthors (e.g., De Santi et al. on density-entropy / CVaR-style objectives) supply the component definitions but do not uniqueness-force the complementarity claim, which is established by leave-one-out reward and distribution-shift measurements. The cheap single-pass density estimator is an explicit computational approximation with post-hoc NLL checks, not a circular redefinition of the objective. Normalized reward (Eq. 3) rescales per-target scores to [0,1] using figure-specific min/max for cross-target averaging; this is standard presentation scaling and does not determine method rankings or invent the performance gains. The study is self-contained against external benchmarks. Score 0.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 2 invented entities

The central claim is empirical and rests on standard ML/RL design choices plus a few paper-specific modeling approximations (cheap density estimator, fixed hyperparameter defaults, learned oracles as proxies for expensive assays). No new physical entities are postulated. Free parameters are the usual algorithmic knobs held fixed across ablations rather than fitted to force the headline result.

free parameters (5)
  • CVaR quantile τ (q) = 0.8
    Fixed at 0.8 for all runs; controls how aggressively the upper reward tail is emphasized.
  • Density-entropy weight γ = 1.0
    Fixed at 1.0; scales the model-debiasing penalty in the shaped log-reward.
  • Invalid-SMILES penalty r_inv = −5
    Chosen as −5 after a small sweep; load-bearing for validity under aggressive exploration.
  • Active-loop sizes M, K, G and buffer capacity = M=1000, K=25, G=50, buffer=10000
    M=1000 candidates, K=25 oracle batch, G=50 gradient steps, buffer 10k; held fixed as reference defaults rather than tuned per method.
  • Thompson ensemble size J and architecture = J=10, width=256
    J=10 MLPs, width 256; defines the uncertainty used for acquisition.
axioms (4)
  • domain assumption Learned oracles (FlashAffinity primary; family fitness surrogates; selected Boltz-2 checks) are adequate proxies for ranking molecular quality under the design objective.
    All primary metrics are oracle scores; wet-lab transfer is not demonstrated.
  • ad hoc to paper The single-pass masked-token estimator log p_θ(x0|xt) is a usable stand-in for the intractable model density in Density Entropy Regularization.
    §2.2 explicitly replaces the ELBO / LLaDA estimator for speed; only post-hoc NLL checks support fidelity.
  • domain assumption Holding all non-ablated hyperparameters at shared defaults isolates the causal effect of each knob.
    Stated in Appendix B.7; no per-method peak tuning, so interactions with optimal hyperparameters remain untested.
  • domain assumption Top-1 / top-k reward under a fixed oracle budget is the right primary success criterion for molecular optimization.
    Eq. (1) and evaluation protocol; multi-objective and synthesizability constraints are secondary.
invented entities (2)
  • Composable online-adaptation harness for discrete diffusion (five plug-in knobs + finetuner-agnostic loop) no independent evidence
    purpose: Organize and ablate acquisition, reward shaping, debiasing, replay, and validity control inside one feedback loop.
    Organizational construct rather than a physical entity; components are prior art, composition is the paper's framing.
  • Cheap single-pass density estimator for discrete Density Entropy Regularization no independent evidence
    purpose: Approximate −log p_θ(x) without n_mc extra forwards per step.
    Paper-specific approximation whose gradient fidelity is only partially validated.

pith-pipeline@v1.1.0-grok45 · 30931 in / 3446 out tokens · 27992 ms · 2026-07-12T06:43:23.884006+00:00 · methodology

0 comments
read the original abstract

Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures. At test time, however, the goal is not to sample from this prior, but to use a limited oracle budget to shift generation toward task-specific high-reward molecules. We study this adaptation problem for discrete diffusion models. Each online round couples several choices. The loop must decide which candidates to evaluate, how rewards become model updates, which feedback to reuse, and how far to move beyond the pretrained prior. These choices have mostly been studied in isolation, leaving open whether they complement one another, become redundant, or interfere inside a full online adaptation loop. We conduct controlled studies across six small-molecule binding-affinity tasks and three protein-fitness tasks. We find that acquisition, reward shaping, and model debiasing provide complementary routes to higher reward, especially for small molecules. Replay further stabilizes learning, while validity penalties keep small-molecule exploration on the valid molecular manifold. Together, these findings point to a practical recipe for feedback-efficient molecular optimization: online fine-tuning with acquisition, reward shaping, debiasing, replay, and validity control. This recipe outperforms offline fine-tuning and inference-time search baselines under matched oracle-call budgets and GPU-hour accounting. The gains are largest when high-reward candidates require larger shifts from the pretrained prior.

Figures

Figures reproduced from arXiv: 2607.02834 by Alexander F. G. Goldberg, Ariel Dai, Daniel Khalil, Jason Yang, Maruan Al-Shedivat, Nate Gruver, Pranav Murugan, Riccardo De Santi, Trevor Chen, Wenda Chu, Yisong Yue.

Figure 1
Figure 1. Figure 1: Online test-time feedback loop for molecular optimization. Each round samples a candidate pool Ct of size M from the current discrete diffusion generator pθt , selects a batch St of size K for oracle evaluation, appends the resulting pairs (x, r(x)) to the evaluated set, converts scores into a training reward r˜(x), regularizes the update, and trains pθt into pθt+1 before the next round. identification [e.… view at source ↗
Figure 2
Figure 2. Figure 2: Thompson sampling. Although cA has the higher mean prediction, cB can be selected be￾cause its larger uncertainty gives it nonzero probability of drawing above the greedy threshold. 3 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Reward-shaping design choices. (a) CVaR upper-tail shaping: the right-tail of the reward distribution above the τ -quantile (dashed) is shaded, and the solid red bar marks CVaR1−τ = E[r | r ≥ Qτ ], the expected reward conditional on being in that tail. Optimization concentrates gradient on this tail expectation, not on a hard top-k selection. (b) Density Entropy Regularization (model debiasing): subtractin… view at source ↗
Figure 4
Figure 4. Figure 4: In this study, we explore several tasks across two different domains, optimizing [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Leave-one-out ablations reveal complementary roles most clearly on small-molecule tasks. Each panel removes one component from the full online loop and reports normalized maximum reward after 100, 1k, and 5k oracle calls, for DDPP-LB (top) and VIDD (bottom), averaged over small-molecule targets (left) and protein families (right). At later budgets, small-molecule ablations diverge most strongly, with Densi… view at source ↗
Figure 6
Figure 6. Figure 6: Online adaptation outperforms inference-time search most clearly on small-molecule tasks. Full loop (black) vs. five search baselines at 100/1k/5k calls, over molecule targets (left) and protein families (right). On molecules the full loop leads at every budget and the gap widens with calls, as each gradient step compounds while search only re-selects from a fixed generator. On proteins the baselines reach… view at source ↗
Figure 7
Figure 7. Figure 7: Distribution shifts in pθt explain performance gains for methods on tasks where high-reward molecular design solutions lie outside the original generative prior pθt . Each panel overlays the run-end anchor-NLL density of the full method (blue) against the corresponding leave￾one-out ablation cell; the x-axis (− log ppre(x)) measures how far each sample lies from the original generative prior. Density Entro… view at source ↗
Figure 40
Figure 40. Figure 40: 16 [PITH_FULL_IMAGE:figures/full_fig_p016_40.png] view at source ↗
Figure 8
Figure 8. Figure 8: Unconditional sampling. The simplest baseline: sample from pθ and score every sample under the oracle, with no exploration policy. 19 [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Beam search. At each denoising step, the top-K partial trajectories by score are kept and expanded; the rest are pruned [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Monte Carlo Tree Search. Each node is a partial denoising state; UCB selects the next expansion based on visit counts and value estimates from rollouts [PITH_FULL_IMAGE:figures/full_fig_p020_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Discrete Feynman-Kac correctors (DFKC) / SMC. A particle population is reweighted at each denoising step by a tilt proportional to the (estimated) reward, then resampled when effective sample size drops below threshold. 20 [PITH_FULL_IMAGE:figures/full_fig_p020_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Surrogate screening. A cheap learned surrogate rˆ(xt) scores intermediate states in place of the expensive oracle, reducing per-output oracle cost from O(L) to O(1) [PITH_FULL_IMAGE:figures/full_fig_p021_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Tweedie estimator. The posterior-mean denoised state xˆ0 = E[x0 | xt] is computed in a single denoising pass and scored directly, replacing a full rollout [62, 63]. C.3 Fine-tuning objective and exploration knobs [PITH_FULL_IMAGE:figures/full_fig_p021_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: DDPP-LB loss. Trajectory-balance squared residual matching qθ to the reward-tilted posterior p ⋆ ∝ ppre · r, with a learned partition-function head Zˆ ϕ trained jointly with θ. 21 [PITH_FULL_IMAGE:figures/full_fig_p021_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Reward shapes on a synthetic 25-mode landscape. (a) Prior panchor: 25 roughly Gaussian modes. (b) Each mode is assigned a score in [0, 1]. (c) Vanilla DDPP-LB with raw reward r(x) puts mass on modes proportional to their reward weight (ideal under unlimited sampling, but compute-inefficient under limited oracle calls). (d) CVaR shaping concentrates most mass on the r ≥ 0.3 modes, enabling deeper search of… view at source ↗
Figure 16
Figure 16. Figure 16: Thompson sampling. Acquisition rule: a per-candidate sampled surrogate reward rˆi = µi + εiσi over the ensemble’s mean/disagreement, oracle-evaluating the top-K by rˆi . Different draws of ε favor different molecular regions [PITH_FULL_IMAGE:figures/full_fig_p023_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Replay buffer. Each generated (x, r) pair is stored in a fixed-capacity buffer with priority eviction; mini-batches are drawn from the buffer (with a fixed fresh-fraction) so the gradient sees a time-averaged version of the policy distribution rather than the sharp on-policy one. C.4 Active loop [PITH_FULL_IMAGE:figures/full_fig_p023_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Active loop. The high-level structure shared by every fine-tuning method we evaluate: generate candidates, score them under the oracle, then refine the backbone pθ and/or the surrogate before generating again. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_18.png] view at source ↗
Figure 1
Figure 1. Figure 1: Require: Pretrained ppre, oracle r(·), fine-tuning loss Lft, wall budget Twall Require: Generation batch M, oracle batch K, fine-tuning steps per round G Require: Reward-shape: CVaR quantile q, debiasing γ, invalid penalty rinv; ensemble size J 1: Initialize pθ0 ← ppre; buffer B ← ∅; ensemble {rˆϕj } J j=1 random 2: while wall time < Twall do 3: Ct ← sample M candidates from pθt 4: (µi , σi) ← ensemble mea… view at source ↗
Figure 19
Figure 19. Figure 19: Final comparison. Top-10 mean FA score vs. (a) cumulative oracle calls and (b) GPU hours, comparing the full online active loop (Our method: DDPP-LB + CVaR + debiasing + Thompson + buffer + invalid penalty) against the strongest pure-search baseline (Beam Search) and untreated online finetuning (DDPP Online No-CVaR). The full stack dominates both baselines under matched compute on the FA oracle. 24 [PITH… view at source ↗
Figure 20
Figure 20. Figure 20: Surrogate screening on inference-time search. Beam, MCTS, and DFKC each compared against their surrogate-pre-filter variant (cheap MLP scores all candidates, oracle only sees the top-K). Solid: oracle-on-everything. Dashed: surrogate pre-filter. Seed-averaged (n=5) top-1 best FA against cumulative oracle calls (left) and GPU hours (right). Surrogate screening dominates plain search on both axes across all… view at source ↗
Figure 21
Figure 21. Figure 21: Tweedie vs. completion (inference-time search). Same beam/MCTS/DFKC families as above. Solid: completion (full denoising rollout to score). Dashed: Tweedie posterior-mean estimate xˆ0 = E[x0 | xt] in place of the rollout. Seed-averaged (n=5). The Tweedie estimator’s noise in the discrete setting is reflected directly in score: every Tweedie variant underperforms its completion counterpart, and the gap is … view at source ↗
Figure 22
Figure 22. Figure 22: Compute breakdown for inference-time search, completion vs. Tweedie. Stacked bars are seed-averaged (n=5) wall-clock spend at the 2-hour FA budget; oracle compute is estimated as ncalls × cFA with cFA ≈ 50 ms (matched to the per-phase log on the active runs). Tweedie shifts the per-family balance but does not recover enough wall time to compensate for the score loss in [PITH_FULL_IMAGE:figures/full_fig_p… view at source ↗
Figure 23
Figure 23. Figure 23: Online vs. offline DDPP-LB on top-1 best FA score so far, against cumulative oracle calls (left) and GPU hours (right). Both runs use the same DDPP-LB loss and no reward shaping; the only difference is whether the buffer of (x, r(x)) pairs is collected up front and trained on once (offline) or refreshed continuously by the current qθ between gradient steps (online). Seed-averaged (n=5) bands. Online domin… view at source ↗
Figure 24
Figure 24. Figure 24: Tweedie vs. completion (active hybrids). Same Tweedie/completion comparison but applied inside the online DDPP-LB hybrids of Sec. 4 (beam and MCTS only; DFKC is not paired here). Solid: completion. Dashed: Tweedie. Seed-averaged (n=5). The Tweedie variants underperform their completion counterparts here as well, despite the saved generation passes nominally freeing time for fine-tuning [PITH_FULL_IMAGE:f… view at source ↗
Figure 25
Figure 25. Figure 25: Compute breakdown under a 2-hour wall budget for plain online finetuning (left) versus [PITH_FULL_IMAGE:figures/full_fig_p027_25.png] view at source ↗
Figure 26
Figure 26. Figure 26: Compute breakdown for active hybrids, completion vs. Tweedie. Per-phase wall times pulled from active_loop_log.jsonl; seed-averaged (n=5). Generation dominates for both completion and Tweedie variants, leaving very little wall time for fine-tuning (≤ 0.1 hr) regardless of whether the search uses the Tweedie shortcut or completion. The fine-tuning budget reduction observed for completion-based hybrids in … view at source ↗
Figure 27
Figure 27. Figure 27: Retroactive oracle-cost rescaling on online search hybrids. Same three online methods (ddpp_online_cvar vs. ddpp_beam_cvar vs. ddpp_mcts_cvar) plotted on the GPU-hours axis with the per-call oracle cost rescaled by 1/100 (left) and by 100 (right). Rescaling shifts the GPU￾hours axis by four orders of magnitude in either direction without altering the relative ordering: plain online DDPP-LB matches or exce… view at source ↗
Figure 28
Figure 28. Figure 28: Compute breakdown. Per-method decomposition of wall-clock spend at standard FA oracle cost (left, 1×) and under a retroactive 100× oracle-cost rescaling (right). Pure search spends almost all of its budget on oracle calls, so its absolute compute scales nearly linearly with coracle. Online finetuning amortizes generation and gradient updates across the same oracle budget, so its share of oracle compute is… view at source ↗
Figure 29
Figure 29. Figure 29: CVaR ablation, secondary metrics. Same four cells as [PITH_FULL_IMAGE:figures/full_fig_p029_29.png] view at source ↗
Figure 30
Figure 30. Figure 30: Model-debiasing ablation, secondary metrics. Same two cells as [PITH_FULL_IMAGE:figures/full_fig_p030_30.png] view at source ↗
Figure 31
Figure 31. Figure 31: Replay-buffer ablation, secondary metrics. Same four cells as [PITH_FULL_IMAGE:figures/full_fig_p030_31.png] view at source ↗
Figure 32
Figure 32. Figure 32: Anchor-NLL distribution per leave-one-out ablation cell. Histogram of per-token anchor NLL − log ppre(x0 | xt) at run end, computed over all valid samples (n in title). Top-left is the full method; the other five panels each remove one knob and overlay the resulting distribution against the full-method reference (blue). Removing the density entropy regularization / debiasing knob (bottom-left) collapses m… view at source ↗
Figure 33
Figure 33. Figure 33: CVaR ablation. Four cells on DDPP-LB: solid = CVaR-on, dotted = CVaR-off; green = offline, teal = online. Seed-averaged (n=5) with ±1 std bands. CVaR lifts all three metrics in both regimes; the gain on top-10% mean (bottom row) for the online + CVaR cell — which continues to climb to ∼ 0.5 rather than saturating — confirms that the CVaR-shaped gradient shifts the policy distribution itself rather than on… view at source ↗
Figure 34
Figure 34. Figure 34: Model-debiasing ablation. Two cells, both online DDPP-LB with CVaR off: solid (with debiasing) vs. dotted (without). Seed-averaged (n=5). Holding CVaR off isolates the debiasing knob from the upper-tail reweighting in [PITH_FULL_IMAGE:figures/full_fig_p033_34.png] view at source ↗
Figure 35
Figure 35. Figure 35: Best-so-far oracle curves: cheap estimator vs. LLaDA-Alg. 3 variants (FA/2VT4, 2 hr wall). Each subplot shows best FA score per oracle call (left) and per GPU hour (right) for five density-entropy-regularization variants: single-pass cheap estimator (black) and LLaDA-Alg. 3 with nmc ∈ {1, 4, 8, 16} (colored). Top: DDPP-LB. Bottom: VIDD. The cheap estimator matches or exceeds LLaDA variants on oracle calls… view at source ↗
Figure 36
Figure 36. Figure 36: NLL under the pretrained model for the top-200 discovered molecules, per estimator (FA/2VT4). NLL evaluated post-hoc via LLaDA-Alg. 3 (nmc=32); dashed line = median. Top: DDPP-LB. Bottom: VIDD. The cheap estimator produces distributions shifted at least as far off the prior as the LLaDA variants, validating that the single-pass approximation achieves the intended exploration pressure. 34 [PITH_FULL_IMAGE… view at source ↗
Figure 37
Figure 37. Figure 37: Replay-buffer ablation. Two cells, both online DDPP-LB+CVaR: solid (with replay buffer) vs. dashed (no buffer). Seed-averaged (n=5). Top-1 separates the cells modestly. The clearest signal is in top-10% mean (bottom row), where the buffer-on cell continues to climb past −0.4 while the buffer-off cell plateaus near −0.7. The buffer is shifting where the policy puts mass over the run, not merely re-presenti… view at source ↗
Figure 38
Figure 38. Figure 38: Thompson-sampling ablation. Pink: Thompson on top of online DDPP-LB+CVaR. Cyan: same configuration with deterministic top-K selection by ensemble mean instead of Thompson draws. Seed-averaged (n=5). Top-1 is comparable; top-10% mean (bottom row) is where the contrast is sharpest — Thompson reaches ∼ −0.05 vs. ∼ −0.4 for deterministic selection, a clear distribution shift toward higher-reward modes that co… view at source ↗
Figure 39
Figure 39. Figure 39: Invalid-SMILES penalty sweep rinv ∈ {0, −1, −3, −5, −7}. On top of Approach A + Thompson. Light-to-dark purple = increasing penalty magnitude. Top-1 and top-10 narrow the gap between cells over time; top-10% mean separates them most clearly, with smaller penalties (−1, −3) reaching slightly higher upper-tail FA but at the cost of validity ( [PITH_FULL_IMAGE:figures/full_fig_p037_39.png] view at source ↗
Figure 40
Figure 40. Figure 40: Validity panel paired with top-1 FA for the same penalty sweep as Fig. 39, used in the [PITH_FULL_IMAGE:figures/full_fig_p037_40.png] view at source ↗
Figure 41
Figure 41. Figure 41: Multi-knob ablation: lift-fraction trajectories. Normalized reward (1 = FULL final) vs. cumulative oracle calls, averaged across the six FA-oracle small-molecule targets and five seeds; pair removals dashed, triple removal dotted. Removing additional exploration / shaping knobs degrades lift roughly monotonically the triple sits lowest at every budget on DDPP-LB, with VIDD showing the same ordering at com… view at source ↗
Figure 42
Figure 42. Figure 42: Multi-knob ablation: normalized reward at oracle-call slices. Same normalization as [PITH_FULL_IMAGE:figures/full_fig_p038_42.png] view at source ↗
Figure 43
Figure 43. Figure 43: Anchor-NLL distribution per multi-knob ablation cell (DDPP-LB, 2VT4). Per-token anchor NLL − log ppre(x0 | xt) under the frozen pretrained generator on all oracle-evaluated SMILES (5 seeds), rough estimator with T = 10 random-mask MC averaging; FULL (steelblue) overlaid as reference. Every cell with Density Entropy Regularization removed collapses mass back toward the prior, identifying DE as the dominant… view at source ↗
Figure 44
Figure 44. Figure 44: VIDD harness ablation across all 9 tasks. Mirror of [PITH_FULL_IMAGE:figures/full_fig_p040_44.png] view at source ↗
Figure 45
Figure 45. Figure 45: Per-task ablation, small molecules (set 1: 2VT4, 5SDV, 6CM4). Rows are FA-oracle targets; columns are DDPP-LB ablations (left), VIDD ablations (centre), search baselines (right). The VIDD column is consistently more compressed than DDPP-LB on the same target, matching the cross-finetuner observation in [PITH_FULL_IMAGE:figures/full_fig_p041_45.png] view at source ↗
Figure 46
Figure 46. Figure 46: Per-task ablation, small molecules (set 2: 7BKC, 7C7M, 7YLL). Continuation of [PITH_FULL_IMAGE:figures/full_fig_p042_46.png] view at source ↗
Figure 47
Figure 47. Figure 47: Per-task ablation, protein-fitness targets (GB1, CreiLOV, TrpB). Rows are protein￾fitness families; columns match [PITH_FULL_IMAGE:figures/full_fig_p042_47.png] view at source ↗
Figure 48
Figure 48. Figure 48: DDPP-LB harness ablation across all 9 tasks. Normalized reward (1 = full DDPP-LB on that target) for each leave-one-out ablation of the knobs in [PITH_FULL_IMAGE:figures/full_fig_p043_48.png] view at source ↗
Figure 49
Figure 49. Figure 49: Inference-time search-only baselines across all 9 tasks. Same lift-fraction normalization as [PITH_FULL_IMAGE:figures/full_fig_p044_49.png] view at source ↗
Figure 50
Figure 50. Figure 50: Aggregated ablation grid, oracle-call view. 2 × 3 grid: row 1 averages over all six FA-oracle small molecules, row 2 averages over the three protein-fitness families; columns are DDPP-LB ablations, VIDD ablations, and search-only baselines. Same lift-fraction normalization as [PITH_FULL_IMAGE:figures/full_fig_p045_50.png] view at source ↗
Figure 51
Figure 51. Figure 51: Aggregated ablation grid, GPU-hours view. GPU-hours twin of [PITH_FULL_IMAGE:figures/full_fig_p045_51.png] view at source ↗
Figure 52
Figure 52. Figure 52: DDPP-LB harness ablation, GPU-hours view. GPU-hours twin of [PITH_FULL_IMAGE:figures/full_fig_p046_52.png] view at source ↗
Figure 53
Figure 53. Figure 53: Search-only baselines, GPU-hours view. GPU-hours twin of [PITH_FULL_IMAGE:figures/full_fig_p047_53.png] view at source ↗
Figure 54
Figure 54. Figure 54: VIDD harness ablation, GPU-hours view. GPU-hours twin of [PITH_FULL_IMAGE:figures/full_fig_p048_54.png] view at source ↗
Figure 55
Figure 55. Figure 55: Per-task ablation (mol set 1), GPU-hours view. GPU-hour twin of [PITH_FULL_IMAGE:figures/full_fig_p048_55.png] view at source ↗
Figure 56
Figure 56. Figure 56: Per-task ablation (mol set 2), GPU-hours view. GPU-hour twin of [PITH_FULL_IMAGE:figures/full_fig_p049_56.png] view at source ↗
Figure 57
Figure 57. Figure 57: Per-task ablation (proteins), GPU-hours view. GPU-hour twin of [PITH_FULL_IMAGE:figures/full_fig_p049_57.png] view at source ↗
Figure 58
Figure 58. Figure 58: DDPP-LB and VIDD harness ablation, lift-fraction bar version. Reproduces [PITH_FULL_IMAGE:figures/full_fig_p050_58.png] view at source ↗
Figure 59
Figure 59. Figure 59: Full method vs. inference-time search baselines, lift-fraction bar version. Reproduces [PITH_FULL_IMAGE:figures/full_fig_p050_59.png] view at source ↗
Figure 60
Figure 60. Figure 60: Online vs. offline DDPP-LB — Boltz-2. Companion to [PITH_FULL_IMAGE:figures/full_fig_p051_60.png] view at source ↗
Figure 61
Figure 61. Figure 61: CVaR ablation — Boltz-2. Companion to [PITH_FULL_IMAGE:figures/full_fig_p051_61.png] view at source ↗
Figure 62
Figure 62. Figure 62: Replay-buffer ablation — Boltz-2. Companion to [PITH_FULL_IMAGE:figures/full_fig_p051_62.png] view at source ↗
Figure 63
Figure 63. Figure 63: Thompson-sampling ablation — Boltz-2. Companion to [PITH_FULL_IMAGE:figures/full_fig_p052_63.png] view at source ↗
Figure 64
Figure 64. Figure 64: Surrogate screening on inference-time search — Boltz-2. Companion to [PITH_FULL_IMAGE:figures/full_fig_p052_64.png] view at source ↗
Figure 65
Figure 65. Figure 65: Online search hybrids at standard Boltz-2 oracle cost. Plain online DDPP-LB+CVaR vs. DDPP+beam+CVaR vs. DDPP+MCTS+CVaR. Plain online matches or exceeds both search￾augmented variants on top-1 Boltz-2, mirroring the FA conclusion in Sec. 4. 52 [PITH_FULL_IMAGE:figures/full_fig_p052_65.png] view at source ↗
Figure 66
Figure 66. Figure 66: Retroactive Boltz-2 oracle-cost rescaling. Companion to [PITH_FULL_IMAGE:figures/full_fig_p053_66.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 10 canonical work pages · 1 internal anchor

  1. [1]

    GenMol: A Drug Discovery Generalist with Discrete Diffusion, January 2025

    Seul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu, Danny Reidenbach, Yuxing Peng, Saee Paliwal, Weili Nie, and Arash Vahdat. GenMol: A Drug Discovery Generalist with Discrete Diffusion, January 2025. URL http://arxiv.org/abs/2501.06158. arXiv:2501.06158 [cs]

  2. [2]

    de Almeida, Alexander Rush, Thomas Pierrot, and V olodymyr Kuleshov

    Yair Schiff, Subham Sekhar Sahoo, Hao Phung, Guanghan Wang, Sam Boshar, Hugo Dalla- torre, Bernardo P. de Almeida, Alexander Rush, Thomas Pierrot, and V olodymyr Kuleshov. Simple Guidance Mechanisms for Discrete Diffusion Models. InThirteenth International Conference on Learning Representations, 2025. doi: 10.48550/arXiv.2412.10193. URL http://arxiv.org/a...

  3. [3]

    Diffusion Language Models Are Versatile Protein Learners

    Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, and Quanquan Gu. Diffusion Language Models Are Versatile Protein Learners. InProceedings of the 41st In- ternational Conference on Machine Learning, pages 52309 – 52333, February 2024. doi: 10.1126/sciadv.adj3786. URL http://arxiv.org/abs/2402.18567. arXiv:2402.18567 [cs, q-bio]

  4. [4]

    FlashAffinity: Bridging the Accuracy- Speed Gap in Protein-Ligand Binding Affinity Prediction.bioRxiv, 2025

    Songlin Jiang, Yifan Chen, Ze Cao, and Wengong Jin. FlashAffinity: Bridging the Accuracy- Speed Gap in Protein-Ligand Binding Affinity Prediction.bioRxiv, 2025. doi: 10.64898/ 2025.12.22.695983. URL https://www.biorxiv.org/content/10.64898/2025.12.22. 695983v1. bioRxiv preprint, December 2025

  5. [5]

    Boltz-2: Towards accurate and efficient binding affinity prediction.bioRxiv, 2025

    Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, David Kwabi-Addo, Dominique Beaini, Tommi Jaakkola, and Regina Barzilay. Boltz-2: Towards accurate and efficient binding affinity prediction.bioRxiv, 2025. doi: 10.1101/2025.06.14.659707. URL https://www. bi...

  6. [6]

    Cambridge University Press, 2020

    Tor Lattimore and Csaba Szepesvári.Bandit Algorithms. Cambridge University Press, 2020. doi: 10.1017/9781108571401

  7. [7]

    Con- strained molecular generation via sequential flow model fine-tuning

    Sven Gutjahr, Riccardo De Santi, Luca Schaufelberger, Kjell Jorner, and Andreas Krause. Con- strained molecular generation via sequential flow model fine-tuning. InICML 2025 Generative AI and Biology (GenBio) Workshop, 2025

  8. [8]

    How artificial intelligence is reengineering protein engi- neering.Science, April 2026

    Jennifer Listgarten and Hanlun Jiang. How artificial intelligence is reengineering protein engi- neering.Science, April 2026. doi: 10.1126/science.aec8444. URL https://www.science. org/doi/10.1126/science.aec8444

  9. [9]

    Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review, January 2025

    Masatoshi Uehara, Yulai Zhao, Chenyu Wang, Xiner Li, Aviv Regev, Sergey Levine, and Tommaso Biancalani. Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review, January 2025. URL http://arxiv.org/abs/2501.09685. arXiv:2501.09685 [cs]

  10. [10]

    Wittmann, Frances H

    Jason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo, Bruce J. Wittmann, Frances H. Arnold, and Yisong Yue. Steering Generative Models with Experimental Data for Protein Fitness Optimization. InAdvances in Neural Information Processing Systems, May 2025. doi: https: //doi.org/10.48550/arXiv.2505.15093. URLhttp://arxiv.org/abs/2505.15093

  11. [11]

    Guiding Generative Models for Protein Design: Prompting, Steering and Aligning, November 2025

    Filippo Stocco, Michele Garibbo, and Noelia Ferruz. Guiding Generative Models for Protein Design: Prompting, Steering and Aligning, November 2025. URL http://arxiv.org/abs/ 2511.21476. arXiv:2511.21476 [q-bio]

  12. [12]

    Tseng, Sergey Levine, and Tommaso Biancalani

    Masatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali, Gabriele Scalia, Nathaniel Lee Diamant, Alex M. Tseng, Sergey Levine, and Tommaso Biancalani. Feed- back Efficient Online Fine-Tuning of Diffusion Models, July 2024. URL http://arxiv.org/ abs/2402.16359. arXiv:2402.16359 [cs]

  13. [13]

    Steering Masked Discrete Diffusion Models via Discrete 10 Denoising Posterior Prediction, October 2024

    Jarrid Rector-Brooks, Mohsin Hasan, Zhangzhi Peng, Zachary Quinn, Chenghao Liu, Sarthak Mittal, Nouha Dziri, Michael Bronstein, Yoshua Bengio, Pranam Chatterjee, Alexander Tong, and Avishek Joey Bose. Steering Masked Discrete Diffusion Models via Discrete 10 Denoising Posterior Prediction, October 2024. URL http://arxiv.org/abs/2410.08134. arXiv:2410.08134 [cs]

  14. [14]

    Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design, August 2025

    Xingyu Su, Xiner Li, Masatoshi Uehara, Sunwoo Kim, Yulai Zhao, Gabriele Scalia, Ehsan Hajiramezanali, Tommaso Biancalani, Degui Zhi, and Shuiwang Ji. Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design, August 2025. URL http://arxiv.org/abs/2507.00445. arXiv:2507.00445 [cs]

  15. [15]

    Reward-Guided Iterative Refinement in Diffusion Models at Test- Time with Applications to Protein and DNA Design, February 2025

    Masatoshi Uehara, Xingyu Su, Yulai Zhao, Xiner Li, Aviv Regev, Shuiwang Ji, Sergey Levine, and Tommaso Biancalani. Reward-Guided Iterative Refinement in Diffusion Models at Test- Time with Applications to Protein and DNA Design, February 2025. URL http://arxiv. org/abs/2502.14944. arXiv:2502.14944 [q-bio]

  16. [16]

    Guiding Generative Protein Language Models with Reinforcement Learning

    Filippo Stocco, Maria Artigues-Lleixà, Andrea Hunklinger, Talal Widatalla, Marc Güell, and Noelia Ferruz. Guiding Generative Protein Language Models with Reinforcement Learning. arXiv, 2025. doi: https://doi.org/10.48550/arXiv.2412.12979. URL http://arxiv.org/abs/ 2412.12979

  17. [17]

    CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation, April

    Yanting Li, Zhuoyang Jiang, Enyan Dai, Lei Wang, Wen-Cai Ye, and Li Liu. CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation, April

  18. [18]

    arXiv:2604.11483 [cs]

    URLhttp://arxiv.org/abs/2604.11483. arXiv:2604.11483 [cs]

  19. [19]

    Molecular De Novo Design through Deep Reinforcement Learning, August 2017

    Marcus Olivecrona, Thomas Blaschke, Ola Engkvist, and Hongming Chen. Molecular De Novo Design through Deep Reinforcement Learning, August 2017. URL http://arxiv.org/abs/ 1704.07555. arXiv:1704.07555 [cs]

  20. [20]

    Functional Alignment of Protein Language Models via Reinforcement Learning with Ex- perimental Feedback.bioRxiv, 2025

    Nathaniel Blalock, Srinath Seshadri, Agrim Babber, Sarah Fahlberg, and Philip Romero. Functional Alignment of Protein Language Models via Reinforcement Learning with Ex- perimental Feedback.bioRxiv, 2025. doi: https://doi.org/10.1101/2025.05.02.651993. URL https://www.biorxiv.org/content/10.1101/2025.05.02.651993

  21. [21]

    Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning.Advances in Neural Information Processing Systems (NeurIPS 2025), 2025

    Riccardo De Santi, Marin Vlastelica, and Ya-Ping Hsieh. Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning.Advances in Neural Information Processing Systems (NeurIPS 2025), 2025. doi: 10.48550/arXiv.2511.22640. URL http: //arxiv.org/abs/2511.22640. arXiv:2511.22640 [cs]

  22. [22]

    Efficient tail-aware generative optimization via flow model fine-tuning.arXiv preprint arXiv:2602.16796, 2026

    Zifan Wang, Riccardo De Santi, Xiaoyu Mo, Michael M Zavlanos, Andreas Krause, and Karl H Johansson. Efficient tail-aware generative optimization via flow model fine-tuning.arXiv preprint arXiv:2602.16796, 2026

  23. [23]

    Provable maximum entropy manifold exploration via diffusion models

    Riccardo De Santi*, Marin Vlastelica*, Ya-Ping Hsieh, Zebang Shen, Niao He, and Andreas Krause. Provable maximum entropy manifold exploration via diffusion models. InProc. International Conference on Machine Learning (ICML), June 2025

  24. [24]

    Verifier- Constrained Flow Expansion for Discovery Beyond the Data, February 2026

    Riccardo De Santi, Kimon Protopapas, Ya-Ping Hsieh, and Andreas Krause. Verifier- Constrained Flow Expansion for Discovery Beyond the Data, February 2026. URL http: //arxiv.org/abs/2602.15984. arXiv:2602.15984 [cs]

  25. [25]

    Thompson

    William R. Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples.Biometrika, 25(3-4):285–294, 1933. doi: 10.1093/biomet/25. 3-4.285

  26. [26]

    An empirical evaluation of Thompson sampling

    Olivier Chapelle and Lihong Li. An empirical evaluation of Thompson sampling. InAdvances in Neural Information Processing Systems 24 (NeurIPS), pages 2249–2257, 2011. URL https:// papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling

  27. [27]

    Rusu, Joel Veness, Marc G

    V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Pe- tersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level control through deep reinforcemen...

  28. [28]

    Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design, October 2024

    Chenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang, Tommaso Biancalani, Avantika Lal, Tommi Jaakkola, Sergey Levine, Hanchen Wang, and Aviv Regev. Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design, October 2024. URLhttp://arxiv.org/abs/2410.13643. arXiv:2410.13643 [cs]

  29. [29]

    Diamant, Alex M

    Masatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali, Gabriele Scalia, Nathaniel L. Diamant, Alex M. Tseng, Tommaso Biancalani, and Sergey Levine. Fine-tuning of continuous- time diffusion models as entropy-regularized control, 2024. URL https://arxiv.org/abs/ 2402.15194

  30. [30]

    A tutorial on Thompson sampling.F oundations and Trends in Machine Learning, 11(1):1–96, 2018

    Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen. A tutorial on Thompson sampling.F oundations and Trends in Machine Learning, 11(1):1–96, 2018. doi: 10.1561/2200000070. URLhttps://arxiv.org/abs/1707.02038

  31. [31]

    Tyrrell Rockafellar and Stanislav Uryasev

    R. Tyrrell Rockafellar and Stanislav Uryasev. Optimization of conditional value-at-risk.Journal of Risk, 2(3):21–41, 2000. doi: 10.21314/JOR.2000.038

  32. [32]

    Tyrrell Rockafellar and Stanislav Uryasev

    R. Tyrrell Rockafellar and Stanislav Uryasev. Conditional value-at-risk for general loss distribu- tions.Journal of Banking and Finance, 26(7):1443–1471, 2002. doi: 10.1016/S0378-4266(02) 00271-6

  33. [33]

    Large language diffusion models.arXiv preprint arXiv:2502.09992, 2025

    Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. Large language diffusion models.arXiv preprint arXiv:2502.09992, 2025

  34. [34]

    Prioritized experience replay

    Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. Prioritized experience replay. InInternational Conference on Learning Representations (ICLR), 2016

  35. [35]

    RDKit: Open-source cheminformatics software

    Greg Landrum and RDKit Contributors. RDKit: Open-source cheminformatics software. https://www.rdkit.org, 2024. Software release; canonical Zenodo DOI: https://doi. org/10.5281/zenodo.10099869

  36. [36]

    ODesign: A World Model for Biomolecular Interaction Design, October 2025

    Odin Zhang, Xujun Zhang, Haitao Lin, Cheng Tan, Qinghan Wang, Yuanle Mo, Qiantai Feng, Gang Du, Yuntao Yu, Zichang Jin, Ziyi You, Peicong Lin, Yijie Zhang, Yuyang Tao, Shicheng Chen, Jack Xiaoyu Chen, Chenqing Hua, Weibo Zhao, Runze Ma, Yunpeng Xia, Kejun Ying, Jun Li, Yundian Zeng, Lijun Lang, Peichen Pan, Hanqun Cao, Zihao Song, Bo Qiang, Jiaqi Wang, Pe...

  37. [37]

    Tadmor, Richard G

    Cheng Zeng, Jirui Jin, Connor Ambrose, George Karypis, Mark Transtrum, Ellad B. Tadmor, Richard G. Hennig, Adrian Roitberg, Stefano Martiniani, and Mingjie Liu. PropMolFlow: property-guided molecule generation with geometry-complete flow matching.Nature Computa- tional Science, pages 1–10, January 2026. ISSN 2662-8457. doi: 10.1038/s43588-025-00946-y. URL...

  38. [38]

    Structure-based Drug Design with Equivariant Diffusion Models, October 2022

    Arne Schneuing, Yuanqi Du, Charles Harris, Arian Jamasb, Ilia Igashov, Weitao Du, Tom Blundell, Pietro Lió, Carla Gomes, Max Welling, Michael Bronstein, and Bruno Correia. Structure-based Drug Design with Equivariant Diffusion Models, October 2022. URL http: //arxiv.org/abs/2210.13695. arXiv:2210.13695 [cs, q-bio]

  39. [39]

    3D Equivariant Diffusion for Target-Aware Molecule Generation and Affinity Prediction, March

    Jiaqi Guan, Wesley Wei Qian, Xingang Peng, Yufeng Su, Jian Peng, and Jianzhu Ma. 3D Equivariant Diffusion for Target-Aware Molecule Generation and Affinity Prediction, March

  40. [40]

    arXiv:2303.03543 [q-bio]

    URLhttp://arxiv.org/abs/2303.03543. arXiv:2303.03543 [q-bio]

  41. [41]

    Batey, Mike Tyers, Michał Koziarski, and Cheng-Hao Liu

    Andrei Rekesh, Miruna Cretu, Dmytro Shevchuk, Vignesh Ram Somnath, Pietro Liò, Robert A. Batey, Mike Tyers, Michał Koziarski, and Cheng-Hao Liu. SynCoGen: Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling, July 2025. URL http: //arxiv.org/abs/2507.11818. arXiv:2507.11818 [cs]

  42. [42]

    Keir Adams, Kento Abeywardane, Jenna Fromer, and Connor W. Coley. ShEPhERD: Diffusing shape, electrostatics, and pharmacophores for bioisosteric drug design, March 2025. URL http://arxiv.org/abs/2411.04130. arXiv:2411.04130 [q-bio]. 12

  43. [44]

    Why is tanimoto index an appropriate choice for fingerprint-based similarity calculations?Journal of Cheminformatics, 7:20, 2015

    Dávid Bajusz, Anita Rácz, and Károly Héberger. Why is tanimoto index an appropriate choice for fingerprint-based similarity calculations?Journal of Cheminformatics, 7:20, 2015. doi: 10.1186/s13321-015-0069-3

  44. [45]

    Richard Bickerton, Gaia V

    G. Richard Bickerton, Gaia V . Paolini, Jeremy Besnard, Sorel Muresan, and Andrew L. Hopkins. Quantifying the chemical beauty of drugs.Nature Chemistry, 4(2):90–98, 2012. doi: 10.1038/ nchem.1243

  45. [46]

    Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions.Journal of Cheminfor- matics, 1:8, 2009

    Peter Ertl and Ansgar Schuffenhauer. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions.Journal of Cheminfor- matics, 1:8, 2009. doi: 10.1186/1758-2946-1-8

  46. [47]

    Test-time scaling of diffusion models via noise trajectory search, 2025

    Vinith Ramesh et al. Test-time scaling of diffusion models via noise trajectory search, 2025. URLhttps://arxiv.org/abs/2506.03164

  47. [48]

    TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion, September 2025

    Sophia Tang, Yuchen Zhu, Molei Tao, and Pranam Chatterjee. TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion, September 2025. URL http://arxiv. org/abs/2509.25171. arXiv:2509.25171 [cs]

  48. [49]

    Discrete Feynman-Kac Correctors, January

    Mohsin Hasan, Viktor Ohanesian, Artem Gazizov, Yoshua Bengio, Alán Aspuru-Guzik, Roberto Bondesan, Marta Skreta, and Kirill Neklyudov. Discrete Feynman-Kac Correctors, January

  49. [50]

    arXiv:2601.10403 [cs]

    URLhttp://arxiv.org/abs/2601.10403. arXiv:2601.10403 [cs]

  50. [51]

    Feynman-Kac Correctors in Diffusion: Annealing, Guidance, and Product of Experts, June 2025

    Marta Skreta, Tara Akhound-Sadegh, Viktor Ohanesian, Roberto Bondesan, Alán Aspuru- Guzik, Arnaud Doucet, Rob Brekelmans, Alexander Tong, and Kirill Neklyudov. Feynman-Kac Correctors in Diffusion: Annealing, Guidance, and Product of Experts, June 2025. URL http://arxiv.org/abs/2503.02819. arXiv:2503.02819 [cs]

  51. [52]

    Bronstein, Martin Steinegger, Emine Kucukbenli, Arash Vahdat, and Karsten Kreis

    Kieran Didi, Zuobai Zhang, Guoqing Zhou, Danny Reidenbach, Zhonglin Cao, Sooyoung Cha, Tomas Geffner, Christian Dallago, Jian Tang, Michael M. Bronstein, Martin Steinegger, Emine Kucukbenli, Arash Vahdat, and Karsten Kreis. Scaling atomistic protein binder design with generative pretraining and test-time compute. InThe F ourteenth International Conference...

  52. [53]

    A General Framework for Inference-time Scaling and Steering of Diffu- sion Models, January 2025

    Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, and Rajesh Ranganath. A General Framework for Inference-time Scaling and Steering of Diffu- sion Models, January 2025. URL http://arxiv.org/abs/2501.06848. arXiv:2501.06848 [cs]

  53. [54]

    Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding, October 2024

    Xiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia, Gokcen Eraslan, Surag Nair, Tommaso Biancalani, Aviv Regev, Sergey Levine, and Masatoshi Uehara. Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding, October 2024. URLhttp://arxiv.org/abs/2408.08252. arXiv:2408.08252 [cs]

  54. [55]

    Unlocking Guidance for Discrete State-Space Diffusion and Flow Models

    Hunter Nisonoff, Junhao Xiong, Stephan Allenspach, and Jennifer Listgarten. Unlocking Guidance for Discrete State-Space Diffusion and Flow Models. In13th International Conference on Learning Representations, 2025. doi: https://doi.org/10.48550/arXiv.2406.01572. URL http://arxiv.org/abs/2406.01572. arXiv:2406.01572 [cs]

  55. [56]

    Frey, Tim G

    Nate Gruver, Samuel Stanton, Nathan C. Frey, Tim G. J. Rudner, Isidro Hotzel, Julien Lafrance- Vanasse, Arvind Rajpal, Kyunghyun Cho, and Andrew Gordon Wilson. Protein Design with Guided Discrete Diffusion. InAdvances in Neural Information Processing Systems 36, May

  56. [57]

    URL http://arxiv.org/abs/2305

    doi: https://doi.org/10.48550/arXiv.2305.20009. URL http://arxiv.org/abs/2305. 20009. arXiv:2305.20009 [cs, q-bio]

  57. [58]

    Oltrogge, David F

    Junhao Xiong, Ishan Gaur, Maria Lukarska, Hunter Nisonoff, Luke M. Oltrogge, David F. Savage, and Jennifer Listgarten. ProteinGuide: On-the-fly property guidance for protein sequence generative models, January 2026. URL http://arxiv.org/abs/2505.04823. arXiv:2505.04823 [cs]. 13

  58. [59]

    Gao, and Wing Hung Wong

    Wenhui Sophia Lu, Xiaowei Zhang, Luis Santiago Mille-Fragoso, Haoyu Dai, Xiaojing J. Gao, and Wing Hung Wong. ProV ADA: Generation of Subcellular Protein Variants via Ensemble- Guided Test-Time Steering, July 2025. URL https://www.biorxiv.org/content/10. 1101/2025.07.11.664238v1. ISSN: 2692-8205 Pages: 2025.07.11.664238 Section: New Results

  59. [60]

    Dynamic Search for Inference-Time Alignment in Diffusion Models, June 2025

    Xiner Li, Masatoshi Uehara, Xingyu Su, Gabriele Scalia, Tommaso Biancalani, Aviv Regev, Sergey Levine, and Shuiwang Ji. Dynamic Search for Inference-Time Alignment in Diffusion Models, June 2025. URLhttp://arxiv.org/abs/2503.02039. arXiv:2503.02039 [cs]

  60. [61]

    Krishnapriyan

    Yue Jian, Curtis Wu, Danny Reidenbach, and Aditi S. Krishnapriyan. General Binding Affinity Guidance for Diffusion Models in Structure-Based Drug Design.Journal of Chemical Infor- mation and Modeling, 66(3):1342–1352, February 2026. ISSN 1549-9596, 1549-960X. doi: 10.1021/acs.jcim.5c01166. URL http://arxiv.org/abs/2406.16821. arXiv:2406.16821 [cs]

  61. [62]

    Hit and Lead Discovery with Explorative RL and Fragment-based Molecule Generation

    Soojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu, and Sung Ju Hwang. Hit and Lead Discovery with Explorative RL and Fragment-based Molecule Generation.Advances in Neural Information Processing Systems 34 (NeurIPS 2021), 2021. doi: 10.48550/arXiv.2110.01219. URLhttp://arxiv.org/abs/2110.01219. arXiv:2110.01219 [cs]

  62. [63]

    PepTune: De Novo Generation of Ther- apeutic Peptides with Multi-Objective-Guided Discrete Diffusion, December 2024

    Sophia Tang, Yinuo Zhang, and Pranam Chatterjee. PepTune: De Novo Generation of Ther- apeutic Peptides with Multi-Objective-Guided Discrete Diffusion, December 2024. URL http://arxiv.org/abs/2412.17780. arXiv:2412.17780 [q-bio]

  63. [64]

    Emmanuel Noutahi, Cristian Gabellini, Michael Craig, Jonathan S. C. Lim, and Prudencio Tossou. Gotta be SAFE: A New Framework for Molecular Design, December 2023. URL http://arxiv.org/abs/2310.10773. arXiv:2310.10773 [cs]

  64. [65]

    Rush, Yair Schiff, Justin T

    Subham Sekhar Sahoo, Marianne Arriola, Aaron Gokaslan, Edgar Mariano Marroquin, Alexan- der M. Rush, Yair Schiff, Justin T. Chiu, and V olodymyr Kuleshov. Simple and effective masked diffusion language models. InAdvances in Neural Information Processing Systems (NeurIPS),

  65. [66]

    URLhttps://arxiv.org/abs/2406.07524

  66. [67]

    Tweedie’s formula and selection bias.Journal of the American Statistical Association, 106(496):1602–1614, 2011

    Bradley Efron. Tweedie’s formula and selection bias.Journal of the American Statistical Association, 106(496):1602–1614, 2011. doi: 10.1198/jasa.2011.tm11181

  67. [68]

    McCann, Marc L

    Hyungjin Chung, Jeongsol Kim, Michael T. McCann, Marc L. Klasky, and Jong Chul Ye. Diffusion posterior sampling for general noisy inverse problems. InInternational Conference on Learning Representations (ICLR), 2023. URLhttps://arxiv.org/abs/2209.14687. 14 A Declaration of LLM Usage LLMs were used to write code, interpret results, assist with writing, and...