{"id":"adf20bdf-6de1-4f78-a9d2-ce570a8956c8","arxiv_id":"2607.06421","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":6,"one_line_summary":"GB-FESO backpropagates a KL-divergence loss through a frozen conditional diffusion model's sampling trajectory to optimize system parameters so the generated ensemble matches a target free-energy surface.","lead":"This paper introduces GB-FESO, a method that optimizes the input parameters of a frozen diffusion model to make its generated molecular ensembles match a target free-energy landscape. It matters because designing molecules for specific thermodynamic behaviors—rather than just static structures—could enable better catalysts, drugs, and materials.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Optimization success is evaluated entirely within the diffusion surrogate; no ground-truth MD validation of optimized conditions is performed, leaving the extrapolation claim untested.","rationale":"The reader correctly identifies initialization dependence and the rough loss landscape as practical concerns, and the CONDITIONAL verdict is appropriate. However, the reader's weakest_assumption focuses on optimization navigability rather than the more fundamental issue: the evaluation metric itself is surrogate-internal. The paper never validates that optimized conditions reproduce target FES under ground-truth MD. This is the single most load-bearing gap because it undermines the central claim's empirical foundation — 'successfully optimizes the interaction parameters to reproduce target free-energy landscapes' is only demonstrated within the surrogate, not against ground truth. The Gaussian proof-of-concept (§III.A) is exempt from this concern because the surrogate is trivially exact for Gaussian sampling, but the toy peptide results (§III.B) are not. The paper is honest about being a 'first step' and provides code/data, which is commendable. The concern is addressable with a modest computational cost (a handful of LAMMPS runs at optimized ε values), and until that check is performed, the extrapolation claim in particular should be treated as unsupported. The CONDITIONAL verdict should remain, with the added requirement of ground-truth MD validation of at least a subset of optimized conditions.","tokens_in":14024,"tokens_out":2187,"duration_ms":205069,"concrete_test":"Select 5–10 optimized ε pairs from successful runs (including at least 3 where the optimized ε lies outside the training domain). For each, run independent LAMMPS MD simulations at those ε values, compute the ground-truth FES, and compare to the target FES using the same KL divergence metric. If the ground-truth KL divergences are within a factor of 2 of the in-surrogate values, the surrogate is faithful and the method works end-to-end. If they diverge systematically (especially for extrapolated conditions), the surrogate introduces systematic errors that the current evaluation masks.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that GB-FESO optimizes interaction parameters to reproduce target free-energy landscapes. However, both the optimization objective and the success criterion operate entirely through the frozen diffusion model: the KL divergence is computed between diffusion-generated ensembles and target ensembles (the latter also originally from MD, but the comparison is surrogate-to-target, not ground-truth-to-target). The paper explicitly states (§V.C) that success is defined by 'agreement between the generated and target distributions' — i.e., the surrogate's output — not by running MD at the optimized ε values and checking the resulting FES. This creates a circularity risk: if the diffusion model's conditional distribution p_θ(x|ε) deviates from the true p(x|ε) at the optimized ε values (particularly for out-of-training-domain conditions, which the paper claims to handle), the in-surrogate KL loss could be low while the true FES at those ε values differs substantially from the target. Figure 3b shows the model reproduces distributions for two specific unseen ε pairs, but this is a spot-check, not a systematic validation at the optimized conditions. The extrapolation claim in the conclusion ('targets whose conditioning parameters lie outside the training region') is especially vulnerable because the surrogate's fidelity at extrapolated ε values is precisely where errors would accumulate, and no ground-truth check exists to catch them. The reader's concern about initialization dependence (25% failure from high-high starts) is real but secondary — it addresses optimization robustness, not whether the success metric itself is valid.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The manuscript introduces Gradient-Based Free Energy Surface Optimization (GB-FESO), a framework in which a trained, frozen conditional diffusion model serves as a differentiable surrogate for an ensemble distribution. The conditioning variables (e.g., Lennard-Jones interaction parameters) are optimized via backpropagation of a KDE-based KL-divergence loss through a deterministic DDIM sampling trajectory so that the generated ensemble matches a prescribed target free-energy surface. The method is validated on a 1D Gaussian proof-of-concept (both continuous and Gumbel-relaxed discrete conditions) and on a 4-particle Lennard-Jones toy peptide with three metastable states, including optimization in both full internal-coordinate (5D) and reduced HLDA collective-variable (2D) representations. Code and data are publicly available.","tokens_in":14169,"tokens_out":2065,"duration_ms":116545,"significance":"The idea of treating a frozen conditional diffusion model as a differentiable surrogate for an ensemble distribution—and optimizing physical design variables through it—is novel and conceptually appealing. The Gaussian proof-of-concept is clean and well-controlled, and the demonstration that discrete conditioning variables can be handled via Gumbel relaxation is a useful contribution. The release of reproducible code and LAMMPS simulation scripts is a strength. The framework, if validated, would be a meaningful step toward ensemble-level inverse design.","major_comments":[{"comment":"§V.C and Conclusion: The central claim—that GB-FESO optimizes interaction parameters to 'reproduce target free-energy landscapes'—is evaluated entirely within the diffusion surrogate. Success is defined as 'agreement between the generated and target distributions' (§V.C), i.e., the KL divergence between the diffusion-generated ensemble and the target ensemble. No ground-truth MD simulation is run at the optimized ε values to verify that the true FES at those parameters actually matches the target. Since the authors have LAMMPS available and used it to generate all training data, this validation would be straightforward to perform. Without it, the paper cannot distinguish between cases where the optimizer found ε values that genuinely reproduce the target FES and cases where the surrogate p_θ(x|ε) is simply inaccurate at the optimized ε, yielding a low in-surrogate KL but a poor true FES.","section":null},{"comment":"Conclusion: The paper claims that GB-FESO succeeds 'including targets whose conditioning parameters lie outside the training region of the diffusion model.' Figure 3b provides a spot-check of model fidelity at two unseen ε pairs, but this is not a systematic validation of surrogate accuracy at the optimized (potentially extrapolated) ε values. The extrapolation claim is especially vulnerable because surrogate errors are expected to be largest out-of-distribution, and no ground-truth check exists to detect them. This claim should either be supported by MD validation at extrapolated optimized conditions or substantially softened.","section":null},{"comment":"Table I: The initialization dependence is significant—runs starting from the high–high ε region fail to populate all target-active basins in 10/40 runs (25%), compared to 0/40 (2.5%) for low–low starts. The median ΔF error for high–high starts (0.89 k_BT in 5D) is roughly 4.5× larger than for low–low starts (0.20 k_BT). The paper attributes this to 'highly localized distributions associated with strong Lennard-Jones interactions' but does not analyze whether the failures reflect surrogate inaccuracy at large ε (where distributions are sharply peaked and harder to learn) or genuine multimodality of the KL loss landscape. This distinction is load-bearing: if the former, the method's reliability is bounded by surrogate quality; if the latter, it is a fundamental optimization challenge. A brief analysis (e.g., comparing surrogate-generated vs. ground-truth distributions at the failed ε)would","section":null}],"minor_comments":[{"comment":"§V.C: The two-stage optimization threshold (the KL value below which stage 2 is triggered) is not specified. Please state it explicitly.","section":null},{"comment":"§V.C: The stochastic perturbation scheme used in stage 1 is not described. What distribution, amplitude, and schedule are used?","section":null},{"comment":"§V.C: The KDE bandwidth h=0.5 for the toy peptide is large relative to the scale of the internal coordinates (bond distances ~1, angles in radians). No sensitivity analysis is provided. A brief study of how results depend on h would strengthen the paper.","section":null},{"comment":"Eq. (9): The reverse KL term weight λ is set to 0 for the toy peptide. The paper states 'exclusion of the symmetric term does not significantly affect performance,' but no quantitative comparison is given. A table or figure comparing λ=0 vs. λ>0 would help.","section":null},{"comment":"Figure 4: The axis labels and colorbar are difficult to read. Please increase font sizes and ensure all panels are clearly labeled.","section":null},{"comment":"§III.A: The claim that the Gaussian DDPM 'demonstrates the ability to extrapolate beyond the training domain' would be strengthened by reporting the optimized condition values and their distance from the nearest training condition.","section":null},{"comment":"Table I: The 'Avg. KL successful' column mixes 5D and 2D values in a way that is initially confusing. Consider separating into two tables or using clearer formatting.","section":null},{"comment":"§II.B, last paragraph: The sentence beginning 'One possible approach is binary logit relaxation[36]...' is missing a period.","section":null}],"recommendation":"major_revision","confidential_remarks":"The lack of ground-truth MD validation at optimized conditions is the key issue. The authors have LAMMPS and the simulation infrastructure readily available, so this should be feasible within a revision cycle. I would view the addition of even a small number of MD validations at optimized ε values (especially for extrapolated conditions) as substantially strengthening the paper. The initialization dependence is also worth investigating more carefully—it may reveal whether the surrogate or the optimizer is the binding constraint."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The core methodological contribution—using a frozen conditional diffusion model as a differentiable surrogate for ensemble-level inverse design—is correctly identified, and the referee's concerns about surrogate validation are well-taken. We address each major comment below.","responses":[{"response":"The referee is correct. This is the most important gap in the current manuscript. Success is currently defined entirely within the surrogate, and without ground-truth MD validation at the optimized ε values, we cannot rule out the scenario where the surrogate p_θ(x|ε) is inaccurate at the optimized conditions, producing a low KL divergence against the target while the true FES at those parameters is poor. We will address this by running LAMMPS simulations at the optimized ε values for a representative set of successful and failed optimization runs across all initialization regions (low–low, low–high, high–low, high–high), and comparing the ground-truth FES to both the surrogate-generated FES and the target FES. This will include both 5D and 2D (HLDA) optimization cases. We will report the ground-truth ΔF errors alongside the in-surrogate metrics in a revised Table I and add a new figure showing target vs. surrogate vs. ground-truth FES comparisons. This validation is straightforward given our existing LAMMPS infrastructure and will be incorporated into the revised manuscript.","revision_made":"yes","referee_comment":"No ground-truth MD simulation is run at the optimized ε values to verify that the true FES at those parameters actually matches the target. The paper cannot distinguish between genuine reproduction of the target FES and surrogate inaccuracy yielding low in-surrogate KL but poor true FES."},{"response":"The referee's concern is valid. The extrapolation claim in the Conclusion is currently supported only by the two-point spot-check in Figure 3b and by the Gaussian proof-of-concept, neither of which constitutes systematic validation of surrogate fidelity at extrapolated optimized conditions. We will take two corrective actions. First, as part of the MD validation described in our response to the first comment, we will explicitly identify which optimized ε values fall outside the training domain and report ground-truth FES comparisons for those cases specifically. This will allow us to either support or qualify the extrapolation claim with direct evidence. Second, if the ground-truth validation reveals that surrogate accuracy degrades at extrapolated conditions, we will soften the claim accordingly and discuss the limitation explicitly. We agree that the current wording is stronger than the evidence supports.","revision_made":"yes","referee_comment":"The extrapolation claim ('including targets whose conditioning parameters lie outside the training region') is not systematically validated. Figure 3b is only a spot-check at two unseen ε pairs. Surrogate errors are expected to be largest out-of-distribution, and no ground-truth check exists."},{"response":"This is a fair point. The current manuscript offers a plausible but untested explanation for the initialization dependence. We will perform the analysis the referee suggests: for the failed runs (primarily from high–high starts), we will compare the surrogate-generated distribution at the optimized ε against the ground-truth MD distribution at the same ε. If the surrogate and ground truth agree at the failed ε values but both differ from the target, this indicates genuine multimodality or local minima in the KL loss landscape—an optimization challenge. If the surrogate and ground truth disagree at the failed ε values, this indicates surrogate inaccuracy at large ε, where distributions are sharply peaked and harder to learn—a model quality limitation. We expect both effects may contribute, and the analysis will allow us to disentangle them. We will add this analysis as a dedicated subsection or paragraph in the revised manuscript and update the discussion in §V.C and the Conclusion accordingly.","revision_made":"yes","referee_comment":"The initialization dependence (25% failure from high–high starts vs. 2.5% from low–low) is attributed to 'highly localized distributions associated with strong Lennard-Jones interactions' but not analyzed. The distinction between surrogate inaccuracy at large ε and genuine multimodality of the KL loss landscape is load-bearing and should be analyzed."}],"tokens_in":13858,"tokens_out":1146,"duration_ms":76708,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Short version: the core idea is genuinely new and the paper ships working code. The main gap is that the stress-test concern lands — success is measured entirely within the diffusion surrogate, with no MD validation at the optimized epsilon values. That said, the methodological contribution is real and the paper is honest about being a first step. It deserves a serious referee who pushes the authors to close this loop before publication, but the framework itself is worth engaging with. The Gaussian proof-of-concept is clean and the discrete relaxation via Gumbel-Softmax is a nice touch showing the method handles categorical design variables, not just continuous ones. The toy peptide results show the method works in a physically motivated setting with multiple metastable states, and the HLDA-based collective variable reduction is a sensible scaling strategy. Code and data are on Zenodo, which is good practice. The stress-test concern is the real issue. Both the optimization objective and the success criterion operate through the frozen diffusion model. The paper defines success as agreement between generated and target distributions — both surrogate outputs — rather than running MD at the optimized epsilon values and checking the actual FES. Figure 3b shows the model reproduces distributions for two unseen epsilon pairs, but this is a spot-check, not systematic validation at optimized conditions. The extrapolation claim in the conclusion is especially vulnerable because surrogate fidelity at out-of-training-domain conditions is exactly where errors would accumulate, and nothing in the paper catches that. The reader's concern about initialization dependence (25% failure from high-high starts) is real but secondary. It tells us the loss landscape is rough, which the authors acknowledge. More concerning is that the two-stage optimization threshold values are not stated, and there is no hyperparameter sensitivity analysis. These are fixable issues. The central methodological claim — that a frozen conditional diffusion model can serve as a differentiable surrogate for ensemble-level inverse design — holds up as a proof of concept. The framework is not circular in the formal sense: the model is trained on external MD data and the condition variables are optimized against an independent target. The problem is that the evaluation never leaves the surrogate. Who is this for? Methodologists in molecular design and enhanced sampling who are interested in pushing inverse design from structural to thermodynamic objectives. The paper is a credible first step that needs one critical addition before it can be trusted: run MD at a few optimized epsilon values and show the true FES matches the target. Without that, the extrapolation claim is unsupported. Recommend serious peer review with that experiment as a required revision.","headline":"Novel differentiable-surrogate framework for inverse FES design; validation is entirely in-surrogate with no ground-truth MD check at optimized conditions.","tokens_in":14978,"tokens_out":585,"would_cite":false,"duration_ms":190358,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["02.50.-r","87.15.A-","87.15.hp"],"model":"glm-5.2","headline":"Frozen diffusion model backprops gradients to design target free-energy landscapes","keywords":[],"falsifier":"Apply GB-FESO to a molecular system with more than a handful of degrees of freedom and multiple metastable states. If the KDE-based KL loss landscape becomes so rough that gradient descent fails to converge regardless of initialization, or if the frozen diffusion model cannot generalize to the optimized conditioning values and produces physically meaningless ensembles, the method's utility would be confined to toy systems.","tokens_in":14106,"feed_emoji":"🧬","tokens_out":882,"duration_ms":139398,"temperature":0.7,"pith_summary":"The paper introduces GB-FESO (Gradient-Based Free Energy Surface Optimization), a method that repurposes a trained conditional diffusion model as a differentiable surrogate for a molecular ensemble distribution. After training, the model's neural network weights are frozen. The design variables that condition the model — such as Lennard-Jones interaction parameters — are then treated as optimization targets. A loss function measuring the Kullback-Leibler divergence between the diffusion model's generated ensemble and a prescribed target free-energy surface is backpropagated through the deterministic sampling trajectory to update those conditioning variables. The paper validates the approach on one-dimensional Gaussian distributions (both continuous and relaxed-discrete conditioning) and then on a four-particle Lennard-Jones toy peptide with three metastable states. In the peptide system, GB-FESO recovers target free-energy landscapes in the majority of test cases, including targets whose optimal parameters lie outside the model's training domain. The paper also demonstrates that optimization can be performed in a reduced collective-variable representation (two bond angles rather than the full five-dimensional internal-coordinate space) with only minor performance loss.","feed_headline":"Frozen diffusion model designs molecules to match target energy landscapes","feed_subtitle":"GB-FESO backpropagates through a diffusion sampler to tune interaction parameters so generated ensembles reproduce prescribed free-energysur","key_machinery":"Conditional diffusion model (DDPM) with deterministic DDIM sampling; KDE-based KL-divergence loss; FiLM-conditioned MLP backbone; two-stage optimization (high learning rate with stochastic perturbations, then low learning rate); HLDA collective-variable reduction","core_discovery":"The central discovery is that a frozen conditional diffusion model, originally trained to generate conformational ensembles given system parameters, can be inverted: by backpropagating a distribution-level KL-divergence loss through the deterministic sampling path back to the conditioning inputs, one can optimize those inputs so that the generated ensemble matches a target free-energy surface. The model thus functions as a differentiable map from design parameters to ensemble distributions, and gradient descent on that map recovers parameters reproducing desired thermodynamic landscapes. The paper shows this works on a physically motivated system — a Lennard-Jones four-bead peptide — and not","pith_inferences":[],"forward_implications":["If the method scales, molecular engineers could specify desired thermodynamic behavior (metastable state populations, barrier heights) and automatically recover interaction parameters or chemical compositions that produce it, rather than relying on trial-and-error simulation.","The approach extends diffusion-model-based molecular design beyond single-structure generation to ensemble-level design, where the target is a distribution rather than a conformation.","Reduced-representation optimization via collective variables (demonstrated with HLDA) suggests a path to applying the method to higher-dimensional systems where full-coordinate optimization would be intractable.","The ability to find parameters outside the training domain hints that the diffusion model learns a smooth enough mapping from conditions to distributions to support limited extrapolation, which could reduce the breadth of training data needed."],"fun_headline_variants":["Optimizing system parameters through a frozen diffusion model to match target energy lands","Backpropagating through diffusion samplers to design target free-energy landscapes","GB-FESO: Inverting diffusion models to design prescribed free-energy landscapes","Tuning molecular parameters via differentiable diffusion to match target energy landscapes","Free-energy landscape design via gradient backpropagation through frozen diffusion models"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The method assumes that the KDE-based KL-divergence loss landscape is navigable by gradient descent (with stochastic perturbations) and that the frozen diffusion model generalizes well enough to interpolate to optimized conditioning values. The paper acknowledges the loss landscape is rough and that optimization success depends substantially on initialization — runs starting from high-interaction-strength parameters fail roughly 25% of the time versus 2.5% for low-strength初始化","fun_headline_variants_meta":{"raw":{"variants":["Optimizing system parameters through a frozen diffusion model to match target energy landscapes","Backpropagating through diffusion samplers to design target free-energy landscapes","GB-FESO: Inverting diffusion models to design prescribed free-energy landscapes","Tuning molecular parameters via differentiable diffusion to match target energy landscapes","Free-energy landscape design via gradient backpropagation through frozen diffusion models","Inverting frozen diffusion models to optimize molecular free-energy landscapes","Gradient-based optimization of free-energy landscapes using frozen diffusion models","Designing target free-energy landscapes by backpropagating through diffusion models","Optimizing ensembles to match target energy landscapes via frozen diffusion models"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1233,"prompt_tokens":571,"completion_tokens":662,"prompt_tokens_details":null},"tokens_in":571,"tokens_out":662,"duration_ms":31860,"temperature":1.0,"reasoning_tokens":640,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T06:10:00.340628+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Apply GB-FESO to a molecular system with more than a handful of degrees of freedom and multiple metastable states. If the KDE-based KL loss landscape becomes so rough that gradient descent fails to converge regardless of initialization, or if the frozen diffusion model cannot generalize to the optimized conditioning values and produces physically meaningless ensembles, the method's utility would be confined to toy systems.","supporting_citations":[],"review_version":1}