{"id":"8839c113-1dfb-4158-97ad-ac8c38a6fbde","arxiv_id":"2602.10637","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Coarse-grained Boltzmann generators reweight flow-model samples using a learned potential of mean force, recovering equilibrium statistics from biased or short simulations.","lead":"This paper builds a coarse-grained Boltzmann generator: a flow model proposes reduced-coordinate molecular configurations, and a learned effective energy reweights them into a valid equilibrium ensemble. It aims to make unbiased sampling practical for solvated molecules where atomistic generators are too expensive.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbiasedness hinges on ESFM recovering the true PMF from 10 ns WT-MetaD data, but Prop. 2–3 require an equilibrium biased ensemble; WT-MetaD is time-dependent, so the learned PMF may be biased and reweighting cannot repair it.","rationale":"The reader's weakest assumption—that ESFM from 10 ns WT-MetaD recovers the true PMF—is the same load-bearing step I identify. The paper's own Proposition 2 requires an equilibrium biased ensemble; WT-MetaD is explicitly not such an ensemble, and Proposition 3 is only cited, not proved. If the learned PMF is biased, the importance weights target pη ∝ exp(−βUη) rather than p(R), and no amount of reweighting corrects this. This concern is sharp and concrete: it targets the mechanism that makes CG-BGs 'asymptotically correct,' and it can be tested by comparing the WT-MetaD-trained PMF against one trained under a fixed equilibrium bias. I do not think the concern is fatal or warrants rejection: the reported reweighted ϕ/ψ profiles match MD closely, which is encouraging but not decisive because those are low-dimensional projections. The weight-clipping issue is also real and independently undermines the literal 'asymptotically exact' framing, but the PMF-bias issue is more fundamental and is the one the reader flagged. Since the reader already assigned CONDITIONAL and the concern is addressable rather than refuted, no change to the reader's verdict is needed.","tokens_in":25808,"tokens_out":8650,"duration_ms":98113,"concrete_test":"Train the same MACE PMF architecture on alanine dipeptide using (a) the paper's 10 ns WT-MetaD (γ=9) dataset and (b) a fixed-bias equilibrium dataset of comparable cost, e.g. umbrella sampling or a simulation with the final metadynamics bias held static, covering the same CG region. Evaluate both PMFs by (i) mean-force RMSE against projected forces from the independent 500 ns unbiased MD trajectory and (ii) CG-BG reweighted ϕ/ψ free-energy profiles. If (a) and (b) agree within bootstrap error, the time-dependent-bias concern is benign; if not, PMFB in §4.3 is unvalidated and the 'unbiased statistics' conclusion must be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central unbiasedness claim relies on the PMF Uη being close to the true PMF U. Proposition 2 (§3.2) states that a bias V(R) leaves the conditional distribution pV(r|R) = p(r|R), and Proposition 3 (cited to Chen et al. 2026) then equates the LESFM optimum with the VFM optimum. Both propositions assume the biased data are sampled from the Boltzmann distribution of a fixed bias. The alanine PMFs used for the main reweighted results (PMFB, §4.3) are trained on 10 ns well-tempered metadynamics with γ=9 (§B.2), where Gaussian hills are deposited every 1 ps and the bias is never a fixed potential. WT-MetaD trajectories are not samples from exp[−β(u+V)]/ZV; the ensemble is time-dependent and history-dependent, so Proposition 2 does not strictly apply. If the fast solvent and eliminated degrees have not fully equilibrated to the slowly growing bias, the conditional mean of projected forces—the ESFM regression target—can be systematically biased. Since CG-BG reweighting targets exp(−βUη), any such PMF error is exactly what importance reweighting cannot fix: the self-normalized estimator converges to pη, not to the true marginal p(R). The paper provides no independent validation of Uη beyond reweighted ϕ/ψ projections, which are insensitive to errors in other CG modes. In addition, the implemented estimator is not literally unbiased: without the 1% weight clipping, ESS collapses to ~0 and metrics degrade sharply (Table 6), while clipping removes the asymptotic exactness claimed for the method.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Coarse-Grained Boltzmann Generators (CG-BGs), a framework that performs generative modeling and importance sampling in a coarse-grained coordinate space. A normalizing flow proposes CG configurations, and a neural-network potential of mean force (PMF) learned by enhanced-sampling force matching (ESFM) is used to reweight the proposals via the self-normalized importance sampling estimator. The authors claim that CG-BGs recover unbiased equilibrium statistics of the CG marginal distribution, including solvent-mediated effects, at substantially lower cost than atomistic BGs. Experiments are reported on the Müller–Brown potential and alanine dipeptide in explicit solvent, with comparisons against implicit-solvent MD baselines and a simulation-free PMF benchmarking application.","tokens_in":26220,"tokens_out":2939,"duration_ms":33553,"significance":"If the claims hold, CG-BGs would be a useful step toward scalable equilibrium sampling in reduced coordinates: the idea of combining a learned PMF with flow-based importance reweighting is natural but nonetheless novel in this form, and the explicit-solvent reference set is a strength relative to prior BG work that often treats implicit-solvent MD as ground truth. The paper also provides a neat simulation-free diagnostic for comparing learned CG PMFs. The central asymptotic-exactness claim, however, is only as strong as the accuracy of the learned PMF, and the current evidence does not fully support the abstract's 'asymptotically correct statistics' formulation. The manuscript includes reproducible code, detailed experimental appendices, and computational cost benchmarks, which are all positive features.","major_comments":[{"comment":"The unbiasedness of the learned PMF is load-bearing for the entire method, but Propositions 2 and 3 assume the biased training data are drawn from the Boltzmann distribution of a fixed bias V(R). The PMFB used for the main reweighted alanine results is trained on 10 ns well-tempered metadynamics with γ=9 and Gaussian hills deposited every 1 ps (§B.2). WT-MetaD is a history-dependent, time-dependent process; its instantaneous ensemble is not exp[−β(u+V)]/Z_V. Consequently, the ESFM regression target may not equal the true conditional mean force, and importance reweighting with the learned U_η would converge to p_η(R)∝exp(−βU_η), not to the true marginal p(R). The paper needs either (i) a rigorous extension of Prop. 2 to time-dependent biases, or (ii) empirical evidence that the WT-MetaD conditional force averages have converged to the unbiased mean force, e.g., by comparing a PMF trained","section":"§3.2, Proposition 2–3; §B.2, Table 3"},{"comment":"The implemented estimator is not literally asymptotically exact. Table 6 shows that with 0% weight clipping the ESS collapses to ~0 and JS/PMF errors become very large (e.g., Heavy Atom 0% clipping JS=0.2364, PMF=8.33); the reported results all use 1% clipping. Clipping the top 1% of weights introduces a bias that is not accounted for. The text in §3.3 says the procedure 'gives unbiased estimates under p(R)' and the abstract claims 'asymptotically correct statistics'; these statements are only true if U_η=U and if no clipping is applied. The authors should either use a bias-corrected truncation scheme, report the bias as an explicit error bar, or soften the exactness claims accordingly. As written, the central 'exactness' narrative is stronger than what the algorithm actually computes.","section":"§3.3, §G.2, Table 6"},{"comment":"Proposition 3 (equivalence of LESFM and LVFM optima) is stated without proof and is cited to Chen et al. (2026). This proposition is not a peripheral detail: it is what justifies transferring the unbiasedness of Prop. 2 to the force-matching loss used in training. The paper should include a proof or at least a precise statement of the conditions (model expressivity, data distribution, possible dependence on the bias V) under which the equivalence holds. Without this, the claim that ESFM 'enables accurate PMF learning from rapidly converged data' remains insufficiently supported.","section":"§3.2, Proposition 3"}],"minor_comments":[{"comment":"The abstract says 'asymptotically correct statistics' while §3.3 correctly conditions on U_η approximating the true PMF. Suggest aligning the wording throughout: the method is exact conditioned on an exact PMF, and the learned-PMF case is an approximation whose error is not quantified in the paper.","section":"Abstract and §1"},{"comment":"Typo: 'Jenson-Shannon' should be 'Jensen–Shannon'.","section":"§4, Metrics"},{"comment":"The Heavy Atom mapping is described as 'Fig. 3a' in the text; the caption lists panels (a)–(d), but it would help to label the mapping panel explicitly in the figure itself.","section":"§4.1, Figure 3"},{"comment":"For reproducibility, please clarify whether the 10 ns WT-MetaD datasets used for the CNF and for the PMF are the same trajectory or different runs, and report the initial hill height and deposition stride explicitly in the table or text.","section":"§B.2, Table 3"},{"comment":"It is stated that the divergence is computed exactly with automatic differentiation; please state the computational overhead of this choice relative to Hutchinson's estimator, since scalability is a claimed advantage.","section":"§C.3"}],"recommendation":"major_revision","confidential_remarks":"The central conceptual contribution is worthwhile, but the paper's strongest claim—asymptotically exact equilibrium sampling—is not fully supported because (i) the PMF is learned from time-dependent WT-MetaD data under a fixed-bias assumption that does not hold, and (ii) the implemented estimator uses weight clipping that introduces unquantified bias. These are fixable with additional validation or reworded claims, so I recommend major revision rather than rejection. The authors should also consider whether the learned-PMF target should be described as 'the target distribution' without acknowledging that it is a model-dependent distribution; a more careful distinction between p(R) and p_η(R) would significantly strengthen the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper combines two established ingredients — normalizing-flow Boltzmann generators and force-matched coarse-grained PMFs — into a single pipeline: learn a PMF from biased trajectories via enhanced-sampling force matching, train a flow on biased or unbiased CG data, and reweight the flow samples with that PMF. The specific combination is new as far as the cited literature goes, and the experiments support the practical claims. The Müller-Brown case is a clean test with an exact reference, and the alanine dipeptide results match explicit-solvent MD well after reweighting, even when the flow was trained on 10 ns WT-MetaD data. The reweighted samples also beat implicit-solvent baselines, which is a real improvement over prior BG work that treats implicit solvent as the reference. Code is public.\n\nThe main soft spot is the language about exactness. The target of the importance weights is the learned PMF, so the estimator is unbiased for the model distribution pη, not for the true marginal p. The paper states this conditionally in §3.3, but the abstract and introduction say 'exact' or 'asymptotically correct' without the caveat. That is an overstatement, not a fatal flaw.\n\nTwo other issues are more substantive. Proposition 3 is stated without proof. In the fixed-bias equilibrium case it follows from Proposition 2 plus the usual regression argument that the minimizer is the conditional expectation of the projected force. But the paper never gives that argument, and, more importantly, the main alanine PMFs are trained on well-tempered metadynamics, where the bias is time-dependent. Proposition 2 requires a fixed bias, and a 10 ns WT-MetaD run is not an equilibrium sample from any fixed biased Boltzmann distribution. The paper provides no validation of the learned PMF beyond the φ/ψ dihedral projections — which are exactly the CVs used for biasing — so the unbiasedness claim rests on an assumption that is not checked.\n\nThe weight clipping is a third, smaller concern. Without the 1% clipping, ESS collapses to zero and the metrics degrade badly; with it, the estimator is no longer asymptotically unbiased. The ablation is honest, but it further undercuts the 'exact' framing.\n\nThese are addressable rather than fatal. The method is useful and clearly presented. I'd send it to peer review, asking for a rewritten abstract, a proof or careful attribution of Proposition 3, and at least one PMF validation on a mode not biased during training.","headline":"A solid and useful combination of known pieces, with an overclaimed 'exact' label; the learned PMF and the time-dependent bias make the asymptotic claims weaker than the experiments suggest.","tokens_in":26685,"tokens_out":3629,"would_cite":true,"duration_ms":37874,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Coarse-grained Boltzmann generators can deliver asymptotically exact equilibrium statistics by reweighting flow samples with a learned potential of mean force, even when flow and PMF are trained on rapidly converging biased simulations.","keywords":["coarse-grained Boltzmann generator","importance sampling","normalizing flow","potential of mean force","force matching","enhanced sampling","equilibrium sampling","molecular dynamics"],"falsifier":"Generate two biased equilibrium ensembles of the same system with different CG-only bias potentials and compare, in overlapping CG bins, the mean projected atomistic forces recomputed from the unbiased potential; if the two sets of regression targets differ beyond sampling error, Proposition 2's invariance claim fails and the reweighted estimator is biased. Alternatively, on a system with an exactly known PMF, train Uη via ESFM from a biased ensemble, reweight flow samples, and check whether the recovered marginal matches the exact p(R) within error.","tokens_in":25705,"feed_emoji":"🧪","tokens_out":6045,"duration_ms":64145,"temperature":0.7,"pith_summary":"The paper's thesis is that the exactness guarantee of Boltzmann Generators survives a move to coarse-grained coordinates: generate samples from a normalizing flow in a reduced space, weight them by exp(−βUη(R))/qθ(R) with a learned potential of mean force Uη, and the weighted ensemble converges to the true coarse-grained equilibrium distribution p(R). The key enabling step is estimating Uη by enhanced-sampling force matching, which provides unbiased mean-force targets from short biased simulations, so the method does not need long unbiased molecular dynamics trajectories. On alanine dipeptide in explicit water, reweighted ensembles from flows trained on 10 ns well-tempered metadynamics data match a 500 ns reference free-energy profile, and the method outperforms implicit-solvent baselines. A sympathetic reader would care because it offers a route to unbiased equilibrium sampling of larger molecular systems at reduced computational cost, plus a built-in, simulation-free way to validate any learned coarse-grained potential.","feed_headline":"Coarse-grained Boltzmann generators stay unbiased via reweighting","feed_subtitle":"A learned PMF and importance weights turn fast biased simulations into exact equilibrium samples of larger molecules.","key_machinery":"The central object is the learned coarse-grained PMF Uη(R), which plays the role of the target energy in the importance weights w(R) ∝ exp(−βUη(R))/qθ(R); it is a free energy containing entropic contributions from eliminated degrees of freedom, and has no ground-truth energy labels, only force labels. It is trained by variational force matching against instantaneous atomistic forces projected onto the CG coordinates, and enhanced-sampling force matching is what makes training data cheap: Proposition 2 shows the fiber conditional distribution p(r|R) is unchanged under a CG-only bias, so biased ensembles give unbiased mean-force targets, and Proposition 3 shows the ESFM loss has the same globa","core_discovery":"On the paper's own terms, the discovery is that importance reweighting in coarse-grained space is not just a heuristic fix but a principled correction: the target of a Boltzmann Generator can be the marginal Boltzmann distribution defined by the PMF, and the weights w(R) ∝ exp(−βUη(R))/qθ(R) yield unbiased estimates of p(R) provided Uη is accurate on the support of the flow. The load-bearing result is that Uη can be learned from biased, rapidly converged simulations: because a bias that acts only on CG coordinates leaves the conditional distribution of fine-grained configurations given R invariant, the projected forces used in variational force matching remain unbiased regression targets, an","pith_inferences":["If the learned PMF is imperfect, the reweighting cannot repair it: the method's accuracy is bounded by PMF quality, so for very aggressive coarse-graining, where the conditional force variance grows, we would expect a gradual loss of exactness — a scaling prediction that could be tested by sweeping mapping resolution.","The same reweighting scheme could be applied to any reduced latent space, not just physical CG coordinates, e.g., sampling lattice field theories or glassy systems in a collective-variable representation.","The one-shot, simulation-free evaluation of PMFs suggests a training loop: one could choose the flow model to maximize effective sample size for a fixed PMF, or jointly train flow and PMF to minimize weight degeneracy, potentially improving ESS beyond the ~20% reported.","Weight clipping introduces a bias-variance trade-off; a principled alternative would be to use the flow itself to target high-weight regions, effectively learning the biased sampling distribution that minimizes the variance of the self-normalized estimator."],"forward_implications":["Boltzmann Emulators, which currently train on biased or short trajectories and report biased statistics, can be upgraded to exact samplers by adding the learned PMF reweighting step.","The cost bottleneck of atomistic Boltzmann Generators is reduced because only CG degrees of freedom are generated and the ODE likelihood evaluation (Jacobian trace) scales with the reduced dimension.","Equilibrium statistics can be recovered from 10 ns biased simulations, removing the need for long unbiased MD data that is often the dominant expense.","Solvent-mediated and many-body interactions, which implicit solvent models miss, are captured by a PMF learned from explicit solvent trajectories, giving better accuracy at CG resolution.","The framework doubles as a simulation-free validator for learned PMFs: one set of flow samples and weights estimates observables for any candidate energy model."],"fun_headline_variants":["Coarse-grained Boltzmann generators get unbiased via learned PMF","Reweighted coarse-grained sampling for exact equilibrium stats","Fast coarse-grained generator with importance reweighting","Unbiased sampling of larger molecules via CG-BGs","Learned PMF makes coarse-grained generators exact"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything rests on the claim that enhanced-sampling force matching recovers the true potential of mean force: if the 10 ns biased simulation is not near the biased equilibrium, the model lacks expressivity, or the coarse mapping makes the conditional force noise irreducible, then Uη is biased and the importance weights target the wrong distribution, so the reweighted statistics are no longer asymptotically exact.","fun_headline_variants_meta":{"raw":{"variants":["Coarse-grained Boltzmann generators get unbiased via learned PMF","Reweighted coarse-grained sampling for exact equilibrium stats","Fast coarse-grained generator with importance reweighting","Unbiased sampling of larger molecules via CG-BGs","Learned PMF makes coarse-grained generators exact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":982,"prompt_tokens":681,"completion_tokens":301,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":425,"completion_tokens_details":{"reasoning_tokens":226}},"tokens_in":425,"tokens_out":301,"duration_ms":3667,"temperature":1.0,"reasoning_tokens":226,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T01:00:52.337347+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate two biased equilibrium ensembles of the same system with different CG-only bias potentials and compare, in overlapping CG bins, the mean projected atomistic forces recomputed from the unbiased potential; if the two sets of regression targets differ beyond sampling error, Proposition 2's invariance claim fails and the reweighted estimator is biased. Alternatively, on a system with an exactly known PMF, train Uη via ESFM from a biased ensemble, reweight flow samples, and check whether the recovered marginal matches the exact p(R) within error.","supporting_citations":[],"review_version":1}