REVIEW 4 major objections 4 minor 1 cited by
FES-FM samples the free energy surface directly by learning a transport map in collective-variable space, so generation only requires integrating a low-dimensional ODE.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 15:04 UTC pith:WE3VNBB7
load-bearing objection A plausible reduced flow-matching sampler for free energy surfaces, but the core optimisation claim is unproven and the abstract overpromises. the 4 major comments →
FES-FM: Free Energy Surface Sampling via Reduced Flow Matching
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that there exists a time-dependent velocity field u(y,t) in CV space whose ODE pushes the projected prior distribution onto the target CV density ρ(y), and that this u can be trained without ever computing the high-dimensional target density. The training objective combines a neural velocity u, an auxiliary mean-force network v, and a scalar network c to satisfy the reduced transport equation (20), in which the coefficients are conditional expectations over the CV level set Σy. Combined with the Hessian-informed harmonic prior of Section 5, which gives an E(3)-invariant and physically meaningful starting distribution for many-particle systems, the method claims t
What carries the argument
The reduced transport equation (20) — E_μ[∂tU] + u·E_μ[D] − ∇·u + c = 0, where D is the local mean force (17), E_μ denotes expectation under the time-dependent Boltzmann measure conditioned on the CV level set ξ(x)=y, and c approximates ∂t log Z(t) — is the mathematical engine. The training loss (26) writes the equation as a pointwise residual plus a mean-force regression, and estimates the required expectations by running the coupled non-equilibrium dynamics (8)–(9) and applying the Jarzynski reweighting identity (10). The Hessian-informed harmonic prior (28) is the companion object for many-particle systems: it restricts the prior to the non-trivial Hessian modes at a local minimum, making
Load-bearing premise
The load-bearing premise is that jointly minimizing the mean-force regression ℓ1 and the transport residual ℓ2 lands on the intersection where both are satisfied, so the learned velocity actually transports the prior to the target CV distribution; the paper states this without proof, and a spurious joint optimum would silently break the sampler.
What would settle it
On a low-dimensional system where the conditional expectations in (20) can be computed exactly by numerical quadrature on the CV level sets, train FES-FM and compare the learned u and v to the exact values. If, as λ grows, the residuals do not both vanish or the generated marginal fails to converge to the target ρ(y), the joint-minimization reasoning is false. A second check: feed the trained u into the Jarzynski reweighting estimator; the theory predicts the path weights become constant (zero variance), and a substantial residual weight spread would indicate the transport equation was not sol
If this is right
- Sample generation after training is a d-dimensional Euler integration, so cost per sample depends on the CV dimension, not the full configuration dimension n.
- The free energy surface can be reconstructed directly by histogramming generated CV samples, without projecting or reweighting full-space trajectories.
- The Jarzynski reweighting in training provides an unbiased estimate of the conditional expectations needed for the transport equation, and the method inherits variance reduction from the warm-up drift.
- The E(3)-invariant Hessian prior should make the approach compatible with molecular systems whose potential has a well-defined local minimum and near-harmonic fluctuations.
- On the tested potentials (Müller-Brown, double-well in 50–200 dimensions, two many-particle systems), the method matches or improves accuracy while cutting generation time by roughly an order of magnitude.
Where Pith is reading between the lines
- Because the reduced dynamics are formulated purely in terms of the CV map, the method could be combined with learned, data-driven collective variables to form an end-to-end free energy sampler; the paper does not explore this coupling.
- The training stage still relies on NETS-style high-dimensional non-equilibrium simulation; removing that dependence, as the authors mention in Appendix E, would be the step that makes the method scale to realistic force fields.
- A direct test of the joint-minimization claim: in a system with analytically known conditional mean force, the trained v should match it; if it does not, the ℓ1+λℓ2 objective is being trapped in a spurious minimum.
- The coarea-formula derivation assumes ∇ξ has full rank everywhere in the region of interest; systems where the CV becomes degenerate (e.g., at a saddle) would need a regularized or smoothed CV map.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FES-FM is a reduced flow-matching method for sampling the distribution ρ(y) of collective variables ξ(X) under a Boltzmann target p(x). The authors derive a reduced transport equation (20) for a CV-space velocity u, express the conditional force terms via the local mean force D(x,t), and define a combined loss (26) in which the expensive conditional expectations are reweighted by the NETS/Jarzynski identity (10). Training has a warm-up stage for the full-space drift bθ0 followed by main training of uθ1,cθ1,vθ1; generation solves the d-dimensional ODE (13). The paper also introduces a Hessian-informed harmonic prior with E(3)-invariance. Experiments on Müller-Brown, double-well potentials in n=50–200, and two many-particle systems report lower generation time than projecting full-space NETS samples (NETS-P), with comparable or better 1-Wasserstein error.
Significance. If the central transport claim were rigorously established, FES-FM would be a genuinely useful method: generation cost is decoupled from the ambient dimension, the loss reduction in (26) is elegant and avoids explicitly sampling the CV-conditioned manifolds, and the Hessian-informed prior addresses an important practical issue for molecular systems. The paper makes good use of standard identities (coarea formula, local mean force, Jarzynski reweighting), and the experiments, while limited to synthetic systems, provide initial evidence. The main weaknesses are that the mathematical justification of the training objective is incomplete and one advertised benchmark is absent.
major comments (4)
- [§4.3, Eqs. (21)–(27)] The assertion that minimizing ℓ1+λℓ2 yields uθ1,cθ1,vθ1 satisfying both (22) and (24) is not established. This requires (i) existence of a point in the intersection of the two solution sets and (ii) attainment of zero loss by the chosen neural parameterization. The paper only states that the intersection is non-empty; it does not prove that the reduced Liouville equation (14) [equivalently (20)] admits a solution u, nor that the global minimizer of (26) attains the zero-loss point. Without this, Algorithm 2 is not guaranteed to transport the prior to ρ(y). Please add an existence argument (e.g., construct u from the conditional expectation of the full-space velocity b under μΣy,t) and either prove or empirically verify that the trained u satisfies (20) on held-out (y,t) points.
- [Abstract and §6] The abstract states that the method is evaluated 'including alanine dipeptide in implicit solvent as a molecular benchmark,' but no alanine dipeptide experiment appears in Section 6 or the appendices. The experiments cover only Müller-Brown, high-dimensional double-well potentials, and two synthetic many-particle systems. This is a factual mismatch between the claimed evaluation and the reported results. The claim should be removed or the experiment added.
- [§1 and Table 1] The paper claims that FES-FM 'drastically reduces computational costs,' but the 'Time' reported in Tables 1–2 is only the generation time (solving ODE (13)). Training (Algorithm 1) still requires full-space simulations of (8)–(9) in every epoch, and training wall-clock time is not reported. The abstract's unqualified cost claim is therefore misleading. Please qualify the claim as 'generation cost' and report end-to-end training plus generation time for a fair comparison.
- [§5, Eq. (28)] The Hessian-informed prior U0(x) is defined through the alignment argmin R*0(x)=argmin_{R∈O(3)}∥(I⊗R)x−x0∥. Theorem 2 proves O(3)-invariance but not the regularity needed for the SDE (8), whose drift uses ∇U(x,t)=(1−t)∇U0(x)+t∇Utarget(x). The Kabsch alignment is generally only piecewise smooth, and at points where the argmin is non-unique U0 may be non-differentiable. The paper should either prove that this set has measure zero and the SDE is well-posed, or replace U0 by a smooth surrogate.
minor comments (4)
- [§4.3, Eq. (27)] The notation ∇·u(ξ(x),t) should be explicitly defined as divergence with respect to the y argument of u; otherwise it can be confused with a full-space divergence.
- [§2.1] The phrase 'generates samples for optimization' is unclear; please rephrase to indicate what is optimized.
- [§6.1 / Fig. 1] The caption says 'the contour lines colored by the viridis colormap' but the figure is not colorblind-safe; consider using a labeled contour or a different colormap.
- [Appendix D / Fig. D.5] The caption of Fig. D.5(a) describes FES-FM as using Npre training epochs; clarify that Npre is the number of warm-up epochs for bθ0, not for uθ1.
Circularity Check
No significant circularity; core reduced-transport derivation is self-contained, and same-author citations are non-load-bearing.
full rationale
The derivation chain is not circular. The target CV density ρ(y) is fixed by U and ξ through Eq. (1); the reduced transport equation (20) is obtained by projecting the full Liouville equation via the coarea formula and identities (16)-(18). The training loss (26) is the expectation of the squared residuals of (22) and (24) under νt, with Eνt estimated by the Jarzynski reweighting identity (10), which holds exactly for any drift b̂; no ground-truth CV samples or precomputed free-energy values are used as labels. Thus the learned u would be a genuine solution of (20), not a renamed fit to the target. The joint-minimizer claim in §4.3 is an unproven optimization/representability assumption (finite capacity, nonconvexity), which is a correctness risk, not a circular reduction: it does not presuppose the target distribution or the trained velocity. The only same-author citations (Liu et al. 2025b,c, Appendix C.3) are offered as alternative E(3)-equivariant architecture choices and are not load-bearing for the main derivation; no uniqueness theorem or ansatz is imported from the authors' prior work. Hence no circular step is exhibited, and the score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- λ (loss weight in L1) =
1.0
- ϵ (diffusion coefficient in SDE (8)) =
0.2 (Müller-Brown), 0.1 (DW-50/100/200D), 0.02 (R2-3P, R3-4P)
- Npre (warm-up epochs) =
2000 for all experiments
- Prior hyperparameters for Müller-Brown =
mean (-0.7, 0.7), std 0.3
axioms (7)
- standard math Coarea formula and Liouville equation for ODE marginals (Eq. 1, Eq. 14)
- standard math Jarzynski equality / NETS reweighting identity (Eq. 10-11)
- domain assumption ξ ∈ C^1 and rank(∇ξ)=d in the region of interest (Section 1)
- domain assumption U is twice continuously differentiable, O(3)-invariant, and translation-invariant; H(x0) has rank n-6 (Theorem 1)
- ad hoc to paper The joint minimizer of ℓ1 + λℓ2 lies in the intersection of solution sets of (22) and (24) (Section 4.3)
- ad hoc to paper The reduced ODE (13) with velocity satisfying (20) has a unique solution whose marginal is ρ(y,t)
- ad hoc to paper The Hessian-informed prior U0(x) of Eq. (28) is smooth enough for the SDE (8) despite the alignment argmin
read the original abstract
Sampling the distribution of collective variables (CVs) and estimating the associated free energy surface are crucial problems in statistical physics, as they underpin a better understanding of chemical reactions and conformational transitions. Traditional methods usually rely on simulations in high-dimensional configuration space and project the resulting configurations onto the CV space. To improve sampling speed, we propose FES-FM, a reduced flow matching (FM) method for free energy surface (FES) sampling. We train a dynamical transport map in the CV space, thereby enabling direct sampling of CV distributions and reconstruction of the corresponding free energy surface. For many-particle systems, we construct a prior distribution based on the Hessian at a local minimum of the potential, which ensures both rotation-translation invariance and physically meaningful configurations. We evaluate the proposed method across a variety of potential functions and collective variables, including alanine dipeptide in implicit solvent as a molecular benchmark. Comparative experiments demonstrate that our approach significantly improves sampling speed while maintaining accuracy.
Figures
Forward citations
Cited by 1 Pith paper
-
iSMART: An Iterative Sampling-and-Regression Technique for Solving Martingale-Based PDEs
Solving martingale-based PDEs becomes a sequence of least-squares regressions on SDE path samples, avoiding nested expectations and adversarial training.
Reference graph
Works this paper leans on
-
[3]
arXiv preprint arXiv:2211.01364
An optimal control perspective on diffusion-based generative modeling. arXiv preprint arXiv:2211.01364 . Bortoli, V.D., Hutchinson, M.J., Wirnsberger, P., Doucet, A.,
-
[4]
Target score matching. ArXiv abs/2402.08667. Carbone, D., Hua, M., Coste, S., Vanden-Eijnden, E.,
-
[5]
arXiv preprint arXiv:2410.03282
Neural sampling from boltzmann densities: Fisher-rao curves in the wasserstein geometry. arXiv preprint arXiv:2410.03282 . Chen, J., Richter, L., Berner, J., Blessing, D., Neumann, G., Anandkumar, A.,
-
[6]
arXiv preprint arXiv:2412.07081
Sequential controlled langevin diffusions. arXiv preprint arXiv:2412.07081 . Ciccotti, G., Lelievre, T., Vanden-Eijnden, E.,
-
[7]
arXiv preprint arXiv:2209.03003
Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 . Liu, Z., Zhang, W., Li, T., 2025b. Improving the euclidean diffusion generation of manifold data by mitigating score function singularity, in: The Thirty-ninth Annual Conference on Neural Information Processing Systems. Liu, Z., Zhang, W.,...
-
[9]
Particle denoising diffusion sampler. ArXiv abs/2402.06320. Plainer, M., Wu, H., Klein, L., Günnemann, S., Noé, F.,
-
[11]
arXiv preprint arXiv:2407.07873
Dynamical measure transport and neural pde solvers for sampling. arXiv preprint arXiv:2407.07873 . Tian, Y., Panda, N., Lin, Y.T.,
-
[12]
Xu, Y., Wang, Y., Luo, S., Gao, K., He, T., Liu, C., He, D.,
Iterated energy-based flow matching for sampling from boltzmann densities.arXiv:2408.16249. Xu, Y., Wang, Y., Luo, S., Gao, K., He, T., Liu, C., He, D.,
-
[1955]
Mathematical Proceedings of the Cambridge Philosophical Society 51, 406–413
A generalized inverse for matrices. Mathematical Proceedings of the Cambridge Philosophical Society 51, 406–413. doi:10.1017/S0305004100030401. Phillips, A., Dau, H.D., Hutchinson, M.J., Bortoli, V.D., Deligiannidis, G., Doucet, A.,
-
[2018]
arXiv preprint arXiv:1809.10188
Monge-amp\ere flow for generative modeling. arXiv preprint arXiv:1809.10188 . Zhang, Q., Chen, Y.,
-
[2021]
Forthethree-particlesystemin R2, theparametersin (31)arechosenas α1 = 5000/49, α2 = 5000/49, α3 = 50, r1 = 2 , r2 = 2 , r3 = 2 .4, r4 = 3 .1
and those based on alignment with respect to a given reference configuration (Liu et al., 2025b,c). Forthethree-particlesystemin R2, theparametersin (31)arechosenas α1 = 5000/49, α2 = 5000/49, α3 = 50, r1 = 2 , r2 = 2 , r3 = 2 .4, r4 = 3 .1. For the four-particle system in R3, the parameters in (32) are chosen as α1 = α2 = α3 = α4 = α5 = 5000 /49, α6 = 20...
2000
-
[2022]
arXiv preprint arXiv:2209.15571
Building normalizing flows with stochastic inter- polants. arXiv preprint arXiv:2209.15571 . Albergo, M.S., Vanden-Eijnden, E.,
-
[2024]
Iterated denoising energy matching for sampling from boltzmann densities. ArXiv abs/2402.06121. Albergo, M.S., Vanden-Eijnden, E.,
-
[2025]
Consistent sampling and simulation: Molecular dynamics with energy-based diffusion models. ArXiv abs/2506.17139. Raissi, M., Perdikaris, P., Karniadakis, G.E.,
-
[2048]
The 1-Wasserstein distance is computed using 10,000 points sampled from the generated distribution and 10,000 points sampled from the ground-truth distribution
ODE(4) and ODE (13) are solved via the Euler method with a step size of 0.001. The 1-Wasserstein distance is computed using 10,000 points sampled from the generated distribution and 10,000 points sampled from the ground-truth distribution. Remaining hyperparameters are summarized in Table C.4. Detailed experimental setups are provided in the subsections b...
2000
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.