REVIEW 3 major objections 5 minor 41 references
Neural non-equilibrium Hamiltonian paths, once trained, can be corrected exactly to sample Boltzmann distributions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 22:35 UTC pith:YUWE2643
load-bearing objection The theory is sound and the round-trip kernel is genuinely new, but the empirical case is under-evidenced; the manuscript deserves serious refereeing and a request for code and error bars. the 3 major comments →
Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the recorded dimensionless work W_theta(Gamma) equals the log ratio of the forward path law to a reverse reference path law whose endpoint is Boltzmann-distributed, up to a free-energy constant. Under the theorem's assumptions — normalized conditional densities and invertible, volume-preserving, time-reversible maps — one obtains E[exp(-W)] = Z, so the same identity supplies normalizer estimates, path importance weights, independent path Metropolis acceptance, and a shared-bridge round-trip kernel that preserves the Boltzmann target. Minimizing mean work minimizes path-space KL divergence and bounds endpoint mismatch.
What carries the argument
The central object is the dimensionless generalized recorded work W_theta(Gamma), computed as the log ratio of forward to reverse path probabilities at each stage, with deterministic kick-drift maps chosen to be invertible, volume-preserving, and reversible under momentum flip so that Jacobian terms vanish. The same quantity drives training, evaluation weights, Metropolis acceptance, and the round-trip kernel, where a shared-bridge involution makes the configuration-space transition reversible with respect to the target.
Load-bearing premise
The recorded work is exactly the path log-ratio only if every implemented proposal stage satisfies the theorem's conditions — normalized conditional densities, invertible volume-preserving maps, exact time-reversibility, and correct evaluation of reverse probabilities; if any implemented map is mis-evaluated or non-invertible, the correction is biased rather than exact.
What would settle it
On a target with a known normalizing constant, run path-SNIS with a proposal whose deterministic map has a small, deliberate Jacobian error (e.g., a scaling that is not volume-preserving) and check whether the sample average of exp(-W) deviates from Z beyond Monte Carlo error. Alternatively, apply the molecular internal-coordinate implementation with its non-invertible boundary projections to a target with a known free energy and compare the estimated normalizer to the reference; a systematic bias would show that the exactness claim fails for such implementations.
If this is right
- Minimizing average recorded work during training is equivalent to minimizing a path-space KL divergence to the reverse path, which upper-bounds the KL between proposal endpoints and the target.
- Averaging exp(-W) over forward paths gives an unbiased estimator of the normalizing constant Z, enabling free-energy difference estimation without any endpoint density.
- Path-SNIS weighted endpoint averages converge to the true Boltzmann expectations, even when no analytic endpoint proposal density is available.
- Independent path Metropolis with acceptance 1∧exp(-W' + W) leaves the Boltzmann distribution invariant on path space, and its endpoint marginal is the target.
- The shared-bridge round-trip kernel preserves the Boltzmann distribution on configuration space, with acceptance that reduces to a difference of recorded works.
Where Pith is reading between the lines
- Because the correction relies only on the path probability ratio rather than an explicit endpoint density, the construction should transfer to any proposal whose full path law is tractable, including learned diffusion or flow paths.
- A concrete stress test: on a target with a known normalizing constant, deliberately introduce a small Jacobian error in a map and check whether E[exp(-W)] departs from Z beyond Monte Carlo error; any departure indicates exactness is lost.
- The molecular boundary-projection issue suggests practical deployments need either fully invertible coordinate maps or an explicitly approximate correction whose bias can be quantified on systems with known free energies.
- Symmetry-averaged proposal densities, as discussed in the appendix, could improve overlap and acceptance for symmetric targets without changing the correction identity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct learned Hamiltonian path proposal for unnormalized Boltzmann targets. A forward path law Q_θ and a reverse-reference path law P_θ with Boltzmann endpoint marginal are constructed; the generalized work W_θ is defined in Eq. (10) as the log ratio of the two path densities. Theorem 1 states dP_θ/dQ_θ = exp(−W_θ)/Z, from which the paper derives path-SNIS weights, Jarzynski normalizer estimates, path-IMH acceptance (Theorem 2), and a shared-bridge round-trip Metropolis kernel preserving π on configuration space (Proposition 1). Training minimizes E[W_θ] plus a variance regularizer. Experiments on double-well and finite-volume lattice φ4 targets report corrected observables, ESS, acceptance, and autocorrelation; a molecular internal-coordinate study is explicitly presented as a feasibility study because its boundary projections are non-invertible. The paper is careful to separate the exact-invertibility setting from the molecular feasibility setting and to label single-chain and finite-volume results as such.
Significance. The construction is a coherent and useful unification of non-equilibrium work, path-space importance sampling, and involutive MCMC for learned Hamiltonian path proposals. If the implementation exactly satisfies the theorem’s assumptions, the framework gives a tractable way to correct learned stochastic Hamiltonian proposals and a clean derivation of a configuration-space round-trip kernel. The paper is unusually transparent about limitations: it explicitly excludes the molecular study from the exact-invertibility claims, reports the low-acceptance U(1) pilot as a stress test, and labels single-chain comparisons as representative. It also ships a small numerical verification of inverse-map and momentum-reversal residuals for the U(1) pilot. The main weaknesses are that the central identities are definitional rather than independent fluctuation theorems, and the main DW/lattice experiments do not yet verify the theorem’s implementation assumptions or provide uncertainty quantification. These issues are fixable but currently make the empirical support partial.
major comments (3)
- [§3.1, App. C.6, App. D] Theorem 1 (Eqs. 17–18) is exact only if the implemented maps satisfy the stated assumptions: normalized q^F,q^R,r^F,r^R, invertible volume-preserving g_s, and Φ_t^{-1}=R∘Φ_t∘R. For the DW and φ4 runs, these properties are asserted “by construction,” but the code is not yet released (Appendix D) and no numerical residuals are reported for those systems. The U(1) pilot does verify inverse-map and momentum-reversal residuals below 4e-15 (Appendix C.6), but this verification is absent for the main benchmarks. The molecular study is explicitly excluded because its boundary projections are non-invertible (Appendix B.1, §3.5). Since the exactness of the correction is the central claim, the main experimental evidence does not yet close the gap between theorem assumptions and implementation. Please add residual checks for DW and φ4 (or release the code), and include a direct diagnostic such as E[
- [Table 1, §4.1, App. B.3] The headline numbers are single-checkpoint point estimates. The Table 1 caption itself states that normalizer errors do not include across-training uncertainty. For DW8, |Δ log Z| = 2.45×10^{-4} is reported alongside path ESS of 2.65% and path-IMH acceptance of 12.7%; without error bars or repeated-seed statistics this could be a favorable evaluation rather than a reliable property of the method. The mode-mass L1 error of 3.02×10^{-2} is not negligible. Similarly, the φ4 curves in Figure 3 show no displayed uncertainties even though Appendix B.3 describes jackknife errors. I ask for uncertainty quantification on the central tables/figures, or an explicit statement marking which entries are single-run demonstrations rather than estimates with statistical error bars.
- [Eq. (10), §3.1] The recorded work W_θ is defined in Eq. (10) exactly so that dP_θ/dQ_θ = exp(−W_θ)/Z. Theorem 1 and Corollary 1 are therefore consistency identities that hold by construction; they are not independent physical predictions or new fluctuation theorems. This is not a flaw in the construction, but the paper should state this openly and locate the contribution in the tractable path-ratio construction and the round-trip kernel, rather than presenting the identities as derived results. This framing also clarifies the training objective: minimizing E[W_θ] is the path-space KL divergence by construction, not an approximate or heuristic objective.
minor comments (5)
- [Eq. (10)] In Eq. (10), the term −log q^R_{θ,t}(−\bar p_{t+1} | x_{t+1}) relies on the output momentum \bar p_{t+1} defined in Eq. (5). A brief reminder that the reverse momentum is the negated output momentum would help the reader parse the formula.
- [Figure 3] The φ4 panels would be much more informative with error bars or shaded bands; Appendix B.3 describes contiguous-block jackknife uncertainties, so these quantities are available. Showing them would let the reader judge whether the SNIS-corrected curve agrees with the HMC reference within errors.
- [Appendix D] The reproducibility statement says code “will be released,” but no repository URL, DOI, or commit identifier is provided. For the final version, please include a release artifact or a clear timeline so that the invertibility and reversibility properties of the implemented maps can be independently checked.
- [§4.3] The molecular table uses “score ESS,” which might be misread as an effective sample size for the target Boltzmann distribution. Since the paper emphasizes that these are prior-action scores and not exact Radon–Nikodym weights, consider renaming this column to “prior-score ESS” consistently across the table and text.
- [§3.4] The notation f_θ(a|u) and r_θ(b|x) is introduced in the round-trip section without an explicit connection to the per-stage factors in Eq. (10). A one-line expansion or a forward reference to Appendix A.11 would help readers verify the equality in Eq. (30).
Circularity Check
The recorded work Wθ is defined as the forward/reverse path log-ratio, so Theorem 1 and the Jarzynski normalization are identities by construction rather than independent physical predictions; the round-trip kernel retains independent content.
specific steps
-
self definitional
[Section 2.2 Eq. (10); Section 3.1 Theorem 1 proof (Eqs. 17–20)]
"Wθ(Γ)=log q0(x0)−log γ(xL)+Σ_t[log r^F...−log r^R...+log q^F...−log q^R...]. Dividing Eqs. (8) and (9)... = Z exp[Wθ(Γ)], where π(xL)=γ(xL)/Z and Eq. (10) was used in the second line."
Eq. (10) defines Wθ precisely as the forward-minus-reverse path log-density ratio plus log q0 − log γ. The reverse law Pθ is constructed to start from xL∼π=γ/Z and to use the same stage laws reversed. Therefore dQθ/dPθ is exactly Z exp[Wθ] by construction; Theorem 1 and Eq. (17) restate this definition. Corollary 1's Jarzynski identity E[exp(−Wθ)]=Z and the path-SNIS/IMH target laws follow from this definitional ratio plus the assumed normalization of Pθ, so those 'corrections' are identities inherent in the definition of the recorded work, not independent predictions.
full rationale
The one clear reduction-by-construction is the core work identity: Wθ is not independently measured or derived from dynamics, but is defined as the log ratio of forward and reverse path densities, so Theorem 1 and the Jarzynski normalization are tautological. There is no fitted-input-called-prediction problem: the training objective uses Wθ, and the reported corrected estimates also use Wθ as an importance/acceptance weight, but no parameter is fit directly to the target observables and then renamed a prediction. The only self-citation (Chen et al., 2026b, a stochastic path sampler by co-authors including the present author) appears as related-work context and is not load-bearing; no uniqueness theorem is imported from the authors. The round-trip NHMC-MH kernel and its detailed-balance proof constitute independent content: they use a measure-preserving shared-bridge involution and do not reduce to the definition of Wθ alone. The paper also honestly flags its own limitation in Section 3.5 and Appendix B.1, excluding the molecular study from Theorem 1 because the boundary projections are non-invertible; this is an admitted scope restriction rather than a circular step. The DW/lattice exactness relies on asserted invertibility, volume preservation, and time-reversibility of the implemented maps, which is a verification gap (code not yet released), not a circularity. Overall, the central identity is definitional, justifying a score of 6, while the round-trip construction and numerical checks keep the paper from being fully circular.
Axiom & Free-Parameter Ledger
free parameters (4)
- λ_var (variance regularizer coefficient) =
0.05 for DW runs; 0.01 for φ4 and U(1) runs
- Number of path stages L and leapfrog steps per stage =
DW4 main: 16/4; DW4 round-trip: 32/5; DW8: 20/5; φ4: 16/5; Ala: 3/1; U(1): 20/5 (Table 4)
- Leapfrog step sizes =
Not reported numerically
- Neural network widths/architectures =
MLP width 128 (DW), CNN channels 32/32 (φ4), 48/48 (U(1)); SiLU activations
axioms (5)
- domain assumption The target density is normalizable: 0 < Z = ∫ γ(x) dx < ∞
- domain assumption The deterministic maps g_s and Φ_t are invertible, volume-preserving, and Φ_t is reversible under momentum flip (Φ_t^{-1} = R Φ_t R)
- domain assumption q_0, q^F_t, q^R_t, r^F_t, r^R_t are normalized densities/mass functions
- domain assumption Q_θ and P_θ are mutually absolutely continuous and path weights are finite positive
- standard math Standard probability results: strong law of large numbers, detailed balance, Radon–Nikodym theorem
read the original abstract
Sampling from an unnormalized Boltzmann density requires proposals that move probability mass globally while retaining enough path-probability information for statistical correction. We introduce Neural Non-Equilibrium Hamiltonian Monte Carlo (NHMC), a train-then-correct learned Hamiltonian sampler. Starting from a tractable base distribution, NHMC learns stochastic Hamiltonian-style paths toward the target. Once training is complete, the learned proposal parameters are fixed; the proposal then generates complete paths and endpoint configurations, which are statistically corrected using the recorded non-equilibrium work. This dimensionless generalized work is determined by the probability ratio between the forward proposal path and a reverse reference path. During training, minimizing its mean reduces a path-space KL divergence and controls an upper bound on endpoint mismatch. During evaluation, the same quantity defines weights for self-normalized importance sampling on paths (path-SNIS), estimates normalizing constants or free-energy differences, and gives the acceptance ratio for path-space independent Metropolis-Hastings (path-IMH). The same forward-reverse laws also define a shared-bridge round-trip Metropolis kernel that acts directly on configurations and preserves the Boltzmann target. On double-well and finite-volume lattice $\phi^4$ targets, the NHMC construction gives corrected estimates when path overlap is sufficient; when overlap is poor, weight degeneracy, low acceptance, and long autocorrelation expose proposal failure. We additionally report a molecular internal-coordinate feasibility study using an MD prior and learned-force path proposal.
Figures
Reference graph
Works this paper leans on
-
[1]
Duane, Simon and Kennedy, A. D. and Pendleton, Brian J. and Roweth, Duncan , journal=. 1987 , doi=
1987
-
[2]
and Sohl-Dickstein, Jascha , booktitle=
Levy, Daniel and Hoffman, Matthew D. and Sohl-Dickstein, Jascha , booktitle=. Generalizing. 2018 , doi=
2018
-
[3]
, booktitle=
Foreman, Sam and Jin, Xiao-Yong and Osborn, James C. , booktitle=. Deep Learning. 2021 , doi=
2021
-
[4]
2017 , doi=
Song, Jiaming and Zhao, Shengjia and Ermon, Stefano , booktitle=. 2017 , doi=
2017
-
[5]
Science , volume=
No. Science , volume=. 2019 , doi=
2019
-
[6]
Albergo, M. S. and Kanwar, G. and Shanahan, P. E. , journal=. Flow-Based Generative Models for. 2019 , doi=
2019
-
[7]
Efficient Modeling of Trivializing Maps for Lattice
Del Debbio, Luigi and Marsh Rossney, Joe and Wilson, Michael , journal=. Efficient Modeling of Trivializing Maps for Lattice. 2021 , doi=
2021
-
[8]
and Botev, Aleksandar and Boyda, Denis and Cranmer, Kyle and Hackett, Daniel C
Abbott, Ryan and Albergo, Michael S. and Botev, Aleksandar and Boyda, Denis and Cranmer, Kyle and Hackett, Daniel C. and Kanwar, Gurtej and Matthews, Alexander G. D. G. and Racani. Normalizing Flows for. arXiv preprint arXiv:2305.02402 , year=. doi:10.48550/arXiv.2305.02402 , url=
-
[9]
International Conference on Learning Representations , year=
Flow Annealed Importance Sampling Bootstrap , author=. International Conference on Learning Representations , year=. doi:10.48550/arXiv.2208.01893 , url=
-
[10]
Arbel, Michael and Matthews, Alexander G. D. G. and Doucet, Arnaud , booktitle=. Annealed Flow Transport. 2021 , doi=
2021
-
[11]
and Vanden-Eijnden, Eric , booktitle=
Albergo, Michael S. and Vanden-Eijnden, Eric , booktitle=. 2025 , doi=
2025
-
[12]
Advances in Neural Information Processing Systems , volume=
Stochastic Normalizing Flows , author=. Advances in Neural Information Processing Systems , volume=. 2020 , doi=
2020
-
[13]
and Crooks, Gavin E
Nilmeier, Jerome P. and Crooks, Gavin E. and Minh, David D. L. and Chodera, John D. , journal=. 2011 , doi=
2011
-
[14]
Matthews, Alexander G. D. G. and Arbel, Michael and Rezende, Danilo J. and Doucet, Arnaud , booktitle=. Continual Repeated Annealed Flow Transport. 2022 , doi=
2022
-
[15]
Advances in Neural Information Processing Systems , volume=
Equivariant Flow Matching , author=. Advances in Neural Information Processing Systems , volume=. 2023 , doi=
2023
-
[16]
and Bose, Avishek Joey and Lin, Chen and Klein, Leon and Bronstein, Michael M
Tan, Charlie B. and Bose, Avishek Joey and Lin, Chen and Klein, Leon and Bronstein, Michael M. and Tong, Alexander , booktitle=. Scalable Equilibrium Sampling with Sequential. 2025 , doi=
2025
-
[17]
arXiv preprint arXiv:2509.03726 , year=
Energy-Weighted Flow Matching: Unlocking Continuous Normalizing Flows for Efficient and Scalable Boltzmann Sampling , author=. arXiv preprint arXiv:2509.03726 , year=. doi:10.48550/arXiv.2509.03726 , url=
-
[18]
International Conference on Learning Representations , year=
Denoising Diffusion Samplers , author=. International Conference on Learning Representations , year=. doi:10.48550/arXiv.2302.13834 , url=
-
[19]
Iterated Denoising Energy Matching for Sampling from
Akhound-Sadegh, Tara and Rector-Brooks, Jarrid and Bose, Avishek Joey and Mittal, Sarthak and Lemos, Pablo and Liu, Cheng-Hao and Sendera, Marcin and Ravanbakhsh, Siamak and Gidel, Gauthier and Bengio, Yoshua and Malkin, Nikolay and Tong, Alexander , booktitle=. Iterated Denoising Energy Matching for Sampling from. 2024 , doi=
2024
-
[20]
Transactions on Machine Learning Research , year=
OuYang, RuiKang and Qiang, Bo and Hern. Transactions on Machine Learning Research , year=. doi:10.48550/arXiv.2409.09787 , url=
-
[21]
Training Neural Samplers with Reverse Diffusive
He, Jiajun and Chen, Wenlin and Zhang, Mingtian and Barber, David and Hern. Training Neural Samplers with Reverse Diffusive. Proceedings of the 28th International Conference on Artificial Intelligence and Statistics , series=. 2025 , doi=
2025
-
[22]
Proceedings of the 42nd International Conference on Machine Learning , series=
Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching , author=. Proceedings of the 42nd International Conference on Machine Learning , series=. 2025 , doi=
2025
-
[23]
Physical Review Letters , volume=
Nonequilibrium Equality for Free Energy Differences , author=. Physical Review Letters , volume=. 1997 , doi=
1997
-
[24]
Physical Review E , volume=
Entropy Production Fluctuation Theorem and the Nonequilibrium Work Relation for Free Energy Differences , author=. Physical Review E , volume=. 1999 , doi=
1999
-
[25]
Journal of Computational Physics , volume=
Efficient Estimation of Free Energy Differences from Monte Carlo Data , author=. Journal of Computational Physics , volume=. 1976 , doi=
1976
-
[26]
Statistics and Computing , volume=
Annealed Importance Sampling , author=. Statistics and Computing , volume=. 2001 , doi=
2001
-
[27]
arXiv preprint arXiv:1205.1925 , year=
Hamiltonian Annealed Importance Sampling for Partition Function Estimation , author=. arXiv preprint arXiv:1205.1925 , year=. doi:10.48550/arXiv.1205.1925 , url=
-
[28]
Stochastic Path Sampler for
Chen, Shiyang and Qian, Moxian and Aarts, Gert and Lucini, Biagio and Zhou, Kai , journal=. Stochastic Path Sampler for. 2026 , doi=
2026
-
[29]
Cohn-Gordon, Reuben and Seljak, Uro. Counterdiabatic. arXiv preprint arXiv:2602.21272 , year=. doi:10.48550/arXiv.2602.21272 , url=
-
[30]
International Conference on Learning Representations , year=
Path Integral Sampler: A Stochastic Control Approach for Sampling , author=. International Conference on Learning Representations , year=. doi:10.48550/arXiv.2111.15141 , url=
-
[31]
International Conference on Learning Representations , year=
Transport Meets Variational Inference: Controlled Monte Carlo Diffusions , author=. International Conference on Learning Representations , year=. doi:10.48550/arXiv.2307.01050 , url=
-
[32]
Markov Chain Monte Carlo with Diffusion Paths
Markov Chain Monte Carlo with Diffusion Paths , author=. arXiv preprint arXiv:2607.11631 , year=. doi:10.48550/arXiv.2607.11631 , url=
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2607.11631
-
[33]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Sequential Monte Carlo Samplers , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2006 , doi=
2006
-
[34]
arXiv preprint arXiv:2408.16249 , year=
Iterated Energy-Based Flow Matching for Sampling from Boltzmann Densities , author=. arXiv preprint arXiv:2408.16249 , year=. doi:10.48550/arXiv.2408.16249 , url=
-
[35]
and Koes, David Ryan , booktitle=
Aggarwal, Rishal and Chen, Jacky and Boffi, Nicholas M. and Koes, David Ryan , booktitle=. 2025 , doi=
2025
-
[36]
and McGibbon, Robert T
Eastman, Peter and Swails, Jason and Chodera, John D. and McGibbon, Robert T. and Zhao, Yutong and Beauchamp, Kyle A. and Wang, Lee-Ping and Simmonett, Andrew C. and Harrigan, Matthew P. and Stern, Chaya D. and Wiewiora, Rafal P. and Brooks, Bernard R. and Pande, Vijay S. , journal=. 2017 , doi=
2017
-
[37]
and Dror, Ron O
Lindorff-Larsen, Kresten and Piana, Stefano and Palmo, Kim and Maragakis, Paul and Klepeis, John L. and Dror, Ron O. and Shaw, David E. , journal=. Improved Side-Chain Torsion Potentials for the. 2010 , doi=
2010
-
[38]
Proteins: Structure, Function, and Bioinformatics , volume=
Exploring Protein Native States and Large-Scale Conformational Changes with a Modified Generalized Born Model , author=. Proteins: Structure, Function, and Bioinformatics , volume=. 2004 , doi=
2004
-
[39]
Topology of
L. Topology of. Communications in Mathematical Physics , volume=. 1982 , doi=
1982
-
[40]
Functional Integration: Basics and Applications , series=
Monte Carlo Methods in Statistical Mechanics: Foundations and New Algorithms , author=. Functional Integration: Basics and Applications , series=. 1997 , doi=
1997
-
[41]
Involutive
Neklyudov, Kirill and Welling, Max and Egorov, Evgenii and Vetrov, Dmitry , booktitle=. Involutive. 2020 , publisher=
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.