{"id":"769a31f8-d095-494c-999c-05e185e130f5","arxiv_id":"2504.18911","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"SamAdams, an adaptive-stepsize Langevin sampler with an auxiliary moving-average control variable, achieves larger stable steps and better accuracy than fixed-step BAOAB on several benchmark and neural-network posterior sampling problems.","lead":"This paper introduces a way to make Monte Carlo sampling faster and more stable by automatically shrinking or growing the numerical step size depending on the local steepness of the distribution. The method is inspired by the Adam optimizer used in machine learning, and the authors show it can sample hard test problems and neural network posteriors with larger average steps than fixed-step samplers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 excludes the implemented monitor and all test potentials; the stationary-bias guarantee for the actual SamAdams algorithm is absent.","rationale":"The paper's continuous-time construction is sound: in real time the (x,p) marginal is exactly the original underdamped Langevin dynamics, so the reweighting identity (15) is exact in the continuum. The ZBAOABZ splitting is a natural Strang composition, and the finite-time weak-order argument would be reasonable under the stated smoothness assumptions. The numerical experiments are extensive and internally consistent, which is real evidence. The concern is narrower and precisely located: the theorem's hypotheses exclude the monitor used in all experiments, and no ergodic or stationary-bias analysis fills the gap. Remark 3 shows the authors are aware of the non-Lipschitz monitor, but awareness is not support. Because the headline claims are about long-run averages and stability, the missing guarantee for the implemented algorithm is load-bearing. I considered whether the greater risk is the missing [62]/AdamMCMC baseline or hyperparameter sensitivity; both are real but secondary. The [62] baseline would contextualize the gains, and Section 5.2.4 already documents a wide stable region, so neither undermines the central claim as directly as the theory-practice mismatch. The concrete test on a Gaussian target with the implemented monitor versus a smoothed, theorem-compliant monitor would separate a technicality from a genuine failure: if the non-Lipschitz monitor is needed for the large stable stepsizes and still yields Delta tau^2 bias, the concern is resolved; if bias degrades or the large stepsizes disappear, the paper's central assertion needs qualification. The reader's CONDITIONAL verdict is the right call, so no change is recommended.","tokens_in":36807,"tokens_out":13906,"duration_ms":147383,"concrete_test":"Run ZBAOABZ on a Gaussian target U(x) = x^T C x / 2 in d=100, with the implemented monitor g = ||grad U||^2 = ||C x||^2, fixed hyperparameters (e.g., alpha, Omega, m, M, r as in Section 5.2.4), and Delta tau varying from, say, 0.005 to 0.08. Use very long runs (e.g., 10^9 steps or enough effective samples) and compute reweighted ergodic means of observables like x^2 and ||x||^2, plus the mean stepsize. If the bias relative to the exact Gaussian mean does not decrease as Delta tau^2 (or plateaus instead of converging), the non-Lipschitz monitor breaks the theorem's guarantee. As a control, repeat with g_epsilon = sqrt(||grad U||^2 + epsilon), which satisfies the C^6 bounded-derivative hypotheses for this U; comparing the two bias curves isolates the role of the non-Lipschitz g.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SamAdams/ZBAOABZ 'samples the canonical distribution after reweighting' rests on Theorem 1 (Section 3.2), whose hypotheses require psi, sqrt(psi), grad U, and g to be C^6 with all partial derivatives bounded and psi uniformly bounded. Every numerical experiment uses the monitor g = Omega^{-1} ||grad U||^s (Eq. 26) with s=2 (and s=1 in some funnel runs). For s=2 this g is not globally Lipschitz, and even for s=1 the composition with grad U is not globally Lipschitz unless grad U itself is. The funnel, Beale, and neural-network potentials have unbounded gradients or non-smooth ReLU activations, so the theorem does not apply to any reported experiment. Remark 3 explicitly acknowledges the mismatch and suggests a smoothed variant, but no experiment uses it. Theorem 1 bounds only the finite-time weak error of the ratio E(chi mu)/E(mu); it does not establish a stationary-bias bound for the ergodic averages computed in Section 5, and no such analysis is supplied for the implemented algorithm. Section 5.2.4 further shows the stability threshold depends strongly on alpha and Omega, so the practical efficiency claim is not robustly anchored. The missing stationary-bias guarantee for the non-Lipschitz monitor is the load-bearing gap between theory and the paper's headline claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops an adaptive-stepsize Langevin sampling framework called SamAdams. It augments the phase space with an auxiliary variable ζ that evolves by a relaxation equation driven by a monitor function g(x,p), typically a power of the gradient norm, and uses a bounded Sundman kernel ψ(ζ) to rescale time by dt=ψ(ζ)dτ. The authors show that in continuous time the (x,p) process in real time is ordinary underdamped Langevin dynamics, so canonical averages can be recovered by reweighting samples with ψ(ζ). A symmetric splitting integrator ZBAOABZ is introduced and a finite-time weak second-order error bound is stated in Theorem 1 under strong regularity assumptions. Numerical experiments on toy potentials, a 9D funnel, and MNIST MLP/CNN classifiers report improved stability and larger mean stepsizes than fixed-step BAOAB.","tokens_in":37011,"tokens_out":13882,"duration_ms":144096,"significance":"If the theoretical gap can be closed, this is a genuinely useful contribution: the reweighting construction is simple and does not require modifying the drift, the moving-average mechanism is a principled bridge to Adam, and the numerical study is extensive, with striking stability improvements on the 9D funnel and the CNN. The paper is also honest in Remark 3 that the implemented monitor violates the assumptions of Theorem 1. However, the central claim that SamAdams samples the canonical distribution is currently supported only for the continuous-time dynamics and for the finite-time weak error of a regularized version of the algorithm; no stationary-bias guarantee is provided for the implemented method. This gap is load-bearing and should be addressed before the headline claims are accepted.","major_comments":[{"comment":"Theorem 1 requires ψ, √ψ, ∇U, and g to be C^6 with all partial derivatives bounded and r<ψ<M. The monitor used in every experiment is g=Ω^{-1}||∇U||^s with s=1 or 2 (Eq. (26)); for s=2 this is not globally Lipschitz, and for s=1 it is not differentiable at ∇U=0. The funnel, Beale, and ReLU neural-network potentials also have unbounded or non-smooth derivatives. Remark 3 acknowledges this and proposes a smoothed variant, but no experiment uses that smoothed variant. Thus Theorem 1 cannot be invoked for any of the reported numerical claims.","section":"Section 3.2, Remark 3, Eq. (26), Section 5"},{"comment":"The paper proves only a finite-time weak-error bound for the reweighted ratio (Eq. (18)); the constant C depends on T and the bound does not control the limit n→∞ with fixed ∆τ. Section 2.1 gives the exact continuous-time reweighting identity, but the asymptotic bias of ergodic averages computed by the discrete ZBAOABZ is not analyzed; the text explicitly defers this to future work. Since the abstract and conclusion claim that the method samples the canonical distribution after reweighting, a stationary-bias theorem for the implemented algorithm is required, or the claims must be weakened.","section":"Section 3.2 and Section 2.1"},{"comment":"For the displayed kernels ψ(1) and ψ(2), one has ψ(0)=M/m and ψ(∞)=m, not ψ(0)=M as stated in Section 4. Consequently the claimed stepsize bounds ∆t∈[m∆τ,M∆τ] do not follow from the formulas, and the repeated assertion that ζ0=0 gives ∆t0=M∆τ (e.g., Section 5 and Fig. 5) is algebraically inconsistent for m<1. Either the kernels or the bound statements are misprinted; as written, all reported values of ∆tmax and the initialization argument need to be revised.","section":"Eq. (25), Section 4, and Section 5 initialization"},{"comment":"Fig. 11 shows strong dependence of the stability threshold on α and Ω, with a 'highly unstable' region and an optimal zone near α≈1, Ω≈100. This is in tension with the statement in Section 5 that detailed hyperparameter tuning is usually unnecessary. The authors should provide guidance on how to choose (α,Ω) for a new problem, or explicitly state the admissible region and its dependence on the force scale.","section":"Section 5.2.4, Fig. 11"},{"comment":"The claimed 400% efficiency improvement in the 9D funnel rests on ESS estimates computed from non-uniformly spaced samples after an interpolation step that is not described. Table 1 also shows non-negligible variation in mean log posterior across stepsizes (−10.55 to −10.42), so the accuracy comparison is not fully quantified. Please report the ESS estimation procedure and provide effective sample size per gradient evaluation for both samplers.","section":"Section 5.3.1, Table 1 and Fig. 14"}],"minor_comments":[{"comment":"The uniform lower bound r<ψ<M is used to justify division by E(ψ), but it is only introduced later in Theorem 1; state it when the reweighting formula is first presented.","section":"Section 2.1"},{"comment":"The pseudocode does not define burn-in or how nmeas interacts with the reported averages; specify that the first n_burn iterations are discarded and that 'collect' means storing (x_n,p_n,µ_n) for reweighting.","section":"Algorithm 1"},{"comment":"The sentence 'using the mean adaptive stepsize for BAOAB gives the benefit of the doubt to the fixed-stepsize method' needs a caveat, because the mean uses future information from the full trajectory that would not be available when selecting a fixed stepsize in practice; the direction of the resulting bias should be stated explicitly.","section":"Section 5.3.2"},{"comment":"The notation q(n∆t) is ambiguous for the adaptive scheme; clarify whether n∆t refers to real time for the fixed-step scheme and to the corresponding rescaled time for ZBAOABZ.","section":"Section 5, Eq. (27)"},{"comment":"The paper does not state whether code is available; for a methods paper with extensive experiments, a reproducibility statement or code release would be helpful.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"None of these issues suggest scientific misconduct; the continuous-time construction is sound and the numerical observations are plausible. My main concern is the mismatch between Theorem 1 and the implemented algorithm, which the authors themselves flag. I recommend major revision rather than rejection because the gap is identifiable and potentially fixable: either prove a stationary-bias estimate under weaker conditions, or run the experiments with the smoothed monitor of Remark 3 and restrict the claims accordingly. The Eq. (25) bound error should also be corrected, as it affects all reported stepsize ranges."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is real and well-engineered: an auxiliary variable zeta that accumulates a moving average of a gradient-norm monitor, a Sundman transform psi(zeta) that adapts the timestep, and a reweighting by psi to recover canonical averages. That combination is not in the prior work I know, including the configuration-dependent scheme in [62] and AdamMCMC. The continuous-time reweighting argument is self-consistent, and the ZBAOABZ splitting is a natural, cheap way to wrap any fixed-step Langevin integrator. The numerical story is also consistent across several benchmarks: the method stays stable where fixed-step BAOAB blows up, and the stepsize histograms show the mechanism doing what the prose says it does. I believe the stability gains are real, and the MNIST CNN results are suggestive even if not a full BNN evaluation.\n\nThe soft spots are exactly where the reader's report and the stress-test note put them. Theorem 1 assumes psi, sqrt(psi), grad U, and g are C^6 with bounded derivatives and psi uniformly bounded. The implemented monitor g = ||grad U||^2 is not globally Lipschitz, and neither the funnel, Beale, nor ReLU networks satisfy the smoothness hypotheses. Remark 3 says this plainly, which is honest, but no experiment uses the smoothed variant that would satisfy the theorem. So the headline claim that SamAdams 'samples the canonical distribution after reweighting' is not actually proved for the method being sold. The paper also states that stationary-bias analysis is future work, which is fine, but it weakens the central assertion. Two additional issues are practical: no code or data is provided, and the closest adaptive-stepsize baseline [62] is never compared empirically, so it is hard to say whether the moving-average mechanism earns its keep over a simpler configuration-dependent stepsize. The hyperparameter sensitivity documented in Section 5.2.4 is not fatal, but it means the efficiency gains are not plug-and-play.\n\nNone of this is fatal. The method is promising, the exposition is clear, and the limitations are acknowledged rather than hidden. The right path is a major revision that either extends the theory to the non-Lipschitz monitor or switches the experiments to the smoothed variant, adds a stationary-bias analysis or at least a careful numerical bias study, and includes code plus a direct comparison to [62] and AdamMCMC. I would send it to peer review with that expectation. The paper deserves referee time; it just needs to earn the canonical-sampling claim.","headline":"A genuinely new adaptive-stepsize Langevin scheme whose experiments are more convincing than its theory; worth a serious referee, but the authors need to close the gap between the proved result and the implemented algorithm.","tokens_in":37632,"tokens_out":1634,"would_cite":true,"duration_ms":19372,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60J22","65C05","65C30"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adaptive-stepsize Langevin sampling with an Adam-style moving average of gradient norms can run at mean stepsizes far above the fixed-step stability limit while preserving canonical averages through reweighting.","keywords":["sampling methods","computational statistics","Adam","Langevin dynamics","adaptive stepsize","Bayesian sampling","neural network"],"falsifier":"On a target with exactly known marginals and unbounded gradient norm—for example a high-dimensional Neal funnel with heavier tails—measure the stationary bias of reweighted observables under ZBAOABZ at decreasing $\\Delta\\tau$: if the bias fails to shrink at order $\\Delta\\tau^2$, or if the mean-step advantage over BAOAB disappears when the monitor is replaced by a globally Lipschitz smoothed version, the central efficiency claim is refuted.","tokens_in":36491,"feed_emoji":"🎲","tokens_out":7505,"duration_ms":70471,"temperature":0.7,"pith_summary":"The paper claims that the biggest practical constraint on Langevin-based MCMC—the need to pick one small step size that is safe in the steepest part of the log-posterior—can be relaxed by letting the step size adapt to recent force magnitudes, in the spirit of the Adam optimizer. It introduces SamAdams, which adds an auxiliary scalar $\\zeta$ whose relaxation equation accumulates a moving average of a monitor $g=\\|\\nabla U\\|^s$, and uses a bounded Sundman transform $\\psi(\\zeta)$ to rescale time so that $\\Delta t = \\psi(\\zeta)\\Delta\\tau$. The paper shows that canonical expectations are still obtained by reweighting samples with $\\psi(\\zeta)$, and its symmetric integrator ZBAOABZ is proved to have weak order two. In experiments, SamAdams runs at mean step sizes two to four times larger than the fixed-step BAOAB stability threshold, with roughly 400% efficiency gain on the 9D Neal funnel and more stable, more accurate Bayesian neural network training on MNIST.","feed_headline":"Adam-style step control speeds MCMC up ~400 percent","feed_subtitle":"A time-rescaled Langevin sampler stays stable at mean step sizes that break fixed-step integrators.","key_machinery":"The load-bearing object is the Sundman time-rescaled underdamped Langevin SDE in augmented phase space, $$dx=\\psi(\\zeta)p\\,d\\tau,\\quad dp=-\\psi(\\zeta)\\nabla U\\,d\\tau-\\gamma\\psi(\\zeta)p\\,d\\tau+\\sqrt{2\\gamma\\$beta^{{-1}}$\\psi(\\zeta)}\\,dW,\\quad d\\zeta=(-\\$\\alpha$\\zeta+g(x,p))\\,d\\tau,$$ with $dt=\\psi(\\zeta)d\\tau$. The monitor $g=\\|\\nabla U\\|^s$ (typically $s=1$ or $2$, scaled by $\\Omega$) drives the moving average, and the bounded filter $\\psi(\\zeta)= (m\\zeta^r+M)/(\\zeta^r+m)$ keeps $\\Delta t$ inside $[m\\Delta\\tau,M\\Delta\\tau]$. The integrator ZBAOABZ wraps the standard BAOAB splitting with half-steps of the $\\zeta$-flow, and the reweighting identity $\\mathbb{E}_{\\pi_\\beta}(\\phi)=\\mathbb{E}_{\\Pi_\\tau}(\\phi\\psi)/\\mathbb{E}_{\\Pi_\\tau}(\\psi)$ is what converts rescaled-time samples into canonical averages.","core_discovery":"On its own terms, the central claim is that a time-rescaled Langevin process with an Adam-style monitor is a practical, provably second-order sampling method whose adaptive step size is automatically reduced in regions of steep change of the log posterior and increased on plateaus. The invariant measure of the rescaled dynamics is not the Boltzmann measure, but the paper proves the reweighting identity $\\mathbb{E}_{\\pi_\\beta}(\\phi) = \\mathbb{E}_{\\Pi_\\tau}(\\phi\\,\\psi)/\\mathbb{E}_{\\Pi_\\tau}(\\psi)$ and uses it to compute canonical averages. The numerical heart is the claim that ZBAOABZ remains stable and accurate at mean step sizes $\\langle\\Delta t\\rangle$ that would make BAOAB unstable, because small step sizes are needed only in rare steep regions. The paper supports this with a weak-order-two convergence theorem and experiments on the star potential, asymmetric double well, entropic barrier, Beale potential, Neal's funnel, and Bayesian neural networks on MNIST.","pith_inferences":["A monitor based on minibatch gradient variance rather than the gradient norm would make the step size track the injected noise; the paper's logistic-regression appendix suggests this could automatically adjust for changing batch size, replacing hand-tuned learning-rate schedules.","The scalar $\\zeta$ discards Adam's per-coordinate step sizes; a vector-valued $\\zeta$ could add anisotropy and help ill-conditioned targets, but the paper gives no implementation or bias analysis for that variant.","Because the step size reacts to the gradient norm, the method should transfer to molecular dynamics with repulsive cores or other force fields with rare large forces, where fixed-step integrators are throttled by the worst-case force."],"forward_implications":["ZBAOABZ can be run at mean step sizes above the stability threshold of fixed-step BAOAB; stability is governed by local, not global, curvature.","On the 9D Neal funnel the method yields about 400% more effective samples per unit cost at equal accuracy to the best fixed-step run.","On the MNIST MLP and CNN, the adaptive step prevents early-training loss spikes, and posterior-averaged test accuracies are higher and have far smaller variance than BAOAB at comparable cost.","Any fixed-stepsize Langevin integrator can be wrapped in the two Z half-steps, and the same reweighting identity then provides canonical expectations without Metropolis correction."],"supporting_citations":[{"why":"introduced adaptive stepsize Langevin dynamics via a bounded Sundman kernel, the foundation SamAdams extends with the auxiliary variable","marker":"[62]"},{"why":"defines the BAOAB splitting integrator that ZBAOABZ wraps with symmetric Z half-steps","marker":"[56]"},{"why":"supplies the weak second-order convergence result for splitting methods with multiplicative noise on which Theorem 1 relies","marker":"[103]"},{"why":"shows the Adam optimizer is the Euler discretization of an ODE, motivating the choice of monitor and transform kernel","marker":"[23]"},{"why":"introduced the Adaptive Verlet idea of treating the time-rescaling factor as a dependent variable","marker":"[44]"},{"why":"defines the Adam optimizer whose moving-average step control inspires the SamAdams monitor and kernel","marker":"[54]"}],"fun_headline_variants":["Adam-inspired step control accelerates MCMC","Langevin sampler adapts steps like Adam optimizer","Time-rescaled MCMC with Adam-style adaptive step","Adaptive-step Langevin from Adam's gradient intuition","MCMC stability via Adam-infused step scheduling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The efficiency claim rests on the assumption that one scalar moving average of the gradient norm, combined with a bounded filter, keeps the implemented integrator stable and nearly unbiased in high dimensions; the paper's convergence theorem requires smoother monitors than the $\\|\\nabla U\\|^2$ actually used.","fun_headline_variants_meta":{"raw":{"variants":["Adam-inspired step control accelerates MCMC","Langevin sampler adapts steps like Adam optimizer","Time-rescaled MCMC with Adam-style adaptive step","Adaptive-step Langevin from Adam's gradient intuition","MCMC stability via Adam-infused step scheduling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000848,"raw_usage":{"total_tokens":3701,"prompt_tokens":966,"completion_tokens":2735,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":2662}},"tokens_in":582,"tokens_out":2735,"duration_ms":22153,"temperature":1.0,"reasoning_tokens":2662,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:06:54.738660+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a target with exactly known marginals and unbounded gradient norm—for example a high-dimensional Neal funnel with heavier tails—measure the stationary bias of reweighted observables under ZBAOABZ at decreasing $\\Delta\\tau$: if the bias fails to shrink at order $\\Delta\\tau^2$, or if the mean-step advantage over BAOAB disappears when the monitor is replaced by a globally Lipschitz smoothed version, the central efficiency claim is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"introduced adaptive stepsize Langevin dynamics via a bounded Sundman kernel, the foundation SamAdams extends with the auxiliary variable"},{"cited_title":"Rational construction of stochastic numerical methods for molecular sampling","cited_arxiv_id":null,"evidence_quote":"defines the BAOAB splitting integrator that ZBAOABZ wraps with symmetric Z half-steps"},{"cited_title":"Weak second order multirevolution composition methods for highly oscillatory stochastic differential equations with additive or multiplicative noise","cited_arxiv_id":null,"evidence_quote":"supplies the weak second-order convergence result for splitting methods with multiplicative noise on which Theorem 1 relies"}],"review_version":1}