{"id":"4a6b1454-6120-44cc-857a-15c340142d91","arxiv_id":"2607.13757","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Girsanov reweighting of AMS-sampled reactive trajectories propagates machine-learned interatomic potential parameter uncertainty to committor probabilities and, under extra assumptions, to reaction rates.","lead":"Rare-event simulations with machine-learned force fields are fast but uncertain about how model error changes reaction probabilities. This paper shows how to reweight one set of reactive trajectories to many parameter values, turning potential uncertainty into confidence intervals on committor probabilities and rates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Butane validation does not test the variance-control premise: the POPS posterior is trained in metastable basins, not on reactive paths, so Girsanov weight degeneracy could invalidate the central efficiency claim.","rationale":"The Girsanov estimator itself is derived carefully and the toy-model validations are suggestive. The mathematical unbiasedness argument for p̂_G(θ) is standard: the AMS path-distribution estimator is unbiased, and the Girsanov weight transforms the path measure. The first-order cumulant estimator p̂_C(θ) is a practical approximation, and the paper is explicit that it is used for the butane application. The real soft spot is not the algebra but the applicability of the variance-control assumption to realistic MLIP uncertainty. The butane UQ posterior is constructed from configurations in the reactant and product basins only; it does not quantify misspecification in the barrier region that reactive trajectories actually visit. Since Girsanov weights are exponential in path-integrated forces/descriptors along those reactive paths, a modest misspecification near the barrier can make the weights degenerate. The reader identified exactly this as the weakest load-bearing assumption. My independent reading agrees with that assessment. The paper does not report any diagnostic of weight degeneracy, such as effective sample size or the distribution of L(X, θ_L, θ), for the butane system. The proposed ESS check would settle whether the premise holds. If it fails, the central claim that one AMS run at θ_L and a score pass is sufficient for full parameter-posterior propagation would be unsupported for realistic MLIPs, though the theoretical framework would remain valid. Since the reader's verdict was already CONDITIONAL and the concern does not move the verdict, I recommend UNCHANGED.","tokens_in":21404,"tokens_out":6548,"duration_ms":66809,"concrete_test":"Using the released AMS trajectories and POPS samples from the GitHub repository, compute the full Girsanov log-weights L(X_i, θ_L, θ_k) for both MACE surrogates for N_θ = 2000 posterior draws θ_k (Eq. 25), and the normalized weights w_i ∝ exp(L_i). Report the median and 5th percentile of the effective sample size ESS = (Σw_i)² / Σw_i² over θ_k. If the median ESS is below ~10 (or the 5th percentile below ~1), the §2.4 variance-control premise fails for this system, and the butane results cannot support the efficiency claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is the §2.4 claim that Girsanov reweighting variance stays controlled because reactive trajectories are short and potential differences are 'sufficiently bounded' over the configurational space explored by the reactive ensemble. The butane experiment does not establish this. Step 1 of §4.2.2 fine-tunes the MACE surrogates on 6400 configurations sampled from trajectories started in metastable basins A and B; the POPS posterior πPOPS is therefore calibrated on basin configurations, not on the transition-state region traversed by the AMS reactive paths. Girsanov weights (Eq. 23, Eq. 28) are exponential in path integrals of ∇D·ξ and ∇D·∇Dᵀ along exactly those paths. If the surrogate potential or force is misspecified near the barrier, the effective sample size of the weights can collapse, and neither p_G nor the cumulant approximation p_C (Eq. 31) can be trusted. Since the paper's headline pipeline uses p_C on butane, this is a direct gap in the validation of the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a rare-event kinetic uncertainty-propagation framework for machine-learned interatomic potentials (MLIPs). The authors combine Adaptive Multilevel Splitting (AMS) at a reference parameter θ_L with Girsanov path reweighting to estimate the averaged committor probability p(θ) for nearby potentials without rerunning AMS for each θ. For linear-in-parameter potentials, they derive an unbiased full estimator p̂_G (Eqs. 19–28) and a cheaper first-order cumulant estimator p̂_C (Eq. 31), then reconstruct a log-normal conditional law p|θ and mix over a parameter posterior π(θ). Validation is performed on a dimer in a WCA solvent, a rugged Müller–Brown potential, a 1D model, and butane conformational transitions using MACE foundation models with POPS posteriors. Under an explicit assumption on the exit flux φ_A, the framework is extended to reaction rates via Hill's relation.","tokens_in":21788,"tokens_out":12615,"duration_ms":126981,"significance":"The theoretical contribution is valuable: the unbiasedness argument for p̂_G via the path-measure estimator γ̂ is clean, and the linear-parametrization trick makes score/FIM computations tractable for architectures like MACE. The toy tests are well designed and show good agreement with direct AMS over a range of parameter perturbations. The paper also ships code and data, which is a strength. The main open risk is whether the approximations and variance-control assumptions survive in the real butane application, where only p̂_C is used and no direct validation at sampled θ is provided. If the suggested diagnostics are added, the framework could be a solid methodological contribution to MLIP uncertainty propagation for rare-event kinetics.","major_comments":[{"comment":"The butane pipeline uses p̂_C, the first-order-in-(θ−θ_L) cumulant estimator, but there is no evidence that this truncation is accurate for the POPS posterior actually sampled. The toy models validate p̂_C for single-parameter perturbations of about ±10% around θ0; the butane posterior lives in a high-dimensional parameter space and the effective width in the score direction is not reported. Because p̂_C discards all second-order and higher terms in (θ−θ_L), the reported 'coverage' of the reference MPA value could simply reflect a wide posterior. Please add a calibration check: compare p̂_C to p̂_G on a subset of the 2000 drawn θ, or report the distribution of (θ−θ_L)ᵀs̄ and the magnitude of the neglected quadratic term, ideally with direct AMS at a few θ points. This is load-bearing for the central claim that the framework 'successfully recovers reference rare event probabilities' in a","section":"§4.2.2, Eq. (31)"},{"comment":"The variance-control premise for the full Girsanov estimator is not tested in the butane application. Section 2.4 asserts that reactive trajectories are short and potential differences are 'sufficiently bounded across the configurational space explored by the reactive ensemble,' but the POPS posterior in §4.2.2 step 1 is calibrated on configurations from metastable basins A and B, not from the transition-state region traversed by the AMS reactive paths. For p̂_G in Eq. (28), the weights are exponential in path integrals of s and I; if the surrogate potential or force is misspecified near the barrier, the effective sample size of those weights can collapse. Please report an effective-sample-size or variance diagnostic for the p̂_G weights on the butane reactive trajectories, or explicitly restrict the central variance-controlled claim to p̂_C and state that p̂_G is validated only on the t","section":"§2.4, §4.2.2 step 1"},{"comment":"The unbiasedness proof in Eq. (20) uses the exact normalization Z(θ_L,θ). In practice, Eq. (29) estimates Z with K samples, and plugging 1/Ẑ into Eq. (28) introduces a finite-sample bias; E[1/Ẑ] ≠ 1/E[Ẑ]. The text in §3.2 describes p̂_G as 'although unbiased,' which is not exactly true for the implemented estimator with estimated Z. Please clarify this distinction, and either use an unbiased estimator of the ratio or quantify the bias (e.g., by reporting K and the variance of Ẑ). This is a correctness caveat for the central estimator.","section":"§3.2, Eqs. (23), (28), (29)"}],"minor_comments":[{"comment":"The phrase 'first-order cumulant' is potentially misleading. The first cumulant κ₁(L) contains a term proportional to E[I](θ−θ_L)², so Eq. (31) is a first-order-in-(θ−θ_L) truncation, not a first-cumulant truncation. Please state explicitly which small parameter is being used for the truncation.","section":"Eqs. (30)–(31)"},{"comment":"The notation s̄ʲ and the aggregation over n_real realizations are unclear. It appears the score is averaged over all trajectories of all realizations, while the AMS probability estimates are averaged separately. Please define each term and provide the explicit delta-method variance formula used to produce confidence intervals.","section":"Eq. (33)"},{"comment":"The confidence intervals and violin plots are not fully specified. State how the log-normal fit is performed (moment matching on which scale), how the reported intervals are derived from the delta-method variance, and how many θ samples N_θ are used in each figure.","section":"Figs. 7, 9, 10"},{"comment":"The assumption that φ_A is essentially θ-independent is supported only by three point estimates of mean exit times. Please add statistical uncertainties on φ_A⁻¹ to quantify the support for this assumption, and note explicitly that Table 1 compares discrete models rather than the POPS posterior around θ_L.","section":"Table 1, §4.3"},{"comment":"The statement that FIM computation requires 'P=256 automatic differentiation passes' should clarify whether P is the number of final-layer parameters of MACE-OMAT-0 small and MACE-MP-0a small, and whether this number is architecture-specific.","section":"§4.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a promising methodological contribution, and the core derivation appears sound. My main concern is that the real-world butane validation relies on the approximate estimator p̂_C without a direct check of its truncation error, and the variance-control premise for the full estimator is not examined on reactive paths. These are fixable with additional diagnostics rather than fundamental flaws. I would support publication after the authors provide the suggested calibration checks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: the central estimator is right, and the toy tests earn their keep. The unbiasedness argument in Eq. (20) is correct, and the linear-potential score/FIM form is a genuinely new trick for propagating MLIP parameter uncertainty to committor probabilities without rerunning AMS. The cumulant shortcut (p_C) is also a nice practical contribution: it drops the expensive FIM and still tracks the full estimator on the two toy models. I would send this to a serious referee.\n\nSecond: the butane section is the weak spot, and it is weaker than the authors let on. The stated variance-control premise in §2.4 — that reactive paths are short and potential differences bounded, so Girsanov weight variance is tame — is exactly the thing the real application should validate. It does not. The POPS posterior is fit on 6400 configurations drawn from the metastable basins (step 1 of §4.2.2), not from the transition-path ensemble. The Girsanov weights are exponential in path integrals of forces along the reactive paths, which pass through the barrier region. If the surrogate is misspecified there, the weights can degenerate, and neither p_G nor p_C is trustworthy. The paper doesn't report effective sample sizes or weight variance for the butane trajectories, so we cannot tell whether the good coverage in Fig. 9 is method or luck. That gap doesn't invalidate the derivation, but it means the headline claim 'efficient route for MLIP uncertainty to rare event kinetics' is not closed by this paper.\n\nThe other soft spots are smaller. The first-order cumulant truncation is uncontrolled in general; the toy agreement is nice evidence, but there is no diagnostic to tell a user when p_C is safe. The rate extension in §4.3 assumes the exit frequency φ_A is well-determined, and the authors say so themselves, so that is an honest limitation, not a concealed one. The log-normal assumption for p|θ is an arbitrary but harmless modeling choice.\n\nBottom line: the math is sound, the code is released, the toy validations are convincing, and the real-system demo is suggestive. The load-bearing missing piece is a characterization of the Girsanov weight distribution on the butane reactive ensemble. That is a tractable addition — an ESS analysis and maybe a barrier-region POPS refit — and a good referee should ask for it. Yes, send it to peer review.","headline":"The estimator is correctly derived and convincingly validated on toy models, but the butane application does not test the variance-control assumption the practical pipeline depends on.","tokens_in":22172,"tokens_out":2630,"would_cite":true,"duration_ms":27663,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Girsanov reweighting turns one rare-event simulation into a full posterior over kinetic observables, propagating machine-learned potential uncertainty without resampling.","keywords":["machine-learning interatomic potentials","Girsanov reweighting","rare-event kinetics","committor probability","Adaptive Multilevel Splitting","uncertainty quantification","path-space importance sampling","Hill relation"],"falsifier":"For butane or a similar MLIP system, draw θ samples from a POPS posterior estimated from configurations sampled on the transition path (the barrier region) rather than only the metastable basins, then compare p_G(θ) and p_C(θ) against direct AMS reruns at the same θ. If the distribution of per-trajectory weights M(X,θ_L,θ) develops heavy tails or the reweighted confidence intervals systematically miss the direct estimates, the bounded-variance premise fails.","tokens_in":21358,"feed_emoji":"⚛️","tokens_out":9036,"duration_ms":88534,"temperature":0.7,"pith_summary":"This paper aims to propagate the parameter uncertainty of machine-learned interatomic potentials (MLIPs) to rare-event kinetic observables—the averaged committor probability p and, under an extra assumption, the reaction rate k_AB. Its central insight is that instead of rerunning expensive rare-event sampling for every parameter realization, one can sample reactive trajectories once at a reference potential and reweight them in path space using Girsanov's theorem. For potentials linear in their parameters, the reweighting factor is an exponential linear-quadratic form in the parameter displacement, built from a trajectory score and Fisher information matrix; the paper derives a full unbiased estimator and a cheaper first-order cumulant estimator. It validates both on dimers in solvent, a rugged Müller-Brown potential, and butane with MACE foundation models, reporting about 10^3x speedups and uncertainty-aware distributions that cover the reference rare-event probability. If the variance-boundedness assumption holds, this gives a practical route to UQ for kinetics with MLIPs, where previously only equilibrium or pointwise uncertainties were accessible.","feed_headline":"One AMS run maps MLIP uncertainty to rare-event kinetics","feed_subtitle":"Girsanov reweighting turns a single rare-event simulation into a posterior over committor probabilities and reaction rates.","key_machinery":"The central object is the Girsanov likelihood ratio for linear potentials V(x,θ)=θᵀD(x): M(X,θ_L,θ) = exp[(θ−θ_L)ᵀs(X) − ½(θ−θ_L)ᵀI(X)(θ−θ_L)] / Z(θ_L,θ). Here s(X) is a path score accumulating vector-Jacobian products of the descriptor field along the trajectory, I(X) is the path Fisher information matrix, and Z normalizes the exit-distribution ratio. This factorization converts path reweighting into a cheap parameter-space tilt: one AMS run at θL yields predictions for every θ. The cumulant estimator p_C keeps only the score term, avoiding the costly I(X) computation that would need roughly 256 automatic-differentiation passes for a MACE-style network. The unbiasedness of p_G derives from","core_discovery":"On its own terms, the paper establishes that for potentials linear in their parameters, the path-space likelihood ratio between two potentials is an exponential linear-quadratic form in the parameter displacement, built from a trajectory score s(X) and Fisher information matrix I(X). This makes the Girsanov estimator p_G(θ) in Eq. (28) an unbiased estimator of the committor probability p(θ) using only trajectories sampled at the reference parameter θL, and the first-order cumulant estimator p_C(θ) in Eq. (31) a cheap approximation sharing its Taylor expansion. From these, one constructs the uncertainty-aware distribution P(p|π) over any parameter posterior π, and—under the assumption that th","pith_inferences":["The variance-boundedness premise is the Achilles heel; a stress test with a POPS posterior built from transition-path configurations (rather than basin configurations) would reveal whether the weights M(X,θ_L,θ) remain well-behaved where the reaction actually occurs.","The log-normal ansatz for p|θ is an extra modeling assumption; a nonparametric conditional distribution or a beta fit could change the tails of P(p|π) even with identical first two moments.","The same reweighting machinery could, in principle, correct for estimator bias by reweighting from a cheap surrogate to a high-fidelity reference potential, not just among posterior samples, as long as the potential difference stays small along reactive paths.","For long reactive trajectories (protein folding, large conformational changes), Girsanov variance will grow exponentially with duration; extensions using control variates or learned ratio estimation would be needed to keep the method practical."],"forward_implications":["A single AMS run at the MLE parameter plus one score pass replaces a per-parameter resampling loop; the paper reports a speedup of about 10^3x for butane.","The cumulant estimator p_C reproduces the full estimator's accuracy in all tested systems while costing only scores, making the method tractable for high-dimensional MLIP parameter spaces.","The framework cleanly separates algorithmic sampling noise from model-misspecification uncertainty, producing narrow, covering intervals for accurate surrogates and wide, covering intervals for biased ones.","Under the stated basin-accuracy assumption, uncertainty bounds on the reaction rate k_AB follow from Hill's relation with exit frequencies estimated once at θL.","Because it relies only on an unbiased reactive path distribution, the method is agnostic to the splitting sampler (FFS, SMC, AMS) and to the parameter posterior model."],"fun_headline_variants":["Girsanov reweighting propagates MLIP uncertainty to rare-event kinetics","One AMS run yields MLIP kinetic uncertainty via Girsanov reweighting","MLIP uncertainty quantified for rare events without resampling trajectories","Efficient path-space reweighting maps MLIP error to reaction-rate bounds"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that Girsanov path-reweighting variance remains bounded over the reactive ensemble—short reactive trajectories and small potential perturbations under MLIP parameter uncertainty—along with the assumption that the exit-frequency φA is essentially independent of θ; the butane validation, whose POPS posterior is built from basin configurations, does not directly test the variance bound near the transition barrier.","fun_headline_variants_meta":{"raw":{"variants":["Girsanov reweighting propagates MLIP uncertainty to rare-event kinetics","One AMS run yields MLIP kinetic uncertainty via Girsanov reweighting","MLIP uncertainty quantified for rare events without resampling trajectories","Efficient path-space reweighting maps MLIP error to reaction-rate bounds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1463,"prompt_tokens":790,"completion_tokens":673,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":593}},"tokens_in":534,"tokens_out":673,"duration_ms":6385,"temperature":1.0,"reasoning_tokens":593,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T03:47:28.726111+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For butane or a similar MLIP system, draw θ samples from a POPS posterior estimated from configurations sampled on the transition path (the barrier region) rather than only the metastable basins, then compare p_G(θ) and p_C(θ) against direct AMS reruns at the same θ. If the distribution of per-trajectory weights M(X,θ_L,θ) develops heavy tails or the reweighted confidence intervals systematically miss the direct estimates, the bounded-variance premise fails.","supporting_citations":[],"review_version":1}