{"id":"7da49b32-e171-49bc-9346-edd172d3bb13","arxiv_id":"2505.02589","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A neural-network surrogate for the log-likelihood gradients makes Hamiltonian Monte Carlo trajectories 30 times faster than relative-binning gradients and recovers LVK-consistent posteriors for two binary neutron star mergers.","lead":"DeepHMC uses a neural network to approximate the expensive likelihood gradients that slow down Hamiltonian Monte Carlo sampling for binary neutron star gravitational waves. On the real events GW170817 and GW190425 it recovers posteriors similar to the LVK public results while running in hours on a laptop.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Phase III paper trail does not confirm the Metropolis step uses the true endpoint likelihood, so the posterior-recovery claim remains conditional.","rationale":"The central speed claim is well-supported by the trajectory timing measurements and is not the weak point. The weak point is the validation of posterior recovery. The reader identifies approximate-gradient bias as the weakest assumption. Strictly, HMC with an approximate force remains correct if (a) the proposal map is volume-preserving and time-reversible and (b) the Metropolis acceptance uses the exact target Hamiltonian at the endpoints. The leapfrog with a learned force satisfies (a) even if the force is not a gradient, because each half-step is a shear with unit determinant and the same network is used forward and backward. Thus the DNN approximation per se need not bias the stationary distribution; it mainly affects acceptance rates and ergodicity. What would bias it is omitting the exact endpoint correction, or using a stochastic network (dropout active) that breaks reversibility. The manuscript's Phase III bullet list is silent on this. Given that the only posterior validation is against LVK samples produced under different settings (fixed sky for GW170817; unimodal-only comparisons for GW190425), the reported overlap cannot confirm that the DeepHMC chain targets the true posterior. A definitive check is to instrument the acceptance step and run a same-setup reference sampler. This justifies keeping the CONDITIONAL verdict: the method is promising and the speed numbers are concrete, but the central posterior-recovery claim needs an explicit exact-endpoint implementation statement and a direct reference comparison.","tokens_in":27630,"tokens_out":11684,"duration_ms":165258,"concrete_test":"Instrument Phase III to log the true relative-binning log-likelihood at the start and end of every proposed trajectory and the resulting Metropolis acceptance ratio; verify that the acceptance decision uses these exact values, not the DNN-evaluated Hamiltonian. Then run a same-setup comparison for both events: identical waveform (IMRPhenomD-NRTidal), priors, PSDs, and 128s data, for DeepHMC versus a reference HMC using exact numerical gradients (or Bilby's sampler) with equal computational budget, and compare marginal posteriors with two-sample Anderson-Darling and a multivariate energy statistic. If endpoint correction is absent or posteriors differ beyond Monte Carlo error, the consistency claim fails; if they match, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The least secure link in the central claim is the Phase III target distribution. The generic HMC description includes a Metropolis accept/reject step, and the introductory description of the earlier shadow-potential method says the end-point Hamiltonian is recomputed with the true potential. But the DeepHMC Phase III specification (Sec. V.B) only says to \"run N trajectories using the DNN approximation for the gradients\" and never states that the endpoint Hamiltonian is evaluated with the true likelihood for the accept/reject decision. If the endpoint correction is omitted, the chain samples the DNN shadow posterior, not the true posterior; large R^2 on in-region test gradients (Figs. 10-11) does not bound the stationary distribution, especially for GW190425 where tidal gradients reach only R2≈0.90. If the endpoint correction is present, DNN gradient error is largely an efficiency issue, not a bias issue, because the approximate-force leapfrog is still volume-preserving and time-reversible. The paper provides no same-setup reference sampler or diagnostic to distinguish these two cases, and the LVK comparisons are not apples-to-apples (fixed sky for GW170817; unimodal-only table for GW190425). Therefore the posterior-recovery claim is conditional on an implementation detail the manuscript does not explicitly document.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DeepHMC, a Hamiltonian Monte Carlo algorithm for binary neutron star parameter estimation in which the expensive log-likelihood gradients are replaced by gradients predicted by a fully connected deep neural network. The algorithm is structured in three phases: Phase I collects trajectory points and numerical gradients; Phase II trains the DNN and updates the mass-matrix scales from the Phase I covariance; Phase III runs production trajectories using the DNN gradients. The authors report that DNN-gradient trajectories are about 30 times faster than relative-binning-gradient trajectories and about 7000 times faster than naive-gradient trajectories, and that for GW170817 they obtain 5000 statistically independent samples in about two hours on a MacBook Pro, while for GW190425 they obtain 17,500 independent samples in about 2.5 days. They validate the resulting posteriors by comparing medians and 90% credible intervals with LVK public samples for both events and find broad agreement for many intrinsic parameters.","tokens_in":27873,"tokens_out":5723,"duration_ms":81513,"significance":"If the method samples the true posterior, the result is practically significant: a laptop-scale, sub-minute-per-sample HMC for 12-dimensional BNS inference, compatible with Bilby, would be a useful tool for current and future GW analysis, especially as signals become longer. The paper's concrete timing measurements, the explicit three-phase pipeline, and the use of real LVK public data for two BNS events are strengths. The posterior medians for intrinsic parameters largely overlap the LVK public values, and the speedup claims are internally consistent. However, the central posterior-recovery claim is conditional on an implementation detail: the manuscript does not explicitly state that the Phase III Metropolis acceptance step uses the true Hamiltonian rather than the DNN shadow Hamiltonian. The validation comparisons are also not same-setup comparisons, so the evidence for unbiased posterior sampling is currently qualitative. These gaps are fixable but load-bearing.","major_comments":[{"comment":"The Phase III specification, as written, does not state whether the Metropolis-Hastings acceptance at the end of each trajectory is evaluated with the true likelihood/true Hamiltonian or with the DNN-shadow Hamiltonian. The generic HMC algorithm in Sec. II.B includes a Metropolis step using H(q,p), and the earlier shadow-potential method described in Sec. I explicitly states that one 'moved back to the true potential' at the endpoint. In contrast, Sec. V.B only says to 'run N trajectories using the DNN approximation for the gradients.' If the endpoint Hamiltonian is not recomputed with the true likelihood, the chain samples the DNN shadow posterior, not the target posterior, and the R^2 scores in Figs. 10-11 do not bound the stationary distribution. This is the central correctness point of the paper, and it must be documented and verified explicitly, for example by stating the exact code path for the acceptance probability or by demonstrating with a same-setup reference sampler that the posterior is unbiased.","section":"Sec. V.B (Phase III) and Sec. II.B"},{"comment":"The validation against LVK public samples is not an apples-to-apples comparison. For GW170817, the LVK public samples fix the sky position using the electromagnetic counterpart and fix the time of coalescence, whereas DeepHMC samples these parameters freely; the LVK columns for alpha, delta, and tc in Table I are constants, not credible intervals from a free inference. For GW190425, Table I explicitly presents 'unimodal posterior distributions only,' and the sky maps in Fig. 9 show appreciably different localization areas between the LVK samples and DeepHMC. Consequently, the median overlap on intrinsic parameters is encouraging but does not establish that DeepHMC samples the same posterior as a reference sampler on the same data, waveform, and priors. I recommend adding a same-setup comparison—for example, running a standard sampler on the identical data, waveform, prior, and fixed versus free parameter choices—and reporting a quantitative distributional comparison for all sampled parameters, not only medians and intervals.","section":"Sec. VI, Table I, Figs. 7-9"},{"comment":"The stopping threshold for the DNN is event-dependent (R^2 > 0.99 for GW170817 but R^2 > 0.9 for GW190425), and for GW190425 the tidal-parameter gradients only reach R^2 ≈ 0.90, with the text noting that the tidal parameters are consistently the slowest to converge. A 10% unexplained variance in exactly the tidal directions is not by itself a demonstration of unbiased sampling; whether it matters depends on whether the endpoint correction uses the true Hamiltonian and on the energy-error statistics of the accepted trajectories. The manuscript provides no diagnostic, such as a comparison of Phase III acceptance energy errors or a posterior-bias check against a reference sampler on an injection with known parameters, to show that the approximate-gradient leapfrog does not shift the stationary distribution. This should be added before the claimed posterior-recovery result is taken as established.","section":"Sec. V.A and Appendix B, Fig. 11"}],"minor_comments":[{"comment":"The text contains several typos and grammatical slips: 'acclerated' in the title, 'produces produces' in the abstract, 'LIKEKIHOOD' in the Sec. V heading, 'strucutre' in Sec. V.B, and 'very 10^4 trajectories' in the Phase III bullet list.","section":"Throughout"},{"comment":"In the sampler-space coordinate definition, the text writes 'M = n η^{3/5}' where the symbol n appears to be a typo for the total mass m; this should be corrected for clarity.","section":"Sec. II.C"},{"comment":"The notation (μ/M)^{5/2} = η is used without immediately tying it to the standard symmetric mass ratio; a brief reminder near Eq. (42) would help the reader follow the Jacobian transformation.","section":"Sec. IV.A and Appendix A"},{"comment":"Several panels in the R^2 plots use poorly formatted scientific tick labels and different axis ranges, which makes the 'diagonal line' agreement harder to judge; reformatting the ticks and using a common scale across panels would improve readability.","section":"Figs. 10 and 11"},{"comment":"The entries for alpha, delta, and tc for GW170817 are written as point values with zero-width intervals, which is misleading in a table of credible intervals; the caption should state explicitly that these were fixed in the LVK comparison.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"The key risk to the paper is the Phase III acceptance documentation gap identified in my major comment 1. If the implementation does use the true endpoint Hamiltonian, the core method is likely sound and the remaining issues are local; if it does not, the posterior recovery claims would need substantial rewriting. I would encourage the editor to request the same-setup reference-sampler comparison as a condition of acceptance, since the current LVK public comparisons are not sufficient to certify unbiased posterior sampling. The paper fits the journal's scope and builds on a coherent line of the authors' prior work; I see no scope or attribution problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is the first version of this HMC line that replaces hand-built gradient surrogates with a DNN, applies it to real BNS events with spins and tides, and measures the speedup rather than claiming it. The trajectory timings are concrete: 0.06 s per 200-step trajectory with DNN gradients, 1.8 s with relative binning, 7 minutes naive. That is a genuine result, and the prior-boundary and Jacobian treatments show careful implementation.\n\nThe soft spots are real but addressable. The biggest is the Phase III target distribution. The generic HMC description includes a Metropolis-Hastings accept/reject with the true Hamiltonian. The earlier shadow-potential method explicitly says the endpoint Hamiltonian is recomputed with the true potential. The DeepHMC Phase III spec only says 'run N trajectories using the DNN approximation for the gradients' and never states that the endpoint Hamiltonian is evaluated with the true likelihood. Without that, the chain samples the DNN shadow posterior, not the true posterior. The R^2 values (even 0.99) don't bound the stationary distribution. This is an implementation detail, not a fatal conceptual problem, but it must be spelled out.\n\nThe validation also isn't apples-to-apples: for GW170817 the LVK comparison used a fixed sky position, and for GW190425 only unimodal results are tabulated. There is no same-setup run of a reference sampler with the same waveform, priors, and data, and no bias diagnostic. The R^2 threshold is set per event after seeing the result, and the code isn't released. Individually each is minor; together they mean the posterior-recovery claim is plausible but not nailed down.\n\nWho gets value? Anyone working on fast GW parameter estimation, and method developers in amortized inference. The paper deserves a serious referee. The acceleration claims are measurable and the technical content is sound enough that a revision with a same-setup reference comparison and a clear statement of the Phase III accept/reject step would settle my main concern.","headline":"A genuine acceleration technique with measured speedups and plausible posteriors, but the paper leaves the exact target distribution of Phase III underspecified and the validation is not apples-to-apples.","tokens_in":28421,"tokens_out":2530,"would_cite":true,"duration_ms":30198,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DeepHMC replaces expensive gravitational-wave likelihood gradients with a trained neural network, making Hamiltonian Monte Carlo practical for binary neutron star parameter estimation.","keywords":["Hamiltonian Monte Carlo","deep neural network gradient surrogate","gravitational-wave parameter estimation","binary neutron star","relative binning","GW170817","GW190425","posterior sampling"],"falsifier":"Run DeepHMC and an exact-gradient HMC sampler with the same waveform, priors, and data on a low-signal-to-noise or multimodal event, and check whether posterior means and credible intervals differ by more than the sampling error. A more targeted version: take a GW190425-like injection, measure the DNN gradient error in low-likelihood regions, and test whether increasing that error changes the inferred tidal parameters.","tokens_in":27364,"feed_emoji":"🌌","tokens_out":5856,"duration_ms":70072,"temperature":0.7,"pith_summary":"This paper tries to remove the main obstacle to using Hamiltonian Monte Carlo for gravitational-wave parameter estimation: the cost of computing gradients of the log-likelihood at every leapfrog step. The proposed algorithm, DeepHMC, first runs a short phase of ordinary HMC with numerical gradients, then trains a deep neural network to predict those gradients from the sampled positions, and finally uses the network's gradients to integrate long HMC trajectories. The paper argues this preserves the posterior while cutting trajectory cost by a factor of 30 relative to relative-binning gradients and about 7000 relative to naive likelihood gradients. Applied to the 128-second public data for the binary neutron star mergers GW170817 and GW190425, DeepHMC produces posteriors that agree with the public catalog results, obtaining 5000 statistically independent samples for GW170817 in about two hours on a laptop. If true, this would make full 12-dimensional neutron-star binary inference a desktop-scale problem and open the door to much longer signals from future detectors.","feed_headline":"Neural-net gradients cut neutron-star inference cost 7000x","feed_subtitle":"DeepHMC turns Hamiltonian Monte Carlo into a laptop-scale sampler, matching public results for GW170817 and GW190425.","key_machinery":"The load-bearing object is the DNN gradient surrogate: a fully connected network with three hidden layers of 10D, 10D, and 100D neurons plus a 25% dropout layer, trained by mean-squared-error regression on Phase I data. It plays the role of a shadow potential: Phase III leapfrog steps use the DNN's predicted gradients instead of waveform-derived gradients, while the start and end of each trajectory are evaluated with the true potential to compute the Metropolis acceptance. Also load-bearing are the coordinate choice of log chirp mass, log reduced mass, dimensionless spins, tidal parameters, log distance, sky position, inclination, and log duration with phase marginalized; reflective boundary handling for priors and the equal-mass line; and a sinusoidal regularization of the spin-prior gradient singularity.","core_discovery":"The central claim is that the gradient bottleneck of HMC can be bypassed with a learned surrogate without breaking the sampler: Hamiltonian trajectories in the final sampling phase may be integrated using DNN-approximated gradients, provided the network was trained on positions and numerical gradients collected during an earlier information-gathering phase. The endpoint Metropolis-Hastings acceptance still uses the true Hamiltonian, so the chain remains an approximate HMC sampler whose target is the true posterior. On the two real events tested, the algorithm reproduces the public posterior medians and credible intervals for intrinsic parameters, with autocorrelation times around 10 lags for GW170817 and a few hundred for GW190425, translating into roughly one independent sample per second and one per 26 seconds respectively.","pith_inferences":["Inference: the same Phase I/II/III surrogate idea could be applied with more expressive gradient emulators or to precessing waveforms, but the paper only demonstrates aligned spins and a single waveform family.","Inference: the reported 30x and 7000x trajectory speedups are measured within this pipeline against relative-binning and naive likelihood gradients; the end-to-end gain is smaller once Phase I and Phase II training are included, so the fairest user-facing metric is wall-clock time to a fixed number of independent samples under identical settings.","Inference: using DNN gradients inside the leapfrog integration makes the sampler's stationary distribution only approximately the target; a natural testable extension is to use the DNN only as a proposal generator and correct with an exact Metropolis step, which would remove the bias at some computational cost.","Inference: the lower R-squared scores and slower autocorrelations for GW190425 suggest the method's advantage is signal-to-noise dependent, so a bias diagnostic comparing DeepHMC chains against an exact-gradient chain on a subset of runs should become a standard validation step."],"forward_implications":["Binary neutron star parameter estimation with spins and tides can run on a laptop in hours rather than days or weeks for high-signal-to-noise events.","Because no a priori classification of unimodal versus multimodal parameters is required, the same pipeline handles both the unimodal GW170817 posterior and the multimodal GW190425 posterior.","The per-trajectory speedup makes HMC competitive for the longer-duration signals expected from third-generation detectors, where random-walk samplers are expected to slow down.","DeepHMC can be pointed at new waveform models without re-deriving hand-tuned gradient approximations, since the network learns gradients from data.","The Phase I low-latency localization, giving a sky area and distance estimate in about 20 minutes, suggests the method could support early warning or rapid follow-up."],"supporting_citations":[{"why":"Supplies the three-phase HMC structure with a fitted gradient surrogate that DeepHMC extends to neutron-star sources.","marker":"[29]"},{"why":"Provides the previous cubic-OLUTs HMC for binary neutron stars that failed on multimodal posteriors and motivates replacing hand-tuned gradient approximations with a DNN.","marker":"[30]"},{"why":"Introduces the relative binning likelihood approximation used for Phase I gradients and as the baseline for the 30x trajectory speedup.","marker":"[20]"},{"why":"Introduces Hamiltonian Monte Carlo itself, the sampling framework DeepHMC is built on.","marker":"[33]"},{"why":"Supplies the leapfrog integration, step-size tuning, and autocorrelation-time diagnostics used to measure statistically independent samples.","marker":"[13]"},{"why":"Defines the IMRPhenomD-NRTidal waveform model that supplies the gravitational-wave templates for both test events.","marker":"[44]"},{"why":"Provides the public 128-second GW170817 data set used for validation.","marker":"[36]"},{"why":"Provides the public GW190425 data set used to test the algorithm on a low-signal-to-noise, multimodal source.","marker":"[37]"},{"why":"Supplies the Python likelihood and prior infrastructure with which DeepHMC is made compatible.","marker":"[11]"}],"fun_headline_variants":["Deep learning speeds neutron-star inference 7000x","Neural-net gradients trim HMC cost for neutron stars","DeepHMC: laptop-scale sampling for binary neutron stars","DNN-accelerated HMC: <1s per independent sample for GW170817","Neural surrogate gradients make HMC 7000x faster for GW170817"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sampler's final distribution is still the true posterior even though the leapfrog trajectories are integrated with approximate neural-network gradients instead of the exact likelihood gradients; the paper supports this with fit quality and overlap with public samples, but not with a same-setup comparison to a reference sampler or a bias diagnostic.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning speeds neutron-star inference 7000x","Neural-net gradients trim HMC cost for neutron stars","DeepHMC: laptop-scale sampling for binary neutron stars","DNN-accelerated HMC: <1s per independent sample for GW170817","Neural surrogate gradients make HMC 7000x faster for GW170817"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001003,"raw_usage":{"total_tokens":4235,"prompt_tokens":932,"completion_tokens":3303,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":3211}},"tokens_in":548,"tokens_out":3303,"duration_ms":27968,"temperature":1.0,"reasoning_tokens":3211,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:47:16.654725+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run DeepHMC and an exact-gradient HMC sampler with the same waveform, priors, and data on a low-signal-to-noise or multimodal event, and check whether posterior means and credible intervals differ by more than the sampling error. A more targeted version: take a GW190425-like injection, measure the DNN gradient error in low-likelihood regions, and test whether increasing that error changes the inferred tidal parameters.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Python likelihood and prior infrastructure with which DeepHMC is made compatible."},{"cited_title":"Bayesian inference for binary neutron star inspirals using a Hamiltonian Monte Carlo Algorithm","cited_arxiv_id":"1810.07443","evidence_quote":"Provides the previous cubic-OLUTs HMC for binary neutron stars that failed on multimodal posteriors and motivates replacing hand-tuned gradient approximations with a DNN."},{"cited_title":"Canizares, S","cited_arxiv_id":null,"evidence_quote":"Introduces Hamiltonian Monte Carlo itself, the sampling framework DeepHMC is built on."},{"cited_title":"Duane, A","cited_arxiv_id":null,"evidence_quote":"Defines the IMRPhenomD-NRTidal waveform model that supplies the gravitational-wave templates for both test events."},{"cited_title":"Roulet, J","cited_arxiv_id":null,"evidence_quote":"Provides the public GW190425 data set used to test the algorithm on a low-signal-to-noise, multimodal source."}],"review_version":1}