{"id":"b251ebce-1522-42b6-a893-527ad1a148aa","arxiv_id":"2509.06924","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Automatic differentiation through the Abeles reflectivity model gives exact gradients for neutron reflectometry, making Hamiltonian Monte Carlo and variational inference fast, practical alternatives to slow MCMC for thin-film uncertainty quantification.","lead":"Neutrons are bounced off layered materials, and the reflected pattern is used to infer each layer's thickness and density. Exact gradients through the physics now make Bayesian uncertainty analysis much faster, and the authors share an open-source Python library for it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline speed/efficiency claims rest on ESS-per-iteration on one dataset and uncalibrated wall-clock timing, so 'order-of-magnitude' and 'real-time' may not survive a cost-adjusted multi-dataset benchmark.","rationale":"The paper's core contribution—automatic differentiation through the Abeles reflectivity model—is real and well-supported by the open-source refjax library; I do not question its correctness. The central claim, however, is not just that gradients exist but that they make Bayesian inversion dramatically faster. That claim is currently justified by a single ESS-per-iteration curve and one 20 s VI timing. The reader's condition 1 flags the same single-dataset/single-workstation basis, though the reader's stated weakest assumption is the Gaussian likelihood. I see the likelihood assumption as a standard, disclosed modelling choice rather than the most load-bearing risk; the scale of the speed claim is more directly tied to the benchmark methodology. A cost-adjusted, multi-dataset benchmark would either confirm the order-of-magnitude claim or show it is an artifact of the chosen comparison metric. I therefore keep the reader's CONDITIONAL verdict rather than raising to REJECT or lowering to ACCEPT.","tokens_in":16463,"tokens_out":11492,"duration_ms":129820,"concrete_test":"Re-run the reported comparison on 5-10 synthetic NR datasets spanning different layer counts, smearing widths, and signal-to-noise ratios. For each dataset, run NUTS and a representative gradient-free MCMC (e.g., adaptive Metropolis or ensemble sampler) to a target ESS of at least 200 per parameter on the same workstation, and record median wall-clock time and ESS/second with error bars. Also record the ELBO of each of the 100 VI restarts using a large Monte Carlo sample (e.g., 10^4 draws) before selecting the best; if the ranking changes, the quoted 20 s result is affected. If NUTS is not an order of magnitude ahead in ESS/second, or VI is not substantially faster than converged MCMC, the headline claims should be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing empirical claim is that exact AD gradients yield a step-change in inference speed (Abstract; Section 3.1 bullet). Figure 6 compares NUTS vs SA-MCMC as ESS per iteration, not ESS per wall-clock second; the text admits NUTS 'per-sample computational complexity is several times higher' (Section 3.1) but never quantifies this. An 'order of magnitude' gain in ESS per sample can be erased by a comparable per-sample cost increase, and the reported VI 'seconds rather than hours' (Section 3.2) has no same-hardware MCMC baseline, reports a single 20 s run on a 64-core Threadripper, and selects among 100 ELBO restarts by maximum noisy ELBO, which can bias the selected variational solution. The single quartz dataset and a few chains provide no error bars over datasets or posterior geometries, so the central quantitative claims are less secure than the wording suggests.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents refjax, an open-source JAX library that computes specular neutron reflectivity from an arbitrary slab model and, via automatic differentiation, exact gradients with respect to model parameters. These gradients are used to run gradient-based optimisation (Adam), Hamiltonian Monte Carlo (NUTS), and variational inference on NR inverse problems. The method is demonstrated on a crystalline quartz film on silicon (D17 instrument) and on a joint fit of four OLED devices (59 parameters), with an additional lipid bilayer benchmark in the Appendix. The authors report a posterior-mean χ² of 1.22 for quartz (better than the published 1.32), an OLED co-fit via VI in under 20 seconds on a 64-core workstation, and an ESS-per-iteration comparison indicating that NUTS converges where a sample-adaptive MCMC sampler appears to stall. The central claim is that exact gradients through the physical forward model provide a step-change in inference speed and sample efficiency compared with gradient-free MCMC, and that VI can deliver approximate UQ on the order of seconds.","tokens_in":16637,"tokens_out":7589,"duration_ms":78285,"significance":"If the central claim holds, the work is a significant practical contribution to NR analysis: it removes the need for surrogate models, provides a public code base (refjax), and opens the door to gradient-based Bayesian methods for a community that still largely uses Levenberg-Marquardt or gradient-free samplers. The paper is generally honest: it discloses the non-physical quartz oxide SLD, notes the mode-seeking variance underprediction of VI, and includes an external benchmark against refnx. The use of exact gradients through the Abeles model is conceptually clean and the released code makes the approach reproducible. However, the quantitative evidence for the headline speed/efficiency gains is currently incomplete. The ESS comparison is not cost-adjusted, and the VI timing has no same-hardware baseline, so the 'order of magnitude' and 'seconds rather than hours' claims are not yet established. The paper would be strengthened by a careful cost-adjusted benchmark on multiple datasets and posterior geometries.","major_comments":[{"comment":"The sample-efficiency comparison is reported as ESS per iteration, not ESS per wall-clock time. The text acknowledges that NUTS's per-sample cost is 'several times higher' than SA-MCMC, but no ratio is given. Because the Abstract claims 'order of magnitude gains in sample-efficiency' and the Introduction claims a 'step-change in inference speed', the comparison must be normalised by computational cost. Please report ESS per second (or per cost-normalised gradient evaluation) on the same hardware, and state the per-sample cost ratio. If the ratio is, say, 5–10×, the ESS-per-iteration advantage shown in Fig. 6 may disappear.","section":"Section 3.1, Figure 6"},{"comment":"The timing claim ('<20 seconds', 'seconds rather than hours') is based on a single run of VI on a 64-core Threadripper with no same-hardware comparison to HMC/NUTS or to a gradient-free sampler on the same 59-parameter problem. The 'hours' appears to come from typical literature experience rather than from measurement. Please provide a side-by-side wall-clock comparison (including total time for all 100 restarts) and report the elapsed time for each method on the same hardware and data. In addition, selecting the best of 100 restarts by maximum noisy ELBO can bias the chosen surrogate posterior; please report the spread of ELBO values across restarts and/or use a common random seed or a deterministic ELBO evaluation for the selection.","section":"Section 3.2, OLED VI"},{"comment":"The reported posterior widths and the VI-versus-HMC comparison are conditional on a Gaussian likelihood with variances taken from the count data and on the assumed slab model (layer count, Gaussian roughness, smearing kernel). The paper does not validate this likelihood against the data. For instance, the quartz posterior mean assigns a native-oxide SLD of 0.323 Å^-2, far from the physical value (~3.4), which the authors attribute to slab-model slack. This indicates sensitivity to the model class. Please add a posterior predictive check (e.g., simulating from the posterior and comparing to the observed data) and, ideally, a simulation-based calibration study to assess whether the reported posterior widths are trustworthy.","section":"Section 1.2, Eq. (25), Table 1"},{"comment":"The ESS plot is described as 'the effective sample size of a single parameter in a single chain' and no R-hat diagnostics are reported for the full parameter vector. The sample-efficiency claim should be supported by summary ESS and R-hat across all parameters and all chains, not just one trace. Please specify which parameter(s) are plotted, report the convergence diagnostics for the entire model, and ensure the plotted chain/parameter is representative rather than cherry-picked.","section":"Section 3.1, Figure 6 and convergence diagnostics"}],"minor_comments":[{"comment":"'complex multiplayer structures' should be 'multilayer structures'.","section":"Introduction"},{"comment":"'numypro' is a typo for 'numpyro'.","section":"Section 3.1"},{"comment":"The background initial value '5.0' is inconsistent with its prior U(e^-20, 1×10^-7); presumably the intended value is 5.0×10^-7. Please correct.","section":"Table 4"},{"comment":"'D20' and 'H20' should be 'D2O' and 'H2O'.","section":"Appendix B"},{"comment":"'variation inference' should be 'variational inference'.","section":"Section 1.4"},{"comment":"The caption is ambiguous: it says 'each line is the effective sample size of a single parameter in a single chain'—please clarify which parameters/chains are shown and how many lines appear.","section":"Figure 6 caption"},{"comment":"The claim 'for the first time in NR, exact gradients through the reflectivity are computed' is stated as a fact. The authors may wish to soften this to 'to the best of our knowledge' and ensure the literature search covers adjacent fields (X-ray reflectometry, optical thin films) where AD has been used.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The novelty claim of 'first exact gradients in NR' is hard to verify; I recommend a careful check by a domain expert for prior AD-based reflectometry work. The main gap is the benchmarking: if the authors provide cost-adjusted ESS-per-second and same-hardware VI-versus-MCMC timing on multiple datasets, the paper could become acceptable. The likelihood-calibration concern is secondary but would strengthen the Bayesian claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading. The core move is simple and genuinely new in this field: push exact gradients through the Abeles matrix model with JAX, then use NUTS and variational inference for Bayesian inversion. The authors ship refjax, an open-source library, and demonstrate it on three datasets. The external checks are real: the quartz posterior-mean chi-squared (1.22) beats the published 1.32, and the OLED co-fit lands the silicon SLD at 2.09, close to the physical 2.07. The reporting is also honest. Appendix C openly discusses VI's variance underprediction, and Section 3.1 admits the native-oxide SLD is non-physical because it is just absorbing slab-model slack. That is the right way to write this up.\n\nWhere the paper oversells itself is the magnitude of the speed claims. The abstract promises a 'step-change' and 'order of magnitude' gains. Figure 6 is ESS per iteration, not per wall-clock second, and the text in Section 3.1 concedes that NUTS 'per-sample computational complexity is several times higher' than SA-MCMC without ever quantifying that factor. If the per-sample cost is high enough, the ESS gain evaporates. The VI 'seconds rather than hours' also lacks a same-hardware MCMC baseline; the single 20-second run is on a 64-core Threadripper, and the chosen posterior is the best of 100 ELBO restarts, which biases the selected solution. These are fixable but the claims as written are not yet supported.\n\nSmaller issues: Table 4 gives the shared background initial value as 5.0, which sits outside its prior (1e-20 to 1e-7). The lipid benchmark in Appendix B uses 17 free parameters versus 13 in refnx, so the comparison is not apples-to-apples. None of these damage the central feasibility result, but they should be cleaned up before publication.\n\nThis is a solid methods paper. I would send it to peer review, with a request that the authors either produce cost-adjusted wall-clock comparisons over multiple datasets with error bars, or soften 'order of magnitude' to 'up to an order of magnitude' where honest. I'd bring it to a reading group; it is a good example of exact gradients making physical-model inversion practical, and the self-critical appendices are a model of tone. If I worked in the area, I'd cite it.","headline":"Worth reading: exact AD gradients through the Abeles model make HMC and VI practical for NR, and the paper is honest about its limits; but the headline speed gains are not yet secured by the evidence.","tokens_in":17268,"tokens_out":3587,"would_cite":true,"duration_ms":34780,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By threading automatic differentiation through the exact reflectivity model, this paper makes gradient-based Bayesian inference — Hamiltonian Monte-Carlo and variational inference — practical for neutron reflectometry, cutting uncertainty q","keywords":["neutron reflectometry","Bayesian inference","Hamiltonian Monte Carlo","variational inference","automatic differentiation","inverse problems","uncertainty quantification","thin-film analysis"],"falsifier":"Fit a deliberately misspecified slab model, such as too few layers or a fixed smearing width, to a simulated dataset with known ground truth: if NUTS or VI then report narrow, confident posteriors incompatible with the true SLD profile, the Gaussian-likelihood/slab-model assumption — not the gradient machinery — limits the method. A cheaper check is to re-analyse the quartz data with a likelihood that explicitly models correlations from the known pointwise smearing and see whether the reported posterior widths change materially.","tokens_in":16284,"feed_emoji":"⚛️","tokens_out":11567,"duration_ms":103403,"temperature":0.7,"pith_summary":"The paper aims to make Bayesian uncertainty quantification for neutron reflectometry practical by computing exact gradients of the forward specular reflectivity model — the Abeles matrix formalism that maps layer thicknesses, scattering length densities, and roughnesses to the measured reflectivity — with respect to every unknown parameter. This is claimed to be the first time exact gradients through the reflectivity are available, unlocking gradient-based inference schemes that gradient-free MCMC cannot use. On a benchmark quartz film, the posterior-mean fit reaches a chi-squared of 1.22 versus a published 1.32, and the Hamiltonian Monte-Carlo sampler (NUTS) builds effective sample size steadily where a gradient-free adaptive MCMC scheme stalls. On a 59-parameter co-fit of four organic LED devices, variational inference returns an approximate posterior in under 20 seconds per run, putting UQ on the timescale of fast kinetic experiments. The authors argue this makes surrogate machine-learning models unnecessary for fast NR inversion and release an open-source library so the community can build on the approach.","feed_headline":"Cut neutron reflectometry fitting from hours to seconds","feed_subtitle":"Exact gradients through the physical reflectivity model power the speed-up — no machine-learned surrogates needed.","key_machinery":"The load-bearing object is the exact gradient of the specular reflectivity R̂(Q, θ) obtained by automatic differentiation through the Abeles matrix formalism, the closed-form multilayer model that maps thicknesses, scattering length densities, and interfacial roughnesses to reflected intensity. That gradient flows into any error function or into the evidence lower bound (ELBO) of a variational surrogate q(θ; φ), turning NR inversion into the sort of optimisation problem that ADAM, the NUTS Hamiltonian sampler, and reparameterised variational inference already solve efficiently. The single-sample reparameterisation estimate of the ELBO gradient is what makes VI fast, and symplectic leapfrog i","core_discovery":"For the first time in neutron reflectometry, exact gradients of the specular reflectivity with respect to all slab and instrument parameters are computed, by differentiating through the Abeles matrix formalism and its smearing kernels. Because the gradients are exact and cheap, the full gradient-based Bayesian toolbox applies directly to the physical forward model: Hamiltonian Monte-Carlo (specifically the NUTS variant) converges to well-mixed posteriors within about 2000 samples per chain where a gradient-free sample-adaptive MCMC scheme stalls, and variational inference fits a 59-parameter joint posterior over four organic LED devices in under 20 seconds per run, with all predictive fits b","pith_inferences":["If the gradient route holds up, it plausibly transfers to any indirect scattering technique with a differentiable forward model — small-angle scattering, ellipsometry, grazing-incidence scattering — since nothing in the method is specific to reflectometry beyond the kernel; a direct test is re-running the same VI/HMC pipeline on those forward models.","The seconds-scale runtime opens the door to closed-loop process control and active experimental design at the beamline without surrogates, for example using the VI posterior of one measurement to choose settings for the next; the paper gestures at this via Bayesian optimisation but does not demonstrate it.","Because the surrogate q is chosen to be Gaussian, VI will understate uncertainty whenever the true posterior is multimodal or skewed, which the phase-loss degeneracy of reflectometry makes likely; a testable extension is a more flexible surrogate family such as mixtures or normalising flows to recover tails while keeping the runtime short.","The posterior differences between OLED devices, such as the thicker oxide inferred for the 140°C-annealed sample, come from a single measurement each; replicate devices would reveal whether those differences are genuine sample variability or slack in the slab-model degrees of freedom."],"forward_implications":["Gradient-based optimisation with ADAM on exact gradients fits the quartz benchmark to chi-squared 1.302, slightly below the published best of 1.32, with all layer and instrument parameters free.","NUTS reaches well-converged effective sample sizes within roughly 2000 samples per chain where a gradient-free sample-adaptive MCMC scheme loses effective sample size, making HMC a practical replacement for MCMC when accurate UQ is required.","Variational inference produces a full approximate posterior over 59 joint parameters, co-fitting four OLED devices with shared instrument parameters, in under 20 seconds per run, bringing UQ to the timescale of fast kinetic NR experiments.","VI's mode-seeking tendency to understate posterior variance is acknowledged and demonstrated on a lipid bilayer benchmark against HMC, giving practitioners a concrete picture of the speed-versus-fidelity trade-off.","The library makes the kernels, smearing options, gradients, just-in-time compilation, and GPU parallelism available so the demonstrated optimisation and inference schemes can be reproduced and extended."],"supporting_citations":[{"why":"Supplies the Abeles matrix formalism, the exact forward reflectivity model that the paper differentiates.","marker":"[23]"},{"why":"Supplies the JAX automatic-differentiation ecosystem on which the gradient computation and GPU parallelism are built.","marker":"[34]"},{"why":"Supplies the no-U-turn sampler, the adaptive HMC variant used for the quartz posterior and the ESS comparison.","marker":"[29]"},{"why":"Supplies the numpyro implementations of NUTS and variational inference used in both case studies.","marker":"[37]"},{"why":"Supplies the ADAM optimiser used both for gradient-descent fitting of the quartz data and for optimising the ELBO in VI.","marker":"[24]"},{"why":"Supplies the quartz benchmark dataset and the published chi-squared of 1.32 that the paper's fits are compared against.","marker":"[38]"},{"why":"Supplies the refnx benchmark lipid bilayer system used in Appendix B to compare VI against HMC.","marker":"[9]"},{"why":"Supplies the reparameterisation trick that makes the single-sample ELBO gradient estimate used by VI valid and computable.","marker":"[33]"}],"fun_headline_variants":["Neutron reflectometry: exact gradients speed Bayesian inversion from hours to seconds","Exact gradients cut neutron reflectometry fitting from hours to seconds","Gradient-based Bayesian inversion for neutron reflectometry: hours to seconds","Neutron reflectometry UQ: surrogate-free, exact-gradient inversion in seconds","Speeding neutron reflectometry inversion with exact gradients, no surrogates"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Everything rests on the assumption that the chosen slab model, with a Gaussian likelihood whose variances are taken directly from the neutron count data, is the true generator of the measured reflectivity; if the reduced data carry correlated errors or the layer model is misspecified, the posterior widths and speed comparisons inherit that error.","fun_headline_variants_meta":{"raw":{"variants":["Neutron reflectometry: exact gradients speed Bayesian inversion from hours to seconds","Exact gradients cut neutron reflectometry fitting from hours to seconds","Gradient-based Bayesian inversion for neutron reflectometry: hours to seconds","Neutron reflectometry UQ: surrogate-free, exact-gradient inversion in seconds","Speeding neutron reflectometry inversion with exact gradients, no surrogates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1304,"prompt_tokens":753,"completion_tokens":551,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":455}},"tokens_in":497,"tokens_out":551,"duration_ms":6616,"temperature":1.0,"reasoning_tokens":455,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:53:40.509528+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit a deliberately misspecified slab model, such as too few layers or a fixed smearing width, to a simulated dataset with known ground truth: if NUTS or VI then report narrow, confident posteriors incompatible with the true SLD profile, the Gaussian-likelihood/slab-model assumption — not the gradient machinery — limits the method. A cheaper check is to re-analyse the quartz data with a likelihood that explicitly models correlations from the known pointwise smearing and see whether the reported posterior widths change materially.","supporting_citations":[{"cited_title":"Abelès.La théorie générale des couches minces","cited_arxiv_id":null,"evidence_quote":"Supplies the Abeles matrix formalism, the exact forward reflectivity model that the paper differentiates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the no-U-turn sampler, the adaptive HMC variant used for the quartz posterior and the ESS comparison."},{"cited_title":"Gutfreund, T","cited_arxiv_id":null,"evidence_quote":"Supplies the quartz benchmark dataset and the published chi-squared of 1.32 that the paper's fits are compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the refnx benchmark lipid bilayer system used in Appendix B to compare VI against HMC."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the reparameterisation trick that makes the single-sample ELBO gradient estimate used by VI valid and computable."}],"review_version":1}