{"id":"c0b84259-294d-4816-ad35-2e0280597fb7","arxiv_id":"2608.10350","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An AI surrogate plus Bayesian inference maps neutron-scattering data into uncertainty over spin-Hamiltonian parameters and picks the next measurement angle that most reduces that uncertainty.","lead":"This paper builds a machine-learning surrogate for neutron-scattering spectra and uses it to quantify which spin-Hamiltonian parameters a given measurement can actually resolve, with uncertainty. It demonstrates on the quantum magnet NiPS3 that powder and single-crystal data can be combined adaptively, and that the estimated observation geometry predicts where remaining parameter ambiguity will be.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The posterior used for the uncertainty claim is built from a non-likelihood robust score (Supp. Note 2), so the 'parameter uncertainty' that observation geometry predicts is not demonstrated to be calibrated; the uncertainty-aware central claim lacks a coverage check.","rationale":"The reader's weakest_assumption is surrogate accuracy, and their conditions include calibration. I think the calibration issue is the more load-bearing single concern. Even with a perfect surrogate, the posterior from Eqs. S4–S8 has no demonstrated probabilistic calibration: the robust score discards the Poisson likelihood normalization in a way that makes posterior widths depend on min-max rescaling across particles. The paper transparently labels this a generalized posterior, which is a credit, but it means the abstract claim of propagating posterior uncertainty is not yet supported. The geometry-versus-uncertainty correlation in Fig. 3(e) would be conclusive only if the vertical axis were a calibrated uncertainty. Simulation-based calibration is a standard, inexpensive check that directly settles this. I do not see an internal inconsistency in the metric construction itself; the equal-weight L2 metric is a modeling choice, and Fig. 3(e) provides some empirical support. The surrogate training-range/prior mismatch for Ax is a secondary concern but less central than calibration. Verdict remains CONDITIONAL: the paper is a credible methodological contribution, but the uncertainty quantification needs calibration evidence. No change to reader's verdict.","tokens_in":27188,"tokens_out":8109,"duration_ms":81312,"concrete_test":"Run simulation-based calibration on the synthetic benchmark: draw 100 θ* from the prior, generate Poisson-noised single-crystal and powder spectra as in Fig. 3, run the exact SMC/robust-score pipeline, and record whether θ* falls inside the 68% and 95% highest-posterior-density or quantile regions of each marginal. Report empirical coverage and also compare posterior widths to the actual root-mean-square error of point estimates. If coverage deviates substantially from nominal, or widths do not track RMSE, the posterior is not calibrated and the uncertainty-aware claim fails; if coverage is close, the robust-score concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assertion is that the observation geometry predicts the parameter uncertainty remaining after inference. The posterior used for that inference is not a calibrated Bayesian posterior: Supplementary Note 2 replaces the Poisson log-likelihood with (i) a per-particle softmax over measured points, (ii) a per-point min-max normalization across particles, and (iii) a mean log-score, and then states that the result 'should be interpreted as an effective (generalized) posterior' rather than the posterior of the raw Poisson likelihood. No coverage or calibration check is reported anywhere in the paper. Consequently the posterior widths plotted in Fig. 3(e) have no demonstrated relationship to the actual remaining ambiguity in θ; they are a function of the arbitrary normalization choices in the robust score. If those widths are not calibrated, Fig. 3(e) only shows that the sloppiness score correlates with the width of an auxiliary score-shaped distribution, not with true parameter uncertainty. The same uncalibrated posterior feeds the BOED utility Eq. (10), so the experimental design and the uncertainty-aware claim inherit the problem. This is more load-bearing than surrogate MSE: even a perfect surrogate cannot turn an uncalibrated score into a calibrated uncertainty.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for uncertainty-aware Hamiltonian inference and adaptive experimental design in quantum magnets, demonstrated on NiPS3 with multimodal inelastic neutron scattering (INS). The central idea is to equip Hamiltonian-conditioned neural surrogates with an observation geometry: the pullback metric G_mn(θ) = <∂_m S, ∂_n S>_Ω (Eq. 3) computed from surrogate derivatives, whose inverse diagonal entries define parameter sloppiness scores. These scores are claimed to predict, a priori from the forward model, the posterior uncertainty remaining after Bayesian inference (Fig. 3e). The framework is demonstrated with simulated sequential single-crystal measurements, with and without powder-informed initialization, and with experimental powder and single-crystal INS data, where posterior samples are checked against independent Sunny/LSWT spectra (Fig. 4b). The paper also introduces a Poisson-derived robust particle score (Supplementary Note 2) that defines an 'effective (generalized) posterior', and a Bayesian experimental design utility based on predictive spectral variance (Eq. 10).","tokens_in":27456,"tokens_out":8056,"duration_ms":74057,"significance":"If the central claims hold, the paper offers a practical and general methodology for a real problem: determining which Hamiltonian parameters are identifiable from a given scattering experiment and which additional measurements best resolve the remaining ambiguity. The use of a differentiable surrogate to compute the observation geometry, the sequential powder-then-single-crystal workflow, and the posterior-predictive validation against independent Sunny calculations (Fig. 4b) are genuine strengths. The paper also provides detailed architecture and data-generation descriptions, which is valuable for reproducibility. However, the load-bearing uncertainty claims currently rest on an uncalibrated generalized posterior and on surrogate derivatives that are not validated against the physics forward model. These gaps are fixable but are central to the paper's stated contribution.","major_comments":[{"comment":"The particle weights that define the posterior behind Fig. 3(e) and the BOED utility in Eq. (10) are updated with the min-max normalized 'robust score' ℓ_n defined in Eqs. (S4)–(S8), and the text states that the result 'should be interpreted as an effective (generalized) posterior' rather than a likelihood-based posterior. No coverage or calibration check is reported anywhere in the paper. Because the central claim is that the observation geometry predicts the parameter uncertainty remaining after inference, the posterior widths in Fig. 3(e) must be shown to have a well-defined relationship to actual remaining ambiguity; as it stands, they are a function of the particular normalization choices in the score. I would require a calibration or coverage validation on simulated data, for example empirical coverage of posterior credible intervals over many ground-truth draws under Poisson noise, or a comparison with an exact Poisson-likelihood posterior for a subset of parameters, before the uncertainty-aware claims can be assessed.","section":"Supplementary Note 2; Methods, 'Bayesian inference and experimental design'"},{"comment":"The observation metric G, the sloppiness scores S_m, and the stiff/sloppy eigen-directions in Fig. 2 are all computed from derivatives of the neural surrogate, yet the paper reports only the validation MSE (0.01319) and gives no held-out spectral error or derivative error against Sunny. If the surrogate is biased in any region of parameter space, the metric, the sloppiness ordering, and the posterior widths all inherit that bias. Please provide a quantitative comparison of surrogate predictions and surrogate derivatives with Sunny/LSWT at held-out Hamiltonian parameters, for example relative L2 error of ∂S/∂θ_m as a function of parameter location, and show that the sloppiness ordering and the stiff/sloppy directions are stable under this validation.","section":"Eqs. (3)–(5); Methods, 'Neural surrogate model'"},{"comment":"The positive correlation in Fig. 3(e) is computed between two quantities generated by the same surrogate and the same effective likelihood, so it is a self-consistency check rather than an independent validation of the 'a priori predictor' claim; moreover, it is based on only nine points, one per Hamiltonian parameter. Please report the correlation coefficient with its uncertainty, state explicitly that this is an internal consistency result, and ideally re-evaluate the sloppiness scores and posterior widths using independent Sunny spectra and a calibrated posterior for at least a subset of the benchmark cases.","section":"Fig. 3(e); 'Posterior-guided multimodal Hamiltonian inference'"},{"comment":"The experimental powder spectrum used for inference is the output of source-separation and feature-enhancement vision transformers whose details are deferred to a separate publication, and the note itself states that the result 'should be interpreted as an effective magnetic-spectrum estimate' and that 'quantitative differences in the inferred Hamiltonian may depend on the specific preprocessing procedure.' The demonstration with experimental powder data is therefore conditional on an unpublished, model-dependent preprocessing pipeline. Either the preprocessing must be described and validated against alternative approaches, or the main-text claims about experimental multimodal validation should be correspondingly softened.","section":"Supplementary Note 10; 'Validation with experimental multimodal INS data'"}],"minor_comments":[{"comment":"The predictive-variance utility is described as an approximation to expected information gain, but the text does not discuss the conditions under which this approximation is accurate or its known biases; a clarifying sentence about its heuristic status would be helpful.","section":"Eq. (10)"},{"comment":"Please report the numerical correlation coefficient, and preferably a rank correlation, for the nine plotted points, since visual inspection of a nine-point scatter is not sufficient to support the claimed predictive relation.","section":"Fig. 3(e)"},{"comment":"The training ranges θ_i ∼ U(0, 2θ_ref,i), or the reflected version, are broader than the prior bounds in Eq. (14); please clarify whether any prior-supported region lies outside the training support and whether surrogate extrapolation is a concern there.","section":"Methods, 'Training data generation'"},{"comment":"Because the sloppiness scores are normalized independently within each modality, the cross-modality comparison in the text is only about relative rankings; this is stated in the caption but should also be stated in the main text.","section":"Fig. 2(c) and main text"},{"comment":"The scaling argument that a nine-dimensional grid with the same forward-model budget gives 'about three grid points along each parameter dimension' assumes a uniform Cartesian grid; a brief clarification that other sampling strategies may fare differently would improve the discussion.","section":"Results, 'Validation with experimental multimodal INS data'"}],"recommendation":"major_revision","confidential_remarks":"This is a promising methods paper with a clear central idea and a physically meaningful demonstration. The two main gaps I see are the calibration of the generalized posterior and the lack of derivative-level validation of the surrogate; both are load-bearing for the paper's central uncertainty predictions and are fixable with additional experiments and analysis. I would also flag the reliance on an unpublished vision-transformer preprocessing pipeline for the experimental powder data when the revision is discussed with the authors. No concerns about authorship or citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing here is the integration: observation geometry (the sloppiness-style metric G from Eq. 3), a FiLM-SIREN surrogate, SMC Bayesian inference, and BOED on crystal orientation, all applied to a nine-dimensional NiPS3 Hamiltonian using real powder and single-crystal INS data. Each ingredient is established, but the combined multimodal demo is new. The paper also does something right that many ML inverse papers skip: it validates posterior samples against independent Sunny spectra (PCC scatter in Fig. 4b) and is explicit about what the framework does and does not claim.\n\nThe soft spots are real. The surrogate is validated only by training/validation MSE (0.01319). No held-out spectral or derivative error against Sunny is reported, and the derivatives define the metric G and the likelihood uses the predictions, so surrogate bias would propagate everywhere. That is fixable but needs doing.\n\nThe bigger issue, which the stress-test note gets right, is that the posterior is not a calibrated Bayesian posterior. Supp Note 2 replaces the Poisson log-likelihood with a per-particle softmax, min-max normalization across particles, and a mean log-score, and the authors call the result an effective generalized posterior. No coverage or calibration check is reported anywhere. Consequently, the posterior widths in Fig. 3e have no demonstrated relationship to true remaining ambiguity; Fig. 3e is partly a self-consistency check because the same surrogate and the same effective likelihood generate both the sloppiness scores and the posterior widths. The BOED utility inherits the same issue. Even a perfect surrogate cannot turn an uncalibrated score into calibrated uncertainty.\n\nOne more, more minor: the finite-Q-binning analysis shifts parameter estimates noticeably (e.g., J3b from 13.52 to 14.86 in Table S2), and it is not clearly stated which posterior corresponds to Fig. 4. The paper says the conclusions are robust; the numbers should be reconciled or the choice stated.\n\nThe observation-geometry idea itself holds up as a diagnostic. The paper deserves a serious referee. My recommendation: send it to peer review, and require (1) a coverage/calibration check of the robust score, on simulated data if necessary; (2) held-out surrogate validation against Sunny including derivatives; (3) clarity on the binning. With those, the claim of uncertainty-aware inference would be supported. As it stands, it is a nice demonstration with an unquantified uncertainty calibration.","headline":"Worth a serious look, but the uncertainty claim needs a calibration check before I'd trust the posterior widths.","tokens_in":28037,"tokens_out":2411,"would_cite":true,"duration_ms":23570,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A spectral geometry predicts which spin-Hamiltonian parameters a neutron experiment can resolve before any data are collected.","keywords":["observation geometry","sloppiness analysis","Hamiltonian inference","Bayesian experimental design","implicit neural representations","inelastic neutron scattering","NiPS3","quantum magnets"],"falsifier":"Generate a held-out set of about 2,000 parameter vectors drawn uniformly from the prior, compute the spin-wave spectra and parameter derivatives with the true forward solver, and compare them with the surrogate's spectra, derivatives, sloppiness scores, and posterior widths obtained by re-running the particle filter with the true solver for a subset of particles; if the surrogate–solver disagreement in any region is comparable to the spectral separation between stiff and sloppy directions, the predicted identifiability ranking and the reported correlation would not survive re-computation.","tokens_in":26952,"feed_emoji":"🧲","tokens_out":12159,"duration_ms":103467,"temperature":0.7,"pith_summary":"This paper tries to prove a practical claim: before any neutron beam time is spent, the forward model alone can tell you which spin-Hamiltonian parameters a proposed measurement will determine and which it will leave ambiguous. The authors define an observation metric on Hamiltonian parameter space from derivatives of a neural surrogate of the dynamical structure factor, so that large-eigenvalue stiff directions are well constrained and small-eigenvalue sloppy directions are not. They show, on simulated and experimental inelastic neutron scattering from the quantum magnet NiPS3, that the sloppiness score computed from this metric correlates with the posterior uncertainty left after sequential Bayesian inference, and that a powder measurement can reshape the choice of subsequent single-crystal orientations. A sympathetic reader would care because it offers a route from which Hamiltonian fits the data to which Hamiltonian the data can actually see, and because the same surrogate-plus-metric machinery transfers to any differentiable forward model.","feed_headline":"Sloppiness maps predict what a neutron experiment can resolve","feed_subtitle":"AI surrogates rank which neutron measurements trim Hamiltonian uncertainty before any beam time is spent.","key_machinery":"The load-bearing object is the observation metric $G_{mn}(\\theta)=\\langle\\partial_{\\theta_m} S,\\partial_{\\theta_n} S\\rangle_\\Omega$, a Gram matrix of spectral derivatives integrated over the accessible momentum–energy domain and approximated with Monte Carlo samples; its eigenvalue decomposition separates stiff parameter combinations, which produce large spectral response, from sloppy ones, which are nearly invisible, and the diagonal of its inverse, after rescaling by parameter magnitudes, gives a per-parameter sloppiness score. The computational enabler is a Hamiltonian-conditioned neural surrogate: a shared embedding network feeding feature-wise linear modulation (FiLM) into modality-specific coordinate networks with periodic activations, returning differentiable spectra $S(\\mathbf{Q},\\omega;\\theta)$ and $S(|\\mathbf{Q}|,\\omega;\\theta)$ almost instantly. A sequential Monte Carlo particle filter supplies the posterior, and a predictive-variance utility selects the next crystal orientation.","core_discovery":"The central claim is that the intensity-induced $L^2$ pullback metric $G_{mn}(\\theta)=\\langle\\partial_{\\theta_m} S,\\partial_{\\theta_n} S\\rangle_\\Omega$, computed by Monte Carlo from surrogate derivatives, defines the local observation geometry of neutron scattering, and that this geometry's stiff and sloppy eigendirections predict the identifiability of microscopic interactions before data are acquired. Concretely, the paper asserts that the diagonal of the inverse scaled metric is an effective a priori predictor of the remaining posterior uncertainty after inference, and demonstrates the correlation for the nine-parameter Hamiltonian of NiPS3 across powder and single-crystal modalities. It further claims that sequential Bayesian inference coupled with this geometry lets a powder posterior initialize and reshape the adaptive selection of single-crystal orientation, and that the resulting joint posterior contains Hamiltonians that reproduce both experimental spectra while explicitly exhibiting which parameters, notably $A_x$, $A_z$, $J_{2a}$, and $J_{2b}$, remain poorly resolved.","pith_inferences":["A direct robustness test would be to recompute the sloppiness ranking with derivatives from the spin-wave solver itself, or from a second independent emulator, on a modest grid; if the ranking and the reported correlation survive, the surrogate is not the source of the geometry.","One could replace SMC-based Bayesian experimental design with a derivative-only criterion, for example selecting the setting that maximizes the volume or smallest eigenvalue of $G$ over the accessible region; the paper's equations make that alternative testable with the same machinery.","Because the experimental powder spectrum enters through an extracted magnon estimate rather than raw counts, the posterior's dependence on that preprocessing remains an open question; the metric provides a quick way to predict which parameters would shift under an alternative extraction.","If surrogate uncertainty were propagated through ensembles or Bayesian networks, the same geometry could flag regions where the emulator itself is untrustworthy, turning model mismatch into a diagnosed quantity rather than a silent bias."],"forward_implications":["Given only the forward model, one can pre-compute which Hamiltonian parameters a proposed single-crystal orientation or powder configuration will constrain, without collecting data.","Powder and single-crystal INS can be joined through the shared Hamiltonian parameter space: powder inference shrinks the posterior before single-crystal beam time, and the subsequent adaptive orientation sequence shifts to target remaining sloppy directions.","Posterior samples, not just a best fit, can be checked against both modalities; the NiPS3 analysis shows many Hamiltonians fit the powder spectrum indistinguishably while single-crystal data discriminate more sharply among them.","Because the metric requires only derivatives of a differentiable forward map, the same observation-geometry diagnostics and design loop extend to other spectroscopic probes and to control variables such as incident energy, magnetic field, or pressure."],"supporting_citations":[{"why":"Supplies the reference NiPS3 Hamiltonian, the single-crystal data, and the seven-parameter model that the paper relaxes to nine parameters.","marker":"[3]"},{"why":"Introduces the sloppiness concept that motivates identifying flat spectral-discrepancy directions as under-constrained parameters.","marker":"[5]"},{"why":"Extends sloppiness analysis to physical models, justifying the interpretation of stiff and sloppy directions.","marker":"[6]"},{"why":"Provides the information-geometric metric framework that the paper's observation metric instantiates.","marker":"[7]"},{"why":"Is the linear spin-wave solver used to generate the 20,000 training spectra for both modalities.","marker":"[23]"},{"why":"Introduces implicit neural representations with periodic activations, the base architecture for the surrogate coordinate networks.","marker":"[43]"},{"why":"Demonstrates INR-based surrogate steering and underlies the predictive-variance utility used to choose the next measurement.","marker":"[45]"},{"why":"Supplies the simulation-based Bayesian optimal experimental design formalism used for orientation selection.","marker":"[48]"},{"why":"Provides the particle-rejuvenation resampling used in the sequential Monte Carlo posterior.","marker":"[57]"},{"why":"Introduces feature-wise linear modulation, the conditioning mechanism that connects Hamiltonian parameters to the coordinate networks.","marker":"[58]"}],"fun_headline_variants":["AI predicts which neutron measurements tighten Hamiltonian bounds","Surrogate geometry forecasts resolution before beam time","Data-driven map ranks neutron experiments by uncertainty cut","Sloppy directions reveal which spin parameters stay hidden","Geometry-based AI selects optimal crystal orientations for neutron shots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument depends on the neural surrogate being an accurate and smooth stand-in for the spin-wave spectra across the entire nine-dimensional parameter range, including the derivatives that build the metric and the likelihood, yet the paper reports only an overall validation error and does not compare held-out surrogate spectra or gradients against the true solver.","fun_headline_variants_meta":{"raw":{"variants":["AI predicts which neutron measurements tighten Hamiltonian bounds","Surrogate geometry forecasts resolution before beam time","Data-driven map ranks neutron experiments by uncertainty cut","Sloppy directions reveal which spin parameters stay hidden","Geometry-based AI selects optimal crystal orientations for neutron shots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000714,"raw_usage":{"total_tokens":3185,"prompt_tokens":895,"completion_tokens":2290,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":2219}},"tokens_in":511,"tokens_out":2290,"duration_ms":12528,"temperature":1.0,"reasoning_tokens":2219,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:22:23.706600+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a held-out set of about 2,000 parameter vectors drawn uniformly from the prior, compute the spin-wave spectra and parameter derivatives with the true forward solver, and compare them with the surrogate's spectra, derivatives, sloppiness scores, and posterior widths obtained by re-running the particle filter with the true solver for a subset of particles; if the surrogate–solver disagreement in any region is comparable to the spectral separation between stiff and sloppy directions, the predicted identifiability ranking and the reported correlation would not survive re-computation.","supporting_citations":[{"cited_title":"Training data generation","cited_arxiv_id":null,"evidence_quote":"Supplies the reference NiPS3 Hamiltonian, the single-crystal data, and the seven-parameter model that the paper relaxes to nine parameters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the sloppiness concept that motivates identifying flat spectral-discrepancy directions as under-constrained parameters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends sloppiness analysis to physical models, justifying the interpretation of stiff and sloppy directions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the information-geometric metric framework that the paper's observation metric instantiates."},{"cited_title":"Blosser, N","cited_arxiv_id":null,"evidence_quote":"Is the linear spin-wave solver used to generate the 20,000 training spectra for both modalities."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces implicit neural representations with periodic activations, the base architecture for the surrogate coordinate networks."},{"cited_title":"Capturing dynamical correlations using implicit neural representations","cited_arxiv_id":"2304.03949","evidence_quote":"Demonstrates INR-based surrogate steering and underlies the predictive-variance utility used to choose the next measurement."},{"cited_title":"AIMS: an AI experimentalist turns uncertainty into quantum matter discovery","cited_arxiv_id":"2607.16544","evidence_quote":"Provides the particle-rejuvenation resampling used in the sequential Monte Carlo posterior."},{"cited_title":"Liu and M","cited_arxiv_id":null,"evidence_quote":"Introduces feature-wise linear modulation, the conditioning mechanism that connects Hamiltonian parameters to the coordinate networks."}],"review_version":1}