{"id":"6b661fc1-7d49-4100-b20f-bc71d750b0bd","arxiv_id":"1908.09102","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Adaptive simulation-based selection of the next measurement angle reduces the required number of SANS measurements by 50-65 percent in virtual experiments.","lead":"The authors propose two simulation-based strategies that decide which scattering angle to measure next in a SANS experiment, and show in computer simulations that both strategies cut the number of measurements by about half compared with random sampling. The work is a proof-of-concept that adaptive, data-driven sampling can save expensive neutron beamtime.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own noiseless assumption (§3) is the load-bearing weak point: real SANS cost is dominated by counting statistics, and the decision rules in (20) and (24) are sensitive to noise, so the 50–65% duration reduction is not established for actual experiments.","rationale":"I agree with the reader that the noiseless, same-distribution setting is the weakest point. I considered the alternative concern that hyperparameters (M=3, K''=12) were tuned on the same 100 samples; that is real but secondary, because the performance gap is large and would likely survive a validation split. The noise issue is more load-bearing: it connects the decision rule directly to the physical quantity (measurement time) the paper claims to reduce, and the authors' own escape ('noise can be suppressed by extending measurement duration') conflicts with the duration claim. A noise-rescaled experiment is the minimal check. The reader's CONDITIONAL verdict is appropriate; adding noise could either confirm or overturn the magnitude of the claimed saving, so no verdict change is needed.","tokens_in":13015,"tokens_out":8627,"duration_ms":90179,"concrete_test":"Simulate Poisson counting statistics on the same 100 virtual samples: for a fixed exposure time t per point, draw observed counts N(q) ~ Poisson(t·I(q)) (optionally plus background), set I_obs(q)=N(q)/t, and rerun Method 1, Method 2, MV, GP, and random sampling. Record the number of measurements and the total exposure time needed to reach the discrepancy that random sampling achieves with 40 noiseless measurements. Vary t over a realistic SANS range (e.g., 10^2 to 10^6 counts at the high-q end). If Method 2's saving drops below ~50% or the selected q sequence changes substantially, the noiseless benchmark is the source of the claimed acceleration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 states 'Throughout this work, for simplicity, we neglect the noise associated with measurements,' and the Section 5 benchmark inherits that assumption. The headline claim, however, is about experimental duration: 'the experimental duration required to achieve satisfactory accuracy was reduced by 50–65%.' In real SANS the time needed at each q is set by counting statistics, so suppressing noise by extending measurements is in direct tension with reducing duration. The adaptive rules are built on pointwise intensity comparisons: Method 1's similarity score (20) and Method 2's acquisition function (24) both divide intensity differences by a reference intensity I0(q). With Poisson noise, these quantities contain a statistical component whose relative size grows where I0(q) is small (typically at high q). The 'disagreement' |I^(1)-I^(2)|/I0 in (24) can then be dominated by fluctuations rather than by model uncertainty, so the selected next q need not be informative. The reported 2–3x savings are thus a property of the noiseless synthetic benchmark, not a demonstrated property of SANS experiments. The authors acknowledge real-experiment tests are future work, but the abstract's duration-reduction framing invites the stronger reading.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript addresses adaptive measurement planning for small-angle neutron scattering (SANS). The authors assume the scattered intensity is given by I(q) = ∫ p(r) F(q,r)^2 dr with a spherical form factor, and they propose two simulation-based strategies. Method 1 precomputes a greedy optimal measurement plan for each virtual sample, stores the plans in a database, and during an experiment retrieves the plan of the M most similar database samples. Method 2 stores only simulated intensities, forms two independent predictions of unmeasured intensities—one from a database nearest-neighbor average and one from an indirect Fourier transform (IFT) of the measured data—and measures next at the q of maximum normalized disagreement. The methods are benchmarked on 100 virtual samples drawn from the same generator as the database, against random sampling, a Gaussian-process-based fixed ordering, and a maximal-variance baseline. The authors report that Method 1 and Method 2 reach the discrepancy of 40 random measurements in about 20 and 13–15 measurements, respectively, which they translate into a 50–65% reduction in experimental duration.","tokens_in":13246,"tokens_out":8186,"duration_ms":82176,"significance":"If the results survive more realistic noise and distribution-shift tests, this is a useful contribution to adaptive experimental design in scattering science. The numerical comparison is careful in several respects: all methods share the same IFT evaluation pipeline, the database and test samples are drawn from a common generative model, 100 virtual samples are averaged, and standard deviations are reported for the main metric. The paper is also honest that this is a numerical study and that real SANS/SAXS tests are future work. The central limitation is that the benchmark is noiseless and in-distribution; because the abstract and conclusions state the outcome in terms of 'experimental duration,' the practical significance as written exceeds what the simulations can establish.","major_comments":[{"comment":"The paper's headline claim of a 50–65% reduction in 'experimental duration' is not supported by the noiseless benchmark. Section 3 explicitly says noise is neglected and that in practice noise can be suppressed by extending measurement duration, but duration is exactly the resource the abstract claims to save. Table 1 counts noiseless measurements, not time. Under counting statistics, the time needed at each q is set by the required signal-to-noise ratio, and the acquisition rules in Eqs. (20) and (24) divide intensity differences by a reference intensity I0(q); wherever I0(q) is small, generally at high q, the statistical component of these ratios can dominate the model-disagreement signal. The reported factor 2–3 is therefore a property of the noiseless synthetic benchmark, not a demonstrated property of SANS experiment duration. I ask for either a simulation with Poisson noise and an explicit integration-time model, or a revised abstract and conclusions that claim only a reduction in the number of noiseless measurements.","section":"Section 3; Eq. (20); Eq. (24); Table 1"},{"comment":"Hyperparameters M=3 for Method 1 and K''=12 for the MV baseline are reported as empirically best, apparently on the same 100 virtual samples used for the benchmark. The text states 'Empirically, we found that M = 3 was most effective' and 'we tested K'' = 3, 6, 12, 24 and found that K'' = 12 performed the best.' If these values were selected on the same test samples, the averaged curves in Figure 6 and the entries in Table 1 are optimistically biased, and the comparison with the MV baseline is not on equal footing. Please clarify whether a separate validation set or a nested evaluation was used, or report the sensitivity of the results to M and K''.","section":"Section 5.1; Section 5.2"},{"comment":"The database and the 100 test samples are generated from the same distribution family, so the similarity searches in Method 1 and Method 2 are evaluated entirely in distribution. Real SANS samples need not lie in this family, and the performance gain depends on the database containing close neighbors. A robustness experiment in which the database and test distributions differ (for example, log-normal or Schulz size distributions, or a perturbed version of Eq. (25)) would substantially strengthen the claim that the methods accelerate actual SANS experiments. As written, the conclusion that SANS experiments can be 'sped up by a factor of 2–3' should be scoped to same-distribution virtual samples.","section":"Section 5.1, Eqs. (25)–(27); Section 6"}],"minor_comments":[{"comment":"The notation 'K = 102' and 'K > 102' appears to mean 10^2; as printed, K = 102 is a specific integer and the following sentence is confusing. Please typeset the powers of ten correctly.","section":"Section 5.1"},{"comment":"The reference intensity I0(q) is used in the similarity measure and in the acquisition function but is never specified. Please state how I0(q) is chosen in the numerical experiments, since the behavior of both methods depends on it.","section":"Eqs. (20), (24); Figure 5"},{"comment":"The author name 'Spalzzi' in Reference [19] appears misspelled; the usual spelling is 'Spalazzi.' Please check.","section":"Reference [19]"},{"comment":"The toy problem labels the two approaches 'deductive' and 'inductive,' and the text then says that Method 1 for SANS is akin to the toy's inductive method while Method 2 for SANS is hybrid. This reversal is easy to misread; a short mapping table or a change of terminology would improve clarity.","section":"Section 2; Section 4"},{"comment":"Please state explicitly whether the shaded bands are standard deviations over the 100 virtual samples or standard errors, and describe how the per-sample number of measurements to reach the threshold discrepancy is computed (for example, by interpolation between integer measurement counts).","section":"Figure 6; Table 1"}],"recommendation":"major_revision","confidential_remarks":"The main risk is that the abstract and conclusions overstate the experimental relevance of a noiseless, in-distribution simulation. If the authors add a noise model or rescale the claims, the paper could be suitable for publication. I would consider sending the revised version to an experimental SANS practitioner as a check on the duration argument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a careful numerical demonstration, not a proven experimental speedup. The simulation work is internally consistent and the methods are described well enough to reproduce the ideas, but the abstract's \"experimental duration reduced by 50–65%\" framing invites a stronger reading than the evidence supports.\n\nWhat is genuinely new: applying adaptive, database-driven measurement planning to SANS appears to be a new application, and the problem formulation is clear. Method 2 — comparing an IFT-based prediction against a nearest-neighbor database average and measuring the next q at their maximal normalized disagreement — is simple but sensible. The benchmark is also fairly constructed: the IFT setup is common to all methods, 100 virtual samples are averaged, standard deviations are reported, and the GP and MV baselines give the proposed methods a real run for their money. The authors also deserve credit for being explicit about the noiseless assumption and about the computational cost of building the Method 1 database.\n\nThe soft spots are the usual ones, but they are not minor. First, Section 3 states that noise is neglected \"for simplicity,\" and the whole benchmark inherits that. Real SANS is dominated by counting statistics, and both the similarity score (20) and the acquisition rule (24) compare pointwise intensities, dividing by a reference intensity. In the presence of Poisson noise, these quantities develop a statistical component that can dominate the disagreement signal, especially where the intensity is small. So the 2–3× savings are a property of the noiseless, same-generator synthetic benchmark, not a demonstrated property of SANS experiments. The authors note that real tests are future work, but the conclusion leans on the stronger claim.\n\nSecond, the hyperparameters M = 3 for Method 1 and K'' = 12 for the MV baseline were selected by peeking at the same 100 virtual samples used for the headline comparison, with no independent validation split. That inflates the reported gains relative to what one should expect on a genuinely new sample. This is fixable.\n\nThird, no code or data were released, so the simulation itself cannot be independently checked. That matters less than the noise issue, but it is a real obstacle.\n\nBottom line: the central simulation claim — that these adaptive methods outperform the stated baselines in the described setting — holds up. The extension to real experiments does not yet. This paper is for people working on SANS/SAXS measurement planning or on active learning for physical measurements generally. It deserves serious refereeing, but a referee should send it back for either a noise-robust simulation or a real experiment, plus a proper validation split. I would not cite it as evidence for experimental speedups yet.","headline":"A clean simulation study of adaptive SANS sampling whose 50–65% savings claim outruns the evidence, because the benchmark is noiseless and the hyperparameters are tuned on the same test set.","tokens_in":13799,"tokens_out":1738,"would_cite":false,"duration_ms":20073,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adaptive simulation-based sampling cuts SANS measurements by half or more.","keywords":["small-angle neutron scattering","SANS","adaptive sampling","sequential experiment design","indirect Fourier transform","simulation-based machine learning","materials informatics","scattering intensity database"],"falsifier":"Run the same adaptive-sampling benchmark with Poisson counting noise added to the simulated intensities - for example, setting the detector counts so that the relative error at each q matches typical SANS statistics - and measure the number of measurements needed to reach a fixed L1 distance from the true size distribution; if the reduction versus random sampling drops markedly or disappears, the noiseless result is an artifact of the idealization.","tokens_in":1382,"feed_emoji":"🧪","tokens_out":3327,"duration_ms":53289,"temperature":0.7,"pith_summary":"This paper proposes that small-angle neutron scattering (SANS) experiments can be accelerated by using a precomputed database of simulated scattering intensities to decide adaptively where to measure next. The authors claim that, in numerical simulations of 100 virtual materials, their two methods reduce the experimental duration needed to reach a given accuracy by 50-65% compared with random sampling, amounting to a factor of 2-3 speedup. The best method, which combines an indirect Fourier transform estimate with a database of similar intensities, reached the accuracy of 40 random measurements in about 13-15 measurements. If this holds in real experiments, it would make sequential SANS measurements substantially cheaper and could extend to small-angle X-ray scattering and other sequential inverse problems.","feed_headline":"Simulation-guided sampling halves SANS measurement time","feed_subtitle":"Database-driven adaptive sampling reaches the same accuracy with 13-15 measurements where random sampling needs 40.","key_machinery":"The central machinery is a database of simulated scattering intensities, computed from the forward model I(q) = integral of p(r) F(q,r)^2 dr for a large set of virtual size distributions, together with a normalized similarity measure and the indirect Fourier transform (IFT) that reconstructs p(r) from measured intensities. Method 2's decision rule, q_{n+1} = argmax |I^(1)(q)-I^(2)(q)|/I_0(q), selects the next measurement as the point of maximal disagreement between a database-based intensity guess and an IFT-based intensity guess, thereby targeting the region where the current model is most uncertain. Method 1 instead stores an entire optimized measurement plan per virtual sample and reuses plans from the most similar database entries.","core_discovery":"The paper's central claim is that the sequential measurement problem for SANS, where the goal is to reconstruct a size distribution p(r) from scattering intensities I(q) measured at a sparse set of q values, can be solved far more efficiently by adaptive, simulation-based sampling than by non-adaptive random or Gaussian-process baselines. The authors propose two adaptive methods: Method 1 retrieves a precomputed measurement plan for the most similar virtual samples and follows it, while Method 2 predicts future intensities in two independent ways - one from a database of simulated intensities and one from an indirect Fourier transform of the current data - and measures next at the q where the two predictions most strongly disagree. In their noiseless simulations, Method 2 achieved around a 60-65% reduction in the number of measurements required to match random sampling at 40 measurements, and Method 1 achieved roughly a 50% reduction, as measured by both L1 distance and Kullback-Leibler divergence of the reconstructed size distribution.","pith_inferences":["In real SANS experiments, Poisson counting noise will likely erode the reported savings: the entire benchmark neglects measurement noise, and both the IFT reconstruction and the pointwise intensity comparison are sensitive to noisy intensities, so the 50-65% reduction is an upper bound for practical settings.","The database approach assumes that the new sample is drawn from the same family of size distributions used to generate virtual samples; out-of-distribution samples such as anisotropic or interacting scatterers would probably degrade both methods, and the paper itself leaves non-spherical scatterers to future work.","The generic principle - measure where two independent forward-model-based predictions diverge most - applies to any inverse problem y = F(x) + n with a known forward model, so the method could be ported to other sequential characterization techniques beyond scattering.","Method 1's dependence on a large precomputed database means its net benefit is only realized after many experiments amortize the database construction cost; Method 2's cheaper database makes it the more practical first choice."],"forward_implications":["SANS experiments could be completed in roughly one-third to one-half the time for the same accuracy, since the tested methods reduce the number of measurements needed by 50-65% relative to random sampling.","Method 2 reached the accuracy of 40 random measurements with about 13.4 measurements by L1 distance and 15.2 by KL divergence, suggesting substantial savings accumulate already in early measurements.","Because small-angle X-ray scattering obeys the same forward relation between size distribution and intensity, the same adaptive planning procedure is expected to accelerate SAXS as well.","The two non-adaptive baselines, Gaussian-process-prioritized sampling and maximal-variance sampling, plateau at low accuracy after an initial fast drop, indicating that fixed or variance-based orderings cannot match the information gain of a forward-model-informed adaptive choice.","Once a database is built, it can be reused across many experiments; the authors note the database for Method 1 costs tens of hours to prepare while Method 2's database takes under an hour."],"supporting_citations":[{"why":"Provides the standard reference for small-angle scattering theory and data interpretation.","marker":"[12]"},{"why":"Provides the form factor for homogeneous spheres and the relation between intensity and size distribution used in the forward model.","marker":"[13]"},{"why":"Introduces the indirect Fourier transform (IFT) method used to reconstruct the size distribution from measured intensity.","marker":"[15]"},{"why":"Supplies the case-based planning paradigm that Method 1 adapts to SANS measurement planning.","marker":"[19]"},{"why":"Provides the multi-armed bandit framework that motivates the exploration/exploitation balance in Method 1's selection of multiple similar plans.","marker":"[20]"},{"why":"Defines Gaussian process regression and automatic relevance determination used in the GP baseline method.","marker":"[22]"}],"fun_headline_variants":["SANS experiments 60% faster with simulation-based sampling","Adaptive SANS sampling: 15 measurements match random 40","Machine learning cuts SANS measurements in half","Simulation-guided SANS: smarter sampling, fewer runs"],"cache_read_input_tokens":15872,"weakest_assumption_plain":"The benchmark assumes that measurements are noiseless and that the real sample's size distribution comes from the same generative family as the simulated virtual samples, so the measured savings of 50-65% may not transfer to noisy, out-of-distribution experimental conditions.","fun_headline_variants_meta":{"raw":{"variants":["SANS experiments 60% faster with simulation-based sampling","Adaptive SANS sampling: 15 measurements match random 40","Machine learning cuts SANS measurements in half","Simulation-guided SANS: smarter sampling, fewer runs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1246,"prompt_tokens":840,"completion_tokens":406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":456,"completion_tokens_details":{"reasoning_tokens":340}},"tokens_in":456,"tokens_out":406,"duration_ms":4789,"temperature":1.0,"reasoning_tokens":340,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:21:33.538254+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same adaptive-sampling benchmark with Poisson counting noise added to the simulated intensities - for example, setting the detector counts so that the relative error at each q matches typical SANS statistics - and measure the number of measurements needed to reach a fixed L1 distance from the true size distribution; if the reduction versus random sampling drops markedly or disappears, the noiseless result is an artifact of the idealization.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the standard reference for small-angle scattering theory and data interpretation."},{"cited_title":"Small-Angle Scattering","cited_arxiv_id":"1901.07353","evidence_quote":"Provides the form factor for homogeneous spheres and the relation between intensity and size distribution used in the forward model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the indirect Fourier transform (IFT) method used to reconstruct the size distribution from measured intensity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the case-based planning paradigm that Method 1 adapts to SANS measurement planning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the multi-armed bandit framework that motivates the exploration/exploitation balance in Method 1's selection of multiple similar plans."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines Gaussian process regression and automatic relevance determination used in the GP baseline method."}],"review_version":1}