{"id":"d47b08d7-54e8-4378-b8a2-3b7bc8ea916f","arxiv_id":"2508.19470","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A lensing-inspired binning method that selects one measurement per bin to thin dense nuclear cross section data, yet its toy-model tests show biased parameter estimates.","lead":"The paper introduces a data-thinning method that selects a smaller, representative subset of nuclear cross section measurements before statistical fitting. The method is tested on synthetic data and neutron total cross sections, but its toy-model tests show statistically significant bias in the fitted parameters.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Toy experiment in Table I shows lensed fits deviate significantly from true slope/intercept, undercutting the claim of statistically comparable outcomes.","rationale":"The reader's weakest assumption—that the selection criterion is sufficient to preserve statistical content—is the same load-bearing point I identify, and it is directly contradicted by the paper's own Table I. The paper reports t-tests with P≈0.01 for the intercept in all three variants, yet concludes reliability solely from Chow F tests that compare lensed data to their own bootstrap. This is an internal inconsistency, not a disagreement with consensus: the paper's own falsifiable experiment fails the central claim. A rejection is warranted because the proposed method, as described, does not deliver 'statistically comparable outcomes' even in its simplest validation. The real-data demonstration is qualitative and lacks a quantitative full-data comparison, so it cannot override the toy failure. My read does not change the reader's verdict, so verdict_should_be is UNCHANGED.","tokens_in":10994,"tokens_out":3931,"duration_ms":38031,"concrete_test":"Rerun the simulated-data validation of Sec. IV with N_rep=1000 independent draws from Eqs. (20)–(22), applying each lensing variant with the paper's choice of Nb, h, and retention ratio; fit the linear model to each lensed subset, form 95% confidence intervals using the reported Cramér–Rao covariance, and record empirical coverage of the true slope/intercept (2,2). If any variant's coverage is substantially below nominal, the thinning procedure does not preserve statistical comparability. As a secondary check, fit the full 1000-point sample and compare its parameter confidence region to the lensed subset's region for the same simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that lensing selects a subset yielding 'statistically comparable outcomes' (Conclusion) is contradicted by the paper's own validation. In Table I, the fitted intercept deviates from the true value of 2 for all three variants: percentile 2.62±0.16 (t=3.85, P=0.01), KDE 2.20±0.05 (t=4.07, P=0.01), KDEσ 2.21±0.06 (t=3.46, P=0.01); the percentile slope also deviates (0.98±0.31 vs. 2; t=3.33, P=0.02). These are not borderline results, and the deviations are in a consistent direction (positive intercept), suggesting the selection rule—matching measurements in bin b to percentiles of neighboring bin b+1 (Eqs. 1–2)—systematically biases the reduced sample. The paper's response is to test only whether the lensed data agree with their own bootstrap (Chow F, P>0.1) and to conclude 'the method therefore gives reliable results.' That test is not aimed at the claim: bootstrap self-consistency of a biased selector does not establish that the subset preserves the statistical content of the full sample. Consequently the strongest claim, both in the abstract and conclusion, lacks support from the paper's own experiment. The real-data section (Sec. V) is qualitative and does not provide a quantitative comparison of the lensed fit to the full-data fit, so it cannot repair this defect.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a data-thinning method for nuclear cross-section datasets, inspired by optical lensing. The method bins the data in energy, characterizes each bin's cross-section distribution by percentiles or by a kernel density estimate, and then selects, from each bin, the measurement most similar to the neighboring bin's distribution. This selected subset is intended to be used as a pre-processing step before Bayesian assimilation methods such as FBET and before optical-model fitting with CoH3. The authors validate the method on a simulated linear model with known slope and intercept and on total cross-section data for neutron-induced reactions on 124Sn, 144Sm, 143Nd, and 150Nd, reporting speedups relative to a classical thinning method. The central claim, stated in the abstract and conclusion, is that the lensing procedure preserves critical information and yields a representative subset with statistically comparable outcomes.","tokens_in":11270,"tokens_out":2821,"duration_ms":28246,"significance":"If the central claim were valid, the method would be a practically useful pre-processing tool for nuclear data evaluation, where dense, correlated datasets make full Bayesian assimilation expensive. The lensing idea is intuitive, and the reported computational speedups over the classical thinning baseline are substantial. The paper provides a machine-readable presentation of the algorithm and a transparent toy experiment. However, the toy experiment is the only quantitative test against a known truth, and its results contradict the central claim: the fitted intercepts differ significantly from the true value for all three lensing variants, and the percentile-based slope also differs significantly. The bootstrap-based checks reported in the paper are self-referential and cannot detect systematic selection bias. The real-data section is qualitative and therefore cannot repair this defect. Because the main claim is not supported by the paper's own validation, the significance of the contribution as presented is not established.","major_comments":[{"comment":"The quantitative test against the known linear truth y = 2x + 2 shows statistically significant deviations for all three lensing variants: the percentile-based intercept is 2.62 ± 0.16 (t = 3.85, P = 0.01) and its slope is 0.98 ± 0.31 (t = 3.33, P = 0.02); the KDE and KDEσ intercepts are 2.20 ± 0.05 (P = 0.01) and 2.21 ± 0.06 (P = 0.01), respectively. These results directly contradict the claim in Section VI that the method 'allows analysts to work with a representative subset that yields statistically comparable outcomes.' The deviations are all in the same direction (positive intercept bias), suggesting a systematic effect of the selection rule, not mere sampling noise.","section":"Section IV, Table I"},{"comment":"The Chow F-test and bootstrap distributions compare the lensed dataset only with resamples of itself, not with the full simulated dataset or with the known generating model. A biased selection rule can be perfectly self-consistent under bootstrap resampling while still destroying the statistical content of the original data. The sentence 'The method therefore gives reliable results for all three lensing variants' therefore does not follow from the reported test. The correct comparison, already available in Table I through the t-tests against the true parameters, shows the opposite conclusion, and the paper's discussion omits this discrepancy.","section":"Section IV, Chow F-test paragraph"},{"comment":"The real-data application is purely qualitative. Figures 5 and 6 show lensed and thinned fits overlaid on FBET-smoothed full-data expectations, but no quantitative metric is reported: there is no fitted parameter table, no chi-square for the lensed-data fits versus the full-data fit, and no comparison of uncertainty bands. The text asserts closer agreement of the lensing approach with the FBET-smoothed expectations, but this is a visual impression. A quantitative comparison on at least one nucleus, e.g., n+124Sn, would be needed to substantiate the claim that the reduced subset preserves the statistical content of the full dataset for model fitting.","section":"Section V, Figs. 5 and 6"}],"minor_comments":[{"comment":"In the definition of the KDE, the Gaussian is written as exp(-(y - y_i)^2 / 2h^2) while the sum runs over x_j in the bin; the kernel should be a function of the cross-section values y_j, so the notation appears to contain a typo (y_i should be y_j).","section":"Section II.B, Eq. (3)"},{"comment":"References [20] and [21] are the same publication (Ochotta et al., Quarterly Journal of the Royal Meteorological Society 131, 3427 (2005)) but are cited as if they were distinct works; this duplication should be corrected.","section":"References [20] and [21]"},{"comment":"The prediction intervals are described as sqrt(N_s + 1) times the estimated parameter error and sqrt(N_B + 1) times the lensed-subsample parameter error, but the derivation or citation for this particular scaling is not given; the reader cannot verify that these intervals have the claimed coverage.","section":"Section IV, prediction intervals"},{"comment":"The figures comparing lensed subsets and bootstrap estimates would be easier to interpret if the number of retained points per bin were stated explicitly, since the claimed reduction ratio is central to the method's motivation.","section":"Section V, Figs. 3 and 4"},{"comment":"References [37] and [38] list only the author and 'in prep.' without titles or dates; if these are intended to support the claim of follow-up work, more bibliographic information should be provided.","section":"References [37] and [38]"}],"recommendation":"reject","confidential_remarks":"The reader's stress-test concern is confirmed by the manuscript text: Table I contains statistically significant deviations from the known true parameters, and the paper's discussion of that table substitutes a self-referential bootstrap check for the needed comparison with the truth. The central claim of 'statistically comparable outcomes' is therefore not supported by the paper's own experiment. The idea is not without merit, and a revised version that either modifies the selection rule to eliminate the bias or honestly reframes the method as a fast heuristic for visualization rather than statistically equivalent thinning could be reconsidered. As submitted, however, the mismatch between the validation and the claims is load-bearing and not fixable with minor edits."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing you should know: the toy experiment in Table I undercuts the paper's main claim. Every lensing variant gives a statistically significant bias in the fitted intercept (percentile 2.62±0.16, t=3.85, P=0.01; KDE 2.20±0.05, t=4.07, P=0.01; KDEσ 2.21±0.06, t=3.46, P=0.01), and the percentile's slope is off too (0.98±0.31 vs. 2, t=3.33, P=0.02). The authors only report a Chow F against their own bootstrap, which is self-consistent but doesn't tell you whether the subset preserves the statistical content of the full sample. So the conclusion that the method yields \"statistically comparable outcomes\" is not supported by their own validation.\n\nWhat's new and worth credit: the specific selection rule—matching a point in bin b to the percentiles or KDE of the neighboring bin—is not in the cited Ochotta thinning, and the application to nuclear cross-section data is a reasonable idea. The method is cheap, and the real-data figures show it produces a plausible subset for FBET preprocessing. That's a legitimate starting point.\n\nThe soft spots, in proportion: Table I is the load-bearing flaw, and it's not a small one. The real-data section is qualitative, with no quantitative comparison of the lensed fit to the full-data fit. The free parameters (number of bins, KDE bandwidth, retention ratio) are never set or tested. No code is shipped. There are minor issues: references [20] and [21] are the same paper, and a few typos.\n\nWho gets value from this: nuclear data evaluators wondering whether they can thin before FBET would find the idea appealing, but they can't trust the current validation. The paper deserves a serious referee—there's enough mechanism here that a good reviewer could push it into something usable—but as it stands I would not cite it, and I would not accept it.\n\nRecommendation: send to peer review rather than desk reject, because the flaw is technical and possibly fixable, but expect a major revision or rejection.","headline":"The toy validation contradicts the central claim: the method is a plausible heuristic but has not been shown to preserve statistical comparability.","tokens_in":11836,"tokens_out":5986,"would_cite":false,"duration_ms":51823,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A lensing-inspired data-thinning method reduces nuclear datasets while keeping fit quality.","keywords":["data thinning","optical lensing","nuclear cross sections","Full Bayesian Evaluation Technique","kernel density estimation","statistical reaction model","resonance region","bootstrap uncertainty"],"falsifier":"Run the full Bayesian evaluation on a dataset small enough to be tractable, then lens it with each variant and compare the fitted parameter posteriors to the full-data posteriors; if the lensed posteriors exclude the full-data parameter region, the preservation claim fails. A cheap version already exists in the paper's toy linear model, where the percentile-only variant recovered a slope of $0.98\\pm0.31$ instead of the true value $2$.","tokens_in":10769,"feed_emoji":"🔭","tokens_out":11995,"duration_ms":105801,"temperature":0.7,"pith_summary":"The paper introduces a data-thinning algorithm inspired by optical lensing that compresses dense nuclear cross-section datasets before expensive Bayesian data assimilation. The selection rule keeps, for each energy bin, the measurement whose cross-section value best matches the distribution of measurements in the neighbouring bin, either through min/median/max percentiles or through kernel-density estimates that can weight measurement errors. The authors show on simulated linear data and on total cross sections for neutron-induced reactions on tin-124, samarium-144, neodymium-143, and neodymium-150 that the reduced subsets lead to Full Bayesian Evaluation Technique (FBET) smoothed results and reaction-model fits statistically comparable to the full dataset, while cutting computation time by orders of magnitude relative to classical thinning. If the claim holds, evaluators can work with a representative subset instead of the full dataset, making data assimilation practical for resonance-rich and other dense experimental data.","feed_headline":"New 'lensing' method thins nuclear data while keeping fit quality","feed_subtitle":"A bin-to-bin matching rule keeps representative points for Bayesian fits and beats classical thinning on speed.","key_machinery":"The selection engine is a per-bin likelihood $L_b(y,\\sigma_y)=\\prod_{p\\in\\{0,50,100\\}} e^{-\\frac12((y-y_{p,b})/\\sigma_y)^2}$ that scores a candidate measurement in bin $b$ by how well it reproduces the minimum, median, and maximum of the cross-section distribution in the neighbouring bin; the algorithm sets the neighbouring bin to $b+1$ and takes the argmax over all measurements in bin $b$. A second variant replaces the three percentile values with a Gaussian kernel density estimate $K_h(y\\mid h,b')$ of the neighbouring bin, and a third uses per-point error bars as kernel bandwidths. This matching rule is what carries the argument: it turns 'keep the informative points' into a concrete optimization that preserves the local distribution shape of the data while discarding redundant measurements.","core_discovery":"The central claim is that bin-local distribution matching is a sufficient data-reduction criterion for later model fitting. Concretely, the algorithm bins the data by energy, computes for each bin a target signature—the min/median/max percentiles, or a kernel density estimate optionally using per-point uncertainties—and for every bin picks the measurement from that bin that best matches the signature of the next bin. Repeating this pass removes one point per bin until the desired subset size is reached. The paper argues that this procedure preferentially keeps points near the average cross-section and discards points inside resonance peaks, and it demonstrates that the lensed subsets, after Bayesian smoothing, support reaction-model fits that agree with fits to the full dataset, at a small fraction of the cost of iterative thinning.","pith_inferences":["Inference: because the algorithm preserves local distribution shape rather than extremes, it is best suited to model calibration against average behaviour; analyses whose scientific target is the resonances themselves should check whether lensing removes the signal they need.","Inference: the same bin-to-neighbour distribution matching could transfer to any binned physical dataset whose assimilation cost scales steeply with size, such as opacity tables or equation-of-state point sets.","Inference: a decisive diagnostic is to compare lensed-subset fits with fits on equal-sized random subsets; if random selection performs as well, the method's advantage is purely computational, while if lensing wins, the distribution-matching criterion is doing genuine statistical work."],"forward_implications":["Lensing can serve as a cheap pre-processing step before Bayesian data assimilation, making fits feasible for datasets that would otherwise strain memory and compute budgets.","For resonance-rich nuclei, the lensing subset deliberately avoids resonance peaks and preserves the average cross-section behaviour that the reaction-model parameters respond to.","The error-weighted kernel variant carries the measurement uncertainties into the selection, so the thinned subset remains representative when data quality varies sharply across the energy range.","Runtime for the percentile variant is orders of magnitude below classical thinning: about 0.03 seconds versus 1500 seconds for the tin-124 test case in the paper.","The selected subset is stable under bootstrap resampling, giving analysts a practical route to uncertainty estimates on the thinned data."],"supporting_citations":[{"why":"Defines the Full Bayesian Evaluation Technique that the lensed subsets feed into before model fitting.","marker":"[14]"},{"why":"The iterative thinning algorithm used as the baseline for speed and selection comparisons.","marker":"[21]"},{"why":"The global optical potential whose parameters are adjusted in the code's demonstration fits.","marker":"[22]"},{"why":"The damped least-squares optimization scheme used to fit the model parameters.","marker":"[23]"},{"why":"Bootstrap resampling used to quantify variability of the lensed subsets and fitted parameters.","marker":"[25]"},{"why":"Provides the experimental total-cross-section datasets for the four demonstration nuclei.","marker":"[12]"},{"why":"The reaction-model code in which the optical-model fits are performed.","marker":"[11]"}],"fun_headline_variants":["Lensing-inspired thinning cuts nuclear data, preserves Bayesian fits","Bin-wise matching thins nuclear data faster, keeps fit accuracy","Optical-lensing data thinning speeds nuclear cross-section fitting","Lensing rule picks key nuclear points for fast fitting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that matching a bin's min/median/max or kernel-density profile to the neighbouring bin preserves the statistical content a later model fit needs; the paper's own toy simulation, in which the percentile-only variant recovered a slope of $0.98\\pm0.31$ instead of the true $2$, shows this premise is not automatic.","fun_headline_variants_meta":{"raw":{"variants":["Lensing-inspired thinning cuts nuclear data, preserves Bayesian fits","Bin-wise matching thins nuclear data faster, keeps fit accuracy","Optical-lensing data thinning speeds nuclear cross-section fitting","Lensing rule picks key nuclear points for fast fitting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000292,"raw_usage":{"total_tokens":1620,"prompt_tokens":781,"completion_tokens":839,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":772}},"tokens_in":397,"tokens_out":839,"duration_ms":8565,"temperature":1.0,"reasoning_tokens":772,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:52:18.763497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full Bayesian evaluation on a dataset small enough to be tractable, then lens it with each variant and compare the fitted parameter posteriors to the full-data posteriors; if the lensed posteriors exclude the full-data parameter region, the preservation claim fails. A cheap version already exists in the paper's toy linear model, where the percentile-only variant recovered a slope of $0.98\\pm0.31$ instead of the true value $2$.","supporting_citations":[{"cited_title":"Cardinali, Observation influence diagnostic of a data assimilation system, in Data Assimilation for Atmo- spheric, Oceanic and Hydrologic Applications (Vol","cited_arxiv_id":null,"evidence_quote":"The iterative thinning algorithm used as the baseline for speed and selection comparisons."},{"cited_title":"Ochotta, C","cited_arxiv_id":null,"evidence_quote":"The damped least-squares optimization scheme used to fit the model parameters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Bootstrap resampling used to quantify variability of the lensed subsets and fitted parameters."},{"cited_title":"The well-known problem in such data aggre- gation is that it neglects experimental correlations, such as those between experiments [13]","cited_arxiv_id":null,"evidence_quote":"Provides the experimental total-cross-section datasets for the four demonstration nuclei."}],"review_version":1}