{"id":"30dd0702-3444-421b-9da1-efa319c0a461","arxiv_id":"2412.18872","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A residual CNN trained on neutron monitor data reconstructs AMS-02 proton flux from 1 to 100 GV at daily and hourly resolution, spanning 2011 to 2024.","lead":"This paper uses neural networks to turn ground-based neutron monitor counts into daily and hourly cosmic-ray proton flux estimates, filling gaps in AMS-02 satellite measurements and extending the record to 2024. A generalist might read it because it offers a machine-learning route to continuous, high-resolution cosmic-ray data that spacecraft do not provide.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Random day-split test R2 likely overstates out-of-time generalization; post-2019 daily flux is validated only via coarser monthly BR bins, leaving the central 2011-2024 daily record under-tested.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the model trained on 2011-2019 data is assumed valid after 2019, with validation only via monthly BR data whose wider bins and temporal aggregation can mask errors. This is indeed the single most consequential risk to the central claim of a continuous daily proton flux record, because the headline R2 is computed on a random split of an autocorrelated time series and therefore does not demonstrate predictive skill on unseen epochs. I considered other possible objections: the hourly product is explicitly unverified (Section III E), but it is presented as a secondary contribution and the paper admits the limitation; the lack of code/data release is a reproducibility concern but does not by itself invalidate the scientific claim; the wavelet physics conclusions are consistent with prior AMS results and are not load-bearing for the main reconstruction claim. The temporal generalization issue is the linchpin: if the model cannot predict the final year of AMS daily data accurately when trained only on earlier years, then the post-2019 extension and the gap-filling claims collapse. The proposed temporal-split test is a direct, feasible check that would settle the issue. Since this concern is addressable and the current evidence is suggestive but not definitive, the reader's CONDITIONAL verdict is appropriate; my analysis does not move it.","tokens_in":14953,"tokens_out":2974,"duration_ms":29278,"concrete_test":"Perform a strict temporal split: train on daily data from May 2011 to December 2017, validate on 2018, and test on the daily AMS data from January to October 2019 (the final available daily period). Report per-rigidity-bin R2 and relative errors on this held-out daily period, and also aggregate the model's daily predictions to BR intervals to compare against the AMS monthly data for 2019. If the out-of-time daily R2 remains close to 0.9984 with no systematic rigidity-dependent bias, the concern is resolved. If the R2 drops materially (e.g., below 0.99) or high-rigidity bins show relative errors exceeding a few percent, the 2011-2024 daily flux record is not supported by the current evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the model produces continuous daily AMS-like proton flux from 2011 to 2024 with high accuracy. The reported R2 = 0.9984 is computed on a test set formed by randomly splitting individual days (Section II B 2). Because the target time series is strongly autocorrelated at the solar-cycle timescale, a random split places test days interleaved with training days, allowing the model to interpolate the smooth trend rather than predict genuinely unseen periods. The only out-of-time validation is against monthly Bartels-rotation AMS data (Section III C), which aggregates over 27-day intervals and uses wider rigidity bins, so daily and rigidity-resolved errors can cancel in the averaging. The paper explicitly states in Section III E that hourly accuracy cannot be verified, and the post-2019 daily extension has no daily ground-truth check. Therefore the strongest claim—a continuous, accurate daily flux record independent of AMS availability—rests on an assumption of temporal stationarity that is never tested at daily resolution after 2019. The conservative error estimate in Section III C does not cure the mismatch: it reports a worst-case deviation after subtracting AMS uncertainties, but the monthly BR comparison in Figure 7 cannot detect errors that are periodic within a Bartels rotation or localized to specific rigidity bins.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a deep residual convolutional neural network that maps daily count rates from 18 neutron monitor (NM) stations to AMS-02 daily proton flux in 30 rigidity bins from 1 to 100 GV. The NM data are first cleaned with IQR outlier removal and cross-station event checks, then imputed with a self-attention imputation model (SAITS). The flux model is trained on the 2011-2019 AMS daily dataset and is reported to achieve R2 = 0.9984 on a randomly held-out test set. The paper then uses the model to produce continuous daily flux from 2011 to 2024, validates the post-2019 period against monthly Bartels-rotation AMS data, applies wavelet analysis to the 2014 solar polar-field-reversal period, and introduces hourly proton flux products for Forbush decrease studies.","tokens_in":15039,"tokens_out":5360,"duration_ms":54473,"significance":"If the central claim holds, the method would provide a continuous, rigidity-resolved cosmic-ray proton flux record that bridges AMS data gaps and extends beyond the published daily AMS interval, which is valuable for solar modulation and space-weather studies. The paper makes good use of public NMDB and AMS data, applies careful preprocessing with physical cross-checks for outlier retention, and gives an explicit conservative error-estimation procedure for the monthly validation. The wavelet analysis and Forbush-decrease application indicate potentially useful scientific output. The main contribution is therefore a data-product/method paper whose value depends on the credibility of the generalization claims.","major_comments":[{"comment":"The single reported value R2 = 0.9984 is not sufficient to support the central accuracy claim. The paper does not state whether this R2 is computed by pooling all 30 rigidity bins and all test days, or per bin. Because the mean proton flux decreases by several orders of magnitude from 1 GV to 100 GV, a model that only reproduces the average rigidity spectrum can achieve a very high pooled R2 while failing to capture day-to-day or bin-to-bin variations. In addition, the random day-level split (80/10/10) interleaves test days with training days; since daily flux is strongly autocorrelated on the solar-cycle timescale, the model can effectively interpolate the smooth trend rather than predict a genuinely unseen period. Please report per-rigidity-bin R2 or relative RMSE, and repeat the evaluation with a chronological or block holdout (for example, train on 2011-2016 and test on 2017-2019) so that the out-of-time generalization is measured directly.","section":"II B 2 and III B, Eq. (3)"},{"comment":"The post-2019 daily extension is validated only against AMS monthly (Bartels-rotation) data with wider rigidity bins. Aggregating over 27-day intervals can cancel errors that oscillate within a Bartels rotation, and the wider rigidity bins can wash out errors localized in a single daily bin. Therefore, the monthly comparison shown in Figure 7 does not establish that the daily, per-bin flux after 2019 is accurate. Please quantify the aggregation effect by applying the same BR binning and rigidity rebinning to the 2011-2019 period where daily ground truth exists: compare the model's daily per-bin error against its BR-aggregated error. If the aggregation substantially reduces the apparent error, the post-2019 daily record should be described as an unvalidated extrapolation rather than as a measurement.","section":"III C"},{"comment":"The hourly proton flux is advertised in the abstract and conclusion as a first-time product, but the model is trained on daily data and the paper explicitly states that hourly accuracy cannot be verified. Fourier-filtering the diurnal NM cycle changes the input distribution, and there is no evidence that the daily NM-to-flux relationship transfers to sub-daily timescales. Please either provide some external check for a known event (for example, comparison with any available sub-daily space-based data, even if for a different species or rigidity range) or reframe the hourly output as an illustrative, unvalidated demonstration and remove the unqualified 'for the first time' claim from the abstract and conclusion.","section":"III E and Conclusion"}],"minor_comments":[{"comment":"The definition of the effective relative error epsilon is ambiguous: it is not clear whether the maximum is taken only over time points where |F_pred - F_AMS| exceeds sigma_AMS, or whether negative differences are clipped to zero. Please specify the exact computation and state explicitly that this is a conservative upper bound.","section":"Eq. (5)"},{"comment":"The caption describes the histogram as representing 'AMS time-dependent uncertainties' while the blue dots are individual daily errors; please clarify the units of the x-axis and how the histogram is constructed, since a histogram of uncertainties is unusual and the current description is confusing.","section":"Figure 3, lower panel"},{"comment":"The sentence describing interpolation of the post-2019 BR-derived errors onto daily rigidity bins does not specify the interpolation method; please state whether the interpolation is linear in rigidity, linear in log rigidity, or another monotonic scheme.","section":"III C"},{"comment":"The caption states that during SEP events the proton flux below 3 GV is excluded to match AMS reporting; please explain how this exclusion is applied consistently in training, testing, and the final 2011-2024 product, since the daily AMS dataset already excludes those measurements and the model is expected to predict all 30 bins.","section":"Figure 5 caption"},{"comment":"Reference [21] is the same AMS publication as reference [5] but with additional links; please differentiate the two references, for example by citing the original paper once and, if needed, a separate reference for the extended monthly data table.","section":"References"},{"comment":"The abstract uses 'simulate the relationship' while the conclusion uses 'establish a correlation' and later 'accurately captures the relationship'; please use consistent terminology for what is an empirical regression model, not a physical simulation.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of astro-ph.SR and addresses a useful gap, but the advertised data products (daily 2011-2024 record and hourly record) rest on generalization claims that are not yet convincingly demonstrated. The requested changes - per-bin and out-of-time metrics, an aggregation-smoothing check for the BR validation, and a re-scoped hourly claim - are feasible within the manuscript's scope and should be made before publication. I would also encourage the editor to ask the authors to release the trained model and the continuous flux data products, since reproducibility is essential for a paper whose main contribution is a derived dataset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what's new: this is the first attempt I know of to give hourly, rigidity-resolved proton flux from neutron monitors, and it also produces a continuous daily record across the AMS gap. If the method holds up, that is a genuinely useful data product for Forbush decrease and periodicity studies. The imputation stage is more careful than most ML papers: cross-station consistency checks for solar events, IQR with a conservative threshold, and a comparison of four imputation methods is honest work.\n\nThe central worry is the validation protocol. R2=0.9984 comes from a random split of individual days. Daily proton flux is highly autocorrelated at solar-cycle timescales, so random splits let the model interpolate the smooth trend. The paper would be much stronger with a temporal split (e.g., train on 2011-2017, test on 2018-2019, or leave out contiguous blocks). The monthly Bartels-rotation comparison is a reasonable attempt at out-of-time validation, and it does cover post-2019 to 2022, but 27-day averaging can hide errors that are periodic within a rotation or confined to narrow rigidity bins. So the claim of a continuous 2011-2024 daily record is plausible but not yet demonstrated at daily resolution after 2019.\n\nThe hourly product is explicitly unverified. The authors say they couldn't find any hourly proton flux publication to compare against. That is an honest statement, but it means the headline 'first hourly proton flux' is a claim without a ground-truth check. A basic sanity check against NM-based FD profiles or against AMS daily data aggregated from hourly predictions would at least show the hourly output is not wildly off.\n\nAlso, the paper positions itself as a community resource but no code or reconstructed dataset is released. That is a fixable problem and matters for a paper whose main output is a dataset.\n\nThe wavelet analysis is not the main risk: the finding that periodicities during the polar reversal match AMS's earlier results is a useful consistency check rather than new physics.\n\nBottom line: the core idea is sound, the implementation looks careful, and the limitations are mostly addressable. I'd send it to peer review, but the referee should ask for a temporal-split validation, a release of code/data, and at least one external check of the hourly product. If those come through, the paper earns its place; without them, the high R2 and the 2011-2024 continuity claim are over-sold.","headline":"A useful reconstruction of AMS-02 proton flux from neutron monitors, but the headline R2 and the hourly product both rest on validation gaps that need tightening.","tokens_in":15756,"tokens_out":2445,"would_cite":false,"duration_ms":22154,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep residual CNN trained on 18 imputed neutron monitor stations predicts daily AMS proton flux from 1 to 100 GV with $R^2 = 0.9984$ and extends the record to 2024 with hourly resolution.","keywords":["cosmic ray proton flux","neutron monitor","AMS-02","deep residual network","time series imputation","SAITS","solar modulation","Forbush decrease"],"falsifier":"Compare the model's post-2019 daily predictions with a future AMS daily release, or compare its hourly output with independent sub-daily rigidity-resolved spacecraft measurements during a well-observed Forbush decrease; if the predicted flux drifts systematically over time or the hourly depth and timing of the decrease differ by more than the stated total uncertainty, the generalization claim is falsified.","tokens_in":14568,"feed_emoji":"☀️","tokens_out":10962,"duration_ms":92022,"temperature":0.7,"pith_summary":"The paper sets out to show that ground-based neutron monitor (NM) count rates carry enough information to reconstruct the rigidity-resolved cosmic-ray proton flux that the Alpha Magnetic Spectrometer (AMS) measures in space. It trains a deep residual convolutional network on 18 pre-processed and imputed NM stations to predict daily AMS proton flux in 30 rigidity bins from 1 to 100 GV, reporting $R^2 = 0.9984$ on a held-out test set. If the claim holds, the result is a continuous daily proton-flux record from 2011 to 2024 that fills AMS operational gaps, plus a first hourly rigidity-resolved proton-flux product for studying short-time solar activity. The argument depends on the NM-to-proton relationship learned from daily 2011-2019 data remaining valid after 2019 and at hourly timescales.","feed_headline":"Deep network reconstructs AMS proton flux from 18 neutron monitors","feed_subtitle":"It reaches R-squared 0.9984 on held-out daily data and extends the record to 2024 with hourly resolution.","key_machinery":"The load-bearing object is a residual-block convolutional neural network: a fully connected layer maps the 1-by-18 vector of daily NM count rates to 64 features, six residual blocks (each with two fully connected layers, batch normalisation, and Gaussian Error Linear Unit activations) extract the nonlinear NM-to-flux relationship, and a final layer emits a 1-by-30 vector of rigidity-binned proton fluxes. Residual connections are what allow the network to converge stably despite missing AMS labels. The other essential piece is SAITS, a self-attention time-series imputer, which reconstructs missing station values so that the NM input is continuous; for hourly output a Fourier filter first removes the ground-based diurnal cycle that has no counterpart in space.","core_discovery":"The paper's central claim is that a residual-block CNN can serve as a surrogate for direct space-based proton measurement: given one day's count rates from 18 neutron monitor stations whose gaps have been filled by a self-attention imputer, it outputs the daily AMS proton flux in 30 rigidity bins spanning 1 to 100 GV. On the held-out daily test set the model reaches $R^2 = 0.9984$. The trained model is then applied to the years after AMS daily data end, validated only against coarser monthly AMS data binned by 27-day solar-rotation intervals, and to hourly NM data after the diurnal cycle is removed with a Fourier filter, producing hourly rigidity-resolved proton fluxes that the paper states cannot currently be verified against any published measurement.","pith_inferences":["An extension the paper leaves implicit: the same pipeline could be applied to AMS helium fluxes or DAMPE electron fluxes to build continuous multi-species rigidity-resolved records over the same solar cycle.","A testable implication the authors do not pursue: compare the hourly product with independent sub-daily spacecraft measurements during a well-observed Forbush decrease to check the hourly transfer assumption.","A robustness question the paper does not address: how much of the accuracy depends on having all 18 stations, since training on station subsets would separate learned physics from network redundancy."],"forward_implications":["The method fills the AMS data gaps of September-November 2014 and July 2018-October 2019 and extends the daily proton-flux record through August 2024.","Wavelet analysis of the reconstructed continuous record across the Solar Cycle 24 polar field reversal reproduces the AMS periodicity pattern: 27-day dominance at low rigidity and 13.5- and 9-day periodicities becoming significant at high rigidity.","The hourly product resolves the structure of ICME-driven Forbush decreases, such as the two events on 16 and 17 March 2015, which daily AMS sampling cannot distinguish.","Total uncertainties are assigned conservatively as the maximum of the AMS measurement error, the pre-2019 model error, and the post-2019 model error estimated from monthly bins."],"supporting_citations":[{"why":"This is the daily AMS proton-flux dataset in 30 rigidity bins that serves as the training target and the main pre-2019 comparison baseline.","marker":"[5]"},{"why":"This supplies the neutron monitor database from which the 18 stations' pressure-corrected count rates are drawn.","marker":"[9]"},{"why":"SAITS is the self-attention imputation algorithm chosen to fill missing NM values after it outperformed the alternative imputers in the comparison.","marker":"[12]"},{"why":"This provides the unified time-series imputation framework, including the modified transformer variant, used to compare and run the imputation models.","marker":"[11]"},{"why":"iTransformer is the strongest imputation competitor to SAITS, and its modified variant is evaluated in the imputation comparison.","marker":"[14]"},{"why":"Residual learning is the architectural mechanism that lets the network train stably and reach the reported accuracy.","marker":"[18]"},{"why":"AMS monthly solar-rotation-binned proton fluxes, which extend beyond the daily dataset, provide the post-2019 validation target.","marker":"[21]"},{"why":"This supplies the wavelet-analysis method used to extract time-frequency periodicities from the reconstructed flux record.","marker":"[23]"},{"why":"This documents the diurnal variation of ground NM rates, which motivates the Fourier filtering step before producing hourly fluxes.","marker":"[24]"}],"fun_headline_variants":["AI maps neutron monitor counts to AMS proton flux with R²=0.9984","Deep learning predicts cosmic proton flux from ground neutron monitors only","Neural imputer + CNN yields hourly proton rigidity spectra from ground data","From 18 neutron monitors to hourly cosmic proton flux via AI","Proton flux reconstruction with deep learning matches space detector to 0.9984 R²"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model is trained only on daily AMS proton flux from 2011 to 2019, and the paper assumes the learned NM-to-proton relationship continues to hold for later years and for hourly timescales, even though no daily or hourly space-based measurements are available to check those periods.","fun_headline_variants_meta":{"raw":{"variants":["AI maps neutron monitor counts to AMS proton flux with R²=0.9984","Deep learning predicts cosmic proton flux from ground neutron monitors only","Neural imputer + CNN yields hourly proton rigidity spectra from ground data","From 18 neutron monitors to hourly cosmic proton flux via AI","Proton flux reconstruction with deep learning matches space detector to 0.9984 R²"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001496,"raw_usage":{"total_tokens":5963,"prompt_tokens":862,"completion_tokens":5101,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":5002}},"tokens_in":478,"tokens_out":5101,"duration_ms":31657,"temperature":1.0,"reasoning_tokens":5002,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:21:24.101006+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the model's post-2019 daily predictions with a future AMS daily release, or compare its hourly output with independent sub-daily rigidity-resolved spacecraft measurements during a well-observed Forbush decrease; if the predicted flux drifts systematically over time or the hourly depth and timing of the decrease differ by more than the stated total uncertainty, the generalization claim is falsified.","supporting_citations":[{"cited_title":"Therefore, we align the imputed NM data with AMS data by date, thereby creating a paired NM-AMS dataset that serves as inputs and outputs for training a calculated model","cited_arxiv_id":null,"evidence_quote":"This is the daily AMS proton-flux dataset in 30 rigidity bins that serves as the training target and the main pre-2019 comparison baseline."},{"cited_title":"Moraal and P","cited_arxiv_id":null,"evidence_quote":"This supplies the neutron monitor database from which the 18 stations' pressure-corrected count rates are drawn."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SAITS is the self-attention imputation algorithm chosen to fill missing NM values after it outperformed the alternative imputers in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This provides the unified time-series imputation framework, including the modified transformer variant, used to compare and run the imputation models."},{"cited_title":"Moraal, A","cited_arxiv_id":null,"evidence_quote":"iTransformer is the strongest imputation competitor to SAITS, and its modified variant is evaluated in the imputation comparison."},{"cited_title":"At lower rigidity levels, the proton flux predominantly ex- hibits a dominant 27-day periodicity, which corresponds to solar rotation","cited_arxiv_id":null,"evidence_quote":"This documents the diurnal variation of ground NM rates, which motivates the Fourier filtering step before producing hourly fluxes."}],"review_version":1}