Pith. sign in

REVIEW 4 major objections 5 minor 5 references

Morphological Fingerprints of Forbush Decreases and Their Relation to Geomagnetic Storm Severity

T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read The authors claim that Forbush-decrease events leave reproducible network fingerprints in multi-station neutron-monitor graphs, and that these fingerprints carry measurable information about geomagnetic-storm severity and Forbush magnitude.

desk verdict The storm-intensity classification results are worth reading; the drop-regression R²=0.350 is likely an artifact of target–feature overlap and needs a leakage check before it is believed. read the letter →

arxiv 2602.16128 v3 pith:NXBXM2RY submitted 2026-02-18 astro-ph.IM

classification astro-ph.IM
keywords Forbushdecreaseneutronmonitorgraphfingerprintminimumspanningtreegeomagneticstormspaceweatherleave-one-outvalidationnetworkscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a Forbush decrease leaves a reproducible 'fingerprint' in the graph of how different neutron-monitor stations respond together, and that this fingerprint carries usable information: it can sort storms by NOAA intensity class (G3/G4/G5), flag severe storms, and partly predict the Forbush drop size. The authors build an event graph from pairwise dissimilarities between station count-rate series, prune it to a minimum-spanning-tree backbone, summarize it with compact geometric and topological descriptors, and validate with leave-one-event-out. A sympathetic reader would care because it offers a unified, interpretable way to compare FD morphology across heterogeneous station networks, potentially aiding storm triage and physical understanding of heliospheric drivers.

What carries the argument

The central object is the event-level graph built from pairwise dissimilarities between station response time series in a common event window. Edges carry transformed dissimilarities, and a minimum spanning tree (MST) provides a sparse, connected, density-controlled backbone with exactly N−1 edges per event, making graphs comparable across events of different station coverage. From each graph, compact fingerprints aggregate global efficiency, spectral summaries (Estrada index, Laplacian summary), mesoscopic structure (modularity), mixing (assortativity), centrality aggregates, and complexity descriptors. The machinery's role is to convert heterogeneous, multivariate time series into low-dime

What would settle it

A decisive control experiment: normalize each station's event-window count-rate series to zero mean and unit variance before computing pairwise dissimilarities, then re-run the LOEO drop regression. If R² drops to near baseline, the regression signal is amplitude leakage; if R² persists, the fingerprint captures morphology beyond amplitude. A second check: replace the NMDB-derived drop labels with independent catalog magnitudes (e.g., from solar-wind/Dst-based storm indices) and see whether the graph fingerprints still predict them.

Watch

Extended reading notes

Core claim

Under strict leave-one-event-out validation, graph fingerprints derived from multi-station neutron-monitor responses carry reproducible signal for three tasks: (i) moderate multi-class classification of storm intensity (G3/G4/G5) with errors dominated by adjacent categories (macro-F1 ≈ 0.575); (ii) stronger binary screening of severe storms (≥G4 vs. G3) with high sensitivity to severe events (true-positive rate 0.87); and (iii) partial prediction of Forbush drop magnitude via partial least squares (R² = 0.350, above a fold-wise mean baseline). The fingerprints that dominate intensity classification—average Katz centrality, Estrada index, Laplacian summaries, entropy, modularity—point to glob

Load-bearing premise

The load-bearing premise is that the drop-regression target is not trivially contained in the graph distances; but because the drop is measured from the same neutron-monitor count-rate series that generate those distances, the reported R² may partly reflect amplitude leakage rather than morphology.

Editorial extensions

If this is right

  • If correct, FD morphology can be characterized quantitatively via compact graph fingerprints, enabling event-to-event comparison beyond summary curves or pairwise metrics.
  • The binary severity screening result suggests graph fingerprints could serve as an operational triage signal, flagging ≥G4 storms with high sensitivity even when multi-class separation is imperfect.
  • Rigidity-conditioned node-role analysis indicates that cutoff rigidity systematically shapes station roles in the event backbone, connecting network structure to physical shielding.
  • The regression result implies that global network organization contains quantitative information about event magnitude, though with regression-to-the-mean at extremes.
  • The MST-backbone construction offers a density-controlled representation that could be extended with cyclic structure to test whether mesoscopic loops add predictive or physical value.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A likely caveat the paper does not fully address: because the drop target is computed from the same NMDB count-rate series used to build the dissimilarities, the regression R² may partly reflect amplitude leakage (larger dips produce larger pairwise distances) rather than an independent morphological property; this concern does not affect the classification tasks, whose labels come from external G
  • A natural extension would be to recompute graph fingerprints after normalizing each station's event-window series to zero mean and unit variance, then re-run the LOEO regression; if R² collapses, the regression signal is dominated by amplitude rather than shape.
  • Adding a few shortest non-tree edges (fixed-density graphs) could isolate the incremental value of cycle structure, which is currently absent by MST construction and may suppress mesoscopic descriptors.
  • The same event-graph framework could transfer to other multi-site heliospheric or geophysical monitoring networks—e.g., riometer or magnetometer arrays—where station heterogeneity and coverage gaps complicate event comparison.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a graph-based representation of Forbush decrease (FD) events. Each event is encoded as a network whose nodes are neutron-monitor stations and whose edge weights are pairwise dissimilarities between station count-rate time series; a set of graph-theoretic fingerprints is then used for three predictive tasks: multi-class geomagnetic storm intensity classification (G3/G4/G5), binary severity screening (≥G4 vs. G3), and regression of FD drop magnitude. Validation is performed with leave-one-event-out (LOEO) over a pre-defined pipeline grid. The authors report moderate classification skill (macro-F1 0.575 for G3/G4/G5; sensitivity 0.87 for ≥G4 vs. G3) and positive regression skill (PLS-5 R²=0.350), and interpret these as evidence that FD event morphology leaves reproducible network signatures. The manuscript includes code and processed feature tables, which is a strength for reproducibility.

Significance. If the reported results were fully supported, the paper would offer a compact, interpretable event-level representation of FD morphology that is comparable across events and potentially useful for storm triage and comparative FD characterization. The classification tasks use externally defined G-scale labels, and the LOEO protocol is a reasonable choice for the small event sample. The availability of code and the explicit description of the pipeline grid are also positive. However, two load-bearing issues currently undermine the central claims: (i) the best pipeline is selected using LOEO performance on the same events and then reported without accounting for selection bias; and (ii) the regression target is not independent of the graph features because both derive from the same NMDB count-rate series, and the best regression setting explicitly retains absolute amplitude information. The classification results are less exposed to the second issue, but the regression-based magnitude claim, which is a central part of the paper's stated contribution, is not yet established.

major comments (4)
  1. [§4.2, §4.5; Table 3] The reported performance is obtained by selecting the best configuration from a grid based on LOEO performance and then reporting that same LOEO performance. This is post-hoc selection on the evaluation set; the headline numbers (macro-F1 0.575, binary sensitivity 0.87, regression R²=0.350) are therefore optimistically biased and no confidence intervals or correction for multiple comparisons are provided. The statement that selection criteria were fixed a priori does not address the selection itself. Please use nested LOEO (model selection on training folds only) or report the full distribution of LOEO scores over all pipeline configurations, together with error bars obtained by event-level bootstrap or similar.
  2. [§2, §3.2, §4.4] The drop-regression target is not independent of the input features. The drop is defined in Section 2 as the percentage reduction in the NMDB count rate, taken from the same NMDB records used to construct station matrix X and the pairwise distances in Eq. (1). The best regression pipeline uses 'no transform; no normalization', so absolute count-rate amplitudes enter the distance matrix directly. A larger-magnitude FD will produce larger pairwise distances, and scale-sensitive fingerprints (average closeness, betweenness, assortativity, etc.) can track this overall scale. Thus PLS-5 achieving R²=0.350 may reflect amplitude leakage rather than morphological information. The authors' own discussion in §5 that 'absolute dissimilarity magnitudes can retain information relevant to drop' confirms this mechanism. Please run a control: compute distances after normalizing each station series by it
  3. [§3.4, abstract, §5, Fig. 1] The graph object is inconsistently defined. Section 3.4 states that the graph is complete: 'Edges connect all station pairs (E is complete over the retained stations)'. But the abstract, the figure caption, and the discussion refer to the minimum spanning tree as the event graph or as the 'controlled sparse backbone'. If the fingerprints are computed on the complete graph, then the MST is only a visualization and the abstract is misleading; if the fingerprints are computed on the MST, then Section 3.4 is wrong. This is not a cosmetic point: descriptor values (efficiency, betweenness, assortativity) differ drastically between a complete weighted graph and its MST. Please clarify which graph is used for each reported result. In addition, if MST graphs are used, the claim that the edge count is 'fixed across events' is false whenever the number of retained stations N varies under the covera
  4. [§4.3 vs. Table 3] The binary severity results are reported inconsistently. Section 4.3 states accuracy = 0.758, balanced accuracy = 0.685, macro-F1 = 0.694, while Table 3 reports accuracy = 0.7878, balanced accuracy = 0.7347, macro-F1 = 0.7413 for the same task. Moreover, §4.3 says the binary model uses 'the same best-performing graph-construction setting as above' (log transform, no normalization), whereas Table 3 lists 'LOG Normalización; Decimal-Scaling'. These discrepancies need to be reconciled; as written, the reader cannot tell which numbers are the actual LOEO results.
minor comments (5)
  1. [§3.2, Fig. 1, Table 3] The distance metric is described as ℓp with p=1 and p=2, but Figure 1 labels Minkowski/Euclidean as p=3 and the best configurations say only 'Minkowski adjacency'. Please specify the exact p value used in each result and keep the notation consistent throughout.
  2. [Fig. 1 caption] The caption reads 'Illustrative two event graph' and lists one event as 2023-04-23 (G2) while the text describes a representative FD with drop ≈10.4%. Please correct the caption and make the event identification consistent.
  3. [§4.5, Table 3] Table 3 mixes English and Spanish ('LOG Normalización'). Please use uniform English terminology.
  4. [§3.6, §4.4] The classification uses n=33 events and the regression uses n=34. Please state explicitly why one event is excluded from classification and include the exact event-list/window metadata in the repository.
  5. [§4.1, Fig. 2] The rigidity-stratified node-role analysis is descriptive; no statistical test for group separation is provided. Since this is used to motivate interpretability, please add a simple test (e.g., Mann-Whitney U or Kruskal-Wallis) with multiple-comparison control, or state clearly that the differences are qualitative.

Circularity Check

1 steps flagged · score 5.0 of 10

Drop-regression skill is partially circular: the drop target and the graph distances are both derived from the same NMDB count-rate series, and the winning pipeline retains absolute distance scale.

  1. self definitional [§2 (Data, drop definition) → §3.2 Eq. (1) (pairwise distances) → §4.4 (best regression: no transform, no normalization)]
    "To characterize the magnitude of the event at each station, we used the calculation of the percentage reduction in the cosmic ray count rate with respect to a pre-event reference level which comes from the data metadata from NMDB and then visually confirmed from a curated review of each dataset. ... Event-specific pairwise distances are computed between station responses to form a matrix D ... d(p)_ij = (Σ_t |xi(t) − xj(t)|p)^{1/p} (1) ... The best performance was obtained with Minkowski-based adjacency and no distance transformation or normalization, combined with partial least squares regres"

    The continuous target y (FD drop %) is computed from the same NMDB count-rate records that are the inputs xi(t), xj(t) of the distance matrix in Eq. (1). In the best regression setting, distances are deliberately left untransformed and unnormalized, so the absolute scale of D is preserved. A larger drop therefore inflates the pairwise response differences and any scale-sensitive graph descriptor (closeness, betweenness, Katz centrality, assortativity); PLS-5 can reach R²=0.350 by tracking that amplitude scale rather than by encoding morphology. The paper's own §5 statement that 'absolute dissimilarity magnitudes can retain information relevant to drop' concedes exactly this confound. Thus the magnitude leg of the central claim is not an independent test of morphology; the classification ta

full rationale

The classification results (G3/G4/G5 and ≥G4 vs. G3) use externally defined NOAA G-scale labels, so they are not circular with the graph construction; they provide independent evidence for the severity part of the claim. The regression task, however, is partially circular: the drop label is defined as the percentage reduction of the same NMDB count-rate series that generate Eq. (1)'s pairwise distances, and the winning model intentionally omits distance normalization, so the reported LOEO R²=0.350 may reflect absolute amplitude leakage rather than geometric/topological morphology. This is not an exact Eq. X = Eq. Y identity, but it is a target–feature dependence strong enough to invalidate the 'global network organization carries quantitative information about event magnitude' interpretation in the absence of a scale-invariant control. The only self-citation (Perez-Navarro & Sierra-Porta 2026, §1) is motivational and not load-bearing. Overall score 5 reflects one partially circular prediction amid otherwise independent evaluation.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central predictive claims rest on hand-chosen pipeline parameters (tau=0.5, k=2, distance p, transform/normalization, PLS-5) that are partly selected by LOEO scores, on trust in externally compiled storm labels, and on the assumption that small event graphs with varying station sets are comparable. No new physical entities are introduced; the 'graph fingerprints' are derived descriptors rather than invented entities.

free parameters (6)
  • coverage-filter threshold tau = 0.5
    Fixed exclusion threshold for stations with missing fraction > tau (Section 3.1); no sensitivity analysis, yet it changes which stations enter every event graph.
  • KNN imputation k = 2
    Base-estimator neighbor count for iterative imputation chosen without stated justification (Section 3.1).
  • distance metric p = best: 'Minkowski' (p unclear; Section 3.2 lists p=1,2, Fig. 1 caption says p=3)
    Grid choice; the winning metric is selected by LOEO performance, so it is effectively fit to the evaluation set.
  • distance transform and normalization = best: log/no-norm (classification); none/none (regression)
    Each task's best configuration was chosen from the predefined grid by LOEO performance (Sections 3.3 and 4.5).
  • PLS component count = 5
    Number of PLS components is fixed without reported selection or robustness analysis (Section 4.4).
  • event window length/alignment
    Section 3.1 refers to 'a common analysis window' but never states its length or how it is chosen across events; this changes all pairwise distances and fingerprints.
assumptions (5)
  • domain assumption Neutron-monitor count rates faithfully represent FD morphology at each station
    Used throughout Sections 2-3; if pressure/rigidity/background corrections differ across stations, pairwise distances encode station artifacts rather than physical morphology.
  • domain assumption External storm labels (NOAA G-scale) and FD drop values for the 33/34 events are correct
    Labels are compiled from catalogs and reports (Section 2); no event list or label audit is included, and the drop is derived from the same NMDB series used for the graphs.
  • ad hoc to paper MST/dissimilarity graph is a meaningful, cross-event-comparable representation of FD morphology
    The whole fingerprint approach assumes pairwise response dissimilarity and the MST backbone capture physically relevant coupling (Sections 3.4 and 5).
  • domain assumption Network metrics are comparable across events with different station sets and sizes
    N varies per event due to coverage filtering (Sections 2 and 3.1); many descriptors depend on graph size/density, and this is not corrected.
  • standard math LOEO folds are statistically independent
    Events from the same solar cycle or driver may share heliospheric conditions; the paper treats each event as an independent sample (implicit in Section 3.6).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Morphological Fingerprints of Forbush Decreases and Their Relation to Geomagnetic Storm Severity." pith.science (2026). https://pith.science/paper/NXBXM2RY

@misc{pith2026260216128,
  author       = {Pith},
  title        = {Pith review of: Morphological Fingerprints of Forbush Decreases and Their Relation to Geomagnetic Storm Severity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NXBXM2RY}},
  note         = {Machine review of arXiv:2602.16128}
}
abstract

Forbush decreases (FDs) are transient depressions in the galactic cosmic-ray flux observed by global neutron-monitor networks and are commonly associated with interplanetary disturbances driven by coronal mass ejections and related shocks. Despite extensive observational work, quantitatively comparing FD morphology across events and linking it to storm severity remains challenging due to heterogeneous station responses, coverage gaps, and the multivariate nature of the network. This work introduces a graph-based event representation in which each FD is mapped to an event network constructed from pairwise dissimilarities between station response time series. A controlled sparse backbone is obtained via the minimum spanning tree, enabling comparable event graphs across cases. From each graph, a compact set of geometric/topological fingerprints is computed, including global integration measures, spectral summaries, mesoscopic structure, centrality aggregates, and complexity descriptors. Predictive skill is assessed using strict leave-one-event-out validation over a pre-defined grid of distance metrics and distance-domain transformations, with selection criteria fixed \emph{a priori}. The proposed fingerprints exhibit measurable signal for three tasks: (i) multi-class classification of geomagnetic storm intensity (G3/G4/G5) with moderate but consistent performance and errors dominated by adjacent categories; (ii) stronger binary severity screening ($\ge$G4 vs. G3) with high sensitivity to severe events; and (iii) drop regression with partial least squares achieving positive explained variance relative to a fold-wise mean baseline.

Figures

Figures reproduced from arXiv: 2602.16128 by the authors.

Figure 1
Figure 1. — Illustrative two event graph for 2012-03-08 (G3) and 2023-04-23 (G2) from a representative Forbush decrease. Nodes are stations colored by cutoff rigidity (GV). Edges correspond to the MST computed from the event-specific dissimilarity matrix, yielding a sparse, connected backbone with N − 1 edges that minimizes total distance. Left: Manhattan (p = 1); right: Minkowski/Euclidean (p = 3). While the main station gro… view at source ↗
Figure 3
Figure 3. — LDA projection (LD1–LD2) for the best multi-class intensity pipeline (log distance transform, no normalization and Minkowski adjacency). Colors indicate storm classes (G3/G4/G5). lower (5/10 = 0.50), indicating that a subset of mod￾erate events remains difficult to separate from the severe group. This asymmetry is consistent with the expecta￾tion that stronger disturbances imprint more distinctive global network s… view at source ↗
Figure 4
Figure 4. — PLS (5 components) LOEO regression: predicted vs. observed Forbush drop (%) for the best pipeline (no transform; no normalization; Minkowski adjacency). The dashed line indicates perfect agreement. rather than as a final operational model; larger event sets will enable tighter uncertainty quantification and more definitive benchmarking. TABLE 3 Best-performing configurations (selected by LOEO performance) for each… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

5 extracted references · 1 linked inside Pith

  1. [1]

    Chilingarian, G

    A. Chilingarian, G. Hovsepyan, H. Martoyan, T. Karapetyan, B. Sargsyan, N. Nokolova, H. Angelov, D. Haas, J. Knapp, M. Walter, O. Ploc, J. Shlegl, M. Kakona, and I. Ambrosova, arXiv e-prints , arXiv:2212.13514 (2022), arXiv:2212.13514 [physics.space-ph]. K. Ghosh and P. Raychaudhuri, arXiv e-prints , astro-ph/0701860 (2007), arXiv:astro-ph/0701860 [astro-...

  2. [19]

    Noaa space weather scales (geomagnetic storms g1–g5),

    C. T. Steigies, in AGU Fall Meeting Abstracts, AGU Fall Meeting Abstracts, Vol. 2016 (2016) pp. IN44A–05. H. Mavromichalaki, A. Papaioannou, C. Plainaki, C. Sarlanis, G. Souvatzoglou, M. Gerontidou, M. Papailiou, E. Eroshenko, A. Belov, V. Yanke, E. O. Fl¨ uckiger, R. B¨ utikofer, M. Parisi, M. Storini, K.-L. Klein, N. Fuller, C. T. Steigies, O. M. Rother...

  3. [27]

    Sierra-Porta and A.-R

    D. Sierra-Porta and A.-R. Dom ´ ınguez-Monterroza, Physica A Statistical Mechanics and its Applications 607, 128159 (2022). D. Sierra-Porta, Chaos 34, 023114 (2024). V. Freitas Silva, M. Eduarda Silva, P. Ribeiro, and F. Silva, Data Mining and Knowledge Discovery 39, arXiv:2301.02333 (2025), arXiv:2301.02333 [cs.SI]. V. Freitas Silva, M. Eduarda Silva, P....

  4. [1084]

    Okany, in General Assembly and 5th Annual Conference of the African Astronomical Society(2025) p

    C. Okany, in General Assembly and 5th Annual Conference of the African Astronomical Society(2025) p

  5. [2013]

    p. 012202. B. Y. Yushkov, Advances in Space Research 77, 3549 (2026). T. Laitinen and S. Dalla, ApJ 906, 9 (2021). G. Ihongo, D. Ruffolo, A. Saiz, U. Tortermpun, and A. C. L. Chian, in 36th International Cosmic Ray Conference (ICRC2019), International Cosmic Ray Conference, Vol. 36 (2019) p

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.