{"id":"e9ea6b75-6b71-446e-90e3-dea30ac04e89","arxiv_id":"2412.10894","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An audit of Qiu et al.'s 149 geomagnetic storms finds 29 disagreements with the IKI RAS catalog and 23 missing ICME/sheath events compared with the Richardson and Cane catalog.","lead":"This paper checks the list of 149 magnetic storms from Qiu et al. against two established solar wind event catalogs and finds mismatches in about 20 to 28 percent of cases. It warns that using the unadjusted list can skew studies of how different solar wind types drive geomagnetic storms.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's disagreement rates assume the solar-wind type at the exact Dst minimum is the true storm driver; its own examples 137 and 149 show the disputed CIR ended 50 and 17 h earlier, so the quoted 'error' rates may be timing-convention artifacts rather than identification errors.","rationale":"The reader's weakest-assumption statement is the same issue I find most load-bearing: the paper assigns each storm to the solar-wind type at the exact Dst minimum and treats any other assignment as an identification error. The paper's own examples 137 and 149 make this concrete, because the Yermolaev catalog places those storm minima in SW only after the CIR has ended. If the physical driver is the structure whose southward IMF drives the storm main phase, then Qiu's CIR labels could be correct even when the Yermolaev type at minimum is SW. The disagreement rates then measure a difference in event-association conventions rather than a 19-28% error rate. The paper does present useful, independently checkable comparison statistics, and its recommendation to use established catalogs is reasonable, but the phrase 'identified incorrectly' requires an independent ground truth that is not supplied. This does not change the reader's CONDITIONAL verdict; it sharpens the required revision: either defend the Dst-minimum rule as the physically correct driver-association criterion, or report disagreement rates as convention-dependent and refrain from calling Qiu's identifications false on that basis alone.","tokens_in":18556,"tokens_out":4258,"duration_ms":41439,"concrete_test":"Recompute the comparison using storm-onset anchoring: for each of the 149 storms, take the start of the main phase (first sustained Dst decrease) and define the driver as the Yermolaev/R&C solar-wind type in the interval from that onset to the Dst minimum, choosing the type overlapping the largest portion of the interval. Then recount the 29 discrepancies. If events 137, 149, and similar cases are counted as CIR under this rule and the disagreement rate drops substantially, the paper's headline overstates Qiu error; if the rate remains near 20%, the timing-convention concern is not decisive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The counting part of the paper is probably correct: Qiu et al.'s labels differ from Yermolaev et al. in 29/149 cases and from Richardson and Cane in 23/81 ICME/Sheath cases. The central claim, however, is that these differences make the Qiu list 'incorrect' and that using it 'can lead to false conclusions.' That leap depends on an unstated physical assumption: the interplanetary driver of a storm is the solar-wind type present at the Dst minimum. Section 3.2 and Table 1 apply exactly this rule, and the paper's own examples undermine it. For events 137 and 149 the Yermolaev type is SW only because the CIR ended 50 h and 17 h before the Dst minimum; Qiu labeled these CIR. A compression region that ended before the minimum can still be the structure whose southward field drove the storm main phase. Without a definition of 'driver' tied to storm onset rather than to the minimum epoch, the discrepancy counts are comparisons of two labeling conventions, not a demonstration that Qiu's identifications are false. The paper also treats the Yermolaev catalog as the reference truth and does not independently validate it; the roughly 13% inter-catalog difference shows the reference set itself is not unique. The load-bearing issue is therefore not the arithmetic but the causal-association rule used to convert disagreements into 'incorrect identification.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares the interplanetary driver identifications for 149 moderate and strong geomagnetic storms (Dst ≤ −50 nT, 2009–2019) in the list of Qiu et al. with two reference data sets: the Yermolaev et al. catalog of large-scale solar wind types and the Richardson and Cane ICME catalog. It reports that 29 of 149 events (19.5%) are labeled differently in the Yermolaev et al. catalog, and that 23 of 81 events marked as ICME or Sheath in Qiu et al. are absent from the Richardson and Cane catalog (~28%). The paper concludes that some events in Qiu et al. were identified incorrectly, that using the list can lead to false conclusions in solar–terrestrial studies, and that the community should adopt reference catalogs for solar wind type identification.","tokens_in":18764,"tokens_out":4337,"duration_ms":38418,"significance":"If the main claim were established, the paper would serve a useful cautionary role for users of the Qiu et al. list. The discrepancy statistics are simple, reproducible from public catalogs, and the five figures support the specific reclassifications shown. However, the central claim that the Qiu et al. identifications are 'incorrect' is not supported by the evidence presented: the paper treats agreement with the authors' own catalog as ground truth, and its own examples show that part of the discrepancy is a timing-convention artifact rather than an identification error. The significance is therefore moderate; a revised version that reframes the results as inter-catalog disagreement and validates driver association against independent physical criteria would be more convincing.","major_comments":[{"comment":"The paper's leap from 'disagreement' to 'incorrect identification' depends on the unstated rule that the solar wind type at the exact time of the Dst minimum is the storm's driver. The paper's own examples 137 and 149 contradict this rule: the CIR ended 50 h and 17 h, respectively, before the Dst minimum, yet the paper counts these as Qiu et al. errors because the Yermolaev et al. catalog says SW at the minimum. Without an independent definition of 'driver' tied to the storm's main-phase onset or energy input, the 19.5% and 28% figures are comparisons of two labeling conventions, not a demonstration that Qiu et al.'s identifications are false. The authors should either adopt an explicit, physically motivated association rule or present the numbers only as inter-catalog agreement rates.","section":"§3.1, §3.2, Table 1"},{"comment":"The text admits that whether the shock wave preceding the Ejecta is associated with the Ejecta 'requires additional research and is beyond the scope of this article.' This directly undermines the claim that Qiu et al.'s Sheath label for event 58 is incorrect: if the association is ambiguous, the discrepancy cannot be counted as an identification error. The same caveat applies to other events where a compression region is adjacent to an ICME boundary, such as events 39, 45, 63, and 135.","section":"§3.2, event 58 (Fig. 3)"},{"comment":"The paper treats the Yermolaev et al. catalog as reference truth without independent validation. The paper itself notes that the Yermolaev et al. and Richardson and Cane catalogs differ by about 13% (Section 3.1 and conclusion 3). Unless the Yermolaev et al. catalog is validated against independent ground truth (for example, in situ plasma and magnetic-field signatures or solar source associations), the conclusion that Qiu et al. made errors, rather than that the catalogs merely disagree, is circular.","section":"§1, §4"},{"comment":"The 28% figure (23 of 81 events) counts events labeled ICME or Sheath by Qiu et al. that are absent from the Richardson and Cane catalog. Because Richardson and Cane is an ICME-only catalog with a composite definition that includes sheath, absence may reflect different catalog coverage or definitions rather than a Qiu et al. error. The paper does not check whether the 23 events would qualify as ICMEs under Richardson and Cane criteria; it simply assumes that Qiu et al.'s label obliges their presence. This needs a case-by-case examination or a clear statement that the figure represents a coverage difference rather than an error rate.","section":"§3.1, comparison with Richardson and Cane"},{"comment":"The identification methodology is described only qualitatively; no quantitative thresholds or algorithmic rules are given for assigning a storm to a solar wind type. The paper criticizes Qiu et al. for not providing reproducible criteria, but the comparison presented here relies on the authors' catalog labels without demonstrating their accuracy in this interval. At a minimum, the paper should state how boundary intervals are chosen and how events that straddle two types are handled, since the reported percentages are sensitive to these choices.","section":"§2, §3.1"}],"minor_comments":[{"comment":"There are numerous typos and formatting issues: 'CB CIR' in Section 3.2, 'Ермолаев' in Table 1 column header, 'c Dst' in Fig. 5 caption, and '1x10' appears without exponent superscripts in several figure panels. The Fig. 1 caption says 'August 27 to March 3, 2017' but should read 'August 27 to September 3, 2017.'","section":"Typographical errors"},{"comment":"The meaning of the sixth column (Yermolaev et al. vs. Richardson and Cane comparison after 'coarsening' to a two-class nomenclature) is unclear from the table alone. The note explaining this should appear before the table or in the caption, and the column should explicitly state whether '+' means agreement or simply presence in both catalogs.","section":"Table 1, column 6"},{"comment":"The percentages 19.5% and ~28% are quoted without uncertainty. Given 29/149 and 23/81, the binomial standard errors are about 3.2% and 5.0%, respectively; the authors should report such uncertainties so that readers can judge whether the differences are significant.","section":"Uncertainty estimates"},{"comment":"Reference [31] appears with an incomplete URL in one place ('http://www iki.rssi.ru/pub/omni/catalog/'), and the access dates for the catalogs are not given. The reference list also does not include the Qiu et al. paper's DOI consistently; check all bibliographic entries.","section":"References and URLs"},{"comment":"The phrase '55 events out of 149 in the list of Qiu et al are marked by the authors as type CB CIR' should read 'CIR,' and the count of 10 Yermolaev et al. differences from this subset should be itemized consistently with Table 1 (events 11, 16, 109, 115, 134, 66, 89, 143, 137, 149).","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a critique of a published list; the central counting appears sound, but the interpretation requires a physically motivated definition of 'driver' and an independent validation of the reference catalog. The heavy self-citation pattern is noticeable (many Yermolaev et al. references) but is not itself grounds for rejection. The journal should weigh whether the paper's contribution—a cautionary comparison—merits publication once the interpretation is reframed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read on Lodkina et al. The headline is simple: they compare Qiu et al.'s list of 149 storm drivers for SC24 against the IKI RAS catalog and Richardson & Cane, and the counts are probably correct. 29 of 149 differ from Yermolaev et al. (19.5%), and 23 of the 81 events Qiu labels ICME/sheath don't appear in Richardson & Cane (28%). If you ever use Qiu's list, those are the numbers to know.\n\nThe specific audit is new, even though it's the same group's established approach from refs [23] and [30]. They show five cases with time series and give catalog links for the rest, so the disagreements are checkable from public data. That is real value.\n\nThe soft spot is interpretation. The paper keeps saying Qiu's events are 'identified incorrectly' and that using the list 'can lead to false conclusions.' That leap presupposes the Yermolaev catalog is ground truth, which is not independently validated. And it assumes the storm driver is the solar wind type at the Dst minimum. That assumption is doubtful, and the paper itself provides the counterexamples: events 137 and 149 are classified as SW because the CIR ended 50 h and 17 h before the minimum. A compression region that ended before the minimum can still be the structure that drove the main phase; the minimum is a lagged response. Without a clear definition of 'driver' referenced to storm onset (e.g., the start of the main phase or the peak southward IMF), the discrepancy rates are comparisons of two labeling conventions, not proof that Qiu's labels are false.\n\nAlso missing: uncertainty on the counts, and the Richardson & Cane comparison is coarse because their ICME definition bundles sheath and ejecta; the paper even notes this but doesn't adjust the 28% number. The '13%' agreement between their catalog and R&C is crudely derived. These are fixable.\n\nVerdict: send to a serious referee, but the authors need to rewrite the framing. Report disagreement rates, discuss the timing convention explicitly, and stop claiming error without independent validation. As a citable result, the counts are useful; as a demonstration of incorrectness, it doesn't hold up as written.","headline":"Probably correct counts, but 'incorrect' is overreach: the audit assumes the Dst-minimum type is the driver, and two of the authors' own examples show the disputed structure ended before the minimum.","tokens_in":19377,"tokens_out":3003,"would_cite":true,"duration_ms":27525,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A published list of the drivers of 149 magnetic storms disagrees with one reference catalog in 19.5% of events and with another in about 28% of its CME-related events, so using it uncorrected invites false conclusions.","keywords":["solar wind types","interplanetary drivers","magnetic storms","ICME","CIR","sheath","Dst index","solar cycle 24"],"falsifier":"A reanalysis that assigns each storm's driver from the onset of the main phase, rather than the Dst minimum, and then recomputes the disagreement rates would show whether the 19–28% discrepancy shrinks or disappears; for the two storms whose CIRs ended 50 and 17 hours before the minimum, this reassignment would directly change the count.","tokens_in":18295,"feed_emoji":"🌩","tokens_out":9247,"duration_ms":76477,"temperature":0.7,"pith_summary":"This paper tests a recently published list of the solar wind structures that drove 149 moderate and strong magnetic storms from 2009 to 2019. It compares each storm's assigned driver against two established catalogs: one that distinguishes seven large-scale solar wind types, and one that tracks interplanetary coronal mass ejections. The comparison shows that about one storm in five is assigned a different driver in the tested list, and that nearly a third of the events labeled as CME-related are absent from the CME catalog. The authors conclude that studies using the untested list without correction may draw false conclusions about how the magnetosphere responds to different solar wind types.","feed_headline":"One in five storm drivers mislabeled in published list","feed_subtitle":"A two-catalog check finds 19–28% disagreement, threatening studies that link solar wind types to magnetic storms.","key_machinery":"The comparison rests on the physical signatures that distinguish compression regions from CME bodies. In a sheath or corotating interaction region (CIR), the velocity, density, temperature, plasma beta, and kinetic and thermal pressures rise together; in an ICME body (ejecta or magnetic cloud) these parameters fall. The paper reads the reference catalog's type at the time of each storm's Dst minimum and compares it with the type assigned by the tested list. The other load-bearing element is the reference catalog's coverage of CIRs, which the CME-only catalog lacks, and the distinction between an ICME body and the sheath ahead of it.","core_discovery":"The central claim is that the identification of interplanetary driver types in the published list under examination is unreliable. Of the 149 storms it covers, 29 receive a different driver type in the reference catalog, a disagreement rate of 19.5%; among the 81 events labeled as CME or sheath, 23 do not appear in the CME catalog, a rate of about 28%. The paper argues that these discrepancies are large enough that using the unadjusted list in solar-terrestrial studies can lead to false conclusions, and it recommends that the scientific community adopt reference catalogs for interplanetary event classification.","pith_inferences":["The default to the Dst minimum as the assignment time means the disagreement rates are likely upper bounds; a rule that credits the structure that initiated the storm would probably pull at least the two CIR-ended-early cases back into agreement.","Studies that use the tested list to separate CME-driven from CIR-driven storms inherit a labeling noise floor near 20%, so any reported CME-CIR difference smaller than about that threshold may be an artifact of mislabeling rather than a physical effect.","The same three-way comparison could be applied to other event lists built from automated CME-arrival criteria, giving a quick quality score before such lists are used in downstream statistics.","A stronger test would re-run one published storm-response study using only the events on which the two reference catalogs agree, and check whether the paper's conclusions change."],"forward_implications":["Any study that adopts the tested list without re-checking inherits a roughly 20% driver-mislabeling rate for Solar Cycle 24 storms.","The 28% absence rate for events labeled as CME or sheath means that studies relying on the tested list may mix true CME regions with non-CME solar wind in their storm samples.","The discrepancy between the tested list and the two reference catalogs is larger than the roughly 13% disagreement between the two reference catalogs themselves, marking the tested list as the outlier.","A practical consequence is that the paper's recommendation to use reference catalogs accepted by the community would, if followed, raise the comparability of solar-terrestrial studies."],"supporting_citations":[{"why":"The list of 149 storms and their interplanetary drivers that the paper tests; all disagreement statistics are computed against its labels.","marker":"[32]"},{"why":"Reference catalog of seven large-scale solar wind types, used as the primary benchmark in the comparison.","marker":"[31]"},{"why":"Reference catalog of near-Earth interplanetary coronal mass ejections, used as the second benchmark and source of the 28% absence rate.","marker":"[34]"},{"why":"Provides the method for selecting sheath and ejecta/magnetic-cloud intervals from the CME catalog and comparing them with the other catalog.","marker":"[30]"},{"why":"States the methodological argument that reference catalogs should be used and documents common identification errors.","marker":"[23]"},{"why":"The qualitative CME identification criteria that the tested list's authors rely on; the paper argues they are too vague to reproduce.","marker":"[33]"}],"fun_headline_variants":["20% of storm drivers misidentified in published list","Storm driver mislabeling threatens solar wind studies","Reference catalogs urged: storm lists 20% unreliable","One in five storm sources wrongly typed in catalog"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the correct solar wind type for a storm is the one present at the exact moment of the Dst minimum, rather than the earlier structure that initiated the storm.","fun_headline_variants_meta":{"raw":{"variants":["20% of storm drivers misidentified in published list","Storm driver mislabeling threatens solar wind studies","Reference catalogs urged: storm lists 20% unreliable","One in five storm sources wrongly typed in catalog"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1294,"prompt_tokens":854,"completion_tokens":440,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":378}},"tokens_in":470,"tokens_out":440,"duration_ms":4362,"temperature":1.0,"reasoning_tokens":378,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:30:29.093650+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reanalysis that assigns each storm's driver from the onset of the main phase, rather than the Dst minimum, and then recomputes the disagreement rates would show whether the 19–28% discrepancy shrinks or disappears; for the two storms whose CIRs ended 50 and 17 hours before the minimum, this reassignment would directly change the count.","supporting_citations":[{"cited_title":"The interplanetary origins of geomagnetic storm with Dstmin ≤ 50nT dur ing solar cycle 24 (2009 –2019) // Advances in Space Research, 2022","cited_arxiv_id":null,"evidence_quote":"The list of 149 storms and their interplanetary drivers that the paper tests; all disagreement statistics are computed against its labels."},{"cited_title":"I., Nikolaeva N.S., Lodkina I.G., Yermolaev M.Y","cited_arxiv_id":null,"evidence_quote":"Reference catalog of seven large-scale solar wind types, used as the primary benchmark in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the method for selecting sheath and ejecta/magnetic-cloud intervals from the CME catalog and comparing them with the other catalog."},{"cited_title":"What Solar–Terrestrial Link Researchers Should Know a bout Interplanetary Drivers","cited_arxiv_id":null,"evidence_quote":"States the methodological argument that reference catalogs should be used and documents common identification errors."},{"cited_title":"Wang Z., Pan B., Miao P","cited_arxiv_id":null,"evidence_quote":"The qualitative CME identification criteria that the tested list's authors rely on; the paper argues they are too vague to reproduce."}],"review_version":1}