{"id":"6fe5eb24-9c53-46a8-993b-b40256f851c7","arxiv_id":"1908.01836","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"At Wakkanai, IRI-predicted foF2 agrees with observations, but the proposed BMUF correction for high solar activity is an in-sample fit with no out-of-sample validation.","lead":"This paper compares IRI model predictions with ionosonde measurements at one Japanese station and proposes a correction for the high-sunspot year 2001. The comparison lacks statistical measures, and the correction is fit to the same data it is tested on.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Correction is fitted and evaluated on the same 2001 data, so the reported error reduction is an in-sample artifact and does not establish that IRI's high-sunspot BMUF bias is correctable.","rationale":"The reader's weakest_assumption identifies both data quality and in-sample fitting as fragile premises. I agree that the in-sample fitting is the more decisive issue: even with perfect input data, the paper's correction cannot support a predictive claim because the same data are used for fitting and evaluation. The reader listed data completeness first, but the in-sample artifact is sufficient to reject the central claim on its own, so I focus on it. My concrete hold-out test would directly distinguish a genuine systematic bias from overfitting. Since this concern supports the reader's REJECT verdict, I recommend no change to the verdict.","tokens_in":12428,"tokens_out":3725,"duration_ms":55726,"concrete_test":"Hold out every other hour for each month of 2001: fit the quadratic coefficients (as in Table 2) using only even hours (0, 2, ..., 22) and compute the absolute error on the odd hours (1, 3, ..., 23). If the held-out mean absolute error is close to the Table 3 range (6-13 MHz) rather than the Table 4 range (1-3 MHz), the correction is overfitting the 2001 medians rather than capturing a stable IRI bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central new result is the quadratic correction (coefficients in Table 2, eq. 1) that reduces absolute BMUF errors from roughly 6-13 MHz (Table 3) to 1-3 MHz (Table 4). However, the coefficients are obtained by fitting the predicted-minus-observed error for each month of 2001, and the 'after correction' errors in Table 4 are computed on the same 24 hourly medians used in the fit. With three free parameters per month and 24 data points, a least-squares fit must reduce in-sample residuals; the reported improvement is therefore guaranteed by construction and provides no evidence that IRI's high-sunspot error is a smooth, reproducible function of local time. The paper offers no validation on independent data (different year, station, or held-out hours), no error bars on the fitted coefficients, and no statistical test distinguishing the correction from noise. The abstract's claim that predicted BMUF 'can be corrected' is thus only an interpolation of observed-minus-predicted, not a demonstrated correction capability. The additional concern about missing data completeness/QC for the Wakkanai medians would affect absolute error levels, but the in-sample fitting alone breaks the predictive claim even if the input data were perfect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares observed and IRI-predicted monthly medians of foF2 and BMUF at Wakkanai for 2001 (high SSN) and 2004/2005 (low SSN). It reports generally good agreement for foF2 and for BMUF in low-SSN years, but poor agreement for 2001 BMUF. A quadratic correction (Eq. 1) with month-specific coefficients (Table 2) is fitted to the 2001 data; Table 4 and Figure 9 show that the corrected BMUF values are much closer to observed values. The paper concludes that IRI's high-SSN BMUF predictions can be corrected using this fitting technique.","tokens_in":12630,"tokens_out":3956,"duration_ms":36079,"significance":"If the correction were validated on independent data, it would be of modest practical value for HF frequency planning at mid-latitude stations during high solar activity. However, the paper's central claim is not established: the correction is fitted to the same 2001 data on which it is evaluated, so the error reduction in Table 4 is an in-sample artifact rather than evidence of a stable, reproducible IRI bias. The absence of quantitative agreement metrics, error analysis, and out-of-sample tests means the conclusions are not supported by the evidence presented.","major_comments":[{"comment":"The correction procedure is circular. Equation (1) with the Table 2 coefficients is obtained by fitting the observed-minus-predicted residuals of 2001, and Table 4 and Figure 9 then show the corrected values on the very same 2001 data. With three free parameters per month and 24 hourly points, a least-squares fit necessarily reduces the in-sample residuals, so the improvement from roughly 6-13 MHz (Table 3) to 1-3 MHz (Table 4) is guaranteed by construction. The abstract's claim that predicted BMUF 'can be corrected' is therefore an interpolation statement, not a demonstrated predictive correction. The paper provides no validation on an independent year, station, or held-out hours, and no statistical test distinguishing the correction from a fit to noise.","section":"Section 3, Eq. (1), Tables 2-4, Figure 9"},{"comment":"The claim of 'good correlation' is not quantified anywhere. No correlation coefficients, RMS errors, or confidence intervals are reported, yet the paper draws conclusions about agreement from visual inspection of the figures. For example, Table 3 lists absolute errors as large as 28.5 MHz for December 2001, which is inconsistent with the text's characterization of the 2001 BMUF comparison as merely 'bad correlation' without further detail. A quantitative metric is needed to support the 'good/bad correlation' taxonomy used throughout.","section":"Section 4, Figures 3-8, Table 3"},{"comment":"There is an internal contradiction about the January exception. The abstract states that foF2 shows good correlation 'except for January of years 2001 and 2004,' while Section 4 states that there is good correlation 'for all 12 months and year chosen.' Similarly, the abstract lists months 1, 9, and 12 of 2004 as exceptions for BMUF, but Section 4 says there is good correlation 'for all 12 months when SSN low.' The reader cannot determine which claim is intended, and this ambiguity affects the paper's conclusions about model performance.","section":"Abstract vs. Section 4"},{"comment":"The reliability of the observed data is not established. The analysis uses monthly medians from the Wakkanai ionosonde, but there are no completeness statistics (number of days per month), no quality-control flags, no estimates of observational uncertainty, and no discussion of data gaps. Since the fitted correction and the error tables are built directly on these medians, missing or poor-quality data would directly bias the fitted coefficients and the reported error reductions. The paper should document data availability, median calculation procedures, and a measure of data quality.","section":"Section 2, Data Selection"}],"minor_comments":[{"comment":"Several figure panels are mislabeled: Figure 4 repeats the 'Jul\\2004' label (the second July panel should be August), Figure 6 includes a panel labeled 'Mar\\2001' twice (the fifth panel appears to be May), and Figure 9 repeats 'Jul\\2001' (the second July panel should be August). These mislabels make it difficult to interpret the results.","section":"Figures 4, 6, and 9"},{"comment":"Equation (2) is never defined, referenced, or used in the text; it should be removed or explained.","section":"Section 3, Eq. (2)"},{"comment":"The coefficients are given with inconsistent precision (e.g., -0.006 vs. -2E-05); they should be reported with a common number of significant digits.","section":"Table 2"},{"comment":"The table headers do not state the units of BMUF; the text uses MHz, but the tables should be explicit about units.","section":"Tables 3 and 4"},{"comment":"The reference list is incomplete: the 'Harris (2005)' entry lacks the full author list, and 'Nagar et al. (2015)' appears to contain PACS numbers rather than a journal citation.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper's central result is an in-sample fit presented as a correction method, with no independent validation and with internal contradictions between the abstract and the results. These issues cannot be resolved by minor revision; the study would need to be substantially reworked with out-of-sample tests and quantitative error metrics. The scope and analysis quality are below the standard expected for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the central correction result is an in-sample fit, so the paper doesn't show what it claims. That said, the underlying data comparison is legitimate and the figures are clear enough to follow.\n\nWhat's new: not much. The MUF relation is standard; comparing IRI foF2 and BMUF with ionosonde medians at Wakkanai for 2001, 2004, and 2005 is a routine validation exercise, and the paper cites several earlier IRI validation studies. The only novel piece is the quadratic correction for 2001, and that's where the trouble is.\n\nThe correction coefficients (Table 2) are fitted to the observed-minus-predicted residuals for each month of 2001. Table 4 then reports the absolute error between observed and this corrected prediction for the same 24 hours and same months. With three free parameters per month and 24 points, the error reduction from roughly 6–13 MHz to 1–3 MHz is mathematically guaranteed. It tells you nothing about whether the bias is a stable, correctable function of local time. There is no hold-out year, no different station, not even a held-out set of hours. No error bars on coefficients, no statistical test.\n\nThe 'good correlation' statements are also unsupported—no correlation coefficients, no RMS, no significance tests. And the text contradicts itself: the abstract says foF2 is an exception in January of 2001 and 2004, while the results section says good correlation for all 12 months. Figure 6 has an April panel labelled March, and Figure 9 repeats July for both July and August. These are not fatal to the data, but they signal the manuscript has not been carefully checked.\n\nThe biggest soft spot is the in-sample evaluation. Even if the Wakkanai medians are perfect, the correction is just an interpolation of the residuals. To make the claim stick, the author would need to apply the fitted coefficients to independent data—say 2002 or another station—and show the error stays low.\n\nWho is this for? Someone who wants a quick look at how IRI does at one mid-latitude station during a high-sunspot year. It is not a methods paper and it does not advance the field. My recommendation: desk reject in current form. If the author redoes the validation properly, it could become a minor data report, but as submitted it should not go to review.","headline":"The paper's only new element, a quadratic correction for 2001, is fitted and evaluated on the same data, so the reported error reduction is an in-sample artifact; the rest is a routine IRI validation exercise.","tokens_in":13126,"tokens_out":1920,"would_cite":false,"duration_ms":20399,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fitted, month-specific correction factor brings IRI's predicted maximum usable frequency into line with observations at a mid-latitude station during a sunspot maximum.","keywords":["ionosphere","foF2","M3000F2","BMUF","IRI model","maximum usable frequency","sunspot cycle","HF propagation"],"falsifier":"Apply the published monthly correction coefficients to IRI predictions for a different high-sunspot year (e.g., 2000 or 2002) at Wakkanai, or to a second mid-latitude station; if the corrected predictions no longer match observations, the 2001 agreement was an in-sample fit rather than a discoverable systematic bias.","tokens_in":12198,"feed_emoji":"📡","tokens_out":13919,"duration_ms":115961,"temperature":0.7,"pith_summary":"Basic maximum usable frequency (BMUF), the highest HF radio frequency that can be used between two points via the ionosphere, is commonly estimated from the International Reference Ionosphere (IRI) model. Comparing monthly medians at Wakkanai, Japan, this paper finds that IRI's BMUF tracks observed values well during low solar activity (2004, 2005) but systematically misses them during the high-sunspot year 2001. The paper fits a quadratic correction factor in local time, separately for each month, and shows that applying it reduces the mean absolute BMUF error from about 6–13 MHz to about 1–3 MHz. The claim matters because reliable MUF predictions during solar maxima are exactly what HF link planners need.","feed_headline":"At Wakkanai, a fitted correction fixes IRI's solar-max MUF","feed_subtitle":"Solar-maximum MUF errors at a mid-latitude station fall from ~10 to ~2 MHz after a per-month time fit.","key_machinery":"The load-bearing object is the correction factor $C(t) = a_0 t^2 + a_1 t + a_2$, a month-specific quadratic polynomial in local time $t$, applied as a multiplicative factor to the IRI-predicted BMUF (which itself is $foF2 \\times M(3000)F2$). The paper fits the three coefficients per month to the observed-minus-predicted ratio using all 24 hourly monthly-median values of 2001, then applies the factor across the whole daily curve. Because the observed error pattern is smooth in local time, the quadratic is able to flatten it to roughly 1–3 MHz residual error.","core_discovery":"The paper's central claim is that the discrepancy between IRI-predicted and observed BMUF at Wakkanai in 2001 is not random scatter but a systematic, smoothly varying function of local time, with a different shape in each calendar month. The evidence is that a quadratic polynomial correction--coefficients $a_0, a_1, a_2$ listed for all twelve months--reduces the average absolute error from roughly 6 to 13 MHz to roughly 1 to 3 MHz across the 24 hourly bins. For low-sunspot years 2004 and 2005, no such correction is needed: observed and predicted BMUF tracks already agree except for a few months (January, September, December of 2004). For the F2-layer critical frequency $foF2$, IRI predictions correlate well with observations in all three years except January of 2001 and 2004. The author takes this as evidence that the model's high-sunspot BMUF bias at this mid-latitude station is a stable, correctable error in IRI's F2-layer representation.","pith_inferences":["Because the correction is fitted and evaluated on the same 2001 data, the post-correction errors are likely optimistic; an out-of-sample test on 2000 or 2002 would show how much of the improvement is a real systematic bias rather than in-sample fitting.","The paper does not separate BMUF into its foF2 and M(3000)F2 contributions; if the high-sunspot bias originates mainly in one of these parameters, a more physical correction could target that parameter directly.","The fitted coefficients vary fairly smoothly from month to month, so a future study could parameterize them by sunspot number and provide corrections for arbitrary solar activity levels at mid-latitudes.","Repeating this fitting approach at other ionosonde stations would reveal whether the high-sunspot BMUF bias is local to Wakkanai or a broader feature of the IRI model at mid-latitudes."],"forward_implications":["For low solar activity years, IRI's BMUF predictions at this latitude can be trusted without correction.","For high solar activity years, the BMUF error is systematic in local time and month rather than random, so it can be captured by a compact per-month quadratic function.","Applying such a correction would give HF link planners a simple way to improve frequency selection during sunspot maxima at mid-latitudes."],"supporting_citations":[{"why":"Supplies the comparison of observed and predicted MUF(3000)F2 that this paper extends to a mid-latitude station.","marker":"Athieno et al., 2015"},{"why":"Provides the baseline IRI-versus-MUF comparison during a rising solar cycle that the Wakkanai results build on.","marker":"Malik et al., 2016"},{"why":"Gives the MUF = foF2 x M(3000)F2 relation used to compute BMUF from ionosonde data.","marker":"Fotiadis et al., 2004"},{"why":"Defines M(3000)F2 as the propagation factor, the second ingredient in the BMUF calculation.","marker":"Kouris et al., 2000"},{"why":"Supplies IRI F2-peak parameter context for the model predictions being tested.","marker":"Adeniyi et al., 2003"}],"fun_headline_variants":["IRI solar-max MUF error at Wakkanai fixed by per-month fit","Per-month quadratic fit corrects IRI MUF at high sunspot","High-SSN BMUF: IRI bias corrected with time-dependent fit","Wakkanai MUF: monthly time fit slashes solar-max error"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole correction rests on the Wakkanai monthly medians being accurate and complete, and on the 2001 residual pattern being a stable systematic error rather than noise, because the same year's data are used both to fit the correction and to demonstrate that it works.","fun_headline_variants_meta":{"raw":{"variants":["IRI solar-max MUF error at Wakkanai fixed by per-month fit","Per-month quadratic fit corrects IRI MUF at high sunspot","High-SSN BMUF: IRI bias corrected with time-dependent fit","Wakkanai MUF: monthly time fit slashes solar-max error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1456,"prompt_tokens":1013,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":362}},"tokens_in":629,"tokens_out":443,"duration_ms":5251,"temperature":1.0,"reasoning_tokens":362,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:00:55.012084+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the published monthly correction coefficients to IRI predictions for a different high-sunspot year (e.g., 2000 or 2002) at Wakkanai, or to a second mid-latitude station; if the corrected predictions no longer match observations, the 2001 agreement was an in-sample fit rather than a discoverable systematic bias.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies IRI F2-peak parameter context for the model predictions being tested."}],"review_version":1}