Pith. sign in

REVIEW 2 major objections 4 minor 65 references

Compared over matched spatial areas, an AI weather model matches or beats a 3-km physics-based model for extreme 6-hour precipitation at lead times of 24 hours and beyond, while the physics model wins only at short lead times.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 07:36 UTC pith:2CLFAQ7F

load-bearing objection Useful HiRA + twCRPS integration with sound math and an honest GraphCast vs HRRR comparison; the main fix before publication is a sensitivity test for the ERA5-based extreme thresholds. the 2 major comments →

arxiv 2510.25045 v2 pith:2CLFAQ7F submitted 2025-10-29 physics.ao-ph

Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model

classification physics.ao-ph
keywords extreme precipitationAI weather predictionspatial verificationneighbourhood methodsthreshold-weighted CRPSHiRAGraphCastHRRR
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that how you compare an AI weather model with a conventional high-resolution physics model decides who wins. It combines a spatial verification technique (neighbourhood pseudo-ensembles around rain gauges) with a threshold-weighted proper scoring rule that focuses on extreme precipitation. Applied to 32 months of 6-hour rainfall over the US, it finds that at matched spatial areas the physics-based model (HRRR) only beats the AI model (GraphCast-GFS) on extremes at short lead times; from about 24 hours on, the AI model is competitive or better. The paper also shows the ranking flips with neighbourhood size, so point-to-point verification can mislead. If correct, it means AI models deserve serious consideration for operational extreme-rain forecasting at longer lead times.

Core claim

The central discovery is that the relative skill of an AI weather prediction model and a high-resolution numerical model for extreme precipitation depends strongly on the spatial scale of evaluation and the lead time. Using the HiRA framework to build neighbourhood pseudo-ensembles and the threshold-weighted CRPS with a chaining function v(z)=max(z, q_alpha) to score only the upper tail, the authors find that when the AI model and the 3-km HRRR model are compared over equivalent physical areas (e.g., 63x81 km), HRRR outperforms GraphCast-GFS for 6-hour precipitation above the climatological 99th percentile only at short lead times; at lead times of 24 hours and beyond, GraphCast-GFS is compe

What carries the argument

The framework marries two existing tools. HiRA (High-Resolution Assessment) converts each point observation into a pseudo-ensemble of forecast values from a neighbourhood of grid cells, creating an equal-probability distribution without regridding. The threshold-weighted continuous ranked probability score (twCRPS) evaluates that pseudo-ensemble using a chaining function v(z)=max(z, q_alpha), which gives unit weight only to thresholds above the local climatological extreme (q0.99 or q0.999). A 'fair' correction to the twCRPS accounts for pseudo-ensembles of different sizes, so models with different native grid resolutions can be compared over matched physical areas; the result is a proper, u

Load-bearing premise

The load-bearing assumption is that extreme events are defined by ERA5-reanalysis thresholds rather than by the station observations themselves; since GraphCast is trained on ERA5, this definition may favour the AI model.

What would settle it

Recompute the twCRPS lead-time comparison defining extreme thresholds from the ASOS station climatology (or from a different reanalysis) and check whether GraphCast-GFS still matches or beats HRRR at 24+ hours; if the AI advantage vanishes, the ranking claim depends on the threshold source.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Point-to-point verification can mis-rank models for extremes; spatial scale must be part of any such comparison.
  • At short lead times, the radar-assimilating high-resolution model (HRRR) retains an edge for extreme precipitation; beyond ~24 hours the AI model is at least as good over matched areas.
  • The method lets agencies compare models of different native resolutions without regridding or degrading the higher-resolution model.
  • The threshold weighting makes the score a proper scoring rule, so it rewards honest forecasts and can be used to train or tune post-processing for extremes.
  • The same framework extends to discrimination ability (calibrated DSC), separating whether a model can tell extremes apart from whether it is well calibrated.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the extreme thresholds come from ERA5 — the same reanalysis the AI model was trained on — the long-lead edge could partly reflect a favourable definition of 'extreme'; rerunning with station-based thresholds is a natural stress test.
  • The ranking flip with neighbourhood size means operational model choice depends on the decision's spatial scale (e.g., a single city vs a river catchment), not just the variable and lead time.
  • The approach could generalise to other high-impact thresholds (flash-flood guidance, fire weather) and to multivariate scores, where a single extreme may be less informative than compound events.
  • As limited-area AI models appear at convection-allowing resolutions, this verification design gives a ready template for comparing them against national high-res NWP without the AI model being penalised for smoothness.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper introduces a spatial verification framework that combines the High-Resolution Assessment (HiRA) neighbourhood approach with threshold-weighted continuous ranked probability scores (twCRPS), including a fair-score correction for unequal pseudo-ensemble sizes. The method is demonstrated on 32 months of 00 UTC forecasts from GraphCast-GFS and HRRR v4, verified at ASOS stations for 6-hour precipitation. Extreme events are defined by ERA5 gridpoint climatological quantiles (99th and 99.9th percentiles). The central empirical claims are that model rankings are sensitive to neighbourhood size; that under matched neighbourhood areas HRRR has better overall CRPS at all lead times but better extreme-event twCRPS only at short lead times; and that a CORP-like discrimination decomposition shows HRRR has superior discrimination at 6 h while GraphCast-GFS has slightly better discrimination beyond 24 h. The scoring equations appear correct, the code is made available in the open-source `scores` package, and the statistical testing (Diebold-Mariano with the Hering-Genton modification) is appropriate.

Significance. If the empirical claims hold, the paper makes a useful contribution to the emerging evaluation literature for AI weather prediction models: it provides a re-gridding-free, user-oriented method for comparing models of different resolutions when extremes are the target, and it shows that point-to-point verification may mis-rank AI and NWP models for extreme precipitation. The derivation of the fair twCRPS (Eq. 8) is transparent and correct in structure, and the authors are careful to use proper scoring rules to avoid the forecaster's dilemma. The 32-month, station-based evaluation against a widely used operational NWP model is a practical strength, as is the public availability of the scoring code. The main caveat is that the definition of 'extreme' is tied to ERA5 climatology, and the paper does not demonstrate that the headline ranking is robust to this choice; the discrimination analysis also rests on an in-sample calibration step with unequal neighbourhood sizes. These issues are local rather than fundamental, and they are addressable with additional analysis.

major comments (2)
  1. [Section 2.1 and Section 4.2] The central twCRPS finding—that HRRR outperforms GraphCast for extremes only at short lead times—depends on the definition of an extreme event, but the thresholds q_alpha are derived from ERA5 gridpoint climatology rather than station climatology. The paper acknowledges that these thresholds 'will differ to those derived directly from the station data' but provides no sensitivity test. This is load-bearing: GraphCast is trained on ERA5, and ERA5 has a documented dry bias over CONUS, so the ERA5-based q_0.99 is likely lower than the station-based quantile. As a result, the twCRPS tail in Sec. 4.2 may include events that a station-based user would not regard as extreme, and the lower threshold could disproportionately favour the ERA5-trained model at exactly the longer lead times where the paper claims non-inferiority. The q_0.999 check in Appendix A retains the same ERA5 reference, and th
  2. [Section 7, Eq. (9), Fig. 7] The discrimination (DSC) comparison is based on isotonic regression applied in-sample to each neighbourhood member, and the pseudo-ensemble sizes differ sharply between models: HRRR 21×27 has 567 members versus GraphCast 3×3 has 9, and even the HRRR 7×9 has 63 members versus 1. In-sample monotone calibration has much more flexibility for the larger HRRR neighbourhoods, so the cross-lead ranking of discrimination ability—including the abstract's claim that GraphCast has 'slightly better discrimination ability from a lead time of 24-hours onwards'—may partly reflect overfitting differences rather than true predictive discrimination. Please provide an out-of-sample or cross-validated version of the DSC analysis, or explicitly re-label the comparison as exploratory and note the overfitting risk.
minor comments (4)
  1. [Section 3.2, Eq. (8)] For M = 1, the fair correction term in Eq. (8) has denominator M(M-1) = 0. The point-forecast (1×1) cases are evidently handled separately, but this should be stated explicitly to avoid confusion.
  2. [Section 4.2, text after Fig. 4] Typo: 'neighbohood' should be 'neighbourhood'.
  3. [Section 7] Typo: 'twCPRS' should be 'twCRPS'; 'compoents' should be 'components'. Also, the CORP-decomposition wording is slightly ambiguous about whether the decomposition is applied to the twCRPS or to an average Brier-score style integral.
  4. [Section 8, future research] Typo: 'temeprature' should be 'temperature'.

Circularity Check

0 steps flagged

No significant circularity: core twCRPS/HiRA comparison is self-contained; ERA5-threshold choice and in-sample calibration are limitations, not circular steps.

full rationale

The central derivation chain is self-contained. The twCRPS/HiRA scores in Sec. 4.2 (Figs. 3-4) are computed from standard proper scoring rules: CRPS (Eqs. 1-4, Matheson-Winkler, Gneiting-Raftery, Ferro), twCRPS (Eq. 5, Gneiting-Ranjan 2011), and its ensemble/chaining form (Eqs. 6-8, Allen et al. 2023). These are external results, not fitted to the two models, and no parameter is estimated from the forecast-observation pairs before the central comparison; alpha and neighbourhood sizes are user-selected. The extreme threshold q_alpha is taken from ERA5 gridpoint climatology (Sec. 2.1) as a reference; the paper explicitly notes it 'will differ to those derived directly from the station data.' That is a validity limitation (and could in principle favour a model trained on ERA5), but it is not a circular step: the threshold is not derived from either model's output or from the target ranking. The in-sample isotonic regression in Sec. 7 is labelled 'in sample' and is used only as a diagnostic of potential discrimination; it does not feed the central lead-time ranking and its in-sample nature is a caveat, not a construction that forces the conclusions. The self-citations (Loveday et al. 2024; Pagano et al. 2024; Leeuwenburg et al. 2024) are contextual or software-related, and the scoring equations rest on independent literature. No prediction reduces to a fitted input, no uniqueness claim is imported from the authors, and no ansatz is smuggled in via self-citation. Score 1 reflects only the presence of minor self-references, not any circularity in the derivation.

Axiom & Free-Parameter Ledger

3 free parameters · 8 axioms · 0 invented entities

The core scoring equations are standard or cited; the main assumptions are the ERA5-derived extreme thresholds, station-as-point truth, and the operational-use framing. No free parameters are fitted to the target ranking; alpha and neighbourhood sizes are user choices.

free parameters (3)
  • Extreme threshold quantile alpha = 0.99 (main analysis); 0.999 (Appendix A)
    User-selected quantile of ERA5 climatology that defines 'extreme' in twCRPS; not fitted to the ranking claim, but central to the empirical conclusions.
  • Neighbourhood sizes = HRRR 1x1, 7x9, 21x27; GraphCast 1x1, 3x3 (roughly 3x3 km, 21x27 km, 63x81 km)
    Hand-chosen to compare equivalent physical areas and to illustrate the double-penalty effect; no fitting procedure.
  • Brier decomposition split = 30 mm
    Ad hoc display threshold in Fig. 5; not used in the main CRPS/twCRPS scores.
axioms (8)
  • standard math CRPS can be written as E|X-y| - 1/2 E|X-X'|
    Section 3.2 Eqs. 1-2; cited to Baringhaus/Franz (2004) and Szekely/Rizzo (2005).
  • standard math twCRPS with a non-negative weight and chaining function v is a proper score
    Section 3.2 Eqs. 5-7; cited to Gneiting/Ranjan (2011) and Allen et al. (2023).
  • standard math The fair correction to ensemble twCRPS is unbiased
    Eq. 8; inferred from Ferro (2013) fair scores and not derived in the text.
  • domain assumption ASOS 6-hour accumulations are valid point observations of precipitation
    Section 2.1; only basic world-record QC is applied; gauge-grid representativeness error is not corrected.
  • domain assumption ERA5 1990-2020 gridpoint quantiles define local extreme thresholds
    Section 2.1; may favour ERA5-trained GraphCast; sensitivity to threshold source is not tested.
  • domain assumption Equal-weight pseudo-ensembles from neighbourhood grid cells represent how meteorologists use forecasts
    Section 3.3; the 'user-oriented' framing assumes this specific operational use case.
  • domain assumption Rectangular HRRR neighbourhoods approximate the physical areas of GraphCast grid cells
    Table 1; an approximation because GraphCast's latitude-longitude grid cell size varies with latitude.
  • domain assumption In-sample isotonic regression isolates discrimination ability
    Section 7; calibration is in-sample, so DSC can overstate true prospective discrimination.

pith-pipeline@v1.3.0-alltime-deepseek · 18668 in / 14172 out tokens · 138699 ms · 2026-08-04T07:36:20.263544+00:00 · methodology

0 comments
read the original abstract

Recent advances in AI-based weather prediction have led to the development of artificial intelligence weather prediction (AIWP) models with competitive forecast skill compared to traditional NWP models, but with substantially reduced computational cost. There is a strong need for appropriate methods to evaluate their ability to predict extreme weather events, particularly when spatial coherence is important, and grid resolutions differ between models. We introduce a verification framework that combines spatial verification methods and proper scoring rules. Specifically, the framework extends the High-Resolution Assessment (HiRA) approach with threshold-weighted scoring rules. It enables user-oriented evaluation consistent with how forecasts may be interpreted by operational meteorologists or used in simple post-processing systems. The method supports targeted evaluation of extreme events by allowing flexible weighting of the relative importance of different decision thresholds. We demonstrate this framework by evaluating 32 months of precipitation forecasts from an AIWP model and a high-resolution NWP model. Our results show that model rankings are sensitive to the choice of neighbourhood size. Increasing the neighbourhood size has a greater impact on scores evaluating extreme-event performance for the high-resolution NWP model than for the AIWP model. At equivalent neighbourhood sizes, the high-resolution NWP model only outperformed the AIWP model in predicting extreme precipitation events at short lead times. We also demonstrate how this approach can be extended to evaluate discrimination ability in predicting heavy precipitation. We find that the high-resolution NWP model had superior discrimination ability at short lead times.

Figures

Figures reproduced from arXiv: 2510.25045 by Nicholas Loveday, Tracy Hertneky.

Figure 1
Figure 1. Figure 1: A map of ASOS station locations used in this study. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Graphical illustration of the threshold-weighted continuous ranked probability score (twCRPS) [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: (a) Mean CRPS results aggregated across all stations and timesteps. Lower scores are bet￾ter. (b) Difference between GraphCast-GFS 1×1 and HRRR 1×1 with 99% confidence intervals. (c) Difference between GraphCast-GFS 1×1 and HRRR 7×9 (21 × 27 km equivalent) with 99% confidence intervals. (d) Difference between GraphCast-GFS 3×3 and HRRR 21×27 (63×81 km equivalent)with 99% confidence intervals. In subfigures… view at source ↗
Figure 4
Figure 4. Figure 4: As for Fig. 3 but for the twCRPS with a threshold weight function of [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Brier score decomposition of the CRPS within the HiRA framework. Lower scores are better. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Q-Q plots of observations against forecasts. [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: (a) DSC results. Higher values indicate more discrimination ability for predicting extremes (all thresholds above the climatological 99th percentile). (b) Difference between GraphCast-GFS 1×1 and HRRR 1×1 with 99% confidence intervals. (c) Difference between GraphCast-GFS 1×1 and HRRR 7×9 with 99% confidence intervals. (d) Difference between GraphCast-GFS 3×3 and HRRR 21×27 with 99% confidence intervals. I… view at source ↗
Figure 8
Figure 8. Figure 8: As for Fig. 3 but for the twCRPS with a threshold weight function of [PITH_FULL_IMAGE:figures/full_fig_p017_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

65 extracted references · 16 canonical work pages

  1. [1]

    Smith, Sergey Frolov, Montgomery Flora, and Corey Potvin

    Daniel Abdi, Isidora Jankov, Paul Madden, Vanderlei Vargas, Timothy A. Smith, Sergey Frolov, Montgomery Flora, and Corey Potvin. Hrrrcast: a data-driven emulator for regional weather forecasting at convection allowing scales, 2025. URL https://arxiv.org/abs/2507.05658

  2. [2]

    Building machine learning limited area models: Kilometer-scale weather forecasting in realistic settings, 2025

    Simon Adamov, Joel Oskarsson, Leif Denby, Tomas Landelius, Kasper Hintz, Simon Christiansen, Irene Schicker, Carlos Osuna, Fredrik Lindsten, Oliver Fuhrer, and Sebastian Schemm. Building machine learning limited area models: Kilometer-scale weather forecasting in realistic settings, 2025. URL https://arxiv.org/abs/2504.09340

  3. [3]

    Weighted scoringrules: Emphasizing particular outcomes when evaluating probabilistic forecasts

    Sam Allen. Weighted scoringrules: Emphasizing particular outcomes when evaluating probabilistic forecasts. Journal of Statistical Software, 110 0 (8), 2024. ISSN 1548-7660. doi:10.18637/jss.v110.i08. URL http://dx.doi.org/10.18637/jss.v110.i08

  4. [4]

    Evaluating forecasts for high-impact events using transformed kernel scores

    Sam Allen, David Ginsbourger, and Johanna Ziegel. Evaluating forecasts for high-impact events using transformed kernel scores. SIAM/ASA Journal on Uncertainty Quantification, 11 0 (3): 0 906–940, August 2023. ISSN 2166-2525. doi:10.1137/22m1532184. URL http://dx.doi.org/10.1137/22m1532184

  5. [5]

    Decompositions of the mean continuous ranked probability score

    Sebastian Arnold, Eva-Maria Walz, Johanna Ziegel, and Tilmann Gneiting. Decompositions of the mean continuous ranked probability score. Electronic Journal of Statistics, 18 0 (2), January 2024. ISSN 1935-7524. doi:10.1214/24-ejs2316. URL http://dx.doi.org/10.1214/24-ejs2316

  6. [6]

    Miriam Ayer, H. D. Brunk, G. M. Ewing, W. T. Reid, and Edward Silverman. An empirical distribution function for sampling with incomplete information. The Annals of Mathematical Statistics, 26 0 (4): 0 641–647, December 1955. ISSN 0003-4851. doi:10.1214/aoms/1177728423. URL http://dx.doi.org/10.1214/aoms/1177728423

  7. [7]

    Baringhaus and C

    L. Baringhaus and C. Franz. On a new multivariate two-sample test. Journal of Multivariate Analysis, 88 0 (1): 0 190–206, January 2004. ISSN 0047-259X. doi:10.1016/s0047-259x(03)00079-4. URL http://dx.doi.org/10.1016/s0047-259x(03)00079-4

  8. [8]

    Zied Ben Bouallègue, Mariana C. A. Clare, Linus Magnusson, Estibaliz Gascón, Michael Maier-Gerber, Martin Janoušek, Mark Rodwell, Florian Pinault, Jesper S. Dramsch, Simon T. K. Lang, Baudouin Raoult, Florence Rabier, Matthieu Chevallier, Irina Sandu, Peter Dueben, Matthew Chantry, and Florian Pappenberger. The rise of data-driven weather forecasting: A f...

  9. [9]

    Accurate medium-range global weather forecasting with 3d neural networks

    Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619 0 (7970): 0 533–538, July 2023. ISSN 1476-4687. doi:10.1038/s41586-023-06185-3. URL http://dx.doi.org/10.1038/s41586-023-06185-3

  10. [10]

    Brenowitz, Yair Cohen, Jaideep Pathak, Ankur Mahesh, Boris Bonev, Thorsten Kurth, Dale R

    Noah D. Brenowitz, Yair Cohen, Jaideep Pathak, Ankur Mahesh, Boris Bonev, Thorsten Kurth, Dale R. Durran, Peter Harrington, and Michael S. Pritchard. A practical probabilistic benchmark for ai weather models. Geophysical Research Letters, 52 0 (7), April 2025. ISSN 1944-8007. doi:10.1029/2024gl113656. URL http://dx.doi.org/10.1029/2024gl113656

  11. [11]

    Glenn W. Brier. Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78 0 (1): 0 1–3, January 1950. ISSN 1520-0493. doi:10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2. URL http://dx.doi.org/10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2

  12. [12]

    Charlton-Perez, Helen F

    Andrew J. Charlton-Perez, Helen F. Dacre, Simon Driscoll, Suzanne L. Gray, Ben Harvey, Natalie J. Harvey, Kieran M. R. Hunt, Robert W. Lee, Ranjini Swaminathan, Remy Vandaele, and Ambrogio Volonté. Do ai models produce better weather forecasts than physics-based models? a quantitative evaluation case study of storm ciarán. npj Climate and Atmospheric Scie...

  13. [13]

    An approach to the verification of high-resolution ocean models using spatial methods

    Ric Crocker, Jan Maksymczuk, Marion Mittermaier, Marina Tonani, and Christine Pequignet. An approach to the verification of high-resolution ocean models using spatial methods. Ocean Science, 16 0 (4): 0 831–845, July 2020. ISSN 1812-0792. doi:10.5194/os-16-831-2020. URL http://dx.doi.org/10.5194/os-16-831-2020

  14. [14]

    Timo Dimitriadis, Tilmann Gneiting, and Alexander I. Jordan. Stable reliability diagrams for probabilistic classifiers. Proceedings of the National Academy of Sciences, 118 0 (8), February 2021. ISSN 1091-6490. doi:10.1073/pnas.2016191118. URL http://dx.doi.org/10.1073/pnas.2016191118

  15. [15]

    Mittermaier, Elizabeth E

    Manfred Dorninger, Eric Gilleland, Barbara Casati, Marion P. Mittermaier, Elizabeth E. Ebert, Barbara G. Brown, and Laurence J. Wilson. The setup of the mesovict project. Bulletin of the American Meteorological Society, 99 0 (9): 0 1887–1906, September 2018. ISSN 1520-0477. doi:10.1175/bams-d-17-0164.1. URL http://dx.doi.org/10.1175/bams-d-17-0164.1

  16. [16]

    Dowell, Curtis R

    David C. Dowell, Curtis R. Alexander, Eric P. James, Stephen S. Weygandt, Stanley G. Benjamin, Geoffrey S. Manikin, Benjamin T. Blake, John M. Brown, Joseph B. Olson, Ming Hu, Tatiana G. Smirnova, Terra Ladwig, Jaymes S. Kenyon, Ravan Ahmadov, David D. Turner, Jeffrey D. Duda, and Trevor I. Alcott. The high-resolution rapid refresh (hrrr): An hourly updat...

  17. [17]

    Elizabeth E. Ebert. Fuzzy verification of high‐resolution gridded forecasts: a review and proposed framework. Meteorological Applications, 15 0 (1): 0 51–64, March 2008. ISSN 1469-8080. doi:10.1002/met.25. URL http://dx.doi.org/10.1002/met.25

  18. [18]

    Elizabeth E. Ebert. Neighborhood verification: A strategy for rewarding close forecasts. Weather and Forecasting, 24 0 (6): 0 1498–1510, December 2009. ISSN 0882-8156. doi:10.1175/2009waf2222251.1. URL http://dx.doi.org/10.1175/2009waf2222251.1

  19. [19]

    Edward S. Epstein. A scoring system for probability forecasts of ranked categories. Journal of Applied Meteorology, 8 0 (6): 0 985–987, December 1969. ISSN 0021-8952. doi:10.1175/1520-0450(1969)008<0985:assfpf>2.0.co;2. URL http://dx.doi.org/10.1175/1520-0450(1969)008<0985:assfpf>2.0.co;2

  20. [20]

    C. A. T. Ferro. Fair scores for ensemble forecasts: Fair scores for ensemble forecasts. Quarterly Journal of the Royal Meteorological Society, 140 0 (683): 0 1917–1923, December 2013. ISSN 0035-9009. doi:10.1002/qj.2270. URL http://dx.doi.org/10.1002/qj.2270

  21. [21]

    Fovell and Alex Gallagher

    Robert G. Fovell and Alex Gallagher. An evaluation of surface wind and gust forecasts from the high-resolution rapid refresh model. Weather and Forecasting, 37 0 (6): 0 1049–1068, June 2022. ISSN 1520-0434. doi:10.1175/waf-d-21-0176.1. URL http://dx.doi.org/10.1175/waf-d-21-0176.1

  22. [22]

    Brown, Barbara Casati, and Elizabeth E

    Eric Gilleland, David Ahijevych, Barbara G. Brown, Barbara Casati, and Elizabeth E. Ebert. Intercomparison of spatial forecast verification methods. Weather and Forecasting, 24 0 (5): 0 1416–1430, October 2009. ISSN 0882-8156. doi:10.1175/2009waf2222269.1. URL http://dx.doi.org/10.1175/2009waf2222269.1

  23. [23]

    Making and evaluating point forecasts

    Tilmann Gneiting. Making and evaluating point forecasts. Journal of the American Statistical Association, 106 0 (494): 0 746–762, June 2011. ISSN 1537-274X. doi:10.1198/jasa.2011.r10138. URL http://dx.doi.org/10.1198/jasa.2011.r10138

  24. [24]

    Probabilistic forecasting

    Tilmann Gneiting and Matthias Katzfuss. Probabilistic forecasting. Annual Review of Statistics and Its Application, 1 0 (Volume 1, 2014): 0 125--151, 2014. ISSN 2326-831X. doi:https://doi.org/10.1146/annurev-statistics-062713-085831. URL https://www.annualreviews.org/content/journals/10.1146/annurev-statistics-062713-085831

  25. [25]

    Strictly proper scoring rules, prediction, and estimation

    Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102 0 (477): 0 359–378, March 2007. ISSN 1537-274X. doi:10.1198/016214506000001437. URL http://dx.doi.org/10.1198/016214506000001437

  26. [26]

    Comparing density forecasts using threshold- and quantile-weighted scoring rules

    Tilmann Gneiting and Roopesh Ranjan. Comparing density forecasts using threshold- and quantile-weighted scoring rules. Journal of Business & Economic Statistics, 29 0 (3): 0 411–422, July 2011. ISSN 1537-2707. doi:10.1198/jbes.2010.08110. URL http://dx.doi.org/10.1198/jbes.2010.08110

  27. [27]

    Jordan, and Sebastian Lerch

    Tilmann Gneiting, Tobias Biegert, Kristof Kraus, Eva-Maria Walz, Alexander I. Jordan, and Sebastian Lerch. Probabilistic measures afford fair comparisons of aiwp and nwp model output, 2025. URL https://arxiv.org/abs/2506.03744

  28. [28]

    Ziegel, and Tilmann Gneiting

    Alexander Henzi, Johanna F. Ziegel, and Tilmann Gneiting. Isotonic distributional regression. Journal of the Royal Statistical Society Series B: Statistical Methodology, 83 0 (5): 0 963–993, August 2021. ISSN 1467-9868. doi:10.1111/rssb.12450. URL http://dx.doi.org/10.1111/rssb.12450

  29. [29]

    The era5 global reanalysis

    Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, Andr \'a s Hor \'a nyi, Joaqu \' n Mu \ n oz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al. The era5 global reanalysis. Quarterly journal of the royal meteorological society, 146 0 (730): 0 1999--2049, 2020. doi:https://doi.org/10.1002/qj.3803

  30. [30]

    Evaluation of cold-season precipitation forecasts generated by the hourly updating high-resolution rapid refresh model

    Kyoko Ikeda, Matthias Steiner, James Pinto, and Curtis Alexander. Evaluation of cold-season precipitation forecasts generated by the hourly updating high-resolution rapid refresh model. Weather and Forecasting, 28 0 (4): 0 921–939, July 2013. ISSN 1520-0434. doi:10.1175/waf-d-12-00085.1. URL http://dx.doi.org/10.1175/waf-d-12-00085.1

  31. [31]

    Weatherreal: A benchmark based on in-situ observations for evaluating weather models, 2024

    Weixin Jin, Jonathan Weyn, Pengcheng Zhao, Siqi Xiang, Jiang Bian, Zuliang Fang, Haiyu Dong, Hongyu Sun, Kit Thambiratnam, and Qi Zhang. Weatherreal: A benchmark based on in-situ observations for evaluating weather models, 2024. URL https://arxiv.org/abs/2409.09371

  32. [32]

    Forecasting global weather with graph neural networks, 2022

    Ryan Keisler. Forecasting global weather with graph neural networks, 2022. URL https://arxiv.org/abs/2202.07575

  33. [33]

    Learning skillful medium-range global weather forecasting

    Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Holland, Oriol Vinyals, Jacklynn Stott, Alexander Pritzel, Shakir Mohamed, and Peter Battaglia. Learning skillful medium-range global weather forecasting. Scien...

  34. [34]

    Simon Lang, Mihai Alexe, Matthew Chantry, Jesper Dramsch, Florian Pinault, Baudouin Raoult, Mariana C. A. Clare, Christian Lessig, Michael Maier-Gerber, Linus Magnusson, Zied Ben Bouallègue, Ana Prieto Nemesio, Peter D. Dueben, Andrew Brown, Florian Pappenberger, and Florence Rabier. Aifs -- ecmwf's data-driven forecasting system, 2024 a . URL https://arx...

  35. [35]

    Simon Lang, Mihai Alexe, Mariana C. A. Clare, Christopher Roberts, Rilwan Adewoyin, Zied Ben Bouallègue, Matthew Chantry, Jesper Dramsch, Peter D. Dueben, Sara Hahner, Pedro Maciel, Ana Prieto-Nemesio, Cathal O'Brien, Florian Pinault, Jan Polster, Baudouin Raoult, Steffen Tietsche, and Martin Leutbecher. Aifs-crps: Ensemble forecasting using a model train...

  36. [36]

    Lavers, Adrian Simmons, Freja Vamborg, and Mark J

    David A. Lavers, Adrian Simmons, Freja Vamborg, and Mark J. Rodwell. An evaluation of era5 precipitation for climate monitoring. Quarterly Journal of the Royal Meteorological Society, 148 0 (748): 0 3152–3165, August 2022. ISSN 1477-870X. doi:10.1002/qj.4351. URL http://dx.doi.org/10.1002/qj.4351

  37. [37]

    Ebert, Harrison Cook, Mohammadreza Khanarmuei, Robert J

    Tennessee Leeuwenburg, Nicholas Loveday, Elizabeth E. Ebert, Harrison Cook, Mohammadreza Khanarmuei, Robert J. Taggart, Nikeeth Ramanathan, Maree Carroll, Stephanie Chong, Aidan Griffiths, and John Sharples. scores: A python package for verifying and evaluating models and predictions with xarray. Journal of Open Source Software, 9 0 (99): 0 6889, July 202...

  38. [38]

    Thorarinsdottir, Francesco Ravazzolo, and Tilmann Gneiting

    Sebastian Lerch, Thordis L. Thorarinsdottir, Francesco Ravazzolo, and Tilmann Gneiting. Forecaster’s dilemma: Extreme events and forecast evaluation. Statistical Science, 32 0 (1), February 2017. ISSN 0883-4237. doi:10.1214/16-sts588. URL http://dx.doi.org/10.1214/16-sts588

  39. [39]

    Taggart, Thomas C

    Nicholas Loveday, Deryn Griffiths, Tennessee Leeuwenburg, Robert J. Taggart, Thomas C. Pagano, George Cheng, Kevin Plastow, Elizabeth E. Ebert, Cassandra Templeton, Maree Carroll, Mohammadreza Khanarmuei, and Isha Nagpal. The jive verification system and its transformative impact on weather forecasting operations. Bulletin of the American Meteorological S...

  40. [40]

    Mass, David Ovens, Ken Westrick, and Brian A

    Clifford F. Mass, David Ovens, Ken Westrick, and Brian A. Colle. Does increasing horizontal resolution produce more skillful forecasts? Bulletin of the American Meteorological Society, 83 0 (3): 0 407–430, March 2002. ISSN 1520-0477. doi:10.1175/1520-0477(2002)083<0407:dihrpm>2.3.co;2. URL http://dx.doi.org/10.1175/1520-0477(2002)083<0407:dihrpm>2.3.co;2

  41. [41]

    Matheson and Robert L

    James E. Matheson and Robert L. Winkler. Scoring rules for continuous probability distributions. Management Science, 22 0 (10): 0 1087–1096, June 1976. ISSN 1526-5501. doi:10.1287/mnsc.22.10.1087. URL http://dx.doi.org/10.1287/mnsc.22.10.1087

  42. [42]

    M. P. Mittermaier and G. Csima. Ensemble versus deterministic performance at the kilometer scale. Weather and Forecasting, 32 0 (5): 0 1697–1709, September 2017. ISSN 1520-0434. doi:10.1175/waf-d-16-0164.1. URL http://dx.doi.org/10.1175/waf-d-16-0164.1

  43. [43]

    Mittermaier

    Marion P. Mittermaier. A strategy for verifying near-convection-resolving model forecasts at observing sites. Weather and Forecasting, 29 0 (2): 0 185–204, April 2014. ISSN 1520-0434. doi:10.1175/waf-d-12-00075.1. URL http://dx.doi.org/10.1175/waf-d-12-00075.1

  44. [44]

    Object-oriented verification of tc-jasper rainfall forecasts: Machine learning, 2025

    Hector Morisseau, Hongyan Zhu, Debra Hudson, and Catherine de Burgh-Day. Object-oriented verification of tc-jasper rainfall forecasts: Machine learning, 2025. URL http://www.bom.gov.au/research/publications/researchreports/BRR-106.pdf

  45. [45]

    Regional data-driven weather modeling with a global stretched-grid, 2024

    Thomas Nils Nipen, Håvard Homleid Haugen, Magnus Sikora Ingstad, Even Marius Nordhagen, Aram Farhad Shafiq Salihi, Paulina Tedesco, Ivar Ambjørn Seierstad, Jørn Kristiansen, Simon Lang, Mihai Alexe, Jesper Dramsch, Baudouin Raoult, Gert Mertes, and Matthew Chantry. Regional data-driven weather modeling with a global stretched-grid, 2024. URL https://arxiv...

  46. [46]

    Automated surface observing system (asos) user's guide

    NWS. Automated surface observing system (asos) user's guide. Technical report, NOAA, 1998

  47. [47]

    Do data-driven models beat numerical models in forecasting weather extremes? a comparison of ifs hres, pangu-weather, and graphcast

    Leonardo Olivetti and Gabriele Messori. Do data-driven models beat numerical models in forecasting weather extremes? a comparison of ifs hres, pangu-weather, and graphcast. Geoscientific Model Development, 17 0 (21): 0 7915–7962, November 2024. ISSN 1991-9603. doi:10.5194/gmd-17-7915-2024. URL http://dx.doi.org/10.5194/gmd-17-7915-2024

  48. [48]

    Pagano, Barbara Casati, Stephanie Landman, Nicholas Loveday, Robert Taggart, Elizabeth E

    Thomas C. Pagano, Barbara Casati, Stephanie Landman, Nicholas Loveday, Robert Taggart, Elizabeth E. Ebert, Mohammadreza Khanarmuei, Tara L. Jensen, Marion Mittermaier, Helen Roberts, Steve Willington, Nigel Roberts, Mike Sowko, Gordon Strassberg, Charles Kluepfel, Timothy A. Bullock, David D. Turner, Florian Pappenberger, Neal Osborne, and Chris Noble. Ch...

  49. [49]

    Proper scoring rules for multivariate probabilistic forecasts based on aggregation and transformation

    Romain Pic, Clément Dombry, Philippe Naveau, and Maxime Taillardat. Proper scoring rules for multivariate probabilistic forecasts based on aggregation and transformation. Advances in Statistical Climatology, Meteorology and Oceanography, 11 0 (1): 0 23–58, March 2025. ISSN 2364-3587. doi:10.5194/ascmo-11-23-2025. URL http://dx.doi.org/10.5194/ascmo-11-23-2025

  50. [50]

    Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

    Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. Probabilistic weather forecasting with machine learning. Nature, 637 0 (8044): 0 84–90, December 2024. ISSN 1476-4687. doi:10.1038/s41586-024-08252-9. URL http://d...

  51. [51]

    Radford, Imme Ebert-Uphoff, and Jebb Q

    Jacob T. Radford, Imme Ebert-Uphoff, and Jebb Q. Stewart. A comparison of ai weather prediction and numerical weather prediction models for 1–7-day precipitation forecasts. Weather and Forecasting, March 2025 a . ISSN 1520-0434. doi:10.1175/waf-d-24-0081.1. URL http://dx.doi.org/10.1175/waf-d-24-0081.1

  52. [52]

    Radford, Imme Ebert-Uphoff, Jebb Q

    Jacob T. Radford, Imme Ebert-Uphoff, Jebb Q. Stewart, Kate D. Musgrave, Robert DeMaria, Natalie Tourville, and Kyle Hilburn. Accelerating community-wide evaluation of ai models for global weather prediction by facilitating access to model output. Bulletin of the American Meteorological Society, 106 0 (1): 0 E68–E76, January 2025 b . ISSN 1520-0477. doi:10...

  53. [53]

    Weatherbench 2: A benchmark for the next generation of data‐driven global weather models

    Stephan Rasp, Stephan Hoyer, Alexander Merose, Ian Langmore, Peter Battaglia, Tyler Russell, Alvaro Sanchez‐Gonzalez, Vivian Yang, Rob Carver, Shreya Agrawal, Matthew Chantry, Zied Ben Bouallegue, Peter Dueben, Carla Bromberg, Jared Sisk, Luke Barrington, Aaron Bell, and Fei Sha. Weatherbench 2: A benchmark for the next generation of data‐driven global we...

  54. [54]

    Rodwell, David S

    Mark J. Rodwell, David S. Richardson, Tim D. Hewson, and Thomas Haiden. A new equitable score suitable for verifying precipitation in numerical weather prediction. Quarterly Journal of the Royal Meteorological Society, 136 0 (650): 0 1344–1363, July 2010. ISSN 1477-870X. doi:10.1002/qj.656. URL http://dx.doi.org/10.1002/qj.656

  55. [55]

    Schwartz and Ryan A

    Craig S. Schwartz and Ryan A. Sobash. Generating probabilistic forecasts from convection-allowing ensembles using neighborhood approaches: A review and recommendations. Monthly Weather Review, 145 0 (9): 0 3397–3418, September 2017. ISSN 1520-0493. doi:10.1175/mwr-d-16-0400.1. URL http://dx.doi.org/10.1175/mwr-d-16-0400.1

  56. [56]

    Neighborhood-based ensemble evaluation using the crps

    Joël Stein and Fabien Stoop. Neighborhood-based ensemble evaluation using the crps. Monthly Weather Review, 150 0 (8): 0 1901–1914, August 2022. ISSN 1520-0493. doi:10.1175/mwr-d-21-0224.1. URL http://dx.doi.org/10.1175/mwr-d-21-0224.1

  57. [57]

    Fixing the double penalty in data-driven weather forecasting through a modified spherical harmonic loss function, 2025

    Christopher Subich, Syed Zahid Husain, Leo Separovic, and Jing Yang. Fixing the double penalty in data-driven weather forecasting through a modified spherical harmonic loss function, 2025. URL https://arxiv.org/abs/2501.19374

  58. [58]

    Székely and Maria L

    Gábor J. Székely and Maria L. Rizzo. A new test for multivariate normality. Journal of Multivariate Analysis, 93 0 (1): 0 58–80, March 2005. ISSN 0047-259X. doi:10.1016/j.jmva.2003.12.002. URL http://dx.doi.org/10.1016/j.jmva.2003.12.002

  59. [59]

    Evaluation of point forecasts for extreme events using consistent scoring functions

    Robert Taggart. Evaluation of point forecasts for extreme events using consistent scoring functions. Quarterly Journal of the Royal Meteorological Society, 148 0 (742): 0 306–320, November 2021. ISSN 1477-870X. doi:10.1002/qj.4206. URL http://dx.doi.org/10.1002/qj.4206

  60. [60]

    Scale issues in verification of precipitation forecasts

    Ben Tustison, Daniel Harris, and Efi Foufoula‐Georgiou. Scale issues in verification of precipitation forecasts. Journal of Geophysical Research: Atmospheres, 106 0 (D11): 0 11775–11784, June 2001. ISSN 0148-0227. doi:10.1029/2001jd900066. URL http://dx.doi.org/10.1029/2001jd900066

  61. [61]

    Easy uncertainty quantification (easyuq): Generating predictive distributions from single-valued model output

    Eva-Maria Walz, Alexander Henzi, Johanna Ziegel, and Tilmann Gneiting. Easy uncertainty quantification (easyuq): Generating predictive distributions from single-valued model output. SIAM Review, 66 0 (1): 0 91–122, February 2024. ISSN 1095-7200. doi:10.1137/22m1541915. URL http://dx.doi.org/10.1137/22m1541915

  62. [62]

    Jakob Benjamin Wessel, Christopher A. T. Ferro, Gavin R. Evans, and Frank Kwasniok. Improving probabilistic forecasts of extreme wind speeds by training statistical post-processing models with weighted scoring rules. Monthly Weather Review, April 2025. ISSN 1520-0493. doi:10.1175/mwr-d-24-0151.1. URL http://dx.doi.org/10.1175/mwr-d-24-0151.1

  63. [63]

    Winkler and Allan H

    Robert L. Winkler and Allan H. Murphy. “good” probability assessors. Journal of Applied Meteorology, 7 0 (5): 0 751–758, October 1968. ISSN 0021-8952. doi:10.1175/1520-0450(1968)007<0751:pa>2.0.co;2. URL http://dx.doi.org/10.1175/1520-0450(1968)007<0751:pa>2.0.co;2

  64. [64]

    Guide to hydrological practices

    WMO. Guide to hydrological practices. Technical Report 168, World Meteorological Organization, 1994

  65. [65]

    Numerical models outperform ai weather forecasts of record-breaking extremes, 2025

    Zhongwei Zhang, Erich Fischer, Jakob Zscheischler, and Sebastian Engelke. Numerical models outperform ai weather forecasts of record-breaking extremes, 2025. URL https://arxiv.org/abs/2508.15724