REVIEW 2 major objections 4 minor 65 references
Compared over matched spatial areas, an AI weather model matches or beats a 3-km physics-based model for extreme 6-hour precipitation at lead times of 24 hours and beyond, while the physics model wins only at short lead times.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 07:36 UTC pith:2CLFAQ7F
load-bearing objection Useful HiRA + twCRPS integration with sound math and an honest GraphCast vs HRRR comparison; the main fix before publication is a sensitivity test for the ERA5-based extreme thresholds. the 2 major comments →
Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that the relative skill of an AI weather prediction model and a high-resolution numerical model for extreme precipitation depends strongly on the spatial scale of evaluation and the lead time. Using the HiRA framework to build neighbourhood pseudo-ensembles and the threshold-weighted CRPS with a chaining function v(z)=max(z, q_alpha) to score only the upper tail, the authors find that when the AI model and the 3-km HRRR model are compared over equivalent physical areas (e.g., 63x81 km), HRRR outperforms GraphCast-GFS for 6-hour precipitation above the climatological 99th percentile only at short lead times; at lead times of 24 hours and beyond, GraphCast-GFS is compe
What carries the argument
The framework marries two existing tools. HiRA (High-Resolution Assessment) converts each point observation into a pseudo-ensemble of forecast values from a neighbourhood of grid cells, creating an equal-probability distribution without regridding. The threshold-weighted continuous ranked probability score (twCRPS) evaluates that pseudo-ensemble using a chaining function v(z)=max(z, q_alpha), which gives unit weight only to thresholds above the local climatological extreme (q0.99 or q0.999). A 'fair' correction to the twCRPS accounts for pseudo-ensembles of different sizes, so models with different native grid resolutions can be compared over matched physical areas; the result is a proper, u
Load-bearing premise
The load-bearing assumption is that extreme events are defined by ERA5-reanalysis thresholds rather than by the station observations themselves; since GraphCast is trained on ERA5, this definition may favour the AI model.
What would settle it
Recompute the twCRPS lead-time comparison defining extreme thresholds from the ASOS station climatology (or from a different reanalysis) and check whether GraphCast-GFS still matches or beats HRRR at 24+ hours; if the AI advantage vanishes, the ranking claim depends on the threshold source.
If this is right
- Point-to-point verification can mis-rank models for extremes; spatial scale must be part of any such comparison.
- At short lead times, the radar-assimilating high-resolution model (HRRR) retains an edge for extreme precipitation; beyond ~24 hours the AI model is at least as good over matched areas.
- The method lets agencies compare models of different native resolutions without regridding or degrading the higher-resolution model.
- The threshold weighting makes the score a proper scoring rule, so it rewards honest forecasts and can be used to train or tune post-processing for extremes.
- The same framework extends to discrimination ability (calibrated DSC), separating whether a model can tell extremes apart from whether it is well calibrated.
Where Pith is reading between the lines
- Because the extreme thresholds come from ERA5 — the same reanalysis the AI model was trained on — the long-lead edge could partly reflect a favourable definition of 'extreme'; rerunning with station-based thresholds is a natural stress test.
- The ranking flip with neighbourhood size means operational model choice depends on the decision's spatial scale (e.g., a single city vs a river catchment), not just the variable and lead time.
- The approach could generalise to other high-impact thresholds (flash-flood guidance, fire weather) and to multivariate scores, where a single extreme may be less informative than compound events.
- As limited-area AI models appear at convection-allowing resolutions, this verification design gives a ready template for comparing them against national high-res NWP without the AI model being penalised for smoothness.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a spatial verification framework that combines the High-Resolution Assessment (HiRA) neighbourhood approach with threshold-weighted continuous ranked probability scores (twCRPS), including a fair-score correction for unequal pseudo-ensemble sizes. The method is demonstrated on 32 months of 00 UTC forecasts from GraphCast-GFS and HRRR v4, verified at ASOS stations for 6-hour precipitation. Extreme events are defined by ERA5 gridpoint climatological quantiles (99th and 99.9th percentiles). The central empirical claims are that model rankings are sensitive to neighbourhood size; that under matched neighbourhood areas HRRR has better overall CRPS at all lead times but better extreme-event twCRPS only at short lead times; and that a CORP-like discrimination decomposition shows HRRR has superior discrimination at 6 h while GraphCast-GFS has slightly better discrimination beyond 24 h. The scoring equations appear correct, the code is made available in the open-source `scores` package, and the statistical testing (Diebold-Mariano with the Hering-Genton modification) is appropriate.
Significance. If the empirical claims hold, the paper makes a useful contribution to the emerging evaluation literature for AI weather prediction models: it provides a re-gridding-free, user-oriented method for comparing models of different resolutions when extremes are the target, and it shows that point-to-point verification may mis-rank AI and NWP models for extreme precipitation. The derivation of the fair twCRPS (Eq. 8) is transparent and correct in structure, and the authors are careful to use proper scoring rules to avoid the forecaster's dilemma. The 32-month, station-based evaluation against a widely used operational NWP model is a practical strength, as is the public availability of the scoring code. The main caveat is that the definition of 'extreme' is tied to ERA5 climatology, and the paper does not demonstrate that the headline ranking is robust to this choice; the discrimination analysis also rests on an in-sample calibration step with unequal neighbourhood sizes. These issues are local rather than fundamental, and they are addressable with additional analysis.
major comments (2)
- [Section 2.1 and Section 4.2] The central twCRPS finding—that HRRR outperforms GraphCast for extremes only at short lead times—depends on the definition of an extreme event, but the thresholds q_alpha are derived from ERA5 gridpoint climatology rather than station climatology. The paper acknowledges that these thresholds 'will differ to those derived directly from the station data' but provides no sensitivity test. This is load-bearing: GraphCast is trained on ERA5, and ERA5 has a documented dry bias over CONUS, so the ERA5-based q_0.99 is likely lower than the station-based quantile. As a result, the twCRPS tail in Sec. 4.2 may include events that a station-based user would not regard as extreme, and the lower threshold could disproportionately favour the ERA5-trained model at exactly the longer lead times where the paper claims non-inferiority. The q_0.999 check in Appendix A retains the same ERA5 reference, and th
- [Section 7, Eq. (9), Fig. 7] The discrimination (DSC) comparison is based on isotonic regression applied in-sample to each neighbourhood member, and the pseudo-ensemble sizes differ sharply between models: HRRR 21×27 has 567 members versus GraphCast 3×3 has 9, and even the HRRR 7×9 has 63 members versus 1. In-sample monotone calibration has much more flexibility for the larger HRRR neighbourhoods, so the cross-lead ranking of discrimination ability—including the abstract's claim that GraphCast has 'slightly better discrimination ability from a lead time of 24-hours onwards'—may partly reflect overfitting differences rather than true predictive discrimination. Please provide an out-of-sample or cross-validated version of the DSC analysis, or explicitly re-label the comparison as exploratory and note the overfitting risk.
minor comments (4)
- [Section 3.2, Eq. (8)] For M = 1, the fair correction term in Eq. (8) has denominator M(M-1) = 0. The point-forecast (1×1) cases are evidently handled separately, but this should be stated explicitly to avoid confusion.
- [Section 4.2, text after Fig. 4] Typo: 'neighbohood' should be 'neighbourhood'.
- [Section 7] Typo: 'twCPRS' should be 'twCRPS'; 'compoents' should be 'components'. Also, the CORP-decomposition wording is slightly ambiguous about whether the decomposition is applied to the twCRPS or to an average Brier-score style integral.
- [Section 8, future research] Typo: 'temeprature' should be 'temperature'.
Circularity Check
No significant circularity: core twCRPS/HiRA comparison is self-contained; ERA5-threshold choice and in-sample calibration are limitations, not circular steps.
full rationale
The central derivation chain is self-contained. The twCRPS/HiRA scores in Sec. 4.2 (Figs. 3-4) are computed from standard proper scoring rules: CRPS (Eqs. 1-4, Matheson-Winkler, Gneiting-Raftery, Ferro), twCRPS (Eq. 5, Gneiting-Ranjan 2011), and its ensemble/chaining form (Eqs. 6-8, Allen et al. 2023). These are external results, not fitted to the two models, and no parameter is estimated from the forecast-observation pairs before the central comparison; alpha and neighbourhood sizes are user-selected. The extreme threshold q_alpha is taken from ERA5 gridpoint climatology (Sec. 2.1) as a reference; the paper explicitly notes it 'will differ to those derived directly from the station data.' That is a validity limitation (and could in principle favour a model trained on ERA5), but it is not a circular step: the threshold is not derived from either model's output or from the target ranking. The in-sample isotonic regression in Sec. 7 is labelled 'in sample' and is used only as a diagnostic of potential discrimination; it does not feed the central lead-time ranking and its in-sample nature is a caveat, not a construction that forces the conclusions. The self-citations (Loveday et al. 2024; Pagano et al. 2024; Leeuwenburg et al. 2024) are contextual or software-related, and the scoring equations rest on independent literature. No prediction reduces to a fitted input, no uniqueness claim is imported from the authors, and no ansatz is smuggled in via self-citation. Score 1 reflects only the presence of minor self-references, not any circularity in the derivation.
Axiom & Free-Parameter Ledger
free parameters (3)
- Extreme threshold quantile alpha =
0.99 (main analysis); 0.999 (Appendix A)
- Neighbourhood sizes =
HRRR 1x1, 7x9, 21x27; GraphCast 1x1, 3x3 (roughly 3x3 km, 21x27 km, 63x81 km)
- Brier decomposition split =
30 mm
axioms (8)
- standard math CRPS can be written as E|X-y| - 1/2 E|X-X'|
- standard math twCRPS with a non-negative weight and chaining function v is a proper score
- standard math The fair correction to ensemble twCRPS is unbiased
- domain assumption ASOS 6-hour accumulations are valid point observations of precipitation
- domain assumption ERA5 1990-2020 gridpoint quantiles define local extreme thresholds
- domain assumption Equal-weight pseudo-ensembles from neighbourhood grid cells represent how meteorologists use forecasts
- domain assumption Rectangular HRRR neighbourhoods approximate the physical areas of GraphCast grid cells
- domain assumption In-sample isotonic regression isolates discrimination ability
read the original abstract
Recent advances in AI-based weather prediction have led to the development of artificial intelligence weather prediction (AIWP) models with competitive forecast skill compared to traditional NWP models, but with substantially reduced computational cost. There is a strong need for appropriate methods to evaluate their ability to predict extreme weather events, particularly when spatial coherence is important, and grid resolutions differ between models. We introduce a verification framework that combines spatial verification methods and proper scoring rules. Specifically, the framework extends the High-Resolution Assessment (HiRA) approach with threshold-weighted scoring rules. It enables user-oriented evaluation consistent with how forecasts may be interpreted by operational meteorologists or used in simple post-processing systems. The method supports targeted evaluation of extreme events by allowing flexible weighting of the relative importance of different decision thresholds. We demonstrate this framework by evaluating 32 months of precipitation forecasts from an AIWP model and a high-resolution NWP model. Our results show that model rankings are sensitive to the choice of neighbourhood size. Increasing the neighbourhood size has a greater impact on scores evaluating extreme-event performance for the high-resolution NWP model than for the AIWP model. At equivalent neighbourhood sizes, the high-resolution NWP model only outperformed the AIWP model in predicting extreme precipitation events at short lead times. We also demonstrate how this approach can be extended to evaluate discrimination ability in predicting heavy precipitation. We find that the high-resolution NWP model had superior discrimination ability at short lead times.
Figures
Reference graph
Works this paper leans on
-
[1]
Smith, Sergey Frolov, Montgomery Flora, and Corey Potvin
Daniel Abdi, Isidora Jankov, Paul Madden, Vanderlei Vargas, Timothy A. Smith, Sergey Frolov, Montgomery Flora, and Corey Potvin. Hrrrcast: a data-driven emulator for regional weather forecasting at convection allowing scales, 2025. URL https://arxiv.org/abs/2507.05658
Pith/arXiv arXiv 2025
-
[2]
Simon Adamov, Joel Oskarsson, Leif Denby, Tomas Landelius, Kasper Hintz, Simon Christiansen, Irene Schicker, Carlos Osuna, Fredrik Lindsten, Oliver Fuhrer, and Sebastian Schemm. Building machine learning limited area models: Kilometer-scale weather forecasting in realistic settings, 2025. URL https://arxiv.org/abs/2504.09340
Pith/arXiv arXiv 2025
-
[3]
Weighted scoringrules: Emphasizing particular outcomes when evaluating probabilistic forecasts
Sam Allen. Weighted scoringrules: Emphasizing particular outcomes when evaluating probabilistic forecasts. Journal of Statistical Software, 110 0 (8), 2024. ISSN 1548-7660. doi:10.18637/jss.v110.i08. URL http://dx.doi.org/10.18637/jss.v110.i08
-
[4]
Evaluating forecasts for high-impact events using transformed kernel scores
Sam Allen, David Ginsbourger, and Johanna Ziegel. Evaluating forecasts for high-impact events using transformed kernel scores. SIAM/ASA Journal on Uncertainty Quantification, 11 0 (3): 0 906–940, August 2023. ISSN 2166-2525. doi:10.1137/22m1532184. URL http://dx.doi.org/10.1137/22m1532184
-
[5]
Decompositions of the mean continuous ranked probability score
Sebastian Arnold, Eva-Maria Walz, Johanna Ziegel, and Tilmann Gneiting. Decompositions of the mean continuous ranked probability score. Electronic Journal of Statistics, 18 0 (2), January 2024. ISSN 1935-7524. doi:10.1214/24-ejs2316. URL http://dx.doi.org/10.1214/24-ejs2316
-
[6]
Miriam Ayer, H. D. Brunk, G. M. Ewing, W. T. Reid, and Edward Silverman. An empirical distribution function for sampling with incomplete information. The Annals of Mathematical Statistics, 26 0 (4): 0 641–647, December 1955. ISSN 0003-4851. doi:10.1214/aoms/1177728423. URL http://dx.doi.org/10.1214/aoms/1177728423
arXiv 1955
-
[7]
L. Baringhaus and C. Franz. On a new multivariate two-sample test. Journal of Multivariate Analysis, 88 0 (1): 0 190–206, January 2004. ISSN 0047-259X. doi:10.1016/s0047-259x(03)00079-4. URL http://dx.doi.org/10.1016/s0047-259x(03)00079-4
-
[8]
Zied Ben Bouallègue, Mariana C. A. Clare, Linus Magnusson, Estibaliz Gascón, Michael Maier-Gerber, Martin Janoušek, Mark Rodwell, Florian Pinault, Jesper S. Dramsch, Simon T. K. Lang, Baudouin Raoult, Florence Rabier, Matthieu Chevallier, Irina Sandu, Peter Dueben, Matthew Chantry, and Florian Pappenberger. The rise of data-driven weather forecasting: A f...
2024
-
[9]
Accurate medium-range global weather forecasting with 3d neural networks
Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619 0 (7970): 0 533–538, July 2023. ISSN 1476-4687. doi:10.1038/s41586-023-06185-3. URL http://dx.doi.org/10.1038/s41586-023-06185-3
-
[10]
Brenowitz, Yair Cohen, Jaideep Pathak, Ankur Mahesh, Boris Bonev, Thorsten Kurth, Dale R
Noah D. Brenowitz, Yair Cohen, Jaideep Pathak, Ankur Mahesh, Boris Bonev, Thorsten Kurth, Dale R. Durran, Peter Harrington, and Michael S. Pritchard. A practical probabilistic benchmark for ai weather models. Geophysical Research Letters, 52 0 (7), April 2025. ISSN 1944-8007. doi:10.1029/2024gl113656. URL http://dx.doi.org/10.1029/2024gl113656
-
[11]
Glenn W. Brier. Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78 0 (1): 0 1–3, January 1950. ISSN 1520-0493. doi:10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2. URL http://dx.doi.org/10.1175/1520-0493(1950)078<0001:vofeit>2.0.co;2
-
[12]
Andrew J. Charlton-Perez, Helen F. Dacre, Simon Driscoll, Suzanne L. Gray, Ben Harvey, Natalie J. Harvey, Kieran M. R. Hunt, Robert W. Lee, Ranjini Swaminathan, Remy Vandaele, and Ambrogio Volonté. Do ai models produce better weather forecasts than physics-based models? a quantitative evaluation case study of storm ciarán. npj Climate and Atmospheric Scie...
-
[13]
An approach to the verification of high-resolution ocean models using spatial methods
Ric Crocker, Jan Maksymczuk, Marion Mittermaier, Marina Tonani, and Christine Pequignet. An approach to the verification of high-resolution ocean models using spatial methods. Ocean Science, 16 0 (4): 0 831–845, July 2020. ISSN 1812-0792. doi:10.5194/os-16-831-2020. URL http://dx.doi.org/10.5194/os-16-831-2020
-
[14]
Timo Dimitriadis, Tilmann Gneiting, and Alexander I. Jordan. Stable reliability diagrams for probabilistic classifiers. Proceedings of the National Academy of Sciences, 118 0 (8), February 2021. ISSN 1091-6490. doi:10.1073/pnas.2016191118. URL http://dx.doi.org/10.1073/pnas.2016191118
-
[15]
Manfred Dorninger, Eric Gilleland, Barbara Casati, Marion P. Mittermaier, Elizabeth E. Ebert, Barbara G. Brown, and Laurence J. Wilson. The setup of the mesovict project. Bulletin of the American Meteorological Society, 99 0 (9): 0 1887–1906, September 2018. ISSN 1520-0477. doi:10.1175/bams-d-17-0164.1. URL http://dx.doi.org/10.1175/bams-d-17-0164.1
-
[16]
Dowell, Curtis R
David C. Dowell, Curtis R. Alexander, Eric P. James, Stephen S. Weygandt, Stanley G. Benjamin, Geoffrey S. Manikin, Benjamin T. Blake, John M. Brown, Joseph B. Olson, Ming Hu, Tatiana G. Smirnova, Terra Ladwig, Jaymes S. Kenyon, Ravan Ahmadov, David D. Turner, Jeffrey D. Duda, and Trevor I. Alcott. The high-resolution rapid refresh (hrrr): An hourly updat...
2022
-
[17]
Elizabeth E. Ebert. Fuzzy verification of high‐resolution gridded forecasts: a review and proposed framework. Meteorological Applications, 15 0 (1): 0 51–64, March 2008. ISSN 1469-8080. doi:10.1002/met.25. URL http://dx.doi.org/10.1002/met.25
-
[18]
Elizabeth E. Ebert. Neighborhood verification: A strategy for rewarding close forecasts. Weather and Forecasting, 24 0 (6): 0 1498–1510, December 2009. ISSN 0882-8156. doi:10.1175/2009waf2222251.1. URL http://dx.doi.org/10.1175/2009waf2222251.1
-
[19]
Edward S. Epstein. A scoring system for probability forecasts of ranked categories. Journal of Applied Meteorology, 8 0 (6): 0 985–987, December 1969. ISSN 0021-8952. doi:10.1175/1520-0450(1969)008<0985:assfpf>2.0.co;2. URL http://dx.doi.org/10.1175/1520-0450(1969)008<0985:assfpf>2.0.co;2
-
[20]
C. A. T. Ferro. Fair scores for ensemble forecasts: Fair scores for ensemble forecasts. Quarterly Journal of the Royal Meteorological Society, 140 0 (683): 0 1917–1923, December 2013. ISSN 0035-9009. doi:10.1002/qj.2270. URL http://dx.doi.org/10.1002/qj.2270
doi:10.1002/qj.2270 1917
-
[21]
Robert G. Fovell and Alex Gallagher. An evaluation of surface wind and gust forecasts from the high-resolution rapid refresh model. Weather and Forecasting, 37 0 (6): 0 1049–1068, June 2022. ISSN 1520-0434. doi:10.1175/waf-d-21-0176.1. URL http://dx.doi.org/10.1175/waf-d-21-0176.1
-
[22]
Brown, Barbara Casati, and Elizabeth E
Eric Gilleland, David Ahijevych, Barbara G. Brown, Barbara Casati, and Elizabeth E. Ebert. Intercomparison of spatial forecast verification methods. Weather and Forecasting, 24 0 (5): 0 1416–1430, October 2009. ISSN 0882-8156. doi:10.1175/2009waf2222269.1. URL http://dx.doi.org/10.1175/2009waf2222269.1
-
[23]
Making and evaluating point forecasts
Tilmann Gneiting. Making and evaluating point forecasts. Journal of the American Statistical Association, 106 0 (494): 0 746–762, June 2011. ISSN 1537-274X. doi:10.1198/jasa.2011.r10138. URL http://dx.doi.org/10.1198/jasa.2011.r10138
-
[24]
Tilmann Gneiting and Matthias Katzfuss. Probabilistic forecasting. Annual Review of Statistics and Its Application, 1 0 (Volume 1, 2014): 0 125--151, 2014. ISSN 2326-831X. doi:https://doi.org/10.1146/annurev-statistics-062713-085831. URL https://www.annualreviews.org/content/journals/10.1146/annurev-statistics-062713-085831
-
[25]
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102 0 (477): 0 359–378, March 2007. ISSN 1537-274X. doi:10.1198/016214506000001437. URL http://dx.doi.org/10.1198/016214506000001437
-
[26]
Comparing density forecasts using threshold- and quantile-weighted scoring rules
Tilmann Gneiting and Roopesh Ranjan. Comparing density forecasts using threshold- and quantile-weighted scoring rules. Journal of Business & Economic Statistics, 29 0 (3): 0 411–422, July 2011. ISSN 1537-2707. doi:10.1198/jbes.2010.08110. URL http://dx.doi.org/10.1198/jbes.2010.08110
Pith/arXiv arXiv 2011
-
[27]
Tilmann Gneiting, Tobias Biegert, Kristof Kraus, Eva-Maria Walz, Alexander I. Jordan, and Sebastian Lerch. Probabilistic measures afford fair comparisons of aiwp and nwp model output, 2025. URL https://arxiv.org/abs/2506.03744
Pith/arXiv arXiv 2025
-
[28]
Alexander Henzi, Johanna F. Ziegel, and Tilmann Gneiting. Isotonic distributional regression. Journal of the Royal Statistical Society Series B: Statistical Methodology, 83 0 (5): 0 963–993, August 2021. ISSN 1467-9868. doi:10.1111/rssb.12450. URL http://dx.doi.org/10.1111/rssb.12450
-
[29]
Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, Andr \'a s Hor \'a nyi, Joaqu \' n Mu \ n oz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al. The era5 global reanalysis. Quarterly journal of the royal meteorological society, 146 0 (730): 0 1999--2049, 2020. doi:https://doi.org/10.1002/qj.3803
doi:10.1002/qj.3803 1999
-
[30]
Kyoko Ikeda, Matthias Steiner, James Pinto, and Curtis Alexander. Evaluation of cold-season precipitation forecasts generated by the hourly updating high-resolution rapid refresh model. Weather and Forecasting, 28 0 (4): 0 921–939, July 2013. ISSN 1520-0434. doi:10.1175/waf-d-12-00085.1. URL http://dx.doi.org/10.1175/waf-d-12-00085.1
-
[31]
Weatherreal: A benchmark based on in-situ observations for evaluating weather models, 2024
Weixin Jin, Jonathan Weyn, Pengcheng Zhao, Siqi Xiang, Jiang Bian, Zuliang Fang, Haiyu Dong, Hongyu Sun, Kit Thambiratnam, and Qi Zhang. Weatherreal: A benchmark based on in-situ observations for evaluating weather models, 2024. URL https://arxiv.org/abs/2409.09371
Pith/arXiv arXiv 2024
-
[32]
Forecasting global weather with graph neural networks, 2022
Ryan Keisler. Forecasting global weather with graph neural networks, 2022. URL https://arxiv.org/abs/2202.07575
Pith/arXiv arXiv 2022
-
[33]
Learning skillful medium-range global weather forecasting
Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Holland, Oriol Vinyals, Jacklynn Stott, Alexander Pritzel, Shakir Mohamed, and Peter Battaglia. Learning skillful medium-range global weather forecasting. Scien...
-
[34]
Simon Lang, Mihai Alexe, Matthew Chantry, Jesper Dramsch, Florian Pinault, Baudouin Raoult, Mariana C. A. Clare, Christian Lessig, Michael Maier-Gerber, Linus Magnusson, Zied Ben Bouallègue, Ana Prieto Nemesio, Peter D. Dueben, Andrew Brown, Florian Pappenberger, and Florence Rabier. Aifs -- ecmwf's data-driven forecasting system, 2024 a . URL https://arx...
Pith/arXiv arXiv 2024
-
[35]
Simon Lang, Mihai Alexe, Mariana C. A. Clare, Christopher Roberts, Rilwan Adewoyin, Zied Ben Bouallègue, Matthew Chantry, Jesper Dramsch, Peter D. Dueben, Sara Hahner, Pedro Maciel, Ana Prieto-Nemesio, Cathal O'Brien, Florian Pinault, Jan Polster, Baudouin Raoult, Steffen Tietsche, and Martin Leutbecher. Aifs-crps: Ensemble forecasting using a model train...
Pith/arXiv arXiv 2024
-
[36]
Lavers, Adrian Simmons, Freja Vamborg, and Mark J
David A. Lavers, Adrian Simmons, Freja Vamborg, and Mark J. Rodwell. An evaluation of era5 precipitation for climate monitoring. Quarterly Journal of the Royal Meteorological Society, 148 0 (748): 0 3152–3165, August 2022. ISSN 1477-870X. doi:10.1002/qj.4351. URL http://dx.doi.org/10.1002/qj.4351
doi:10.1002/qj.4351 2022
-
[37]
Ebert, Harrison Cook, Mohammadreza Khanarmuei, Robert J
Tennessee Leeuwenburg, Nicholas Loveday, Elizabeth E. Ebert, Harrison Cook, Mohammadreza Khanarmuei, Robert J. Taggart, Nikeeth Ramanathan, Maree Carroll, Stephanie Chong, Aidan Griffiths, and John Sharples. scores: A python package for verifying and evaluating models and predictions with xarray. Journal of Open Source Software, 9 0 (99): 0 6889, July 202...
-
[38]
Thorarinsdottir, Francesco Ravazzolo, and Tilmann Gneiting
Sebastian Lerch, Thordis L. Thorarinsdottir, Francesco Ravazzolo, and Tilmann Gneiting. Forecaster’s dilemma: Extreme events and forecast evaluation. Statistical Science, 32 0 (1), February 2017. ISSN 0883-4237. doi:10.1214/16-sts588. URL http://dx.doi.org/10.1214/16-sts588
-
[39]
Nicholas Loveday, Deryn Griffiths, Tennessee Leeuwenburg, Robert J. Taggart, Thomas C. Pagano, George Cheng, Kevin Plastow, Elizabeth E. Ebert, Cassandra Templeton, Maree Carroll, Mohammadreza Khanarmuei, and Isha Nagpal. The jive verification system and its transformative impact on weather forecasting operations. Bulletin of the American Meteorological S...
-
[40]
Mass, David Ovens, Ken Westrick, and Brian A
Clifford F. Mass, David Ovens, Ken Westrick, and Brian A. Colle. Does increasing horizontal resolution produce more skillful forecasts? Bulletin of the American Meteorological Society, 83 0 (3): 0 407–430, March 2002. ISSN 1520-0477. doi:10.1175/1520-0477(2002)083<0407:dihrpm>2.3.co;2. URL http://dx.doi.org/10.1175/1520-0477(2002)083<0407:dihrpm>2.3.co;2
-
[41]
James E. Matheson and Robert L. Winkler. Scoring rules for continuous probability distributions. Management Science, 22 0 (10): 0 1087–1096, June 1976. ISSN 1526-5501. doi:10.1287/mnsc.22.10.1087. URL http://dx.doi.org/10.1287/mnsc.22.10.1087
-
[42]
M. P. Mittermaier and G. Csima. Ensemble versus deterministic performance at the kilometer scale. Weather and Forecasting, 32 0 (5): 0 1697–1709, September 2017. ISSN 1520-0434. doi:10.1175/waf-d-16-0164.1. URL http://dx.doi.org/10.1175/waf-d-16-0164.1
-
[43]
Marion P. Mittermaier. A strategy for verifying near-convection-resolving model forecasts at observing sites. Weather and Forecasting, 29 0 (2): 0 185–204, April 2014. ISSN 1520-0434. doi:10.1175/waf-d-12-00075.1. URL http://dx.doi.org/10.1175/waf-d-12-00075.1
-
[44]
Object-oriented verification of tc-jasper rainfall forecasts: Machine learning, 2025
Hector Morisseau, Hongyan Zhu, Debra Hudson, and Catherine de Burgh-Day. Object-oriented verification of tc-jasper rainfall forecasts: Machine learning, 2025. URL http://www.bom.gov.au/research/publications/researchreports/BRR-106.pdf
2025
-
[45]
Regional data-driven weather modeling with a global stretched-grid, 2024
Thomas Nils Nipen, Håvard Homleid Haugen, Magnus Sikora Ingstad, Even Marius Nordhagen, Aram Farhad Shafiq Salihi, Paulina Tedesco, Ivar Ambjørn Seierstad, Jørn Kristiansen, Simon Lang, Mihai Alexe, Jesper Dramsch, Baudouin Raoult, Gert Mertes, and Matthew Chantry. Regional data-driven weather modeling with a global stretched-grid, 2024. URL https://arxiv...
Pith/arXiv arXiv 2024
-
[46]
Automated surface observing system (asos) user's guide
NWS. Automated surface observing system (asos) user's guide. Technical report, NOAA, 1998
1998
-
[47]
Leonardo Olivetti and Gabriele Messori. Do data-driven models beat numerical models in forecasting weather extremes? a comparison of ifs hres, pangu-weather, and graphcast. Geoscientific Model Development, 17 0 (21): 0 7915–7962, November 2024. ISSN 1991-9603. doi:10.5194/gmd-17-7915-2024. URL http://dx.doi.org/10.5194/gmd-17-7915-2024
-
[48]
Pagano, Barbara Casati, Stephanie Landman, Nicholas Loveday, Robert Taggart, Elizabeth E
Thomas C. Pagano, Barbara Casati, Stephanie Landman, Nicholas Loveday, Robert Taggart, Elizabeth E. Ebert, Mohammadreza Khanarmuei, Tara L. Jensen, Marion Mittermaier, Helen Roberts, Steve Willington, Nigel Roberts, Mike Sowko, Gordon Strassberg, Charles Kluepfel, Timothy A. Bullock, David D. Turner, Florian Pappenberger, Neal Osborne, and Chris Noble. Ch...
2024
-
[49]
Romain Pic, Clément Dombry, Philippe Naveau, and Maxime Taillardat. Proper scoring rules for multivariate probabilistic forecasts based on aggregation and transformation. Advances in Statistical Climatology, Meteorology and Oceanography, 11 0 (1): 0 23–58, March 2025. ISSN 2364-3587. doi:10.5194/ascmo-11-23-2025. URL http://dx.doi.org/10.5194/ascmo-11-23-2025
-
[50]
Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. Probabilistic weather forecasting with machine learning. Nature, 637 0 (8044): 0 84–90, December 2024. ISSN 1476-4687. doi:10.1038/s41586-024-08252-9. URL http://d...
-
[51]
Radford, Imme Ebert-Uphoff, and Jebb Q
Jacob T. Radford, Imme Ebert-Uphoff, and Jebb Q. Stewart. A comparison of ai weather prediction and numerical weather prediction models for 1–7-day precipitation forecasts. Weather and Forecasting, March 2025 a . ISSN 1520-0434. doi:10.1175/waf-d-24-0081.1. URL http://dx.doi.org/10.1175/waf-d-24-0081.1
-
[52]
Radford, Imme Ebert-Uphoff, Jebb Q
Jacob T. Radford, Imme Ebert-Uphoff, Jebb Q. Stewart, Kate D. Musgrave, Robert DeMaria, Natalie Tourville, and Kyle Hilburn. Accelerating community-wide evaluation of ai models for global weather prediction by facilitating access to model output. Bulletin of the American Meteorological Society, 106 0 (1): 0 E68–E76, January 2025 b . ISSN 1520-0477. doi:10...
-
[53]
Weatherbench 2: A benchmark for the next generation of data‐driven global weather models
Stephan Rasp, Stephan Hoyer, Alexander Merose, Ian Langmore, Peter Battaglia, Tyler Russell, Alvaro Sanchez‐Gonzalez, Vivian Yang, Rob Carver, Shreya Agrawal, Matthew Chantry, Zied Ben Bouallegue, Peter Dueben, Carla Bromberg, Jared Sisk, Luke Barrington, Aaron Bell, and Fei Sha. Weatherbench 2: A benchmark for the next generation of data‐driven global we...
-
[54]
Mark J. Rodwell, David S. Richardson, Tim D. Hewson, and Thomas Haiden. A new equitable score suitable for verifying precipitation in numerical weather prediction. Quarterly Journal of the Royal Meteorological Society, 136 0 (650): 0 1344–1363, July 2010. ISSN 1477-870X. doi:10.1002/qj.656. URL http://dx.doi.org/10.1002/qj.656
-
[55]
Craig S. Schwartz and Ryan A. Sobash. Generating probabilistic forecasts from convection-allowing ensembles using neighborhood approaches: A review and recommendations. Monthly Weather Review, 145 0 (9): 0 3397–3418, September 2017. ISSN 1520-0493. doi:10.1175/mwr-d-16-0400.1. URL http://dx.doi.org/10.1175/mwr-d-16-0400.1
-
[56]
Neighborhood-based ensemble evaluation using the crps
Joël Stein and Fabien Stoop. Neighborhood-based ensemble evaluation using the crps. Monthly Weather Review, 150 0 (8): 0 1901–1914, August 2022. ISSN 1520-0493. doi:10.1175/mwr-d-21-0224.1. URL http://dx.doi.org/10.1175/mwr-d-21-0224.1
-
[57]
Christopher Subich, Syed Zahid Husain, Leo Separovic, and Jing Yang. Fixing the double penalty in data-driven weather forecasting through a modified spherical harmonic loss function, 2025. URL https://arxiv.org/abs/2501.19374
Pith/arXiv arXiv 2025
-
[58]
Gábor J. Székely and Maria L. Rizzo. A new test for multivariate normality. Journal of Multivariate Analysis, 93 0 (1): 0 58–80, March 2005. ISSN 0047-259X. doi:10.1016/j.jmva.2003.12.002. URL http://dx.doi.org/10.1016/j.jmva.2003.12.002
-
[59]
Evaluation of point forecasts for extreme events using consistent scoring functions
Robert Taggart. Evaluation of point forecasts for extreme events using consistent scoring functions. Quarterly Journal of the Royal Meteorological Society, 148 0 (742): 0 306–320, November 2021. ISSN 1477-870X. doi:10.1002/qj.4206. URL http://dx.doi.org/10.1002/qj.4206
-
[60]
Scale issues in verification of precipitation forecasts
Ben Tustison, Daniel Harris, and Efi Foufoula‐Georgiou. Scale issues in verification of precipitation forecasts. Journal of Geophysical Research: Atmospheres, 106 0 (D11): 0 11775–11784, June 2001. ISSN 0148-0227. doi:10.1029/2001jd900066. URL http://dx.doi.org/10.1029/2001jd900066
-
[61]
Eva-Maria Walz, Alexander Henzi, Johanna Ziegel, and Tilmann Gneiting. Easy uncertainty quantification (easyuq): Generating predictive distributions from single-valued model output. SIAM Review, 66 0 (1): 0 91–122, February 2024. ISSN 1095-7200. doi:10.1137/22m1541915. URL http://dx.doi.org/10.1137/22m1541915
-
[62]
Jakob Benjamin Wessel, Christopher A. T. Ferro, Gavin R. Evans, and Frank Kwasniok. Improving probabilistic forecasts of extreme wind speeds by training statistical post-processing models with weighted scoring rules. Monthly Weather Review, April 2025. ISSN 1520-0493. doi:10.1175/mwr-d-24-0151.1. URL http://dx.doi.org/10.1175/mwr-d-24-0151.1
-
[63]
Robert L. Winkler and Allan H. Murphy. “good” probability assessors. Journal of Applied Meteorology, 7 0 (5): 0 751–758, October 1968. ISSN 0021-8952. doi:10.1175/1520-0450(1968)007<0751:pa>2.0.co;2. URL http://dx.doi.org/10.1175/1520-0450(1968)007<0751:pa>2.0.co;2
-
[64]
Guide to hydrological practices
WMO. Guide to hydrological practices. Technical Report 168, World Meteorological Organization, 1994
1994
-
[65]
Numerical models outperform ai weather forecasts of record-breaking extremes, 2025
Zhongwei Zhang, Erich Fischer, Jakob Zscheischler, and Sebastian Engelke. Numerical models outperform ai weather forecasts of record-breaking extremes, 2025. URL https://arxiv.org/abs/2508.15724
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.