REVIEW 3 major objections 5 minor 3 cited by
HRRRCast: a data-driven emulator for regional weather forecasting at convection allowing scales
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A 23.5-million-parameter diffusion-based neural network, trained on HRRR analyses over the full CONUS, beats the operational HRRR forecast on 20 dBZ composite reflectivity at every lead time out to 48 hours and matches it at 30 dBZ.
desk verdict A credible engineering result—diffusion emulator beats HRRR on light-rain placement over CONUS—but the headline score is measured against the model's own training target, HRRR analysis, so the real skill is probably a bit weaker than advertised. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is ResHRRR: a U-Net-style residual CNN with squeeze-and-excitation channel attention and FiLM time-conditioning, comprising roughly 23.5 million parameters, trained as a Denoising Diffusion Implicit Model (DDIM), an accelerated deterministic sampling variant of diffusion. Its central mechanism is multi-lead training: a single model predicts 1-, 3-, and 6-hour targets from the same HRRR analysis input, is explicitly conditioned on lead time through FiLM, and is rolled out greedily to reach 48 hours, which limits compounding error without altering the diffusion schedule. A second mechanism, GFS-conditioned downscaling, feeds future synoptic-scale states as extra input channels so the model can blend forecasting with downscaling at longer leads.
What would settle it
Re-run the full 4-month evaluation with an independent radar-based composite reflectivity product, rather than HRRR analysis, as ground truth at the 20 and 30 dBZ thresholds; the claimed advantage over HRRR would be falsified if the FSS differences vanish or reverse on that benchmark.
Extended reading notes
Core claim
On the paper's own terms, HRRRCast demonstrates that a generative diffusion emulator can outperform a state-of-the-art operational convection-allowing model at light-to-moderate precipitation. Using ensembles of 3 to 10 members at 6 km resolution over the full CONUS, ResHRRR achieves higher Fractions Skill Scores than HRRR at the 20 dBZ composite-reflectivity threshold at every lead time from 7 to 48 hours, and outperforms HRRR up to 7 hours at 30 dBZ. Grid-based metrics show lower frequency bias and higher success ratios, object-based verification agrees at 20 dBZ, and power spectra of reflectivity match HRRR analysis more closely than HRRR forecast does. The authors attribute the gains to training on analysis rather than forecast fields, multi-lead training with a greedy rollout, and GFS-conditioned downscaling.
Load-bearing premise
The evaluation uses HRRR analysis as the ground truth for reflectivity, and HRRR analysis is also the training target, so the measured edge over the HRRR forecast could partly reflect the emulator reproducing its own training data rather than matching independent observations.
Editorial extensions
If this is right
- A 3-member HRRRCast ensemble already beats the operational HRRR forecast on 20 dBZ reflectivity up to 48 hours, so ensemble size can be traded against computational budget without losing the headline skill.
- Multi-lead training plus greedy rollout extends diffusion-model forecast range beyond the 1-hour autoregressive horizon used in prior work, without modifying the diffusion process.
- GFS conditioning gives the model a downscaling capability: skill beyond 18 hours is sustained when the global input is an analysis, implying improved global forecasts would directly improve regional emulator skill.
- Because HRRR over-predicts reflectivity while HRRRCast under-predicts it, neither system is unbiased; the emulator's better frequency bias suggests it can complement, not just replace, physics-based guidance.
Reading between the lines
- The paper's advantage is measured against the same analysis fields used for training; a head-to-head verification against independent observed radar would quantify how much of the edge is real storm skill.
- At 40 dBZ no model reaches usable skill, so the practical consequence may be a division of labor: cheap emulator ensembles for light-to-moderate threats, physics-based models reserved for severe-thunderstorm thresholds.
- The strong performance of future GFS states as conditioning hints that coupling a data-driven global model to a regional emulator could push useful convective guidance beyond 48 hours, an extension the paper does not test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HRRRCast, a 6-km CONUS data-driven emulator of the operational HRRR model, with two architectures: ResHRRR (SE-ResNet + FiLM + DDIM diffusion) and GraphHRRR (graph-based). ResHRRR is trained on three years of HRRR analysis with GFS synoptic forcing, uses multi-lead-time training (1, 3, and 6 h), and produces probabilistic ensembles of 3-10 members. The central claim is that ResHRRR outperforms operational HRRR forecasts on composite reflectivity at the 20 dBZ threshold out to 48 h and is competitive at 30 dBZ, supported by FSS, contingency-table, object-based, and RMSE metrics, while GraphHRRR underperforms. The evaluation uses HRRR analysis as ground truth, with a single MRMS-based grid check.
Significance. If the headline skill holds against independent observations, the paper would demonstrate that a 23.5M-parameter diffusion emulator can provide fast, full-CONUS, convection-allowing ensemble forecasts with plausible operational value. Strengths of the study include full-CONUS training, multi-lead diffusion conditioning, ensemble spread diagnostics in Appendix A, power-spectrum sharpness analysis, and a preliminary MRMS check, which go beyond a single skill-score comparison. However, the primary verification target is also the training target, so the significance claim currently rests on a partly circular evaluation and needs an independent observational benchmark to be fully persuasive.
major comments (3)
- [§3, Figs. 4 and 7] The headline comparison is evaluated against the same HRRR analysis used as the training target. Section 3 states that "The HRRR analysis data, which served as the training target, is used as the ground truth for evaluation." Because ResHRRR is trained to reproduce the analysis distribution, the FSS advantage at 20 dBZ in Figure 4 can reflect the model's fidelity to the analysis's own biases (including its null patterns) rather than skill against observed reflectivity. Figure 7b, the only MRMS-based check, covers grid-based bias metrics and does not extend to the 4-month FSS curves that support the central claim. The statement that MRMS "would have been more appropriate" acknowledges the limitation but does not quantify its effect. Please repeat the 4-month FSS and object-based evaluation using MRMS reflectivity as ground truth, or explicitly rephrase the headline as agreement with HRRR analysis rather than verified skill.
- [§2.4.2 and §3.2] The claimed benefit of multi-lead-time training is not isolated by an ablation. Section 3.2 notes that "we have not yet conducted a formal ablation study to confirm this" in the context of GFS inputs, and no experiment compares ResHRRR trained on 1 h leads only against the multi-lead version. Since multi-lead training is presented as a key advancement over StormCast and as the mechanism for "reducing cumulative error," the long-lead FSS results in Figure 4 cannot be attributed to this design choice without such a baseline. Add a controlled 1 h-only training run with identical architecture and data, or soften the attribution accordingly.
- [§3.2, Figs. 4 and 5] No uncertainty quantification is provided for the verification metrics. The FSS curves and CSI values are point estimates computed over 488 (or 240) timestamps, yet the text claims "outperforms HRRR at all lead times up to 48 hours" without confidence intervals or a significance test. Lead-time-wise conclusions at threshold 20 dBZ should be accompanied by bootstrap intervals or a paired significance test; otherwise small FSS differences in Figure 4 may not be robust. This is particularly important because the claim of superiority is a central result.
minor comments (5)
- [§2.3.2] The description says GraphHRRR replaces boundary conditions with Dirichlet conditions initialized using a global model, but three paragraphs later it is said to "not ingest synoptic-scale inputs." Please reconcile these statements, since boundary forcing from a global model is a form of synoptic-scale input.
- [Abstract and §2.1] The assertion that StormCast "inadvertently used" +1 h post-analysis data is presented without a reference, appendix, or quantitative demonstration; if this is based on inspection of the StormCast data pipeline, state the evidence or soften the claim.
- [Overall] The paper does not state whether code, trained weights, or data-processing scripts will be released; a reproducibility statement would strengthen the methods contribution.
- [Figure 4 and Appendix C captions] Several captions contain typographical errors, including "upto" and "modesl" (for "models"); a careful proofreading pass is needed.
- [§2.4.1] The loss weights (1.0 for reflectivity, 0.1 for 2-m temperature) are presented without sensitivity analysis; since the paper later attributes reflectivity skill partly to this weight, a short sensitivity check or a discussion of how these weights were chosen would help.
Circularity Check
No significant circularity: the headline skill scores are empirical comparisons, and the analysis-as-ground-truth choice is a data-validity caveat the paper explicitly acknowledges.
full rationale
The paper's central claim is an empirical skill comparison between HRRRCast and the operational HRRR forecast, not a derivation that reduces to its inputs. The training objective (Section 2.4.1) minimizes MSE against HRRR analysis, and HRRR analysis is also used as evaluation ground truth (Section 3: "The HRRR analysis data, which served as the training target, is used as the ground truth for evaluation. While the Multi-Radar Multi-Sensor (MRMS) dataset provides a more accurate source of composite reflectivity... Using MRMS as ground truth would have been more appropriate if the model had been trained directly on MRMS"). This is a legitimate data-validity caveat: a model trained on analysis fields can score well against those same fields for partly distributional reasons, and the 20-dBZ FSS advantage over HRRR may be inflated by learned analysis-specific bias. However, this does not make the comparison circular by construction: ResHRRR is not guaranteed to beat the HRRR forecast on analysis ground truth, and the paper supplements with an MRMS-based grid check (Figure 7b), object-based verification, and power-spectrum comparisons. No fitted parameter is renamed as a prediction, no load-bearing uniqueness theorem or ansatz is imported via self-citation, and no known result is merely renamed. The self-citations (e.g., Flora and Potvin [8], Smith et al. [28]) provide architectural context and prior results rather than the load-bearing justification for the headline claim. The derivation is therefore self-contained; the evaluation-target issue belongs to data validity, not circularity.
Assumptions & free parameters
free parameters (6)
- Composite reflectivity loss weight =
1.0
- 2-meter temperature loss weight =
0.1
- Multi-lead time set =
[1, 3, 6] h
- Diffusion steps / inference steps =
T=200; 30-50 DDIM steps
- Evaluation thresholds and pooling =
20/30/40 dBZ; 6 km FSS window
- Grid subsampling factor =
2 (6 km from 3 km)
assumptions (4)
- domain assumption HRRR analysis is a sufficiently accurate proxy for observed weather, including reflectivity
- domain assumption A 6 km, 12-pressure-level representation preserves enough convective structure for the skill claims
- domain assumption GFS forecast fields are adequate synoptic forcing for emulating HRRR
- domain assumption The diffusion model (DDIM) can represent the conditional distribution of future atmospheric states
Cite this review
Pith. "Pith review of HRRRCast: a data-driven emulator for regional weather forecasting at convection allowing scales." pith.science (2026). https://pith.science/paper/4RIHZKC5
@misc{pith2026250705658,
author = {Pith},
title = {Pith review of: HRRRCast: a data-driven emulator for regional weather forecasting at convection allowing scales},
year = {2026},
howpublished = {\url{https://pith.science/paper/4RIHZKC5}},
note = {Machine review of arXiv:2507.05658}
}
read the original abstract
The High-Resolution Rapid Refresh (HRRR) model is a convection-allowing model used in operational weather forecasting across the contiguous United States (CONUS). To provide a computationally efficient alternative, we introduce HRRRCast, a data-driven emulator built with advanced machine learning techniques. HRRRCast includes two architectures: a ResNet-based model (ResHRRR) and a Graph Neural Network-based model (GraphHRRR). ResHRRR uses convolutional neural networks enhanced with squeeze-and-excitation blocks and Feature-wise Linear Modulation, and supports probabilistic forecasting via the Denoising Diffusion Implicit Model (DDIM). To better handle longer lead times, we train a single model to predict multiple lead times (1h, 3h, and 6h), then use a greedy rollout strategy during inference. When evaluated on composite reflectivity over the full CONUS domain using ensembles of 3 to 10 members, ResHRRR outperforms HRRR forecast at light rainfall threshold (20 dBZ) and achieves competitive performance at moderate thresholds (30 dBZ). Our work advances the StormCast model of Pathak et al. [21] by: a) training on the full CONUS domain, b) using multiple lead times to improve long-range skill, c) training on analysis data instead of the +1h post-analysis data inadvertently used in StormCast, and d) incorporating future GFS states as inputs, enabling downscaling that improves long-lead accuracy. Grid-, neighborhood-, and object-based metrics confirm better storm placement, lower frequency bias, and higher success ratios than HRRR. HRRRCast ensemble forecasts also maintain sharper spatial detail, with power spectra more closely matching HRRR analysis. While GraphHRRR underperforms in its current form, it lays groundwork for future graph-based forecasting. HRRRCast represents a step toward efficient, data-driven regional weather prediction with competitive accuracy and ensemble capability.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 3 Pith papers
-
A Space-Time Transformer for Precipitation Nowcasting
A full space-time attention video transformer recast as 64-class rainfall prediction with log-frequency class weighting won the Weather4Cast 2025 Cumulative Rainfall challenge (CRPS 3.135).
-
Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model
Combining HiRA neighborhood verification with threshold-weighted CRPS shows that AI-vs-NWP rankings for extreme precipitation depend strongly on neighborhood size.
-
CRPS-LAM: Probabilistic Regional Weather Forecasting with Continuous Ranked Probability Score
CRPS-LAM produces 57-hour probabilistic limited-area forecasts on MEPS at diffusion-comparable accuracy with single-forward-pass sampling, roughly 39x faster than Diffusion-LAM.
Reference graph
Works this paper leans on
-
[1]
Adamov, S., Oskarsson, J., Denby, L., Landelius, T., Hintz, K., Christiansen, S., Schicker, I., Osuna, C., Lindsten, F., Fuhrer, O., Schemm, S., 2025. Building machine learning limited area models: Kilometer- 20 scale weather forecasting in realistic settings. URL https://arxiv.org/abs/2504.09340
arXiv 2025
-
[2]
Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., Tian, Q., Jul. 2023. Accurate medium-range global weather forecasting with 3d neural networks. Nature 619 (7970), 533–538. URL https://doi.org/10.1038/s41586-023-06185-3
-
[3]
R., Aittala, M., Kreis, K., Brenowitz, N., Vahdat, A., Mardani, M., Yu, R., 2025
Cachay, S. R., Aittala, M., Kreis, K., Brenowitz, N., Vahdat, A., Mardani, M., Yu, R., 2025. Elucidated rolling diffusion models for probabilistic weather forecasting. URL https://arxiv.org/abs/2506.20024
arXiv 2025
-
[4]
Chase, R. J., Haynes, K., Hoef, L. V., Ebert-Uphoff, I., 2025. Score-based diffusion nowcasting of goes imagery. URL https://arxiv.org/abs/2505.10432
- [5]
-
[6]
Dowell, D. C., Alexander, C. R., James, E. P., Weygandt, S. S., Benjamin, S. G., Manikin, G. S., Blake, B. T., Brown, J. M., Olson, J. B., Hu, M., Smirnova, T. G., Ladwig, T., Kenyon, J. S., Ahmadov, R., Turner, D. D., Duda, J. D., Alcott, T. I., 2022. The high-resolution rapid refresh (hrrr): An hourly updating convection-allowing forecast model. part i:...
work page 2022
-
[7]
(DTC), D. T. C., 2024. METplus User’s Guide: Gen-Ens-Prod Tool. National Center for Atmospheric Research (NCAR), version 5.1.1
work page 2024
-
[8]
Flora, M. L., Potvin, C., 2025. Wofscast: A machine learning model for predicting thunderstorms at watch-to-warning scales. Geophysical Research Letters 52 (10), e2024GL112383, e2024GL112383 2024GL112383. URL https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/2024GL112383
Show all 34 references
-
[9]
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J., 2015. Deep residual learning for image recognition
2015
-
[10]
B., Smirnova, T., Alexander, C., Berner, J., 2019
Jankov, I., Beck, J., Wolff, J., Harrold, M., Olson, J. B., Smirnova, T., Alexander, C., Berner, J., 2019. Stochastically perturbed parameterizations in an hrrr-based ensemble. Monthly Weather Review 147 (1), 153 – 173. URL https://journals.ametsoc.org/view/journals/mwre/147/1...
2019
-
[11]
Elucidating the design space of diffusion-based gener- ative models
Karras, T., Aittala, M., Aila, T., Laine, S., 2022. Elucidating the design space of diffusion-based gener- ative models. URL https://arxiv.org/abs/2206.00364
2022 arXiv
-
[12]
Forecasting global weather with graph neural networks
Keisler, R., 2022. Forecasting global weather with graph neural networks. URL https://arxiv.org/abs/2202.07575
2022 arXiv
-
[13]
Learning skillful medium-range global weather forecasting
Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M., Alet, F., Ravuri, S., Ewalds, T., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer, S., Holland, G., Vinyals, O., Stott, J., Pritzel, A., Mohamed, S., Battaglia, P., 2023. Learning skillful medium-range glob...
2023 doi
-
[14]
Lang, S., Alexe, M., Clare, M. C. A., Roberts, C., Adewoyin, R., Bouall?gue, Z. B., Chantry, M., Dramsch, J., Dueben, P. D., Hahner, S., Maciel, P., Prieto-Nemesio, A., O’Brien, C., Pinault, F., Polster, J., Raoult, B., Tietsche, S., Leutbecher, M., 2024. Aifs-crps: Ensemble f...
2024 arXiv
-
[15]
Diffusion-lam: Probabilistic limited area weather forecasting with diffusion
Larsson, E., Oskarsson, J., Landelius, T., Lindsten, F., 2025. Diffusion-lam: Probabilistic limited area weather forecasting with diffusion. URL https://arxiv.org/abs/2502.07532
2025 arXiv
-
[16]
May 7, 2024 severe weather and tornadoes
National Weather Service, May 2024. May 7, 2024 severe weather and tornadoes. Accessed: 2025-05-29. URL https://www.weather.gov/grr/7 May 2024 SevereWeather
2024
-
[17]
K., Grover, A., 2023
Nguyen, T., Brandstetter, J., Kapoor, A., Gupta, J. K., Grover, A., 2023. Climax: A foundation model for weather and climate. In: International Conference on Machine Learning. URL https://api.semanticscholar.org/CorpusID:256231457
2023
-
[18]
Scaling transformer neural networks for skillful and reliable medium-range weather forecasting
Nguyen, T., Shah, R., Bansal, H., Arcomano, T., Madireddy, S., Maulik, R., Kotamarthi, V., Foster, I., Grover, A., 2023. Scaling transformer neural networks for skillful and reliable medium-range weather forecasting
2023
-
[19]
N., Haugen, H
Nipen, T. N., Haugen, H. H., Ingstad, M. S., Nordhagen, E. M., Salihi, A. F. S., Tedesco, P., Seierstad, I. A., Kristiansen, J., Lang, S., Alexe, M., Dramsch, J., Raoult, B., Mertes, G., Chantry, M., 2024. Regional data-driven weather modeling with a global stretched-grid. URL...
2024 arXiv
-
[20]
Graph-based neural weather prediction for limited area modeling
Oskarsson, J., Landelius, T., Lindsten, F., 2023. Graph-based neural weather prediction for limited area modeling. URL https://arxiv.org/abs/2309.17370
2023 arXiv
-
[21]
Kilometer-scale convection allowing model emulation using generative diffusion modeling
Pathak, J., Cohen, Y., Garg, P., Harrington, P., Brenowitz, N., Durran, D., Mardani, M., Vahdat, A., Xu, S., Kashinath, K., Pritchard, M., 2024. Kilometer-scale convection allowing model emulation using generative diffusion modeling. URL https://arxiv.org/abs/2408.10958
2024 arXiv
-
[22]
K., Carley, J
Potvin, C. K., Carley, J. R., Clark, A. J., Wicker, L. J., Skinner, P. S., Reinhart, A. E., Gallo, B. T., Kain, J. S., Romine, G. S., Aligo, E. A., Brewster, K. A., Dowell, D. C., Harris, L. M., Jirak, I. L., Kong, F., Supinie, T. A., Thomas, K. W., Wang, X., Wang, Y., Xue, M....
2019
-
[23]
R., El-Kadi, A., Masters, D., Ewalds, T., Stott, J., Mohamed, S., Battaglia, P., Lam, R., Willson, M., 2024
Price, I., Sanchez-Gonzalez, A., Alet, F., Andersson, T. R., El-Kadi, A., Masters, D., Ewalds, T., Stott, J., Mohamed, S., Battaglia, P., Lam, R., Willson, M., 2024. Gencast: Diffusion-based ensemble forecasting for medium-range weather
2024
-
[24]
Data-driven medium-range weather prediction with a resnet pretrained on climate simulations: A new model for weatherbench
Rasp, S., Thuerey, N., 2021. Data-driven medium-range weather prediction with a resnet pretrained on climate simulations: A new model for weatherbench. Journal of Advances in Modeling Earth Systems 13 (2), e2020MS002405, e2020MS002405 2020MS002405. URL https://agupubs.onlineli...
2021 doi
-
[25]
Rolling diffusion models
Ruhe, D., Heek, J., Salimans, T., Hoogeboom, E., 2024. Rolling diffusion models. URL https://arxiv.org/abs/2402.09470
2024 arXiv
-
[26]
A., Kossaifi, J., Bonev, B., Choy, C., Kautz, J., Krueger, D., Azizzadenesheli, K., 2024
Siddiqui, S. A., Kossaifi, J., Bonev, B., Choy, C., Kautz, J., Krueger, D., Azizzadenesheli, K., 2024. Exploring the design space of deep-learning-based weather forecasting systems. URL https://arxiv.org/abs/2410.07472
2024 arXiv
-
[27]
Skinner, P., Stratman, D., Kerr, C., Matilla, B., Martin, J., Dowell, D., Jones, T., Flora, M., Guerra, J., Knopfmeier, K., Britt, K., Yussouf, N., ???? Comparing short-term thunderstorm forecasts from the warn-on-forecast system (wofs) and high-resolution rapid refresh (hrrr)
-
[28]
A., Penny, S
Smith, T. A., Penny, S. G., Platt, J. A., Chen, T.-C., 2023. Temporal subsampling diminishes small spatial scales in recurrent neural network emulators of geophysical turbulence. Journal of Advances in Modeling Earth Systems 15 (12), e2023MS003792, e2023MS003792 2023MS003792. ...
2023 doi
-
[29]
Denoising diffusion implicit models
Song, J., Meng, C., Ermon, S., 2022. Denoising diffusion implicit models. URL https://arxiv.org/abs/2010.02502
2022 arXiv
-
[30]
Z., Separovic, L., Yang, J., 2025
Subich, C., Husain, S. Z., Separovic, L., Yang, J., 2025. Fixing the double penalty in data-driven weather forecasting through a modified spherical harmonic loss function. URL https://arxiv.org/abs/2501.19374
2025 arXiv
-
[31]
R., Landolt, S
Xu, M., Thompson, G., Adriaansen, D. R., Landolt, S. D., 2019. On the value of time-lag-ensemble averaging to improve numerical model predictions of aircraft icing conditions. Weather and Forecasting 34 (3), 507 – 519. URL https://journals.ametsoc.org/view/journals/wefo/34/3/w...
2019
-
[32]
This provides the spatial structure that will be preserved in the final PMM field
Compute the ensemble mean field by averaging the reflectivity across all ensemble members at each grid point. This provides the spatial structure that will be preserved in the final PMM field
-
[33]
From this sorted array, every N th value is selected – where N is the number of ensemble members – to produce a downsampled set of values
Pool and sort all reflectivity values from all ensemble members into a single 1D array. From this sorted array, every N th value is selected – where N is the number of ensemble members – to produce a downsampled set of values. This set approximates the typical value distributi...
-
[34]
This ensures the final field has the same spatial pattern as the ensemble mean, but with a value distribution that better reflects the ensemble variability
Reassign the sorted values from step 2 to the grid points of the mean field based on rank ordering: the largest value is assigned to the location with the highest mean, the second-largest to the second-highest, and so on. This ensures the final field has the same spatial patte...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.