Pith. sign in

REVIEW 3 major objections 5 minor 3 cited by

HRRRCast: a data-driven emulator for regional weather forecasting at convection allowing scales

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A 23.5-million-parameter diffusion-based neural network, trained on HRRR analyses over the full CONUS, beats the operational HRRR forecast on 20 dBZ composite reflectivity at every lead time out to 48 hours and matches it at 30 dBZ.

desk verdict A credible engineering result—diffusion emulator beats HRRR on light-rain placement over CONUS—but the headline score is measured against the model's own training target, HRRR analysis, so the real skill is probably a bit weaker than advertised. read the letter →

arxiv 2507.05658 v1 pith:4RIHZKC5 submitted 2025-07-08 physics.ao-ph cs.LG

classification physics.ao-phcs.LG
keywords HRRRmachinelearningweatherpredictionconvection-allowingmodeldata-drivenregionalmodellingAI4NWPdiffusionensembleforecastingcompositereflectivity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

HRRRCast sets out to show that a data-driven emulator can stand in for a high-resolution operational weather model at convection-allowing scales. Its main result is that ResHRRR, a 23.5-million-parameter diffusion-based ResNet trained on three years of HRRR analyses, produces ensemble forecasts of composite reflectivity that beat the operational HRRR forecast at the light-rain threshold (20 dBZ) across the full CONUS at all evaluated lead times up to 48 hours, and stay competitive at 30 dBZ. If the result holds, convective-scale ensemble guidance can be generated at a fraction of the cost of running a physics-based model, with better frequency bias and storm placement at light-to-moderate intensities. The paper also claims that training one model on 1-, 3-, and 6-hour leads, then rolling it out greedily, extends skill without changing the diffusion process.

What carries the argument

The load-bearing object is ResHRRR: a U-Net-style residual CNN with squeeze-and-excitation channel attention and FiLM time-conditioning, comprising roughly 23.5 million parameters, trained as a Denoising Diffusion Implicit Model (DDIM), an accelerated deterministic sampling variant of diffusion. Its central mechanism is multi-lead training: a single model predicts 1-, 3-, and 6-hour targets from the same HRRR analysis input, is explicitly conditioned on lead time through FiLM, and is rolled out greedily to reach 48 hours, which limits compounding error without altering the diffusion schedule. A second mechanism, GFS-conditioned downscaling, feeds future synoptic-scale states as extra input channels so the model can blend forecasting with downscaling at longer leads.

What would settle it

Re-run the full 4-month evaluation with an independent radar-based composite reflectivity product, rather than HRRR analysis, as ground truth at the 20 and 30 dBZ thresholds; the claimed advantage over HRRR would be falsified if the FSS differences vanish or reverse on that benchmark.

Watch

Extended reading notes

Core claim

On the paper's own terms, HRRRCast demonstrates that a generative diffusion emulator can outperform a state-of-the-art operational convection-allowing model at light-to-moderate precipitation. Using ensembles of 3 to 10 members at 6 km resolution over the full CONUS, ResHRRR achieves higher Fractions Skill Scores than HRRR at the 20 dBZ composite-reflectivity threshold at every lead time from 7 to 48 hours, and outperforms HRRR up to 7 hours at 30 dBZ. Grid-based metrics show lower frequency bias and higher success ratios, object-based verification agrees at 20 dBZ, and power spectra of reflectivity match HRRR analysis more closely than HRRR forecast does. The authors attribute the gains to training on analysis rather than forecast fields, multi-lead training with a greedy rollout, and GFS-conditioned downscaling.

Load-bearing premise

The evaluation uses HRRR analysis as the ground truth for reflectivity, and HRRR analysis is also the training target, so the measured edge over the HRRR forecast could partly reflect the emulator reproducing its own training data rather than matching independent observations.

Editorial extensions

If this is right

  • A 3-member HRRRCast ensemble already beats the operational HRRR forecast on 20 dBZ reflectivity up to 48 hours, so ensemble size can be traded against computational budget without losing the headline skill.
  • Multi-lead training plus greedy rollout extends diffusion-model forecast range beyond the 1-hour autoregressive horizon used in prior work, without modifying the diffusion process.
  • GFS conditioning gives the model a downscaling capability: skill beyond 18 hours is sustained when the global input is an analysis, implying improved global forecasts would directly improve regional emulator skill.
  • Because HRRR over-predicts reflectivity while HRRRCast under-predicts it, neither system is unbiased; the emulator's better frequency bias suggests it can complement, not just replace, physics-based guidance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's advantage is measured against the same analysis fields used for training; a head-to-head verification against independent observed radar would quantify how much of the edge is real storm skill.
  • At 40 dBZ no model reaches usable skill, so the practical consequence may be a division of labor: cheap emulator ensembles for light-to-moderate threats, physics-based models reserved for severe-thunderstorm thresholds.
  • The strong performance of future GFS states as conditioning hints that coupling a data-driven global model to a regional emulator could push useful convective guidance beyond 48 hours, an extension the paper does not test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces HRRRCast, a 6-km CONUS data-driven emulator of the operational HRRR model, with two architectures: ResHRRR (SE-ResNet + FiLM + DDIM diffusion) and GraphHRRR (graph-based). ResHRRR is trained on three years of HRRR analysis with GFS synoptic forcing, uses multi-lead-time training (1, 3, and 6 h), and produces probabilistic ensembles of 3-10 members. The central claim is that ResHRRR outperforms operational HRRR forecasts on composite reflectivity at the 20 dBZ threshold out to 48 h and is competitive at 30 dBZ, supported by FSS, contingency-table, object-based, and RMSE metrics, while GraphHRRR underperforms. The evaluation uses HRRR analysis as ground truth, with a single MRMS-based grid check.

Significance. If the headline skill holds against independent observations, the paper would demonstrate that a 23.5M-parameter diffusion emulator can provide fast, full-CONUS, convection-allowing ensemble forecasts with plausible operational value. Strengths of the study include full-CONUS training, multi-lead diffusion conditioning, ensemble spread diagnostics in Appendix A, power-spectrum sharpness analysis, and a preliminary MRMS check, which go beyond a single skill-score comparison. However, the primary verification target is also the training target, so the significance claim currently rests on a partly circular evaluation and needs an independent observational benchmark to be fully persuasive.

major comments (3)
  1. [§3, Figs. 4 and 7] The headline comparison is evaluated against the same HRRR analysis used as the training target. Section 3 states that "The HRRR analysis data, which served as the training target, is used as the ground truth for evaluation." Because ResHRRR is trained to reproduce the analysis distribution, the FSS advantage at 20 dBZ in Figure 4 can reflect the model's fidelity to the analysis's own biases (including its null patterns) rather than skill against observed reflectivity. Figure 7b, the only MRMS-based check, covers grid-based bias metrics and does not extend to the 4-month FSS curves that support the central claim. The statement that MRMS "would have been more appropriate" acknowledges the limitation but does not quantify its effect. Please repeat the 4-month FSS and object-based evaluation using MRMS reflectivity as ground truth, or explicitly rephrase the headline as agreement with HRRR analysis rather than verified skill.
  2. [§2.4.2 and §3.2] The claimed benefit of multi-lead-time training is not isolated by an ablation. Section 3.2 notes that "we have not yet conducted a formal ablation study to confirm this" in the context of GFS inputs, and no experiment compares ResHRRR trained on 1 h leads only against the multi-lead version. Since multi-lead training is presented as a key advancement over StormCast and as the mechanism for "reducing cumulative error," the long-lead FSS results in Figure 4 cannot be attributed to this design choice without such a baseline. Add a controlled 1 h-only training run with identical architecture and data, or soften the attribution accordingly.
  3. [§3.2, Figs. 4 and 5] No uncertainty quantification is provided for the verification metrics. The FSS curves and CSI values are point estimates computed over 488 (or 240) timestamps, yet the text claims "outperforms HRRR at all lead times up to 48 hours" without confidence intervals or a significance test. Lead-time-wise conclusions at threshold 20 dBZ should be accompanied by bootstrap intervals or a paired significance test; otherwise small FSS differences in Figure 4 may not be robust. This is particularly important because the claim of superiority is a central result.
minor comments (5)
  1. [§2.3.2] The description says GraphHRRR replaces boundary conditions with Dirichlet conditions initialized using a global model, but three paragraphs later it is said to "not ingest synoptic-scale inputs." Please reconcile these statements, since boundary forcing from a global model is a form of synoptic-scale input.
  2. [Abstract and §2.1] The assertion that StormCast "inadvertently used" +1 h post-analysis data is presented without a reference, appendix, or quantitative demonstration; if this is based on inspection of the StormCast data pipeline, state the evidence or soften the claim.
  3. [Overall] The paper does not state whether code, trained weights, or data-processing scripts will be released; a reproducibility statement would strengthen the methods contribution.
  4. [Figure 4 and Appendix C captions] Several captions contain typographical errors, including "upto" and "modesl" (for "models"); a careful proofreading pass is needed.
  5. [§2.4.1] The loss weights (1.0 for reflectivity, 0.1 for 2-m temperature) are presented without sensitivity analysis; since the paper later attributes reflectivity skill partly to this weight, a short sensitivity check or a discussion of how these weights were chosen would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline skill scores are empirical comparisons, and the analysis-as-ground-truth choice is a data-validity caveat the paper explicitly acknowledges.

full rationale

The paper's central claim is an empirical skill comparison between HRRRCast and the operational HRRR forecast, not a derivation that reduces to its inputs. The training objective (Section 2.4.1) minimizes MSE against HRRR analysis, and HRRR analysis is also used as evaluation ground truth (Section 3: "The HRRR analysis data, which served as the training target, is used as the ground truth for evaluation. While the Multi-Radar Multi-Sensor (MRMS) dataset provides a more accurate source of composite reflectivity... Using MRMS as ground truth would have been more appropriate if the model had been trained directly on MRMS"). This is a legitimate data-validity caveat: a model trained on analysis fields can score well against those same fields for partly distributional reasons, and the 20-dBZ FSS advantage over HRRR may be inflated by learned analysis-specific bias. However, this does not make the comparison circular by construction: ResHRRR is not guaranteed to beat the HRRR forecast on analysis ground truth, and the paper supplements with an MRMS-based grid check (Figure 7b), object-based verification, and power-spectrum comparisons. No fitted parameter is renamed as a prediction, no load-bearing uniqueness theorem or ansatz is imported via self-citation, and no known result is merely renamed. The self-citations (e.g., Flora and Potvin [8], Smith et al. [28]) provide architectural context and prior results rather than the load-bearing justification for the headline claim. The derivation is therefore self-contained; the evaluation-target issue belongs to data validity, not circularity.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

This is an empirical machine-learning paper, so the ledger's 'free parameters' are training and evaluation choices rather than physical constants. The central claim does not depend on any new physical entity or on a fitted physical constant; it depends on the validity of HRRR analysis as a target, the adequacy of the 6 km representation, and the GFS conditioning, all listed as axioms. Model weights themselves are omitted from the free-parameter list because they are the fitted output of training, not inputs to a derivation.

free parameters (6)
  • Composite reflectivity loss weight = 1.0
    Set by hand in the pressure-weighted loss (Sec 2.4.1); directly shapes how much the model prioritizes reflectivity and therefore the 20/30 dBZ skill claims.
  • 2-meter temperature loss weight = 0.1
    Hand-chosen (Sec 2.4.1); the paper attributes the higher T2M RMSE to this low weight.
  • Multi-lead time set = [1, 3, 6] h
    Chosen training design (Sec 2.4.2); the rollout strategy and long-lead skill depend on this set.
  • Diffusion steps / inference steps = T=200; 30-50 DDIM steps
    Chosen schedule and sampling settings (Sec 2.3.3); affect sample quality and ensemble diversity.
  • Evaluation thresholds and pooling = 20/30/40 dBZ; 6 km FSS window
    Threshold and neighborhood choices define what 'outperforms HRRR' means (Sec 3.2).
  • Grid subsampling factor = 2 (6 km from 3 km)
    Hardware-driven choice (Sec 2.2); the 6 km resolution is a key difference from HRRR and StormCast.
assumptions (4)
  • domain assumption HRRR analysis is a sufficiently accurate proxy for observed weather, including reflectivity
    Used as training target and as primary evaluation ground truth (Sec 3); the paper partially checks with MRMS.
  • domain assumption A 6 km, 12-pressure-level representation preserves enough convective structure for the skill claims
    The model uses subsampled HRRR grid and pressure levels (Sec 2.2, Table 1); aliasing of fine-scale features is acknowledged as a risk.
  • domain assumption GFS forecast fields are adequate synoptic forcing for emulating HRRR
    The model is trained and evaluated with GFS analysis/forecast as conditioning (Sec 2.1); no formal ablation is done for the GFS reflectivity channel (Sec 3.2).
  • domain assumption The diffusion model (DDIM) can represent the conditional distribution of future atmospheric states
    The framework relies on standard DDPM/DDIM assumptions (Sec 2.3.3); ensemble calibration is acknowledged as imperfect in Appendix A.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HRRRCast: a data-driven emulator for regional weather forecasting at convection allowing scales." pith.science (2026). https://pith.science/paper/4RIHZKC5

@misc{pith2026250705658,
  author       = {Pith},
  title        = {Pith review of: HRRRCast: a data-driven emulator for regional weather forecasting at convection allowing scales},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4RIHZKC5}},
  note         = {Machine review of arXiv:2507.05658}
}
read the original abstract

The High-Resolution Rapid Refresh (HRRR) model is a convection-allowing model used in operational weather forecasting across the contiguous United States (CONUS). To provide a computationally efficient alternative, we introduce HRRRCast, a data-driven emulator built with advanced machine learning techniques. HRRRCast includes two architectures: a ResNet-based model (ResHRRR) and a Graph Neural Network-based model (GraphHRRR). ResHRRR uses convolutional neural networks enhanced with squeeze-and-excitation blocks and Feature-wise Linear Modulation, and supports probabilistic forecasting via the Denoising Diffusion Implicit Model (DDIM). To better handle longer lead times, we train a single model to predict multiple lead times (1h, 3h, and 6h), then use a greedy rollout strategy during inference. When evaluated on composite reflectivity over the full CONUS domain using ensembles of 3 to 10 members, ResHRRR outperforms HRRR forecast at light rainfall threshold (20 dBZ) and achieves competitive performance at moderate thresholds (30 dBZ). Our work advances the StormCast model of Pathak et al. [21] by: a) training on the full CONUS domain, b) using multiple lead times to improve long-range skill, c) training on analysis data instead of the +1h post-analysis data inadvertently used in StormCast, and d) incorporating future GFS states as inputs, enabling downscaling that improves long-lead accuracy. Grid-, neighborhood-, and object-based metrics confirm better storm placement, lower frequency bias, and higher success ratios than HRRR. HRRRCast ensemble forecasts also maintain sharper spatial detail, with power spectra more closely matching HRRR analysis. While GraphHRRR underperforms in its current form, it lays groundwork for future graph-based forecasting. HRRRCast represents a step toward efficient, data-driven regional weather prediction with competitive accuracy and ensemble capability.

Figures

Figures reproduced from arXiv: 2507.05658 by the authors.

Figure 1
Figure 1. ResHRRR inputs and architecture: The network accepts 180 channels (variables) with 530 [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Composite reflectivity at forecast initialization time 2024-05-06 23:00 UTC. Shown are the HRRRCast single-member [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Enlarged view of composite reflectivity plots for 10-member HRRRCast ensemble and deterministic models at forecast [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Skill scores of composite reflectivity with different datasets, different thresholds (20, 30, and 40 dBZ) and a 6 km [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Comparison of RMSEs between the HRRR model and HRRRCast using the HRRR analysis as the ground truth. [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Profile of RMSEs for 3D atmospheric variables between HRRR and HRRRCast forecasts of 3h and 18h lead times. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Standard metrics comparing HRRRCast with HRRR. Both grid-based and object-based metrics indicate HRRRCast [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Profile of selected atmospheric variables (horizontal wind components and vertical velocity) at central longitude line [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Power spectrum comparison of different modesl: Left) 1h lead time Right) 6h lead time. Intermittent fields like [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: 500mb geopotential and surface analysis map for our case study on May 07, 2024 00:00 UTC. [PITH_FULL_IMAGE:figures/full_fig_p018_10.png]
Figure 11
Figure 11. Figure 11: Case study: A 3h forecast made by different models at forecast initialization time: 2024-05-06 23:00 UTC. Quantities [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Space-Time Transformer for Precipitation Nowcasting

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A full space-time attention video transformer recast as 64-class rainfall prediction with log-frequency class weighting won the Weather4Cast 2025 Cumulative Rainfall challenge (CRPS 3.135).

  2. Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model

    physics.ao-ph 2025-10 conditional novelty 5.0 of 10

    Combining HiRA neighborhood verification with threshold-weighted CRPS shows that AI-vs-NWP rankings for extreme precipitation depend strongly on neighborhood size.

  3. CRPS-LAM: Probabilistic Regional Weather Forecasting with Continuous Ranked Probability Score

    cs.LG 2025-10 conditional novelty 5.0 of 10

    CRPS-LAM produces 57-hour probabilistic limited-area forecasts on MEPS at diffusion-comparable accuracy with single-forward-pass sampling, roughly 39x faster than Diffusion-LAM.

Reference graph

Works this paper leans on

34 extracted references · 16 canonical work pages · cited by 3 Pith papers

  1. [1]

    Building machine learning limited area models: Kilometer- 20 scale weather forecasting in realistic settings

    Adamov, S., Oskarsson, J., Denby, L., Landelius, T., Hintz, K., Christiansen, S., Schicker, I., Osuna, C., Lindsten, F., Fuhrer, O., Schemm, S., 2025. Building machine learning limited area models: Kilometer- 20 scale weather forecasting in realistic settings. URL https://arxiv.org/abs/2504.09340

  2. [2]

    Bi, K., Xie, L., Zhang, H., Chen, X., Gu, X., Tian, Q., Jul. 2023. Accurate medium-range global weather forecasting with 3d neural networks. Nature 619 (7970), 533–538. URL https://doi.org/10.1038/s41586-023-06185-3

  3. [3]

    R., Aittala, M., Kreis, K., Brenowitz, N., Vahdat, A., Mardani, M., Yu, R., 2025

    Cachay, S. R., Aittala, M., Kreis, K., Brenowitz, N., Vahdat, A., Mardani, M., Yu, R., 2025. Elucidated rolling diffusion models for probabilistic weather forecasting. URL https://arxiv.org/abs/2506.20024

  4. [4]

    J., Haynes, K., Hoef, L

    Chase, R. J., Haynes, K., Hoef, L. V., Ebert-Uphoff, I., 2025. Score-based diffusion nowcasting of goes imagery. URL https://arxiv.org/abs/2505.10432

  5. [5]

    J., 2017

    Clark, A. J., 2017. Generation of ensemble mean precipitation forecasts from convection-allowing en- sembles. Weather and Forecasting 32 (4), 1569 – 1583

  6. [6]

    C., Alexander, C

    Dowell, D. C., Alexander, C. R., James, E. P., Weygandt, S. S., Benjamin, S. G., Manikin, G. S., Blake, B. T., Brown, J. M., Olson, J. B., Hu, M., Smirnova, T. G., Ladwig, T., Kenyon, J. S., Ahmadov, R., Turner, D. D., Duda, J. D., Alcott, T. I., 2022. The high-resolution rapid refresh (hrrr): An hourly updating convection-allowing forecast model. part i:...

  7. [7]

    (DTC), D. T. C., 2024. METplus User’s Guide: Gen-Ens-Prod Tool. National Center for Atmospheric Research (NCAR), version 5.1.1

  8. [8]

    L., Potvin, C., 2025

    Flora, M. L., Potvin, C., 2025. Wofscast: A machine learning model for predicting thunderstorms at watch-to-warning scales. Geophysical Research Letters 52 (10), e2024GL112383, e2024GL112383 2024GL112383. URL https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/2024GL112383

Show all 34 references
  1. [9]

    Deep residual learning for image recognition

    He, K., Zhang, X., Ren, S., Sun, J., 2015. Deep residual learning for image recognition

  2. [10]

    B., Smirnova, T., Alexander, C., Berner, J., 2019

    Jankov, I., Beck, J., Wolff, J., Harrold, M., Olson, J. B., Smirnova, T., Alexander, C., Berner, J., 2019. Stochastically perturbed parameterizations in an hrrr-based ensemble. Monthly Weather Review 147 (1), 153 – 173. URL https://journals.ametsoc.org/view/journals/mwre/147/1...

  3. [11]

    Elucidating the design space of diffusion-based gener- ative models

    Karras, T., Aittala, M., Aila, T., Laine, S., 2022. Elucidating the design space of diffusion-based gener- ative models. URL https://arxiv.org/abs/2206.00364

  4. [12]

    Forecasting global weather with graph neural networks

    Keisler, R., 2022. Forecasting global weather with graph neural networks. URL https://arxiv.org/abs/2202.07575

  5. [13]

    Learning skillful medium-range global weather forecasting

    Lam, R., Sanchez-Gonzalez, A., Willson, M., Wirnsberger, P., Fortunato, M., Alet, F., Ravuri, S., Ewalds, T., Eaton-Rosen, Z., Hu, W., Merose, A., Hoyer, S., Holland, G., Vinyals, O., Stott, J., Pritzel, A., Mohamed, S., Battaglia, P., 2023. Learning skillful medium-range glob...

  6. [14]

    Lang, S., Alexe, M., Clare, M. C. A., Roberts, C., Adewoyin, R., Bouall?gue, Z. B., Chantry, M., Dramsch, J., Dueben, P. D., Hahner, S., Maciel, P., Prieto-Nemesio, A., O’Brien, C., Pinault, F., Polster, J., Raoult, B., Tietsche, S., Leutbecher, M., 2024. Aifs-crps: Ensemble f...

  7. [15]

    Diffusion-lam: Probabilistic limited area weather forecasting with diffusion

    Larsson, E., Oskarsson, J., Landelius, T., Lindsten, F., 2025. Diffusion-lam: Probabilistic limited area weather forecasting with diffusion. URL https://arxiv.org/abs/2502.07532

  8. [16]

    May 7, 2024 severe weather and tornadoes

    National Weather Service, May 2024. May 7, 2024 severe weather and tornadoes. Accessed: 2025-05-29. URL https://www.weather.gov/grr/7 May 2024 SevereWeather

  9. [17]

    K., Grover, A., 2023

    Nguyen, T., Brandstetter, J., Kapoor, A., Gupta, J. K., Grover, A., 2023. Climax: A foundation model for weather and climate. In: International Conference on Machine Learning. URL https://api.semanticscholar.org/CorpusID:256231457

  10. [18]

    Scaling transformer neural networks for skillful and reliable medium-range weather forecasting

    Nguyen, T., Shah, R., Bansal, H., Arcomano, T., Madireddy, S., Maulik, R., Kotamarthi, V., Foster, I., Grover, A., 2023. Scaling transformer neural networks for skillful and reliable medium-range weather forecasting

  11. [19]

    N., Haugen, H

    Nipen, T. N., Haugen, H. H., Ingstad, M. S., Nordhagen, E. M., Salihi, A. F. S., Tedesco, P., Seierstad, I. A., Kristiansen, J., Lang, S., Alexe, M., Dramsch, J., Raoult, B., Mertes, G., Chantry, M., 2024. Regional data-driven weather modeling with a global stretched-grid. URL...

  12. [20]

    Graph-based neural weather prediction for limited area modeling

    Oskarsson, J., Landelius, T., Lindsten, F., 2023. Graph-based neural weather prediction for limited area modeling. URL https://arxiv.org/abs/2309.17370

  13. [21]

    Kilometer-scale convection allowing model emulation using generative diffusion modeling

    Pathak, J., Cohen, Y., Garg, P., Harrington, P., Brenowitz, N., Durran, D., Mardani, M., Vahdat, A., Xu, S., Kashinath, K., Pritchard, M., 2024. Kilometer-scale convection allowing model emulation using generative diffusion modeling. URL https://arxiv.org/abs/2408.10958

  14. [22]

    K., Carley, J

    Potvin, C. K., Carley, J. R., Clark, A. J., Wicker, L. J., Skinner, P. S., Reinhart, A. E., Gallo, B. T., Kain, J. S., Romine, G. S., Aligo, E. A., Brewster, K. A., Dowell, D. C., Harris, L. M., Jirak, I. L., Kong, F., Supinie, T. A., Thomas, K. W., Wang, X., Wang, Y., Xue, M....

  15. [23]

    R., El-Kadi, A., Masters, D., Ewalds, T., Stott, J., Mohamed, S., Battaglia, P., Lam, R., Willson, M., 2024

    Price, I., Sanchez-Gonzalez, A., Alet, F., Andersson, T. R., El-Kadi, A., Masters, D., Ewalds, T., Stott, J., Mohamed, S., Battaglia, P., Lam, R., Willson, M., 2024. Gencast: Diffusion-based ensemble forecasting for medium-range weather

  16. [24]

    Data-driven medium-range weather prediction with a resnet pretrained on climate simulations: A new model for weatherbench

    Rasp, S., Thuerey, N., 2021. Data-driven medium-range weather prediction with a resnet pretrained on climate simulations: A new model for weatherbench. Journal of Advances in Modeling Earth Systems 13 (2), e2020MS002405, e2020MS002405 2020MS002405. URL https://agupubs.onlineli...

  17. [25]

    Rolling diffusion models

    Ruhe, D., Heek, J., Salimans, T., Hoogeboom, E., 2024. Rolling diffusion models. URL https://arxiv.org/abs/2402.09470

  18. [26]

    A., Kossaifi, J., Bonev, B., Choy, C., Kautz, J., Krueger, D., Azizzadenesheli, K., 2024

    Siddiqui, S. A., Kossaifi, J., Bonev, B., Choy, C., Kautz, J., Krueger, D., Azizzadenesheli, K., 2024. Exploring the design space of deep-learning-based weather forecasting systems. URL https://arxiv.org/abs/2410.07472

  19. [27]

    Skinner, P., Stratman, D., Kerr, C., Matilla, B., Martin, J., Dowell, D., Jones, T., Flora, M., Guerra, J., Knopfmeier, K., Britt, K., Yussouf, N., ???? Comparing short-term thunderstorm forecasts from the warn-on-forecast system (wofs) and high-resolution rapid refresh (hrrr)

  20. [28]

    A., Penny, S

    Smith, T. A., Penny, S. G., Platt, J. A., Chen, T.-C., 2023. Temporal subsampling diminishes small spatial scales in recurrent neural network emulators of geophysical turbulence. Journal of Advances in Modeling Earth Systems 15 (12), e2023MS003792, e2023MS003792 2023MS003792. ...

  21. [29]

    Denoising diffusion implicit models

    Song, J., Meng, C., Ermon, S., 2022. Denoising diffusion implicit models. URL https://arxiv.org/abs/2010.02502

  22. [30]

    Z., Separovic, L., Yang, J., 2025

    Subich, C., Husain, S. Z., Separovic, L., Yang, J., 2025. Fixing the double penalty in data-driven weather forecasting through a modified spherical harmonic loss function. URL https://arxiv.org/abs/2501.19374

  23. [31]

    R., Landolt, S

    Xu, M., Thompson, G., Adriaansen, D. R., Landolt, S. D., 2019. On the value of time-lag-ensemble averaging to improve numerical model predictions of aircraft icing conditions. Weather and Forecasting 34 (3), 507 – 519. URL https://journals.ametsoc.org/view/journals/wefo/34/3/w...

  24. [32]

    This provides the spatial structure that will be preserved in the final PMM field

    Compute the ensemble mean field by averaging the reflectivity across all ensemble members at each grid point. This provides the spatial structure that will be preserved in the final PMM field

  25. [33]

    From this sorted array, every N th value is selected – where N is the number of ensemble members – to produce a downsampled set of values

    Pool and sort all reflectivity values from all ensemble members into a single 1D array. From this sorted array, every N th value is selected – where N is the number of ensemble members – to produce a downsampled set of values. This set approximates the typical value distributi...

  26. [34]

    This ensures the final field has the same spatial pattern as the ensemble mean, but with a value distribution that better reflects the ensemble variability

    Reassign the sorted values from step 2 to the grid points of the mean field based on rank ordering: the largest value is assigned to the location with the highest mean, the second-largest to the second-highest, and so on. This ensures the final field has the same spatial patte...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.