Pith. sign in

REVIEW 3 major objections 5 minor 25 references

Scale-aware losses improve machine-learned forecast realism

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 13:14 UTC pith:JDSGMTY5

load-bearing objection A clean derivation and preliminary comparison of graph-localized multivariate scoring losses for global ML weather forecasting, but the scale-awareness claim is confounded by ad hoc weights and small-model experiments. the 3 major comments →

arxiv 2607.19161 v2 pith:JDSGMTY5 submitted 2026-07-21 physics.ao-ph stat.ML

On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules

classification physics.ao-ph stat.ML
keywords probabilistic weather forecastingscoring rulesCRPSenergy scoregraph energy scorescale-aware lossspectral realismmachine-learned ensembles
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether the strong performance of machine-learned probabilistic weather forecasting depends on the particular loss function used for training, or whether alternative scoring rules can achieve the same skill. The authors train variants of a global forecast model with three objectives: a pointwise score (the almost-fair CRPS), a whole-field score (the fair energy score), and a spatially localized graph-based score anchored to the energy score. They find that forecast skill is broadly similar across all three, with the graph score slightly better in the tropics and the global energy score slightly worse. In a second experiment set, they compare twelve loss configurations and find that every explicitly scale-aware loss—whether based on decomposing fields into scale bands or on scoring spectral coefficients—produces forecast spectra closer to the atmosphere's than pointwise scoring does, and that the weights attached to scales and variables matter at least as much as the mechanism. The authors caution that the experiments used a smaller model and a shortened training schedule, with some weights selected by hand, so full-scale conclusions await further testing.

Core claim

Multivariate scores are viable alternatives to pointwise CRPS training for global machine-learned weather forecasting. A graph-based energy score evaluated over local neighbourhoods matches or slightly beats CRPS skill in the tropics, while a global energy score degrades there. Explicit scale awareness—whether through a multi-scale band decomposition or through spectral scores—improves the realism of forecast spectra by constraining small-scale variability. Weighting of scales and variables matters more than the particular scoring mechanism.

What carries the argument

The carrying machinery is a family of fair multivariate scoring rules and their scale-aware extensions. The graph energy score computes a weighted neighbourhood norm on a k-nearest-neighbour graph at each grid node and aggregates over nodes, localizing the energy score without rectangular patches; it is combined with a weak global energy anchor to restore strict propriety. Scale awareness enters in two ways: a multi-scale loss decomposes fields into ordered scale bands via successive smoothing operators (a Laplacian-pyramid-like construction) and scores each band separately, and spectral losses score spherical-harmonic coefficients either as two-dimensional vectors (spectral energy) or by ma

Load-bearing premise

The central conclusions were drawn from a smaller model, a shortened training schedule, and scale and variable weights that were chosen by hand from rough data estimates; if any of these choices drives the observed differences, the claims may not transfer to full-scale operational training.

What would settle it

Re-run the twelve spectral experiments at the operational resolution and full training schedule with the same ad hoc weights, and also with the weights varied by, say, a factor of ten; if the spectral advantage of scale-aware losses disappears, inverts for some variables, or depends entirely on the hand-picked weights, then the claim that explicit scale awareness itself improves realism would be falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Training global probabilistic forecast models does not require a pointwise score; localized multivariate scores can match or slightly improve skill, especially in the tropics.
  • Explicitly scale-aware losses, whether multi-scale or spectral, yield forecast fields whose variance spectra are closer to those of analyses than pointwise CRPS training does.
  • The effective weights applied to different scales and output variables influence spectral fidelity at least as much as the choice of scoring rule, so the weighting scheme is a primary design decision.
  • Graph-based localization of scores applies to irregular grids and sparse observation networks, since it needs only a neighbourhood graph and weights rather than fixed rectangular patches.
  • Some scale-aware configurations slightly overconstrain small-scale variability at early lead times, reducing early tendency variability below the reference.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the tropical advantage of the graph energy score persists at operational resolution, it suggests that spatial coherence is especially valuable where convective-scale variability dominates, and could point toward local graph structure as a design principle for future loss functions.
  • Because weighting matters more than mechanism, a systematic, data-driven procedure for setting scale and variable weights would likely yield larger gains than inventing new score families; the paper's ad hoc weights are a placeholder for this tuning problem.
  • The graph energy score's strict-propriety failure, though mitigated by an anchor, implies a tunable trade-off: too weak an anchor may leave long-range dependence unconstrained, so an explicit sensitivity study of the anchor weight would clarify when the combined score is safe to use.
  • The same multi-scale and spectral scoring ideas could be transferred to other generative tasks with known spectral character (e.g., downscaling, super-resolution, or synthetic imagery), where matching the power spectrum of the target domain is a goal.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript, presented as a preliminary study, compares several probabilistic training losses for the AIFS-CRPS global machine-learned weather forecasting model. It defines a family of graph-based multivariate scores (graph energy, graph variogram, graph edge energy, edge CRPS) with fair variants, and two spectral scores (spectral energy and spectral-magnitude CRPS), and then reports two sets of experiments. In the first experiment, models at about 1 degree resolution are trained with almost-fair CRPS, a fair global energy score, and a fair graph energy score with a weak global anchor; skill is reported as broadly similar, with qualitative tropical differences. In the second, cheaper experiment, twelve loss configurations are compared through accumulated-tendency spectra against ERA5, leading to the claims that scale-aware losses improve spectral fidelity and that any explicit form of scale-awareness improves realism. The authors explicitly acknowledge limitations in Section 5, including smaller models, lower resolution, shortened training, and ad hoc per-scale weighting.

Significance. If the empirical claims can be made robust, the paper would make a useful contribution: it would show that CRPS-based training is not uniquely necessary for global ML ensemble weather prediction, and it would provide flexible graph-based multivariate scores that apply to irregular grids. The score derivations in Section 2 are clear, the fair-score corrections are done carefully, and the implementation in the Anemoi framework is concrete. The practical question of whether scale-aware losses improve spectral realism is of immediate interest to the machine-learned weather forecasting community. However, the current evidence for the spectral claim is confounded by ad hoc weighting and the skill comparison rests on single training runs without uncertainty quantification. The significance is therefore contingent on strengthening the experimental design or substantially softening the claims.

major comments (3)
  1. [Abstract; §3.2, Table 2; §4.2] The central statement that 'any form of explicit scale-awareness improves realism' is not established by the experimental design. Each multi-scale or spectral loss configuration differs from the single-scale CRPS baseline in two ways at once: the presence of a scale decomposition and the introduction of per-scale, per-waveband, and per-variable weights. Section 3.2 states these weights are 'ad hoc and were chosen only to be of the right order of magnitude - estimated from the data - rather than tuned,' and their numerical values are not reported anywhere. Figures 3-14 therefore cannot separate the effect of scale awareness from the effect of reweighting different scales and variables. The paper itself acknowledges this in the abstract, noting that 'the largest differences are likely associated with different effective weights per scale.' To support the claim, a controlled ablation is nee
  2. [§4.1, Fig. 2] The tropical differences — graph energy performing best and global energy showing degradation relative to CRPS — are described as notable and used to support the paper's overall conclusions, but they are based on a single training run per loss and visual inspection of Figure 2. No error bars, confidence intervals, or multiple seeds are provided. Training a neural weather model is stochastic, and differences of this magnitude could easily arise from initialization or optimization noise. The abstract's first conclusion ('multivariate scores are a viable alternative') is comparatively robust because it rests on broad equivalence, but the stronger tropical ranking needs uncertainty quantification, for example by bootstrapping over verification dates or by reporting small-ensemble seed variability.
  3. [§3.2; §5] The spectral and, to a lesser extent, tropical conclusions are obtained with a substantially reduced setup: smaller model, O96 resolution, and a shortened training schedule with only 1,000 optimization steps per rollout length up to eight steps. The limitation statement in Section 5 is candid, but it is in tension with the abstract's wording 'global machine-learned weather forecasting.' The shorter schedule may disadvantage the non-scale-aware baselines differently from the scale-aware losses, and the paper does not provide convergence diagnostics. Please either provide evidence that the compared training runs have converged sufficiently for the conclusions, or explicitly restrict all claims to this reduced training regime and adjust the abstract accordingly.
minor comments (5)
  1. [§2; §3.2] The symbol ℓ is used both as an ensemble-member index in the score definitions and as the total wavenumber for spectral bands in Section 3.2. These uses are easy to confuse; please rename one of them.
  2. [§3.2] The per-scale and per-waveband weighting factors are described only verbally. A table listing the actual weights used for each variable and band, even in an appendix, would make the experiments reproducible and would allow readers to judge whether the conclusions are sensitive to the chosen weights.
  3. [§4.1, Fig. 2] The main skill comparison is presented as a figure without a supporting quantitative table. A compact table of CRPS values, or differences relative to the CRPS baseline, with confidence intervals would make the 'broadly similar' claim much easier to assess.
  4. [§2, graph energy score] The text notes that the graph energy score need not be strictly proper and says that combining it with a positive-weight global energy score 'makes the combined score strictly proper.' This is plausible, but a one-sentence proof or explicit reference to the relevant proposition in [21] would remove ambiguity, since strict propriety of the training loss is a load-bearing property.
  5. [§5, Discussion] The paper would benefit from stating explicitly which parts of the results are new relative to the authors' prior multi-scale loss paper [14]. The discussion mentions operational use of [14] and lists other models with spectral losses, but it does not clearly delineate what the reader should take away as the novel empirical finding beyond 'results are consistent with previous work.'

Circularity Check

0 steps flagged

No load-bearing circularity: empirical loss comparisons are self-contained; minor self-citation and data-estimated scale weights noted.

full rationale

The paper is an empirical ablation, not a derivation. Training objectives (Sec. 2) are defined independently of the verification metrics (fair CRPS, spectra of accumulated tendencies); the skill comparison in Sec. 4.1 evaluates all three models with fair CRPS even though only the baseline was trained with afCRPS, and the multivariate models are not constructed to optimize that score. The spectral comparison in Sec. 4.2 is a measured holdout outcome: the multi-scale/spectral losses decompose training fields into scale bands or spectral modes, while the evaluation computes accumulated-tendency spectra against ERA5 for 2022; these are related but not identical objects, so no equation reduces one to the other. The only mild self-referential elements are (i) the base model and multi-scale loss come from the authors' own refs [13] and [14], and (ii) Sec. 3.2 states the scale/variable weights were 'ad hoc and were chosen only to be of the right order of magnitude - estimated from the data - rather than tuned.' That is a confound for attributing spectral improvements to scale-awareness per se — the paper itself concedes 'the largest differences are likely associated with different effective weights per scale' — but it is not a fitted parameter renamed as a prediction, because the weights were not fitted to the verification metric and the unweighted spectral experiments also constrain small scales. The self-citations are method attributions, not load-bearing uniqueness claims. No circular step meeting the quote-and-reduction standard was found.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 0 invented entities

The central empirical claims depend on training procedures and verification choices rather than analytic derivations. The main free parameters are the almost-fair CRPS mixing angle, the variogram exponent (unreported), and the ad hoc scale/band weighting that directly shapes the spectral results. No new physical entities are introduced.

free parameters (5)
  • almost-fair CRPS mixing parameter alpha = 0.95
    Set for all CRPS-based experiments; controls trade-off between fair and unfair CRPS. Not fitted to validation skill.
  • variogram exponent p = not stated
    The graph variogram score is defined with |z_j - z_n|^p, but the value of p used in the multi-scale variogram experiment is not reported.
  • per-scale weights zeta_i for multi-scale losses = not listed (ad hoc)
    K-scale loss sums scale bands with weights zeta_i; text says weights were chosen ad hoc by order of magnitude from data, but exact values are not given. Acknowledged as influencing spectral results.
  • per-waveband and per-variable weights for spectral losses = not listed (ad hoc)
    Spectral experiments use band- and variable-dependent weighting estimated from data; exact weights omitted; paper says weighting matters as much as the score mechanism.
  • graph neighbourhood size k = 16
    k-NN graph with k=16 chosen for graph scores; a modeling choice, not fitted.
axioms (5)
  • domain assumption Optimizing a proper scoring rule as training loss transfers to forecast skill at longer lead times.
    The experiments initialize with one-step training and then rollout to multiple steps; the paper assumes gradient descent on the loss improves the ensemble forecast distribution beyond the training rollout lengths.
  • domain assumption ERA5 reanalysis is a valid ground truth for both training and spectral verification.
    Used as target and reference; no treatment of analysis errors.
  • domain assumption Graph energy score propriety can be restored by adding a weak global energy anchor with weight 0.1.
    Text acknowledges graph energy score may fail strict propriety; cites [21] for combination preserving propriety, but does not verify for this specific graph construction.
  • domain assumption The 84 initialization dates and T191 truncation give representative spectral diagnostics.
    Spectra averaged over 84 dates; no uncertainty quantification.
  • standard math Standard scoring-rule mathematics (kernel representations, fairness corrections) is correct.
    Assumed background for score definitions.

pith-pipeline@v1.3.0-alltime-deepseek · 10427 in / 12080 out tokens · 124119 ms · 2026-08-01T13:14:58.958004+00:00 · methodology

0 comments
read the original abstract

Probabilistic forecast models can be machine-learned from data using loss functions based on scoring rules such as the Continuous Ranked Probability Score (CRPS). This note summarises a preliminary study comparing versions of AIFS-CRPS, a global weather forecast model, trained with different univariate and multivariate scoring rules that aim to explicitly represent scale-awareness in the loss function. In the first part, we compare the (almost) fair CRPS, a fair global energy score, and a graph energy score based on node neighbourhoods. Across standard verification metrics, forecast skill is broadly similar. In the extratropics we find only small differences, while in the tropics the graph energy score setup performs somewhat better and the global energy score shows some degradation. These results suggest that multivariate scores are a viable alternative to CRPS-based training for global machine-learned weather forecasting. In the second part of the study, we analyse how different scoring rules and scale-aware loss constraints shape the spectra of forecast fields. It is apparent that any form of explicit scale-awareness improves realism. Here, the largest differences are likely associated with different effective weights per scale.

Figures

Figures reproduced from arXiv: 2607.19161 by Martin Leutbecher, Sam Hatfield, Simon Lang.

Figure 1
Figure 1. Figure 1: Different scores for a single destination node [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of the CRPS, energy score, and graph energy score experiments. The top [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Accumulated tendency spectra for geopotential at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Accumulated tendency spectra for geopotential at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Accumulated tendency-spectrum ratios for geopotential at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Accumulated tendency-spectrum ratios for geopotential at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Accumulated tendency spectra for meridional wind at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Accumulated tendency spectra for meridional wind at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Accumulated tendency-spectrum ratios for meridional wind at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Accumulated tendency-spectrum ratios for meridional wind at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p021_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Accumulated tendency spectra for temperature at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p022_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Accumulated tendency spectra for temperature at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p023_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Accumulated tendency-spectrum ratios for temperature at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p024_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Accumulated tendency-spectrum ratios for temperature at 500 hPa. [PITH_FULL_IMAGE:figures/full_fig_p025_14.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references

  1. [1]

    Ander- sson, Jacklynn Stott, R´ emi Lam, Matthew Willson, Alvaro Sanchez-Gonzalez, and Peter W

    Ferran Alet, Ilan Price, Andrew El-Kadi, Dominic Masters, Stratis Markou, Tom R. Ander- sson, Jacklynn Stott, R´ emi Lam, Matthew Willson, Alvaro Sanchez-Gonzalez, and Peter W. Battaglia. Skillful joint probabilistic weather forecasting from marginals, 2025

  2. [2]

    PyTorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation

    Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesen- sky, Bin Bao, Peter Bell, David Berard, Evgeni Burovski, et al. PyTorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation. InPro- ceedings of the 29th ACM International Conference on Architectural Support for Programming L...

  3. [3]

    Learning probabilistic filters with strictly proper scoring rules, 2026

    Eviatar Bach, Ricardo Baptista, Jochen Br¨ ocker, Bohan Chen, and Andrew Stuart. Learning probabilistic filters with strictly proper scoring rules, 2026

  4. [4]

    Collins, Michael S

    Boris Bonev, Thorsten Kurth, Ankur Mahesh, Mauro Bisson, Jean Kossaifi, Karthik Kashinath, Anima Anandkumar, William D. Collins, Michael S. Pritchard, and Alexander Keller. Fourcastnet 3: A geometric approach to probabilistic machine-learning weather fore- casting at scale, 2025

  5. [5]

    Probabilistic neural operators for functional uncertainty quantification, 2025

    Christopher B¨ ulte, Philipp Scholl, and Gitta Kutyniok. Probabilistic neural operators for functional uncertainty quantification, 2025

  6. [6]

    Burt and Edward H

    Peter J. Burt and Edward H. Adelson. The laplacian pyramid as a compact image code.IEEE Transactions on Communications, 31(4):532–540, 1983

  7. [7]

    U-Cast: A surprisingly simple and efficient frontier probabilistic AI weather forecaster, 2026

    Salva R¨ uhling Cachay, Duncan Watson-Parris, and Rose Yu. U-Cast: A surprisingly simple and efficient frontier probabilistic AI weather forecaster, 2026

  8. [8]

    Cristiana Diaconu, Jonas Scholz, Aliaksandra Shysheya, Stratis Markou, Payel Mukhopadhyay, Miles Cranmer, and Richard E. Turner. Otter Weather: Skillful and computationally efficient medium-range weather forecasting, 2026

  9. [9]

    Christopher A. T. Ferro. Fair scores for ensemble forecasts.Quarterly Journal of the Royal Meteorological Society, 140(683):1917–1923, 2014

  10. [10]

    Tilmann Gneiting and Adrian E. Raftery. Strictly proper scoring rules, prediction, and esti- mation.Journal of the American Statistical Association, 102(477):359–378, 2007. 27

  11. [11]

    Hersbach, B

    H. Hersbach, B. Bell, P. Berrisford, et al. The ERA5 global reanalysis.Quarterly Journal of the Royal Meteorological Society, 146:1999–2049, 2020

  12. [12]

    A composite-loss graph neural network for the multivariate post-processing of ensemble weather forecasts.Quarterly Journal of the Royal Meteorological Society, 2026

    M´ aria Lakatos. A composite-loss graph neural network for the multivariate post-processing of ensemble weather forecasts.Quarterly Journal of the Royal Meteorological Society, 2026

  13. [13]

    Simon Lang, Mihai Alexe, Mariana C. A. Clare, Christopher Roberts, Rilwan Adewoyin, Zied Ben Bouall` egue, Matthew Chantry, Jesper Dramsch, Peter D. Dueben, Sara Hahner, Pedro Maciel, Ana Prieto-Nemesio, Cathal O’Brien, Florian Pinault, Jan Polster, Baudouin Raoult, Steffen Tietsche, and Martin Leutbecher. AIFS-CRPS: Ensemble forecasting using a model tra...

  14. [14]

    A multi-scale loss formulation for learning a probabilistic model with proper score optimisation, 2025

    Simon Lang, Martin Leutbecher, and Pedro Maciel. A multi-scale loss formulation for learning a probabilistic model with proper score optimisation, 2025

  15. [15]

    CRPS-LAM: Regional ensemble weather forecasting from matching marginals, 2025

    Erik Larsson, Joel Oskarsson, Tomas Landelius, and Fredrik Lindsten. CRPS-LAM: Regional ensemble weather forecasting from matching marginals, 2025

  16. [16]

    Ensemble size: How suboptimal is less than infinity?Quarterly Journal of the Royal Meteorological Society, 145(S1):107–128, 2019

    Martin Leutbecher. Ensemble size: How suboptimal is less than infinity?Quarterly Journal of the Royal Meteorological Society, 145(S1):107–128, 2019

  17. [17]

    Weyn, Hang Zhang, Yanfei Xiang, Jiang Bian, Weixin Jin, Kit Thambi- ratnam, Qi Zhang, Haiyu Dong, and Hongyu Sun

    Zekun Ni, Jonathan A. Weyn, Hang Zhang, Yanfei Xiang, Jiang Bian, Weixin Jin, Kit Thambi- ratnam, Qi Zhang, Haiyu Dong, and Hongyu Sun. Huracan: A skillful end-to-end data-driven system for ensemble data assimilation and weather prediction, 2025

  18. [18]

    High-resolution proba- bilistic data-driven weather modeling with a stretched-grid, 2025

    Even Marius Nordhagen, H ˚ avard Homleid Haugen, Aram Farhad Shafiq Salihi, Magnus Sikora Ingstad, Thomas Nils Nipen, Ivar Ambjørn Seierstad, Inger-Lise Frogner, Mariana Clare, Simon Lang, Matthew Chantry, Peter Dueben, and Jørn Kristiansen. High-resolution proba- bilistic data-driven weather modeling with a stretched-grid, 2025

  19. [19]

    Adewoyin, Peter Dueben, and Ritabrata Dutta

    Lorenzo Pacchiardi, Rilwan A. Adewoyin, Peter Dueben, and Ritabrata Dutta. Probabilis- tic forecasting with generative networks via scoring rule minimization.Journal of Machine Learning Research, 25:1–64, 2024

  20. [20]

    Andre Perkins, Anna Kwa, Jeremy McGibbon, Troy Arcomano, Spencer K

    W. Andre Perkins, Anna Kwa, Jeremy McGibbon, Troy Arcomano, Spencer K. Clark, Oliver Watt-Meyer, Christopher S. Bretherton, and Lucas M. Harris. HiRO-ACE: Fast and skillful AI emulation and downscaling trained on a 3 km global storm-resolving model, 2026

  21. [21]

    Romain Pic, Cl´ ement Dombry, Philippe Naveau, and Maxime Taillardat. Proper scoring rules for multivariate probabilistic forecasts based on aggregation and transformation.Advances in Statistical Climatology, Meteorology and Oceanography, 11(1):23–58, 2025

  22. [22]

    Michael Scheuerer and Thomas M. Hamill. Variogram-based proper scoring rules for proba- bilistic forecasts of multivariate quantities.Monthly Weather Review, 143(4):1321–1334, 2015

  23. [23]

    Enscale: Temporally-consistent multivariate generative downscaling via proper scoring rules, 2025

    Maybritt Schillinger, Maxim Samarin, Xinwei Shen, Reto Knutti, and Nicolai Meinshausen. Enscale: Temporally-consistent multivariate generative downscaling via proper scoring rules, 2025

  24. [24]

    Philippe Tillet, Hsiang-Tsung Kung, and David D. Cox. Triton: An intermediate language and compiler for tiled neural network computations. InProceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pages 10–19, 2019. 28

  25. [25]

    (Sparse) Attention to the Details: Preserving spectral fidelity in ML-based weather forecasting models, 2026

    Maksim Zhdanov, Ana Lucic, Max Welling, and Jan-Willem van de Meent. (Sparse) Attention to the Details: Preserving spectral fidelity in ML-based weather forecasting models, 2026. 29