REVIEW 3 major objections 3 minor 32 references
Statistical post-processing of operational dual-resolution wind-speed ensemble forecasts
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read High resolution beats ensemble size in wind-speed forecasts
desk verdict Useful operational numbers on mixing ECMWF's 9-km and 36-km wind-speed ensembles, but the headline claim that 'spatial resolution is superior to ensemble size' overreaches the design: the two ensembles are separate operational products that differ in more than resolution and member count. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the truncated-normal EMOS predictive distribution $N_0^\infty(\mu,\sigma^2)$ for wind speed, with location $\mu = a + b_H^2 \bar f_H + b_L^2 \bar f_L$ and variance $\sigma^2 = c^2 + d^2 S^2$, where $\bar f_H$ and $\bar f_L$ are the means of the high- and low-resolution ensemble members and $S^2$ is the variance of the combined ensemble. Parameters are estimated by minimizing the continuous ranked probability score over training data selected regionally, locally, or semi-locally through k-means clustering. This distribution is what converts raw ensemble output into a calibrated probability forecast, and the same scores used to fit it, namely CRPS, quantile score, and Brier score with stationary-bootstrap confidence intervals, are used to compare configurations.
What would settle it
Run a single numerical weather prediction model at 9 km with 50 members and at 36 km with 100 members, keeping perturbation method, physics, and initialization identical, and verify on a full season of wind-speed observations; if the 36-km, 100-member ensemble matches or beats the 9-km, 50-member ensemble on CRPS, the resolution-over-size claim fails. A quicker check is to verify the raw 9-km forecasts against observations upscaled to a 36-km footprint; if the advantage largely disappears, it was representativeness error rather than forecast information.
Extended reading notes
Core claim
On its own terms, the paper shows that for wind speed the 50-member, 9-km ensemble forecast is at least as skillful as the 150-member dual-resolution forecast (50 high-resolution plus 100 low-resolution members) across CRPS, MAE, quantile, Brier, and RMSE scores, and that the 100-member, 36-km forecast is significantly worse on nearly all of them. After local EMOS post-processing every configuration improves, the differences between configurations shrink, and the dual-resolution forecast is significantly better than the high-resolution forecast only for the first two days. Augmenting a 50-member low-resolution ensemble with 1, 2, 4, 8, 16, or 32 high-resolution members helps for every configuration before calibration, with the largest gains coming from the largest number of high-resolution members; after calibration the gain remains significant in CRPS up to about day 4, after which low-resolution-only post-processing catches up.
Load-bearing premise
The whole comparison assumes the two forecast systems differ only in resolution and ensemble size; if hidden differences such as perturbation strategy, model physics, or initialization drive the skill gap, the conclusion that resolution beats size does not follow.
Editorial extensions
If this is right
- For an operational centre, once a high-resolution ensemble has about 50 members, paying for 100 more low-resolution members is unlikely to improve wind-speed forecast skill; spending the same computing budget on resolution rather than low-resolution ensemble size is the better bet.
- For a low-resolution ensemble, adding even one or two high-resolution members improves the raw forecast, and adding more extends the benefit; this is a cheap upgrade path for extended-range systems.
- Statistical post-processing compresses, but does not remove, the configuration gap; after calibration, the choice of resolution and ensemble mix matters mainly for the first few days of the forecast.
- The advantage of high-resolution members in post-processed mixtures is concentrated in CRPS and upper-tail quantiles, and disappears for high wind-speed thresholds, so users should not expect resolution mixing to fix rare strong-wind events.
- For very short lead times (day 1-2), a post-processed dual-resolution forecast is the best choice; beyond about day 4, the low-resolution-only EMOS model is competitive.
Reading between the lines
- Editorial inference: the paper's headline conclusion that resolution beats ensemble size is stated for two operational products that may differ in perturbation strategy, physics, and initialization; a controlled experiment varying only resolution and member count is needed before the rule is treated as general.
- Editorial inference: a direct testable extension is to repeat the same mixture experiment for temperature, precipitation, or data-driven ensemble forecasts; the asymmetry found here may or may not survive.
- Editorial inference: the raw-forecast verification is affected by representativeness error that favours the finer grid, so part of the apparent resolution advantage may be a measurement artifact; verifying against upscaled observations would separate the two.
- Editorial inference: because post-processing reduces the gap, a cost-aware forecast design could use low-resolution ensembles for calibration of long lead times and reserve high-resolution computation for days 1-3.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares raw and EMOS-post-processed ECMWF wind-speed ensemble forecasts at two horizontal resolutions and varying ensemble sizes, using a truncated normal EMOS model with regional, local, and semi-local training-data selection. The main empirical findings are that the 50-member high-resolution (TCO1279) raw ensemble performs comparably to or better than the 150-member dual-resolution (100,50) ensemble, that the 100-member low-resolution (TCO319) raw ensemble is significantly worse on nearly all scores, and that after local EMOS post-processing, adding high-resolution members to a 50-member low-resolution base yields significant CRPS improvements up to about day 4, while adding low-resolution members to a high-resolution base is beneficial only up to day 2. The authors conclude that spatial resolution is superior to ensemble size and that the direction of mixing matters.
Significance. If the results hold in their intended generality, the paper would provide practically useful guidance for designing dual-resolution ensemble systems and for deciding how to allocate computational resources between resolution and ensemble size. The study is careful in several respects: it uses a large station network, applies proper scoring rules (CRPS, BS, QS) together with point-forecast metrics, and accompanies all skill scores with 95% block-bootstrap confidence intervals, allowing significance statements to be checked. The paper also transparently describes the EMOS setup and the configuration search. Its main limitation is that the comparison is not a controlled experiment: the high- and low-resolution systems are distinct operational products that differ in more than resolution and ensemble size, so the headline attribution of skill differences to spatial resolution is not actually established by the data.
major comments (3)
- [Abstract and Section 5] The central claim that "spatial resolution is superior to the ensemble size" is not supported by the design, because the (0,50) and (100,0) systems are not controlled variants of a common ensemble system. Section 2 states that the 50-member medium-range ENS runs at TCO1279 and the 100-member extended-range ENS runs at TCO319; these are separate operational products that plausibly differ in perturbation strategy, stochastic physics, initialization, and model cycle in addition to resolution and membership. The comparison (100,0) vs (0,50) varies both resolution and ensemble size simultaneously, and the "adding low-resolution members" contrast (0,50) vs (100,50) adds members from a different system, so a null or negative result can reflect the inferiority of that system's perturbations rather than the unhelpfulness of low resolution per se. To make the stated conclusion defensible, the paper would need either a controlled experiment in which resolution is varied while holding the ensemble-generation method fixed, or at minimum an explicit acknowledgment and discussion of the confound and a reformulation of the conclusions in terms of the two operational systems actually compared.
- [Section 4.1, first paragraph] The raw-forecast comparisons are affected by representativeness error, which the paper mentions as the cause of the non-monotonic CRPS curves but does not correct. Representativeness error tends to penalize coarser grids because the point observation is compared against a grid-box average or a smooth field, so the finding that raw (100,0) is significantly worse than raw (0,50) in Figures 4–7 may partly reflect this verification artifact rather than an intrinsic forecast-skill difference. The paper should either apply a representativeness correction (e.g., perturbing members as in Ben Bouallègue et al., 2020) or explicitly argue that the magnitude of the effect is too small to change the qualitative ranking. As written, the raw-verification results in Section 4.1 are presented without this necessary caveat.
- [Section 4, configuration selection] The hyperparameter choice (90 clusters, 60-day training window) is selected by minimizing mean CRPS over a validation period from 13 October 2023 to 31 May 2024, and the full verification period used for all reported skill scores is 3 September 2023 to 31 May 2024, which contains the validation period. This means the final evaluation is not a clean out-of-sample test: the hyperparameters were tuned on a subset of the very data used to compute the reported scores. The paper should either report results on a verification period that is disjoint from the tuning period, or present a sensitivity analysis showing that the substantive conclusions are stable across reasonable choices of cluster number and training-window length. Without this, the significance statements in Figures 4–17 may be optimistically biased.
minor comments (3)
- [Eq. (3.1)] The notation (M_L, M_H) is introduced as "M_L members for the low-resolution ENS extended forecasts and M_H for the high-resolution ENS predictions," but the combination notation "(M_L, M_H)" in Section 4 is used as (100,0), (0,50), etc. The order of the two indices is consistent, but it would help to state explicitly in Section 3.1 that the first entry is the number of low-resolution members and the second is the number of high-resolution members, since the later examples (50,32) and (50,1) could otherwise be misread.
- [Figure 2 and Figure 3 captions] The captions refer to "Model" and "Combination" in the legend, but the panels in Figures 2 and 3 show seven curves (raw ensemble plus three post-processing methods for three combinations). The distinction between the raw ensemble and the EMOS models is important; the legends are legible, but a brief note in the caption stating which line corresponds to the raw ensemble would improve readability.
- [Section 4.2.2] The claim in the text that "all models utilizing high-resolution predictions significantly outperform the reference forecast based solely on low-resolution members up to day 4" (based on Figure 13b) is stated before the caveat that this holds only for CRPS and MAE and not for Brier scores at higher thresholds. Since Figures 15–16 show substantial threshold dependence and even negative skill for some quantiles, the sentence in Section 4.2.2 should be qualified to avoid over-generalization.
Circularity Check
No circularity: the forecast skill comparisons are empirical and out-of-sample with respect to the fitted EMOS parameters; the noted caveats are statistical-validity concerns, not circular derivation.
full rationale
This is an empirical forecast evaluation paper. The central comparisons are between raw and EMOS-calibrated ensemble configurations scored against SYNOP observations on a verification period; EMOS coefficients are estimated on rolling training windows preceding each forecast date, and skill differences are assessed with bootstrap confidence intervals. The claims about resolution versus ensemble size are inferences from observed score differences, not consequences of the model equations (3.1). Self-citations (e.g., Lerch and Baran 2017 for semi-local clustering, Baran et al. 2019 for post-processing effects) supply methods and prior context, but the paper re-derives the relevant rankings from its own data. Two caveats are correctness risks, not circularity: (i) the semi-local hyperparameters (90 clusters, 60-day training) were selected on a validation period overlapping the full verification period, so the post-processed skill scores are partly selected rather than purely out-of-sample; and (ii) the headline causal attribution to 'spatial resolution' is confounded because the high- and low-resolution ensembles are separate operational systems differing in more than resolution. Neither caveat makes a derived quantity equal to an input by construction, so no circular step is identified.
Assumptions & free parameters
free parameters (3)
- EMOS coefficients a, b_H, b_L, c, d =
estimated by minimum CRPS per lead time and configuration
- Training window length =
60 days
- Number of clusters in semi-local approach =
90
assumptions (4)
- domain assumption Wind speed predictive distribution is left-truncated normal with location and scale linked affinely to ensemble mean and variance (Eq. 3.1).
- domain assumption Past forecast-observation pairs in the training window are representative of the verification period distribution.
- domain assumption SYNOP station observations are the ground truth and representativeness error is not corrected.
- ad hoc to paper The two ensemble systems differ mainly in resolution and ensemble size for the purpose of interpretation.
Cite this review
Pith. "Pith review of Statistical post-processing of operational dual-resolution wind-speed ensemble forecasts." pith.science (2026). https://pith.science/paper/6XQXLHQV
@misc{pith2026250615578,
author = {Pith},
title = {Pith review of: Statistical post-processing of operational dual-resolution wind-speed ensemble forecasts},
year = {2026},
howpublished = {\url{https://pith.science/paper/6XQXLHQV}},
note = {Machine review of arXiv:2506.15578}
}
read the original abstract
Weather forecasting presents several challenges, including the chaotic nature of the atmosphere and the high computational demands of numerical weather prediction models. To achieve the most accurate predictions, the ideal scenario involves the lowest possible horizontal resolution and the largest ensemble size. This study provides a detailed comparative analysis of the forecast skill of the raw and post-processed medium- and extended-range wind-speed ensemble forecasts of the European Centre for Medium-Range Weather Forecasts issued at 9 km and 36 km horizontal resolutions, respectively, and their various mixtures. We utilized the ensemble model output statistic approach for forecast calibration with three different spatial training data selection techniques. First, we investigate the performance of the 50-member medium-range and 100-member extended-range predictions - referred to as high and low resolution, respectively - and their 150-member dual-resolution combination. Further, we examine whether the performance of raw and post-processed low-resolution forecasts can be improved by incorporating high-resolution ensemble members. Our results confirm that all post-processed forecasts outperform the raw ensemble predictions in terms of probabilistic calibration and point forecast accuracy and that post-processing considerably reduces the differences between the various configurations. We also show that spatial resolution is superior to the ensemble size; augmenting a sufficiently large ensemble of high-resolution forecasts with low-resolution predictions does not necessarily result in a gain in forecast skill. However, our study also highlights the clear benefit of the other direction, namely, incorporating high-resolution members into low-resolution ensemble forecasts, where the most significant gains are observed in configurations with the highest number of high-resolution members.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Baran, S. and Lakatos, M. (2024). Clustering-based spatial interpolation of parametric postprocessing models. Weather Forecast. , 39(11):1591--1604
work page 2024
-
[2]
Baran, S. and Lerch, S. (2015). Log-normal distribution based emos models for probabilistic wind speed forecasting. Q. J. R. Meteorol. Soc. , 141(691):2289--2299
work page 2015
-
[3]
Baran, S., Leutbecher, M., Szab \'o , M., and Ben Bouall \`e gue, Z. (2019). Statistical post-processing of dual-resolution ensemble forecasts. Q. J. R. Meteorol. Soc. , 145(721):1705--1720
work page 2019
-
[4]
Baran, S., Szokol, P., and Szab \'o , M. (2021). Truncated generalized extreme value distribution-based ensemble model output statistics model for calibration of wind speed ensemble forecasts. Environmetrics , 32(6):paper e2678
work page 2021
-
[5]
Ben Bouall \`e gue , Z., Haiden, T., Weber, N. J., Hamill, T. M., and Richardson, D. S. (2020). Accounting for representativeness in the verification of ensemble precipitation forecasts. Mon. Weather Rev. , 148(5):2049--2062
work page 2020
-
[6]
Bentzien, S. and Friederichs, P. (2014). Decomposition and graphical portrayal of the quantile score. Q. J. R. Meteorol. Soc. , 140(683):1924--1934
work page 2014
-
[7]
Bremnes, J. B. (2020). Ensemble postprocessing using quantile function regression based on neural networks and Bernstein polynomials. Mon. Wea. Rev. , 148(1):403--414
work page 2020
-
[8]
Buizza, R. (2018). Ensemble forecasting and the need for calibration. In Vannitsem, S., Wilks, D. S., and Messner, J. W., editors, Statistical Postprocessing of Ensemble Forecasts , pages 15--48. Elsevier
work page 2018
Show all 32 references
-
[9]
Chen, J., Janke, T., Steinke, F., and Lerch, S. (2024). Generative machine learning methods for multivariate ensemble postprocessing. Ann. Appl. Stat. , 18(1):159--189
2024
-
[10]
Clark, M., Gangopadhyay, S., Hay, L., Rajagopalan, B., and Wilby, R. (2004). The schaake shuffle: A method for reconstructing space–time variability in forecasted precipitation and temperature fields. J. Hydrometeorol. , 5(1):243--262
2004
-
[11]
and Hemri, S
Dai, Y. and Hemri, S. (2021). Spatially coherent postprocessing of cloud cover ensemble forecasts. Mon. Weather Rev. , 149(12):3923--3937
2021
-
[12]
IFS Documentation CY49R1 -- Part V: Ensemble Prediction System
ECMWF (2024). IFS Documentation CY49R1 -- Part V: Ensemble Prediction System . ECMWF, Reading
2024
-
[13]
M., Richardson, D
Gasc\'on, E., Lavers, D., Hamill, T. M., Richardson, D. S., Ben Bouall\`egue , Z., Leutbecher, M., and Pappenberger, F. (2019). Statistical postprocessing of dual-resolution ensemble precipitation forecasts across Europe . Q. J. R. Meteorol. Soc. , 145(724):3218--3235
2019
-
[14]
Gneiting, T. (2011). Making and evaluating point forecasts. J. Amer. Statist. Assoc. , 106(494):746--762
2011
-
[15]
and Raftery, A
Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. J. Am. Stat. Assoc. , 102(477):359--378
2007
-
[16]
E., Westveld, A
Gneiting, T., Raftery, A. E., Westveld, A. H., and Goldman, T. (2005). Calibrated probabilistic forecasting using ensemble model output statistics and minimum CRPS estimation . Mon. Weather Rev. , 133 (5):1098--1118
2005
-
[17]
Hemri, S., Scheuerer, M., Pappenberger, F., Bogner, K., and Haiden, T. (2014). Trends in the predictive performance of raw ensemble weather forecasts. Geophys. Res. Lett. , 41(24):9197--9205
2014
-
[18]
Jordan, A., Krüger, F., and Lerch, S. (2019). Evaluating probabilistic forecasts with scoringRules . J. Stat. Softw. , 90(12):1--37
2019
- [19]
-
[20]
and Baran, S
Lerch, S. and Baran, S. (2017). Similarity-based semilocal estimation of post-processing models. J. R. Stat. Soc. , 66 (1):29--51
2017
-
[21]
and Ben Bouall \`e gue, Z
Leutbecher, M. and Ben Bouall \`e gue, Z. (2020). On the probabilistic skill of dual-resolution ensemble forecasts. Q. J. R. Meteorol. Soc. , 146(727):707--723
2020
-
[22]
Murphy, A. H. (1973). Hedging and skill scores for probability forecasts. J. Appl. Meteorol. , 12(1):215--223
1973
-
[23]
Politis, D. N. and Romano, J. P. (1994). The stationary bootstrap. J. Am. Stat. Assoc. , 89(428):1303--1313
1994
-
[24]
and Lerch, S
Rasp, S. and Lerch, S. (2018). Neural networks for postprocessing ensemble weather forecasts. Mon. Weather Rev. , 146(11):3885--3900
2018
-
[25]
L., and Gneiting, T
Schefzik, R., Thorarinsdottir, T. L., and Gneiting, T. (2013). Uncertainty quantification in complex simulation models using ensemble copula coupling. Statist.\ Sci. , 28(4):616--640
2013
-
[26]
M., Bright, J
Song, M., Yang, D., Lerch, S., Xia, X., Yagli, G. M., Bright, J. M., Shen, Y., Liu, B., and Liu, Xingli Mayer, M. J. (2024). Non-crossing quantile regression neural network as a calibration tool for ensemble weather forecasts. Adv. Atmos. Sci. , 41(7):1417--1437
2024
-
[27]
Szab \'o , M., Gasc \'o n, E., and Baran, S. (2023). Parametric postprocessing of dual-resolution precipitation forecasts. Weather Forecast. , 38(8):1313--1322
2023
-
[28]
Taillardat, M. (2021). Skewed and mixture of Gaussian distributions for ensemble postprocessing. Atmosphere , 12(8):paper 966
2021
-
[29]
Thorarinsdottir, T. L. and Gneiting, T. (2010). Probabilistic forecasts of wind speed: Ensemble model output statistics by using heteroscedastic censored regression. J. Roy. Stat. Soc. , 173A(2):371--388
2010
-
[30]
B., Demaeyer, J., Evans, G
Vannitsem, S., Bremnes, J. B., Demaeyer, J., Evans, G. R., Flowerdew, J., Hemri, S., Lerch, S., Roberts, N., Theis, S., Atencia, A., Ben Boual\`egue , Z., Bhend, J., Dabernig, M., De Cruz, L., Hieta, L., Mestre, O., Moret, L., Odak Plenkovi c , I., Schmeits, M., Taillardat, M....
2021
-
[31]
Veldkamp, S., Whan, K., Dirksen, S., and Schmeits, M. (2021). Statistical postprocessing of wind speed forecasts using convolutional neural networks. Mon. Weather Rev. , 149 (4):1141--1152
2021
-
[32]
Wilks, D. S. (2019). Statistical Methods in the Atmospheric Sciences . Elsevier, Amsterdam, 4th edition
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.