{"id":"0ea86c24-d732-44e1-a594-6ded101334bc","arxiv_id":"2506.15578","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For ECMWF wind-speed ensembles, high resolution beats large ensemble size, calibration shrinks the differences between configurations, and injecting high-resolution members into low-resolution forecasts improves skill.","lead":"This paper compares ECMWF wind-speed forecasts at 9 km and 36 km resolution and their mixtures, before and after statistical calibration. It finds that high-resolution forecasts beat larger low-resolution ensembles, and adding high-resolution members to low-resolution forecasts improves skill, while the reverse does not.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim conflates resolution with unmeasured differences between two operational ECMWF ensemble systems; the design never isolates resolution at fixed ensemble size.","rationale":"The reader's weakest_assumption identifies the same fundamental issue: the two ensemble systems are separate operational products, and the causal attribution to resolution is not identified. This concern is more load-bearing than the representativeness error because it affects both raw and post-processed comparisons and undermines the conceptual claim regardless of verification corrections. The paper never compares high- and low-resolution forecasts at fixed ensemble size (e.g., (50,0) vs (0,50)); the primary contrast (100,0) vs (0,50) varies both resolution and member count, and the mixtures in Section 4.2 add members from a different system. The proposed controlled experiment—using a dual-resolution ensemble with shared perturbations, or equivalently regridding high-res members to low-res—would directly test whether the observed skill gap is due to resolution or to other system differences. I keep the CONDITIONAL verdict because the empirical score comparisons are carefully executed and useful as a product comparison, but the abstract's causal statement requires qualification or supporting evidence. UNCHANGED reflects that the reader's conditional verdict is the appropriate classification.","tokens_in":18482,"tokens_out":9413,"duration_ms":109241,"concrete_test":"Obtain or generate a controlled dual-resolution ensemble in which the same set of perturbations is used at both resolutions (e.g., the experimental setup of Leutbecher and Ben Bouallègue, 2020), and re-run the key comparisons: (0,50) vs (100,0), (0,50) vs (100,50), and the Section 4.2 mixtures. If the resolution advantage and the failure of added low-res members to improve high-res forecasts persist in this setup, the operational-product confound is cleared. If they vanish, the paper's causal language is unwarranted and must be downgraded to a descriptive comparison of two operational products.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central causal claim—'spatial resolution is superior to ensemble size' (Abstract; Section 5)—rests on an uncontrolled comparison. The 50-member TCO1279 medium-range ENS and the 100-member TCO319 extended-range ENS are separate operational products (Section 2). They may differ in perturbation strategy, stochastic physics, coupling, and initialization, not just resolution and ensemble size. Crucially, the design never isolates resolution with ensemble size held constant: the headline comparison is (100,0) vs (0,50), varying both factors. The 'adding low-res members to high-res' contrast (0,50) vs (100,50) adds members from a different system, so the null result could reflect the inferiority of that system's perturbations rather than the unhelpfulness of low resolution per se. Similarly, the Section 4.2 gains from adding high-res members to a low-res base conflate resolution with system identity. Without evidence that ENS and ENS extended differ only in TCO truncation and membership, the attribution to 'spatial resolution' is unsupported. The paper does not acknowledge this confound or provide such evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares raw and EMOS-post-processed ECMWF wind-speed ensemble forecasts at two horizontal resolutions and varying ensemble sizes, using a truncated normal EMOS model with regional, local, and semi-local training-data selection. The main empirical findings are that the 50-member high-resolution (TCO1279) raw ensemble performs comparably to or better than the 150-member dual-resolution (100,50) ensemble, that the 100-member low-resolution (TCO319) raw ensemble is significantly worse on nearly all scores, and that after local EMOS post-processing, adding high-resolution members to a 50-member low-resolution base yields significant CRPS improvements up to about day 4, while adding low-resolution members to a high-resolution base is beneficial only up to day 2. The authors conclude that spatial resolution is superior to ensemble size and that the direction of mixing matters.","tokens_in":18669,"tokens_out":2475,"duration_ms":29097,"significance":"If the results hold in their intended generality, the paper would provide practically useful guidance for designing dual-resolution ensemble systems and for deciding how to allocate computational resources between resolution and ensemble size. The study is careful in several respects: it uses a large station network, applies proper scoring rules (CRPS, BS, QS) together with point-forecast metrics, and accompanies all skill scores with 95% block-bootstrap confidence intervals, allowing significance statements to be checked. The paper also transparently describes the EMOS setup and the configuration search. Its main limitation is that the comparison is not a controlled experiment: the high- and low-resolution systems are distinct operational products that differ in more than resolution and ensemble size, so the headline attribution of skill differences to spatial resolution is not actually established by the data.","major_comments":[{"comment":"The central claim that \"spatial resolution is superior to the ensemble size\" is not supported by the design, because the (0,50) and (100,0) systems are not controlled variants of a common ensemble system. Section 2 states that the 50-member medium-range ENS runs at TCO1279 and the 100-member extended-range ENS runs at TCO319; these are separate operational products that plausibly differ in perturbation strategy, stochastic physics, initialization, and model cycle in addition to resolution and membership. The comparison (100,0) vs (0,50) varies both resolution and ensemble size simultaneously, and the \"adding low-resolution members\" contrast (0,50) vs (100,50) adds members from a different system, so a null or negative result can reflect the inferiority of that system's perturbations rather than the unhelpfulness of low resolution per se. To make the stated conclusion defensible, the paper would need either a controlled experiment in which resolution is varied while holding the ensemble-generation method fixed, or at minimum an explicit acknowledgment and discussion of the confound and a reformulation of the conclusions in terms of the two operational systems actually compared.","section":"Abstract and Section 5"},{"comment":"The raw-forecast comparisons are affected by representativeness error, which the paper mentions as the cause of the non-monotonic CRPS curves but does not correct. Representativeness error tends to penalize coarser grids because the point observation is compared against a grid-box average or a smooth field, so the finding that raw (100,0) is significantly worse than raw (0,50) in Figures 4–7 may partly reflect this verification artifact rather than an intrinsic forecast-skill difference. The paper should either apply a representativeness correction (e.g., perturbing members as in Ben Bouallègue et al., 2020) or explicitly argue that the magnitude of the effect is too small to change the qualitative ranking. As written, the raw-verification results in Section 4.1 are presented without this necessary caveat.","section":"Section 4.1, first paragraph"},{"comment":"The hyperparameter choice (90 clusters, 60-day training window) is selected by minimizing mean CRPS over a validation period from 13 October 2023 to 31 May 2024, and the full verification period used for all reported skill scores is 3 September 2023 to 31 May 2024, which contains the validation period. This means the final evaluation is not a clean out-of-sample test: the hyperparameters were tuned on a subset of the very data used to compute the reported scores. The paper should either report results on a verification period that is disjoint from the tuning period, or present a sensitivity analysis showing that the substantive conclusions are stable across reasonable choices of cluster number and training-window length. Without this, the significance statements in Figures 4–17 may be optimistically biased.","section":"Section 4, configuration selection"}],"minor_comments":[{"comment":"The notation (M_L, M_H) is introduced as \"M_L members for the low-resolution ENS extended forecasts and M_H for the high-resolution ENS predictions,\" but the combination notation \"(M_L, M_H)\" in Section 4 is used as (100,0), (0,50), etc. The order of the two indices is consistent, but it would help to state explicitly in Section 3.1 that the first entry is the number of low-resolution members and the second is the number of high-resolution members, since the later examples (50,32) and (50,1) could otherwise be misread.","section":"Eq. (3.1)"},{"comment":"The captions refer to \"Model\" and \"Combination\" in the legend, but the panels in Figures 2 and 3 show seven curves (raw ensemble plus three post-processing methods for three combinations). The distinction between the raw ensemble and the EMOS models is important; the legends are legible, but a brief note in the caption stating which line corresponds to the raw ensemble would improve readability.","section":"Figure 2 and Figure 3 captions"},{"comment":"The claim in the text that \"all models utilizing high-resolution predictions significantly outperform the reference forecast based solely on low-resolution members up to day 4\" (based on Figure 13b) is stated before the caveat that this holds only for CRPS and MAE and not for Brier scores at higher thresholds. Since Figures 15–16 show substantial threshold dependence and even negative skill for some quantiles, the sentence in Section 4.2.2 should be qualified to avoid over-generalization.","section":"Section 4.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical study of two operational ECMWF ensemble products, with careful scoring and bootstrap inference. The main problem is the gap between the descriptive comparisons and the causal-sounding conclusion about \"spatial resolution versus ensemble size.\" Since the two systems differ in many respects, the headline claim is not established, but the paper could be brought to a publishable state by reframing the conclusions as comparisons of specific operational configurations and by either controlling for or openly discussing the confounds. The representativeness-error issue and the overlapping tuning/verification periods are additional load-bearing concerns that should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a careful empirical comparison of ECMWF's 9-km (TCO1279) and 36-km (TCO319) wind-speed ensemble forecasts and their mixtures, with a genuinely new systematic sweep of 1, 2, 4, 8, 16, and 32 high-resolution members added to a 50-member low-resolution base. The bootstrap confidence intervals are good practice, the EMOS implementation is standard and transparent, and the finding that post-processing shrinks differences between configurations is consistent with prior work and well shown. Operational centers will find the concrete CRPSS/MAES numbers useful when deciding how to allocate computational resources.\n\nThe soft spot is the central causal claim. The paper says 'spatial resolution is superior to the ensemble size,' but the design never isolates resolution at fixed ensemble size. The headline comparison (100,0) vs (0,50) varies resolution and member count simultaneously, and the two systems are distinct operational products—medium-range ENS versus extended-range ENS—that may differ in perturbation strategy, stochastic physics, coupling, and initialization. So the superior skill of (0,50) over (100,0) cannot be uniquely attributed to resolution. Likewise, the gains from adding high-res members to a low-res base, and the lack of gains from adding low-res members to a high-res base, could reflect the quality of the extended-range perturbations rather than resolution per se. The paper does not acknowledge this confound or offer evidence that the two systems are otherwise comparable. That is a real limitation, not a minor quibble.\n\nTwo smaller issues: the final configuration (90 clusters, 60-day training) is selected on a validation period that sits inside the verification period, so the reported skill scores are mildly optimistic. And raw verification is affected by representativeness error, which the paper mentions but does not correct; this could favor the finer grid. Both are worth noting but do not undermine the empirical comparisons themselves.\n\nThe paper is for operational forecast centers and researchers working on dual-resolution post-processing. The empirical content is valuable even if the causal framing is too strong. I'd send it to peer review—the referee should ask for a qualification of the conclusion (e.g., 'in these two operational systems, the higher-resolution system outperforms the lower-resolution one despite fewer members') rather than a broad statement about resolution versus ensemble size. The analysis deserves a serious referee; the authors just need to be held to the limits of their design.","headline":"Useful operational numbers on mixing ECMWF's 9-km and 36-km wind-speed ensembles, but the headline claim that 'spatial resolution is superior to ensemble size' overreaches the design: the two ensembles are separate operational products that differ in more than resolution and member count.","tokens_in":19226,"tokens_out":2150,"would_cite":true,"duration_ms":24504,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P12","62M20"],"pacs":[],"model":"deepseek-v4-flash","headline":"High resolution beats ensemble size in wind-speed forecasts","keywords":["ensemble model output statistics","dual-resolution ensemble forecasts","wind-speed forecasting","truncated normal distribution","probabilistic calibration","continuous ranked probability score","ensemble size","spatial resolution"],"falsifier":"Run a single numerical weather prediction model at 9 km with 50 members and at 36 km with 100 members, keeping perturbation method, physics, and initialization identical, and verify on a full season of wind-speed observations; if the 36-km, 100-member ensemble matches or beats the 9-km, 50-member ensemble on CRPS, the resolution-over-size claim fails. A quicker check is to verify the raw 9-km forecasts against observations upscaled to a 36-km footprint; if the advantage largely disappears, it was representativeness error rather than forecast information.","tokens_in":1678,"feed_emoji":"🌬️","tokens_out":5013,"duration_ms":111874,"temperature":0.7,"pith_summary":"The paper sets out to decide, for 10-m wind-speed forecasts up to 15 days ahead, whether running a high-resolution ensemble at 9 km with 50 members, a low-resolution ensemble at 36 km with 100 members, or a 150-member mixture of the two is the better use of computing power. It finds that the raw high-resolution forecast is essentially never worse than the raw mixture, while the raw 100-member low-resolution forecast is clearly worse on nearly every score. Calibrating with a truncated-normal EMOS model shrinks the gaps, and the mixture beats the high-resolution-only forecast only through about day 2. The clear win runs in the other direction: adding high-resolution members to a 50-member low-resolution base improves skill, with CRPS gains that remain significant up to about day 4. The practical consequence is that once an ensemble is large enough, increasing resolution is worth more than adding more low-resolution members.","feed_headline":"High resolution beats ensemble size for wind skill","feed_subtitle":"Adding coarse members to a fine ensemble does not help; adding fine members to a coarse one does.","key_machinery":"The carrying object is the truncated-normal EMOS predictive distribution $N_0^\\infty(\\mu,\\sigma^2)$ for wind speed, with location $\\mu = a + b_H^2 \\bar f_H + b_L^2 \\bar f_L$ and variance $\\sigma^2 = c^2 + d^2 S^2$, where $\\bar f_H$ and $\\bar f_L$ are the means of the high- and low-resolution ensemble members and $S^2$ is the variance of the combined ensemble. Parameters are estimated by minimizing the continuous ranked probability score over training data selected regionally, locally, or semi-locally through k-means clustering. This distribution is what converts raw ensemble output into a calibrated probability forecast, and the same scores used to fit it, namely CRPS, quantile score, and Brier score with stationary-bootstrap confidence intervals, are used to compare configurations.","core_discovery":"On its own terms, the paper shows that for wind speed the 50-member, 9-km ensemble forecast is at least as skillful as the 150-member dual-resolution forecast (50 high-resolution plus 100 low-resolution members) across CRPS, MAE, quantile, Brier, and RMSE scores, and that the 100-member, 36-km forecast is significantly worse on nearly all of them. After local EMOS post-processing every configuration improves, the differences between configurations shrink, and the dual-resolution forecast is significantly better than the high-resolution forecast only for the first two days. Augmenting a 50-member low-resolution ensemble with 1, 2, 4, 8, 16, or 32 high-resolution members helps for every configuration before calibration, with the largest gains coming from the largest number of high-resolution members; after calibration the gain remains significant in CRPS up to about day 4, after which low-resolution-only post-processing catches up.","pith_inferences":["Editorial inference: the paper's headline conclusion that resolution beats ensemble size is stated for two operational products that may differ in perturbation strategy, physics, and initialization; a controlled experiment varying only resolution and member count is needed before the rule is treated as general.","Editorial inference: a direct testable extension is to repeat the same mixture experiment for temperature, precipitation, or data-driven ensemble forecasts; the asymmetry found here may or may not survive.","Editorial inference: the raw-forecast verification is affected by representativeness error that favours the finer grid, so part of the apparent resolution advantage may be a measurement artifact; verifying against upscaled observations would separate the two.","Editorial inference: because post-processing reduces the gap, a cost-aware forecast design could use low-resolution ensembles for calibration of long lead times and reserve high-resolution computation for days 1-3."],"forward_implications":["For an operational centre, once a high-resolution ensemble has about 50 members, paying for 100 more low-resolution members is unlikely to improve wind-speed forecast skill; spending the same computing budget on resolution rather than low-resolution ensemble size is the better bet.","For a low-resolution ensemble, adding even one or two high-resolution members improves the raw forecast, and adding more extends the benefit; this is a cheap upgrade path for extended-range systems.","Statistical post-processing compresses, but does not remove, the configuration gap; after calibration, the choice of resolution and ensemble mix matters mainly for the first few days of the forecast.","The advantage of high-resolution members in post-processed mixtures is concentrated in CRPS and upper-tail quantiles, and disappears for high wind-speed thresholds, so users should not expect resolution mixing to fix rare strong-wind events.","For very short lead times (day 1-2), a post-processed dual-resolution forecast is the best choice; beyond about day 4, the low-resolution-only EMOS model is competitive."],"supporting_citations":[{"why":"Introduces the dual-resolution ensemble design and shows where mixing resolutions helps, the setup this paper transfers to wind speed.","marker":"Leutbecher and Ben Bouallègue (2020)"},{"why":"Earlier study of post-processed dual-resolution ensembles showing that calibration shrinks the gaps, the pattern this paper reproduces.","marker":"Baran et al. (2019)"},{"why":"Introduces EMOS with minimum-CRPS estimation, the calibration method used throughout.","marker":"Gneiting et al. (2005)"},{"why":"Provides the truncated normal EMOS model for wind speed, the predictive distribution in equation (3.1).","marker":"Thorarinsdottir and Gneiting (2010)"},{"why":"Supplies the semi-local clustering-based training-data selection compared in the study.","marker":"Lerch and Baran (2017)"},{"why":"Documents the representativeness error that the paper acknowledges in raw verification.","marker":"Ben Bouallègue et al. (2020)"}],"fun_headline_variants":["Resolution trumps ensemble size in wind forecasts","Adding high-res members boosts coarse wind forecasts","Fine grid beats large ensemble for wind skill","Wind prediction: resolution over size"],"cache_read_input_tokens":21376,"weakest_assumption_plain":"The whole comparison assumes the two forecast systems differ only in resolution and ensemble size; if hidden differences such as perturbation strategy, model physics, or initialization drive the skill gap, the conclusion that resolution beats size does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Resolution trumps ensemble size in wind forecasts","Adding high-res members boosts coarse wind forecasts","Fine grid beats large ensemble for wind skill","Wind prediction: resolution over size"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000228,"raw_usage":{"total_tokens":1508,"prompt_tokens":1011,"completion_tokens":497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":444}},"tokens_in":627,"tokens_out":497,"duration_ms":6402,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:52:34.957387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a single numerical weather prediction model at 9 km with 50 members and at 36 km with 100 members, keeping perturbation method, physics, and initialization identical, and verify on a full season of wind-speed observations; if the 36-km, 100-member ensemble matches or beats the 9-km, 50-member ensemble on CRPS, the resolution-over-size claim fails. A quicker check is to verify the raw 9-km forecasts against observations upscaled to a 36-km footprint; if the advantage largely disappears, it was representativeness error rather than forecast information.","supporting_citations":[{"cited_title":"and Ben Bouall \\`e gue, Z","cited_arxiv_id":null,"evidence_quote":"Introduces the dual-resolution ensemble design and shows where mixing resolutions helps, the setup this paper transfers to wind speed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier study of post-processed dual-resolution ensembles showing that calibration shrinks the gaps, the pattern this paper reproduces."},{"cited_title":"E., Westveld, A","cited_arxiv_id":null,"evidence_quote":"Introduces EMOS with minimum-CRPS estimation, the calibration method used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the truncated normal EMOS model for wind speed, the predictive distribution in equation (3.1)."},{"cited_title":"and Baran, S","cited_arxiv_id":null,"evidence_quote":"Supplies the semi-local clustering-based training-data selection compared in the study."},{"cited_title":"J., Hamill, T","cited_arxiv_id":null,"evidence_quote":"Documents the representativeness error that the paper acknowledges in raw verification."}],"review_version":1}