{"id":"76e6b339-ec8b-4403-95b3-6d61ef7118ff","arxiv_id":"2412.15832","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A stochastic transformer weather model trained with the almost fair CRPS loss produces ensemble forecasts that beat ECMWF's physics-based IFS ensemble on most medium-range variables and match it at subseasonal timescales.","lead":"At ECMWF, researchers trained a machine-learning weather model to simulate many possible futures at once by teaching it to minimize a probabilistic error score. In an eight-month test, the new model beat the physics-based IFS ensemble on most variables in the medium range and matched it at subseasonal time scales.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The subseasonal MJO comparison is confounded: AIFS-CRPS reforecasts are from 2018–2022 while the IFS reforecasts are from 2023 only, so the claimed MJO and subseasonal skill advantage could reflect different MJO years rather than model quality.","rationale":"The reader's CONDITIONAL verdict is appropriate, and the identified IC-perturbation mismatch is a genuine limitation: AIFS-CRPS is trained from the same deterministic ERA5 analysis for all ensemble members, then evaluated from perturbed IFS initial conditions, and Section 6 admits that first tests with a revised perturbation amplitude improve reliability (not shown). This means the reported spread-error relationship is contingent on an operational perturbation design the model never saw during training. That is a valid concern about calibration and about how much of the reported skill is tied to the particular initialization choice. However, as a threat to the strongest_claim, the medium-range comparison is still internally fair because both systems are initialized from the same operational IFS ensemble initial conditions and verified on the same dates. The more load-bearing weakness is the subseasonal comparison in Section 5.3, where the two systems are verified on different years. The MJO is a high-impact, intermittent phenomenon with strong interannual variability, and Figure 11's error bars cannot repair the fact that the verification years are disjoint. This is not an accusation of bad faith; it is a concrete sampling confound that the paper does not address. A matched-year reforecast comparison is the decisive test. The paper also does not release code or weights, but that affects reproducibility rather than the truth of the central claim. In summary, the medium-range claim is reasonably supported, while the subseasonal/MJO claim should be treated as provisional until a matched-sample evaluation is performed; hence the CONDITIONAL verdict should stand.","tokens_in":21571,"tokens_out":5694,"duration_ms":50694,"concrete_test":"Recompute the subseasonal evaluation with matched reforecast years: run the IFS reforecast system initialized on the same 2018–2022 start dates used for AIFS-CRPS (or equivalently generate AIFS-CRPS reforecasts for the 2023 IFS start dates), with 8 members, ERA5-based initial conditions, and the same verification data. Then recompute the MJO bivariate correlation, RMSE, ensemble spread (Figure 11) and the raw/anomaly ΔfCRPSS scorecards (Figure 10). If AIFS-CRPS retains its advantage under matched years, the subseasonal claim is supported; if the advantage shrinks, reverses, or falls within bootstrap uncertainty, the headline conclusion should be weakened to medium-range-only skill.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest_claim includes the statement that AIFS-CRPS “exhibits lower biases and increased Madden-Julian Oscillation (MJO) forecast skill compared to ECMWF’s operational subseasonal forecasting system.” This subseasonal claim rests on Section 5.3, where “the AIFS-CRPS subseasonal reforecast dataset is comprised of 46-day, 8-member ensemble forecasts initialized once per week over the period 2018-2022” but the comparison is against “operational IFS reforecasts produced during 2023, which we subset to use the same ensemble size and start dates as AIFS-CRPS.” Matching calendar dates across different years does not match the MJO state: MJO amplitude, phase, and propagation vary strongly by year, and a calendar-date match cannot remove that confound. The block-bootstrap significance testing and lead-time-dependent climatologies used in Figure 10/11 do not correct for the fact that the verification targets themselves differ between systems. Thus the raw and anomaly-based ΔfCRPSS differences in Figure 10 and the MJO correlations/RMSE in Figure 11 could be driven by the particular MJO events sampled in 2023 versus 2018–2022 rather than by genuine forecast-skill differences. Section 6 explicitly acknowledges the reforecast period is “short” and “out-of-sample,” but shortness alone is not the issue; the asymmetry of the two samples is. For the medium-range claim, the comparison is made over the same dates and initial conditions, so the reader's concern about training on deterministic ERA5 and evaluating with perturbed IFS initial conditions is a robustness/calibration issue rather than a direct threat to the medium-range score comparison. However, the subseasonal claim is a headline conclusion, and it is currently unsupported by a matched-sample comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AIFS-CRPS, a machine-learned global ensemble weather prediction model based on the AIFS architecture, trained by minimizing a newly introduced 'almost fair' CRPS loss (afCRPS) that blends the fair CRPS with the standard CRPS to avoid degeneracy in finite-precision training. The model injects Gaussian noise through conditional layer normalization to generate exchangeable ensemble members. The authors evaluate 15-day medium-range forecasts from two resolutions (O96 and N320) against the operational IFS ensemble, and 46-day subseasonal forecasts against IFS reforecasts. They report that AIFS-CRPS outperforms IFS for most upper-air and surface variables in the medium range, maintains realistic spectral variability without smoothing, and shows lower biases and improved MJO forecast skill at subseasonal lead times, while acknowledging remaining over-dispersion and stratospheric degradation.","tokens_in":21932,"tokens_out":7044,"duration_ms":59795,"significance":"If the results hold, AIFS-CRPS is a significant advance: it demonstrates that a single CRPS-style scoring-rule objective can yield a stochastic ML ensemble that rivals a high-resolution physics-based ensemble at a fraction of inference cost, with stable small-scale variability and useful subseasonal forecasts. The afCRPS formulation is a useful methodological contribution to scoring-rule-based training, as it addresses a real degeneracy of the fair CRPS while remaining simple to implement. The evaluation is unusually thorough for a preprint, including analysis-based and observation-based verification, significance testing in the scorecards, and explicit discussion of limitations such as stratospheric degradation and over-dispersion. The main caveat is that the subseasonal/MJO comparison rests on mismatched reforecast periods and perturbation methods, so those specific claims are not yet established; the medium-range comparison is much cleaner and supports the central skill claim more strongly.","major_comments":[{"comment":"The subseasonal and MJO comparisons are confounded by two asymmetries. The AIFS-CRPS reforecasts are initialized once per week over 2018–2022, while the IFS reforecasts are from 2023 only; matching calendar dates across years does not control for the different MJO states and verification targets sampled. In addition, the AIFS-CRPS subseasonal reforecasts use ERA5-EDA perturbations, whereas the IFS reforecasts use ERA5-EDA plus singular-vector perturbations, so the two systems differ in initial-condition perturbation methodology as well. The block-bootstrap significance testing and lead-time-dependent climatologies do not remove these confounds because the verification data themselves differ between the two samples. The conclusion in Section 7 that AIFS-CRPS exhibits 'increased Madden-Julian Oscillation (MJO) forecast skill compared to ECMWF's operational subseasonal forecasting system' therefore needs either a matched-period reforecast comparison or a substantially tempered claim.","section":"Section 5.3, Figures 10–11"},{"comment":"The model is trained by propagating a small ensemble from the same deterministic ERA5 analysis, but all reported medium-range results are obtained by initializing from perturbed IFS operational initial conditions. Section 6 states that 'First tests with a revised initial perturbation amplitude show improved reliability (not shown),' which indicates that the ensemble calibration, and hence the spread and CRPS results in Figures 5–9, may depend on the perturbation amplitude used at inference. Since that amplitude is an externally chosen parameter not seen in training, the paper should either report the sensitivity of the headline scores to the initial perturbation amplitude or explicitly mark the reliability results as conditional on the current operational perturbation scheme.","section":"Sections 2.3 and 6"},{"comment":"The medium-range skill comparison is based on a single eight-month period (1 February to 30 September 2024). While the scorecards include significance testing, the conclusion in Section 7 is stated without seasonal qualification: 'AIFS-CRPS forecast skill is higher than that of the 9 km physics-based IFS medium-range ensemble for most upper-air fields and surface variables.' This period does not include a Northern Hemisphere winter, so the claim should be restricted to the evaluated season or supported by additional seasons before being stated as a general result.","section":"Section 5.2, Figures 6–7"}],"minor_comments":[{"comment":"Two different equations are both numbered (1): the reference-field truncation update in Section 2.1 and the CRPS definition in Section 2.2. Please renumber to avoid ambiguity.","section":"Sections 2.1 and 2.2"},{"comment":"Section 2.1 states that the O32 grid has approximately 2.5° resolution, while Table 1 lists O32 as approximately 3.0° resolution; these values should be made consistent.","section":"Section 2.1 and Table 1"},{"comment":"The Figure 5 caption contains the typo 'nothern extra-tropics'; it should read 'northern extra-tropics'.","section":"Figure 5 caption"},{"comment":"The choice of α = 0.95 in the afCRPS loss is not accompanied by any sensitivity analysis; a brief study of α values near 1 would help justify the 'almost fair' approximation and show that the results are not sensitive to this hyperparameter.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate for the journal, but I would not accept the subseasonal/MJO claims as stated because of the year mismatch and perturbation-scheme mismatch in the reforecast comparison. The medium-range evaluation is much cleaner and could support the paper's central claim on its own. I recommend asking the authors to either rerun the IFS reforecasts over the same 2018–2022 period (or use a common set of start dates with contemporaneous verification) or remove the year-mismatched MJO comparison from the headline conclusions. The heavy reliance on self-cited ECMWF work is understandable given the operational context, but the paper should more clearly position the novelty of the afCRPS loss and the stochastic-noise architecture relative to the earlier AIFS papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth your time. The core result is that a stochastic transformer trained with the afCRPS loss—a small but sensible modification of the fair CRPS—produces ensemble forecasts that beat the IFS ensemble on most medium-range variables over an 8-month period. That is a real advance, and the evaluation is thorough: same dates, same initial conditions, significance testing, and honest discussion of over-dispersion and stratospheric degradation. The authors also clearly state the caveats, which I appreciate.\n\nThe new piece is the afCRPS loss and its use in training a stochastic weather model. That is a genuine methodological contribution, even if the architecture leans on prior work. The medium-range comparison is the strongest part of the paper and supports the claim that a single CRPS-style objective can produce a competitive ensemble.\n\nThe soft spot is the subseasonal section. The MJO and weekly-mean comparisons use AIFS-CRPS reforecasts from 2018–2022 but IFS reforecasts from 2023 only. That is not a minor issue; MJO activity varies strongly by year, so the claimed subseasonal skill advantage could easily be an artifact of sampling different MJO states. The block-bootstrap significance testing does not fix the mismatch because the verification targets differ between the two systems. The paper acknowledges the period is short, but the problem is not length—it is the asymmetry. The medium-range claim does not rest on this comparison, so the main result survives, but the headline “increased MJO skill” is currently unsupported.\n\nAlso worth noting: the model is trained on deterministic ERA5 analyses but evaluated with perturbed IFS initial conditions. The paper mentions that revised perturbation amplitudes improve reliability, so the current calibration is partly tuned to the wrong perturbations. That is a robustness concern, not a fatal one, but it should be stated more prominently. No code or weights are released, which limits reproducibility for a paper this dependent on specific training details.\n\nWho should read this: anyone working on ML ensemble forecasting or operational NWP. The medium-range results are important and the afCRPS idea is reusable. The subseasonal claims should be treated with caution until a matched-sample comparison is done.\n\nRecommendation: send it to peer review, but require the authors to either re-do the subseasonal evaluation with matched reforecast years or explicitly retract the MJO claim. The medium-range content is solid enough to warrant a serious referee.","headline":"Medium-range skill claims look credible; the subseasonal/MJO comparison is confounded by mismatched reforecast years and should not be taken at face value.","tokens_in":22568,"tokens_out":1856,"would_cite":true,"duration_ms":18167,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training a weather model with a proper score produces a stochastic ensemble that beats a 9-km physics-based system for most fields and lead times.","keywords":["ensemble forecasting","machine learning weather prediction","continuous ranked probability score","proper scoring rule","stochastic neural network","subseasonal forecast","Madden-Julian Oscillation","AIFS"],"falsifier":"Take the same trained AIFS-CRPS and issue forecasts from (a) unperturbed analyses, (b) operational IFS perturbed starts, and (c) starts with the revised perturbation amplitude mentioned in Section 6; if spread-error ratios and CRPS move materially across these, then the reported ensemble skill depends on perturbations the model was never trained on. Alternatively, compute afCRPS in float16 for a case with M-1 members exactly at the observation to see whether the degeneracy it avoids actually returns.","tokens_in":1755,"feed_emoji":"🌦️","tokens_out":1783,"duration_ms":67193,"temperature":0.7,"pith_summary":"The paper introduces AIFS-CRPS, a machine-learned ensemble weather model trained end-to-end to minimize the almost fair Continuous Ranked Probability Score (afCRPS). The central claim is that a single proper scoring objective, applied to a small training ensemble, is enough to teach a stochastic neural network to generate a well-calibrated, arbitrarily large ensemble of exchangeable forecast members. The authors report that this model scores better than the physics-based IFS ensemble for the majority of medium-range upper-air and surface variables, and that it stays competitive into the subseasonal range, with lower biases and better Madden-Julian Oscillation forecasts than the operational subseasonal system. If correct, this means the computationally heavy perturbed physics ensemble can be replaced by one cheap stochastic model evaluation per member.","feed_headline":"One score trains an ML ensemble that beats IFS on most fields","feed_subtitle":"AIFS-CRPS turns a proper scoring rule into a stochastic model competitive from days to weeks.","key_machinery":"The central object is the almost fair CRPS, defined as $\\text{afCRPS}_\\alpha = \\alpha\\,\\text{fCRPS} + (1-\\alpha)\\,\\text{CRPS}$, with $\\epsilon = (1-\\alpha)/M$, interpolating between the standard CRPS and the fair CRPS; it is rearranged as a double sum of non-negative terms so it remains safe under reduced precision. The second engine is noise-conditioned architecture: independent Gaussian noise per member is embedded and injected via conditional layer normalizations, making ensemble members exchangeable and the model stochastic at inference. Reference-field truncation, $x_{t+1} = U(D(x_t)) + f(x_t)$ with downsampling operator $D$ and upsampling operator $U$, prevents the accumulation of small-scale artifacts during autoregressive rollout.","core_discovery":"AIFS-CRPS is a variant of the AIFS transformer-based architecture in which all standard layer normalizations in the processor are replaced by conditional layer normalizations driven by random Gaussian noise. A small ensemble is propagated in parallel during training and scored against the deterministic ERA5 analysis with afCRPS, a convex combination of the fair CRPS and the standard CRPS that removes most of the finite-ensemble-size bias while avoiding the degeneracy that afflicts the fair score in low precision. In inference the model is rolled out autoregressively from each perturbed IFS initial condition, with each member using its own random seed. The paper's central discovery is that this simple training objective produces a stochastic model whose members keep realistic small-scale variability at long lead times and whose ensemble skill exceeds the 9-km IFS ensemble for most variables and lead times in the medium range, with strong subseasonal performance, particularly for the Madden-Julian Oscillation.","pith_inferences":["Extension: the hyperparameter $\\alpha$ in afCRPS offers a tunable knob between bias correction and stability; one could expect the optimal $\\alpha$ to vary with ensemble size, precision, and variable, a recipe transferable to other probabilistic machine-learning forecasters.","Extension: the paper notes that first tests with a revised initial perturbation amplitude improve reliability, which suggests a natural next step is to train with perturbed initial conditions so the model learns initial-condition uncertainty directly instead of aliasing it into model uncertainty.","Extension: the surrogate MJO index used here omits outgoing longwave radiation, yet still yields strong skill; evaluating a full RMM index with OLR would test whether the missing radiative channel matters for MJO prediction in this model.","Extension: since 46-day forecasts remain stable despite training on rolls of at most 72 hours, long-range skill appears to emerge from short-range training, hinting that similar objectives could support seasonal prediction with modest changes."],"forward_implications":["If the reported skill holds, a single afCRPS objective suffices to generate a full ensemble, without singular-vector perturbations or stochastic physics schemes.","Inference cost is one model evaluation per member per 6-hour step, so a 15-day forecast for one member takes about one minute at O96 and four minutes at N320 on an A100 GPU, making large ensembles cheap to produce.","Subseasonal predictions at two to six weeks are competitive with or better than the operational IFS subseasonal system, especially for the Madden-Julian Oscillation, despite training only on short-range rollout steps.","Raising input resolution from O96 to N320 improves most surface variables, suggesting the same training recipe can be pushed to higher resolutions and additional variables.","The system still depends on physics-based analyses for training data and initialization, so it refines rather than replaces the current data-assimilation infrastructure."],"supporting_citations":[{"why":"Introduces ensemble-size-adjusted discrete and continuous ranked probability scores, the starting point for the fair CRPS.","marker":"[Ferro et al., 2008]"},{"why":"Defines the fair CRPS that afCRPS modifies, supplying the bias-correction idea.","marker":"[Ferro, 2013]"},{"why":"Quantifies how suboptimal finite ensemble sizes are and motivates fair-score corrections in the training loss.","marker":"[Leutbecher, 2019]"},{"why":"Describes the AIFS architecture and training setup that AIFS-CRPS inherits and adapts.","marker":"[Lang et al., 2024a]"},{"why":"Provides the ERA5 reanalysis used as training data and verification target.","marker":"[Hersbach et al., 2020]"},{"why":"Documents the operational IFS initial perturbation methodology used to initialize AIFS-CRPS at inference.","marker":"[Lang et al., 2021]"},{"why":"Defines the 9-km IFS ensemble baseline that AIFS-CRPS is compared against.","marker":"[Lang et al., 2023]"},{"why":"Establishes the spread-RMSE reliability framework used to evaluate ensemble calibration.","marker":"[Leutbecher and Palmer, 2008]"}],"fun_headline_variants":["CRPS-trained ML ensemble outperforms IFS on most scores","Stochastic AIFS beats IFS with almost fair CRPS loss","Proper scoring rule yields superior ML weather ensemble","New loss function boosts ML weather ensembles"],"cache_read_input_tokens":24576,"weakest_assumption_plain":"The model is trained on a small ensemble initialized from a single deterministic analysis, but evaluated from perturbed IFS initial conditions it never saw during training; the paper notes that revised perturbation amplitudes change reliability, so the unseen perturbation statistics are load-bearing for the reported calibration and skill.","fun_headline_variants_meta":{"raw":{"variants":["CRPS-trained ML ensemble outperforms IFS on most scores","Stochastic AIFS beats IFS with almost fair CRPS loss","Proper scoring rule yields superior ML weather ensemble","New loss function boosts ML weather ensembles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000625,"raw_usage":{"total_tokens":2906,"prompt_tokens":969,"completion_tokens":1937,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":1874}},"tokens_in":585,"tokens_out":1937,"duration_ms":13208,"temperature":1.0,"reasoning_tokens":1874,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:03:18.381277+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same trained AIFS-CRPS and issue forecasts from (a) unperturbed analyses, (b) operational IFS perturbed starts, and (c) starts with the revised perturbation amplitude mentioned in Section 6; if spread-error ratios and CRPS move materially across these, then the reported ensemble skill depends on perturbations the model was never trained on. Alternatively, compute afCRPS in float16 for a case with M-1 members exactly at the observation to see whether the degeneracy it avoids actually returns.","supporting_citations":[],"review_version":1}