Pith. sign in

REVIEW 3 major objections 3 minor 5 cited by

Numerical models outperform AI weather forecasts of record-breaking extremes

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Against the five leading AI weather models, the operational numerical model HRES still produces smaller errors for record-breaking heat, cold, and wind extremes at nearly all lead times.

desk verdict Plausible, well-scoped benchmark result that AI models trail HRES on record-breaking extremes, but the abstract alone can't rule out metric or baseline artifacts. read the letter →

arxiv 2508.15724 v1 pith:XIKLS42Y submitted 2025-08-21 physics.ao-ph cs.AIstat.AP

classification physics.ao-phcs.AIstat.AP
keywords AIweatherforecastingrecord-breakingextremesnumericalpredictionHRESGraphCastPangu-Weatherextrapolationforecastverification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the supposed AI revolution in weather forecasting does not yet extend to the most dangerous events: record-breaking heat, cold, and wind. Comparing five state-of-the-art AI models (GraphCast, GraphCast operational, Pangu-Weather, Pangu-Weather operational, and Fuxi) against the European Centre's high-resolution numerical model HRES, the authors find that HRES consistently produces smaller forecast errors for record-breaking extremes across nearly all lead times. The AI models also underestimate the frequency and intensity of these records, underpredict hot records and overpredict cold records, with errors growing as the record exceedance gets larger. If true, the result identifies a systematic extrapolation failure: AI models trained on past weather cannot generalize beyond the largest events in their training data. This matters because record-breaking events are becoming more frequent in a warming climate and are exactly the events that trigger early warnings and disaster response.

What carries the argument

The central object is the record-breaking extreme: a weather event that exceeds the previous observed record at a given location and time of year for temperature or wind. The key mechanism is to condition forecast-verification statistics on whether the observation sets a new record, and on how much it exceeds the old record, instead of averaging over all weather. This conditioning exposes the AI models' tendency to underpredict the frequency and intensity of the most extreme events, with errors growing as the record exceedance grows.

What would settle it

A decisive test: initialise HRES, GraphCast, and Pangu-Weather from the same analysis time for a well-documented record event not in the AI training data (e.g., the June 2021 Pacific Northwest heatwave) and compare 2-metre temperature forecast errors at the record location across lead times. The paper's claim predicts the AI models' errors are larger than HRES's and grow with record exceedance; observing the opposite would refute it.

Watch

Extended reading notes

Core claim

Using a unified verification framework, the authors show that when forecasts are evaluated on record-breaking extremes rather than on all weather, the numerical model HRES outperforms all five AI models. The AI models' forecast errors are larger for record-breaking heat, cold, and wind at nearly all lead times; the AI models tend to underforecast both how often records occur and how far records are exceeded; and they show an asymmetric bias: hot records are underpredicted and cold records overpredicted, with errors increasing with the magnitude of the record exceedance. The authors interpret this as evidence that AI models extrapolate poorly beyond their training domain, and conclude that th

Load-bearing premise

The comparison assumes that record-breaking events are defined and identified identically for every model, using the same verification reference and event-detection rules, and that the AI models' error metrics are not inflated by their training climatology; the abstract provides no detail on the event definition, verification data, or uncertainty intervals.

Editorial extensions

If this is right

  • AI weather models should not be deployed as stand-alone forecast systems for extreme-event early warning until their extrapolation behaviour is redesigned and verified.
  • Aggregate benchmark skill (e.g., RMSE over all dates) can conceal a systematic failure on the tail of the distribution, so operational readiness must be evaluated on record-breaking subsets.
  • Hybrid systems that use AI to post-process or emulate numerical models may be safer than pure AI forecasts for extremes, because the numerical component anchors the tail.
  • Because climate change makes hot records more frequent, the AI models' underprediction of hot records will cause more missed warnings as warming continues.
  • Growing error with record exceedance magnitude means the severity of the most extreme events will be underestimated most, which can understate impact.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the asymmetric bias (underpredicting hot records, overpredicting cold records) is consistent with AI models anchoring to the mean of their training distribution; a direct test would compare forecast bias to the climatological anomaly distribution in each model's training data.
  • Beyond the paper: because AI models are trained on historical data, their extrapolation failure may shrink as training sets are updated to include recent record events; the paper's findings may describe the current generation of training data rather than an eternal limitation.
  • Beyond the paper: the result implies a public-safety standard: AI forecasts should be used only where their tail error profile is independently verified, not on the basis of average skill metrics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript addresses whether AI-based weather forecast models can extrapolate to record-breaking extremes, using the numerical HRES model as a benchmark. The abstract claims that HRES consistently outperforms five state-of-the-art AI models (GraphCast, GraphCast operational, Pangu-Weather, Pangu-Weather operational, and Fuxi) for record-breaking heat, cold, and wind across nearly all lead times, and that AI models underestimate both the frequency and intensity of such events, with errors growing for larger record exceedance. The abstract concludes that AI models are currently limited in high-stakes early-warning applications. No verification methodology, event definitions, sample sizes, uncertainty intervals, or lead-time details are presented in the abstract.

Significance. If the central claim holds, the result has substantial practical significance for operational weather forecasting and for the safe deployment of AI emulators, particularly in disaster warning where record-breaking extremes are the most consequential. The paper is valuable for naming specific AI models and for focusing on extrapolation beyond the training distribution, an issue often overlooked in benchmark comparisons. However, the evaluation of significance is impossible from the abstract alone: the claim must be substantiated with a transparent verification protocol. The abstract gives no way to assess event counts, statistical confidence, or whether the result is robust to metric and reference-dataset choices.

major comments (3)
  1. [Abstract] The central claim, that HRES 'consistently outperforms' the AI models across 'nearly all lead times', lacks any supporting verification methodology. The abstract does not define how record-breaking events are identified, what observational or reanalysis reference is used, what error metric is computed, how many events are considered, or what the uncertainty intervals are. Without this information, the result could be an artifact of a particular metric or a small, non-representative sample. The full paper must report this information for the claim to be load-bearing.
  2. [Abstract] The comparison may be confounded by the training data used for AI models. If the reference baseline for defining record-breaking events is a reanalysis (e.g., ERA5) that also served as the training target for these models, then 'record-breaking' means that the verification target lies beyond the maximum label seen during training. The reported pattern—AI models underpredict hot records and overestimate cold records, with growing errors for larger exceedance—is a signature of regression to the mean under an L2 training loss when evaluated selectively on extremes. The authors should state the reference dataset explicitly and demonstrate that the conclusion holds when using an independent observational reference and alternative climatological baselines.
  3. [Abstract] The claim of 'growing errors for larger record exceedance' is reported without quantitative support or confidence intervals. If the number of record-breaking events is small, a few large outliers could dominate the averages. The abstract does not report event counts, the distribution of errors, or any significance testing. The full paper must show that the consistency across models and lead times is statistically robust rather than driven by a handful of events.
minor comments (3)
  1. [Abstract] The term 'consistent' is vague. Does it mean that HRES outperforms on the majority of lead times, on all regions, or with statistical significance? A precise consistency criterion should be defined.
  2. [Abstract] The abstract lists both operational and non-operational versions of GraphCast and Pangu-Weather. The difference between these versions should be clarified, including whether the same initial conditions and lead times are used for fair comparison.
  3. [Abstract] The verification reference is not mentioned. Including a short phrase such as 'verified against ERA5 and station observations' would help the reader interpret the claims.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the claim is an empirical benchmark of model forecasts against external verification data.

full rationale

This is an abstract-only review. The paper's central claim is that the numerical model HRES outperforms several AI weather models for record-breaking extremes. The claimed comparison is between model forecasts and external observations or reanalysis, not between a fitted parameter and its own input. No equations, fitted constants, or self-citations are visible in the abstract. The skeptical concern about the record-breaking baseline being defined from the same reanalysis used to train the AI models is a possible verification-design issue, but it is not exhibited in the text and would require evidence of a specific reduction (e.g., that record exceedance is defined by the AI models' training labels). The abstract's statements about underpredicting hot records and overestimating cold records are empirical findings about model behavior, not definitions in disguise. Therefore, no circular step can be identified from the available text, and the appropriate score is 0.

Assumptions & free parameters 1 free parameters · 1 assumptions · 0 invented entities

The abstract only allows identification of the record-event threshold as a hand-chosen input and the implicit truthfulness of the verification reference. No new entities or fitted constants are mentioned.

free parameters (1)
  • Record-breaking threshold
    Definition of what counts as a record-breaking event (e.g., exceeding a previous climatological maximum) is a hand-chosen criterion that determines which cases enter the comparison; not visible in the abstract.
assumptions (1)
  • domain assumption The observational or reanalysis reference used to compute forecast errors is accurate for extreme heat, cold, and wind.
    The central comparison treats a reference dataset as truth; if that reference is itself uncertain for records, the error differences could be artifacts. This assumption is implicit in the abstract's use of 'forecast errors' for record-breaking events.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Numerical models outperform AI weather forecasts of record-breaking extremes." pith.science (2026). https://pith.science/paper/XIKLS42Y

@misc{pith2026250815724,
  author       = {Pith},
  title        = {Pith review of: Numerical models outperform AI weather forecasts of record-breaking extremes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XIKLS42Y}},
  note         = {Machine review of arXiv:2508.15724}
}
read the original abstract

Artificial intelligence (AI)-based models are revolutionizing weather forecasting and have surpassed leading numerical weather prediction systems on various benchmark tasks. However, their ability to extrapolate and reliably forecast unprecedented extreme events remains unclear. Here, we show that for record-breaking weather extremes, the numerical model High RESolution forecast (HRES) from the European Centre for Medium-Range Weather Forecasts still consistently outperforms state-of-the-art AI models GraphCast, GraphCast operational, Pangu-Weather, Pangu-Weather operational, and Fuxi. We demonstrate that forecast errors in AI models are consistently larger for record-breaking heat, cold, and wind than in HRES across nearly all lead times. We further find that the examined AI models tend to underestimate both the frequency and intensity of record-breaking events, and they underpredict hot records and overestimate cold records with growing errors for larger record exceedance. Our findings underscore the current limitations of AI weather models in extrapolating beyond their training domain and in forecasting the potentially most impactful record-breaking weather events that are particularly frequent in a rapidly warming climate. Further rigorous verification and model development is needed before these models can be solely relied upon for high-stakes applications such as early warning systems and disaster management.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI-boosted rare event sampling to characterize extreme weather

    physics.ao-ph 2025-10 conditional novelty 7.0 of 10

    AI+RES uses AI weather-forecast ensembles as a guide for rare-event simulation, yielding accurate return-period statistics for 1-in-50,000-year heatwaves at roughly 100× lower computational cost.

  2. Integrating GNSS-Derived Zenith Wet Delay into a Weather Foundation Model Improves Precipitation Forecasting

    physics.ao-ph 2026-07 conditional novelty 6.5 of 10

    Integrating GNSS-derived Zenith Wet Delay into Aurora improves six-hour precipitation skill, with an 8.8% ETS gain at the 99th percentile and a more realistic power spectrum.

  3. Assessing Extrapolation of Peaks Over Thresholds with Martingale Testing

    stat.ME 2025-12 conditional novelty 6.0 of 10

    A gambling-style test over the largest observed rainfalls selects the extreme-value threshold that won the EVA2025 challenge — except on one target, where the authors overrode the game and paid for it.

  4. Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model

    physics.ao-ph 2025-10 conditional novelty 5.0 of 10

    Combining HiRA neighborhood verification with threshold-weighted CRPS shows that AI-vs-NWP rankings for extreme precipitation depend strongly on neighborhood size.

  5. Understanding and Utilizing Dynamic Coupling in Free-Floating Space Manipulators for On-Orbit Servicing

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    A coupling metric from an SVD of the base-arm interaction matrix guides trajectory optimization for free-floating space manipulators, aiming for more efficient on-orbit servicing.

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.