Pith. sign in

REVIEW 2 cited by

WeatherReal: A Benchmark Based on In-Situ Observations for Evaluating Weather Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09371 v1 pith:UBCBDGCL submitted 2024-09-14 physics.ao-ph cs.LG

classification physics.ao-phcs.LG
keywords modelsweatherobservationsweatherrealforecastingin-situnumericalai-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, AI-based weather forecasting models have matched or even outperformed numerical weather prediction systems. However, most of these models have been trained and evaluated on reanalysis datasets like ERA5. These datasets, being products of numerical models, often diverge substantially from actual observations in some crucial variables like near-surface temperature, wind, precipitation and clouds - parameters that hold significant public interest. To address this divergence, we introduce WeatherReal, a novel benchmark dataset for weather forecasting, derived from global near-surface in-situ observations. WeatherReal also features a publicly accessible quality control and evaluation framework. This paper details the sources and processing methodologies underlying the dataset, and further illustrates the advantage of in-situ observations in capturing hyper-local and extreme weather through comparative analyses and case studies. Using WeatherReal, we evaluated several data-driven models and compared them with leading numerical models. Our work aims to advance the AI-based weather forecasting research towards a more application-focused and operation-ready approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model

    physics.ao-ph 2025-10 conditional novelty 5.0 of 10

    Combining HiRA neighborhood verification with threshold-weighted CRPS shows that AI-vs-NWP rankings for extreme precipitation depend strongly on neighborhood size.

  2. EPT-2 Technical Report

    cs.LG 2025-07 conditional novelty 4.0 of 10

    EPT-2 is a 9 km global AI weather model reported to beat Aurora, IFS HRES, and the ECMWF ENS mean over 0 to 240 hours on energy-relevant variables.

Pith tools