Pith. sign in

REVIEW 1 cited by

Analyzing and Exploring Training Recipes for Large-Scale Transformer-Based Weather Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19630 v1 pith:W6H2BOM2 submitted 2024-04-30 cs.LG

classification cs.LG
keywords trainingforecastmodelskillexploringinvestigatemodelsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified SwinV2 transformer on ERA5 data, and find that it attains superior forecast skill when compared against IFS. We present some ablations on key aspects of the training pipeline, exploring different loss functions, model sizes and depths, and multi-step fine-tuning to investigate their effect. We also examine the model performance with metrics beyond the typical ACC and RMSE, and investigate how the performance scales with model size.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PEAR: Equal Area Weather Forecasting on the Sphere

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A transformer weather model operating natively on the equal-area HEALPix grid beats an equiangular-grid counterpart at longer lead times with 2.6x fewer parameters.

Pith tools