REVIEW 3 major objections 6 minor 12 references
TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale
T0 review · 3 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read TCBench claims that neural weather models forecast tropical cyclone tracks skillfully, but intensity — especially rapid intensification — only becomes skillful after post-processing or task-specific training, which this benchmark is built t
desk verdict Useful benchmark infrastructure with a load-bearing protocol leak around the postprocessed intensity and RI claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the argument is the benchmark protocol itself. TCBench defines forecasting as predicting the evolution of an already-existing storm, using IBTrACS as ground truth and retaining 6-hourly verification keys; any model that fails to output a storm at a valid key is assigned the persistence forecast. Deterministic track quality is measured by direct, cross-track, and along-track position errors; ensemble quality by a fair version of CRPS with Haversine distances; and rapid intensification is cast as binary classification (24h wind gain >= 30 kt) scored by CSI and PSS. The postprocessing baselines use a storm-centred patch of AI model fields plus the initial storm state
What would settle it
Re-run the postprocessed PanguWeather intensity and RI evaluation with the standard 6-hourly IBTrACS verification grid, using models trained only on 2017-2020 data and validated on 2021-2022, and report where the 48-h RI CSI and overall PSS land; if CSI drops to near zero or PSS turns negative, the paper's central intensity claim is refuted.
Extended reading notes
Core claim
On TCBench's 2023 test set, the paper's central finding is that track and intensity skill split: deterministic AI track forecasts at 1-5 days are competitive with or better than the physics-based GEFS ensemble, while raw AI intensity forecasts are worse than persistence at short lead times and only become useful after 24-48 h. The explicitly post-processed PanguWeather baseline - a neural model paired with a distributional regression of 24h intensification - matches or beats GEFS on absolute wind and pressure error, and it is the only evaluated baseline with positive rapid-intensification skill, with 48-h CSI around 0.12 and overall PSS about +0.05. The paper concludes that neural weather mo
Load-bearing premise
Everything positive about intensity and rapid intensification rests on the post-processed PanguWeather baseline, which the paper's own Fig. 3 caption says deviates from the benchmark protocol - it uses a non-6-hour lead grid and was trained/validated on different years than the other baselines, and those years are never stated; if the postprocessor saw 2023 during training, the claim that only post-processed AI captures RI collapses.
Editorial extensions
If this is right
- Neural weather models can be used as deterministic 1-5 day TC track guidance at skill comparable to physics-based ensembles, at a fraction of computational cost.
- Raw AI intensity output should not be plugged directly into operational warnings; post-processing or task-specific training becomes a recommended step.
- A single post-processed neural model matched GEFS intensity skill and was the only evaluated system with RI skill, so data-driven RI guidance at 48-96 h lead is feasible.
- The benchmark's persistence-filled, IBTrACS-keyed protocol gives future TC forecasting models a standard, reproducible scoring system, removing tracking-algorithm and metric differences as confounders.
- If post-processed RI skill generalizes beyond the 2023 test year, such models could serve as early-warning flags for the most destructive, rapidly intensifying storms.
Reading between the lines
- Because the postprocessor converts coarse AI fields into intensity skill without a high-resolution physics model, the result suggests operational centers could achieve useful intensity guidance with modest compute; the paper leaves this operational implication implicit.
- The strong RI signal from one postprocessed model implies raw AI fields contain latent information about RI that is masked by intensity bias. A testable extension is to apply the same postprocessing to other neural models (AIFS, GenCast) and check whether RI skill appears consistently.
- The persistence-fill convention means a model with poor storm-coverage gets evaluated partly against persistence; the paper's no-fill appendix shows coverage varies by model, so a useful extension would be reporting coverage-adjusted skill to separate tracking recall from forecast accuracy.
- Since TCBench frames forecasting after storm formation, genesis prediction is outside scope; a natural companion benchmark would score how well AI models develop storms at the right time and place, completing the warning chain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TCBench proposes a standardized benchmark for global 1–5-day forecasts of tropical cyclone track and intensity from existing cyclones, using IBTrACS as ground truth, TIGGE ensemble models and several neural weather models as baselines, TempestExtremes for tracking, and a suite of deterministic/probabilistic/rare-event metrics (DPE, ATE, CTE, CRPS, CSI, PSS). On a 2023 test year, the paper reports three headline findings: AI models beat persistence on track forecasts; raw AI intensity forecasts are unskillful until longer leads; and a postprocessed PanguWeather baseline (PANGU_POST_ANN) attains intensity skill comparable to or better than GEFS and is claimed to be the only baseline with rapid-intensification skill.
Significance. If the protocol were airtight, TCBench would be a valuable community resource: it unifies heterogeneous data sources under a storm-relative evaluation design, provides a defined holdout year, releases code and data (subject to the review anonymization), and makes a concrete comparison between physics-based ensembles and neural weather models. The track results are credible and consistent with prior work: all models beat persistence, AIFS is the best deterministic tracker in this sample, and GEFS retains an edge in probabilistic track skill. The intensity/RI finding is the paper's main novel positive claim, but it currently rests on an off-protocol postprocessor whose training years are never stated, and the 'only' RI claim is contradicted by the paper's own scorecard. With those issues repaired, the benchmark would be a useful contribution; as written, the central intensity/RI conclusions are not yet testable under the benchmark's own rules.
major comments (3)
- [§6, Fig. 4, Fig. 9] Load-bearing protocol omission. The Fig. 3 caption states that the post-processing model (e.g., PANGU_POST_ANN) 'deviates from our protocol (non-6 h lead grid; trained/validated on different years)', but the manuscript never states which years were used to train or validate PANGU_POST_ANN, PANGU_POST_MLR, or PANGU_POST_UNET. This directly collides with §3.3, which requires that 'the year 2023 be left for testing' because baselines 'are not trained or tuned on this year'; §D.5's assertion that 'All baselines are trained and evaluated consistently' does not repair the omission. Since the positive intensity and RI results in §6 are carried by these postprocessors, the claims 'Postprocessing AI models yields skillful intensity predictions' and 'Only the postprocessed AI models capture RI' are not verifiable as written. Please report the training/validation years for each postprocessor, demon
- [§4, Eq. (1)] The statement 'Only the postprocessed AI models capture RI' is contradicted by the paper's own scorecard. Fig. 4 reports FNV3 with RI CSI = 0.095 at 48 h, 0.037 at 72 h, and 0.026 at 96 h, and Fig. 9 shows nonzero overall CSI/PSS values for several models. FNV3 is described in §D.3 as fine-tuned on IBTrACS, and the abstract explicitly allows 'task-specific training' as a route to intensity skill. The 'only' claim needs a precise definition of 'skill' (e.g., a threshold, lead-time range, or statistical significance criterion) or should be removed. In addition, the reported CSI values are small and no confidence intervals or significance tests are provided; given the rarity of RI events, sampling uncertainty should be quantified before claiming any model 'captures' RI.
- [§1, §6] Track CRPS is defined by replacing the absolute differences in Eq. (1) with Haversine distances. This is not automatically a proper scoring rule: for the energy-score representation to be proper, the distance must be of negative type. The manuscript provides no citation or proof that Haversine/great-circle distance has this property. Please either cite a result for spherical geodesic distance or use a chordal distance for which negative-definiteness is known, and state what this implies for the probabilistic track comparisons in Fig. 3b.
minor comments (6)
- [§1, §6] The introduction says 'we demonstrate that neural weather models can skillfully forecast TC intensity up to 5 days ahead, particularly when combined with observational data and processed through tools provided in TCBench.' This is at odds with the more careful abstract and §6 framing, where raw AI models are unskillful and only postprocessed models show skill; please rephrase to avoid implying raw AI intensity skill.
- [§6, Fig. 4] The text says 'all models produce forecasts of Vmax and pmin that are worse than the persistence baseline for shorter lead times,' but PANGU_POST_ANN in Fig. 4 already beats persistence at 24 h for both wind and pressure. This inconsistency should be clarified.
- [§3.2, §D.4] The data inclusion criteria in §3.2 require 6-hourly forecast data, yet the postprocessors are admitted with a 'non-6 h lead grid.' Please explain how the postprocessors satisfy the inclusion criteria, or explicitly list them as an exception.
- [Abstract, §D.3] Terminology is inconsistent: the abstract uses 'AIWP' and 'FourCastNet v2,' while the text uses 'neural weather models' and 'FourCastNetV2'; please standardize.
- [D.4] The 50-member postprocessor ensemble is described as sampled from a parametric Gaussian distribution, but the sampling procedure, random seed, and any clipping details are not given. For a benchmark claiming reproducibility, this should be specified.
- [Throughout] Several typos and incomplete sentences should be corrected: 'Gneiting & and, 2007' (Eq. 3 reference), 'focus lesson process-based assessments' (Sec. 1), 'we as provide scores' (Sec. 2), and 'ERA-5' inconsistently styled. A careful proofread is needed.
Circularity Check
Evaluation protocol is externally grounded; the only caveat is a minor self-citation burden in the off-protocol postprocessed intensity baselines, which is a verification concern rather than demonstrated circularity.
full rationale
TCBench is an evaluation benchmark, not a fitted derivation: tracks and intensities come from external physics-based and neural models, tracking is performed with the fixed TempestExtremes parameter table, and the ground truth is the independent IBTrACS record. The 2023 test year is explicitly separated from the recommended training/validation years: 'We require that the year 2023 be left for testing, given that the models we include as baselines are not trained or tuned on this year' (Sec. 3.3). No metric is defined in terms of a model output, and no reported 'prediction' is the fitted value itself. The only self-citation burden is the postprocessing baselines that carry the intensity and RI claims: they 'follow (Gomez et al., 2025)' (Sec. 5) and use 'the details provided for these architectures and associated hyperparameters provided by Gomez et al. (2025)' (App. D.4), while the Fig. 3 caption concedes 'The post-processing model (dotted; e.g., PANGU_POST_ANN) deviates from our protocol (non-6 h lead grid; trained/validated on different years)'. The training years of these postprocessors are never stated, so the headline 'Only the postprocessed AI models capture RI' rests on a baseline whose 2023 holdout status is not documented. This is a legitimate benchmark-integrity and reproducibility caveat, but it is not demonstrated circularity: evaluating a trained postprocessor on a held-out year is not equivalent, by construction, to its training target unless 2023 was in the training set, which the paper does not claim. No equation or definition in the paper reduces a predicted quantity to its own input, so no circular step rises above the minor-self-citation level.
Assumptions & free parameters
free parameters (2)
- TempestExtremes closed-contour/stitch thresholds (MSLP delta 200 Pa; min/max distance 6.5°/5.5°; Δz300,500 = -58.8 m2 s- =
Table 2
- Postprocessor architectures and hyperparameters (MLR/ANN/UNet) =
Not stated; deferred to Gomez et al. (2025, arXiv:2508.17903)
assumptions (6)
- domain assumption IBTrACS (US-agency 1-min sustained wind, MSLP) is a sufficient ground truth for global TC track and intensity evaluation
- domain assumption Tracks and intensities extracted from coarse (0.25-0.5°) global model fields by TempestExtremes are commensurable with best-track values once spurious tracks are filtered by IBTrACS matching
- domain assumption Filling missing model forecasts with the persistence forecast yields a fair inter-model comparison
- ad hoc to paper Postprocessor scores obtained on a non-6 h lead grid and with non-protocol training/validation years are comparable to the other baselines
- ad hoc to paper Track CRPS defined with Haversine distance is a proper scoring rule
- domain assumption Postprocessed intensity distribution is Gaussian (mean + std output sampled to 50 members)
Cite this review
Pith. "Pith review of TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale." pith.science (2026). https://pith.science/paper/BEAF6WAG
@misc{pith2026260123268,
author = {Pith},
title = {Pith review of: TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale},
year = {2026},
howpublished = {\url{https://pith.science/paper/BEAF6WAG}},
note = {Machine review of arXiv:2601.23268}
}
read the original abstract
TCBench is a benchmark for evaluating global, short to medium-range (1-5 days) forecasts of tropical cyclone (TC) track and intensity. To allow a fair and model-agnostic comparison, TCBench builds on the IBTrACS observational dataset and formulates TC forecasting as predicting the time evolution of an existing tropical system conditioned on its initial position and intensity. TCBench includes state-of-the-art physics-based (TIGGE) and Artificial Intelligence Weather Prediction (AIWP) models (AIFS, Pangu-Weather, FourCastNet v2, GenCast, FNV3). If not readily available (e.g., from the NOAA website as is done with TIGGE), TC tracks are consistently derived from model outputs using the TempestExtremes library. TCBench provides deterministic and probabilistic storm-following metrics. On 2023 test cases, AIWP models skillfully forecast TC tracks, while skillful intensity forecasts require additional steps such as post-processing or task-specific training. Designed for accessibility, TCBench helps AI practitioners tackle domain-relevant TC challenges and equips tropical meteorologists with data-driven tools and workflows to improve prediction and TC process understanding. By lowering barriers to reproducible, process-aware evaluation of extreme events, TCBench aims to democratize data-driven TC forecasting.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
interannual to interdecadal variability.Journal of Geophysical Research: Atmospheres, 107 (D24):ACL–26, 2002. Boris Bonev, Thorsten Kurth, Christian Hundt, Jaideep Pathak, Maximilian Baust, Karthik Kashinath, and Anima Anandkumar. Spherical Fourier neural operators: Learning stable dynamics on the sphere, 2023. URLhttp://arxiv.org/abs/2306.03838. P. Bouge...
arXiv 2002
-
[2]
Model-specific track outputs (e.g., TIGGE XML files, neural weather model forecasts in netCDF format) are converted into a uniform CSV format that includes the storm identifier (taken from IBTrACS), position, maximum sustained wind, and minimum sea-level pressure
-
[3]
Accessed May 14, 2025. 11 Tomer Burg and Sam Lillo. TroPYcal, 2019. URL https://github.com/tropycal/ tropycal. Github repository. J.P Cangialosi, E. Blake, M. DeMaria, A. Penny, A. Latto, E. Rappaport, and V . Tallapragada. Recent progress in tropical cyclone intensity forecasting at the National Hurricane Center.Weather and Forecasting, 35:1913–1922, 202...
arXiv 2025
-
[4]
Gridded predictors (e.g., neural weather model forecasts) are stored as multidimensional arrays [samples, time, lat, lon, variables] for use in ML-based experi- ments. C.3 DATAPROCESSING Each dataset undergoes preprocessing to ensure comparability: 17 Data Source Description Website Provided Reanalysis ERA5 European Re-Analysis 5https://cds.climate. coper...
2023
-
[5]
Forecast and reanalysis fields are subset in space (storm-centered region) and time (forecast lead times) using storm initialization from IBTrACS
-
[7]
A TCTrack object is created for each storm, encapsulating observed and predicted values across time steps
-
[9]
Define the observed motion vector from the storm position 12 hours prior to VT to the observed position at VT
-
[10]
Draw a perpendicular line from the forecast position to this vector
Show all 12 references
-
[11]
Find the intersection point of this perpendicular with the observed vector
-
[12]
yes" events from
CTE is the great circle distance between the forecast position and this intersection point. • In theNorthern Hemisphere, a positive CTE indicates the forecast is to therightof the observed path. • In theSouthern Hemisphere, the interpretation is reversed. Note:CTE is undefined...
1950
-
[2021]
URL https://www.star.nesdis.noaa.gov/star/documents/meetings/ 2020AI/presentations/202101/20210128_Slocum.pptx. H. Su, L. Wu, J.H. Jiang, R. Pai, A. Liu, A.J. Zhai, P. Tavallali, and M. DeMaria. Applying satellite observations of tropical cyclone internal structures to rapid i...
2020 doi
-
[2024]
The New York Times, Published July 29,
URL https://www.nytimes.com/interactive/2024/07/29/science/ ai-weather-forecast-hurricane.html . The New York Times, Published July 29,
2024
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.