Pith. sign in

REVIEW 3 major objections 6 minor 12 references

TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale

T0 review · 3 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read TCBench claims that neural weather models forecast tropical cyclone tracks skillfully, but intensity — especially rapid intensification — only becomes skillful after post-processing or task-specific training, which this benchmark is built t

desk verdict Useful benchmark infrastructure with a load-bearing protocol leak around the postprocessed intensity and RI claims. read the letter →

arxiv 2601.23268 v2 pith:BEAF6WAG submitted 2026-01-30 cs.CE

classification cs.CE
keywords tropicalcycloneforecastingbenchmarkdatasetneuralweathermodelsrapidintensificationtrackandintensityverificationpost-processingensembleIBTrACS
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

TCBench is a global benchmark that standardises 1-5 day tropical cyclone track and intensity evaluation on a fixed 2023 test year, with forecasts matched to observed best-track records at 6-hourly keys and missing forecasts filled by persistence. The paper claims that neural weather models produce deterministic track forecasts that rival, and in one case beat, a leading physics-based ensemble (GEFS), but that their raw intensity forecasts are no better than persistence until about 24-48 hours. The load-bearing positive result is that a post-processed PanguWeather model produced intensity errors on par with or better than GEFS and was the only baseline with measurable rapid-intensification skill (positive CSI/PSS at 48-96 h), supporting the paper's conclusion that skillful AI intensity forecasts require post-processing or task-specific training. If the benchmark holds up, it gives the community a reproducible, model-agnostic way to compare TC forecasting systems and a concrete target for improving AI intensity and RI prediction.

What carries the argument

The machinery that carries the argument is the benchmark protocol itself. TCBench defines forecasting as predicting the evolution of an already-existing storm, using IBTrACS as ground truth and retaining 6-hourly verification keys; any model that fails to output a storm at a valid key is assigned the persistence forecast. Deterministic track quality is measured by direct, cross-track, and along-track position errors; ensemble quality by a fair version of CRPS with Haversine distances; and rapid intensification is cast as binary classification (24h wind gain >= 30 kt) scored by CSI and PSS. The postprocessing baselines use a storm-centred patch of AI model fields plus the initial storm state

What would settle it

Re-run the postprocessed PanguWeather intensity and RI evaluation with the standard 6-hourly IBTrACS verification grid, using models trained only on 2017-2020 data and validated on 2021-2022, and report where the 48-h RI CSI and overall PSS land; if CSI drops to near zero or PSS turns negative, the paper's central intensity claim is refuted.

Watch

Extended reading notes

Core claim

On TCBench's 2023 test set, the paper's central finding is that track and intensity skill split: deterministic AI track forecasts at 1-5 days are competitive with or better than the physics-based GEFS ensemble, while raw AI intensity forecasts are worse than persistence at short lead times and only become useful after 24-48 h. The explicitly post-processed PanguWeather baseline - a neural model paired with a distributional regression of 24h intensification - matches or beats GEFS on absolute wind and pressure error, and it is the only evaluated baseline with positive rapid-intensification skill, with 48-h CSI around 0.12 and overall PSS about +0.05. The paper concludes that neural weather mo

Load-bearing premise

Everything positive about intensity and rapid intensification rests on the post-processed PanguWeather baseline, which the paper's own Fig. 3 caption says deviates from the benchmark protocol - it uses a non-6-hour lead grid and was trained/validated on different years than the other baselines, and those years are never stated; if the postprocessor saw 2023 during training, the claim that only post-processed AI captures RI collapses.

Editorial extensions

If this is right

  • Neural weather models can be used as deterministic 1-5 day TC track guidance at skill comparable to physics-based ensembles, at a fraction of computational cost.
  • Raw AI intensity output should not be plugged directly into operational warnings; post-processing or task-specific training becomes a recommended step.
  • A single post-processed neural model matched GEFS intensity skill and was the only evaluated system with RI skill, so data-driven RI guidance at 48-96 h lead is feasible.
  • The benchmark's persistence-filled, IBTrACS-keyed protocol gives future TC forecasting models a standard, reproducible scoring system, removing tracking-algorithm and metric differences as confounders.
  • If post-processed RI skill generalizes beyond the 2023 test year, such models could serve as early-warning flags for the most destructive, rapidly intensifying storms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the postprocessor converts coarse AI fields into intensity skill without a high-resolution physics model, the result suggests operational centers could achieve useful intensity guidance with modest compute; the paper leaves this operational implication implicit.
  • The strong RI signal from one postprocessed model implies raw AI fields contain latent information about RI that is masked by intensity bias. A testable extension is to apply the same postprocessing to other neural models (AIFS, GenCast) and check whether RI skill appears consistently.
  • The persistence-fill convention means a model with poor storm-coverage gets evaluated partly against persistence; the paper's no-fill appendix shows coverage varies by model, so a useful extension would be reporting coverage-adjusted skill to separate tracking recall from forecast accuracy.
  • Since TCBench frames forecasting after storm formation, genesis prediction is outside scope; a natural companion benchmark would score how well AI models develop storms at the right time and place, completing the warning chain.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. TCBench proposes a standardized benchmark for global 1–5-day forecasts of tropical cyclone track and intensity from existing cyclones, using IBTrACS as ground truth, TIGGE ensemble models and several neural weather models as baselines, TempestExtremes for tracking, and a suite of deterministic/probabilistic/rare-event metrics (DPE, ATE, CTE, CRPS, CSI, PSS). On a 2023 test year, the paper reports three headline findings: AI models beat persistence on track forecasts; raw AI intensity forecasts are unskillful until longer leads; and a postprocessed PanguWeather baseline (PANGU_POST_ANN) attains intensity skill comparable to or better than GEFS and is claimed to be the only baseline with rapid-intensification skill.

Significance. If the protocol were airtight, TCBench would be a valuable community resource: it unifies heterogeneous data sources under a storm-relative evaluation design, provides a defined holdout year, releases code and data (subject to the review anonymization), and makes a concrete comparison between physics-based ensembles and neural weather models. The track results are credible and consistent with prior work: all models beat persistence, AIFS is the best deterministic tracker in this sample, and GEFS retains an edge in probabilistic track skill. The intensity/RI finding is the paper's main novel positive claim, but it currently rests on an off-protocol postprocessor whose training years are never stated, and the 'only' RI claim is contradicted by the paper's own scorecard. With those issues repaired, the benchmark would be a useful contribution; as written, the central intensity/RI conclusions are not yet testable under the benchmark's own rules.

major comments (3)
  1. [§6, Fig. 4, Fig. 9] Load-bearing protocol omission. The Fig. 3 caption states that the post-processing model (e.g., PANGU_POST_ANN) 'deviates from our protocol (non-6 h lead grid; trained/validated on different years)', but the manuscript never states which years were used to train or validate PANGU_POST_ANN, PANGU_POST_MLR, or PANGU_POST_UNET. This directly collides with §3.3, which requires that 'the year 2023 be left for testing' because baselines 'are not trained or tuned on this year'; §D.5's assertion that 'All baselines are trained and evaluated consistently' does not repair the omission. Since the positive intensity and RI results in §6 are carried by these postprocessors, the claims 'Postprocessing AI models yields skillful intensity predictions' and 'Only the postprocessed AI models capture RI' are not verifiable as written. Please report the training/validation years for each postprocessor, demon
  2. [§4, Eq. (1)] The statement 'Only the postprocessed AI models capture RI' is contradicted by the paper's own scorecard. Fig. 4 reports FNV3 with RI CSI = 0.095 at 48 h, 0.037 at 72 h, and 0.026 at 96 h, and Fig. 9 shows nonzero overall CSI/PSS values for several models. FNV3 is described in §D.3 as fine-tuned on IBTrACS, and the abstract explicitly allows 'task-specific training' as a route to intensity skill. The 'only' claim needs a precise definition of 'skill' (e.g., a threshold, lead-time range, or statistical significance criterion) or should be removed. In addition, the reported CSI values are small and no confidence intervals or significance tests are provided; given the rarity of RI events, sampling uncertainty should be quantified before claiming any model 'captures' RI.
  3. [§1, §6] Track CRPS is defined by replacing the absolute differences in Eq. (1) with Haversine distances. This is not automatically a proper scoring rule: for the energy-score representation to be proper, the distance must be of negative type. The manuscript provides no citation or proof that Haversine/great-circle distance has this property. Please either cite a result for spherical geodesic distance or use a chordal distance for which negative-definiteness is known, and state what this implies for the probabilistic track comparisons in Fig. 3b.
minor comments (6)
  1. [§1, §6] The introduction says 'we demonstrate that neural weather models can skillfully forecast TC intensity up to 5 days ahead, particularly when combined with observational data and processed through tools provided in TCBench.' This is at odds with the more careful abstract and §6 framing, where raw AI models are unskillful and only postprocessed models show skill; please rephrase to avoid implying raw AI intensity skill.
  2. [§6, Fig. 4] The text says 'all models produce forecasts of Vmax and pmin that are worse than the persistence baseline for shorter lead times,' but PANGU_POST_ANN in Fig. 4 already beats persistence at 24 h for both wind and pressure. This inconsistency should be clarified.
  3. [§3.2, §D.4] The data inclusion criteria in §3.2 require 6-hourly forecast data, yet the postprocessors are admitted with a 'non-6 h lead grid.' Please explain how the postprocessors satisfy the inclusion criteria, or explicitly list them as an exception.
  4. [Abstract, §D.3] Terminology is inconsistent: the abstract uses 'AIWP' and 'FourCastNet v2,' while the text uses 'neural weather models' and 'FourCastNetV2'; please standardize.
  5. [D.4] The 50-member postprocessor ensemble is described as sampled from a parametric Gaussian distribution, but the sampling procedure, random seed, and any clipping details are not given. For a benchmark claiming reproducibility, this should be specified.
  6. [Throughout] Several typos and incomplete sentences should be corrected: 'Gneiting & and, 2007' (Eq. 3 reference), 'focus lesson process-based assessments' (Sec. 1), 'we as provide scores' (Sec. 2), and 'ERA-5' inconsistently styled. A careful proofread is needed.

Circularity Check

0 steps flagged · score 2.0 of 10

Evaluation protocol is externally grounded; the only caveat is a minor self-citation burden in the off-protocol postprocessed intensity baselines, which is a verification concern rather than demonstrated circularity.

full rationale

TCBench is an evaluation benchmark, not a fitted derivation: tracks and intensities come from external physics-based and neural models, tracking is performed with the fixed TempestExtremes parameter table, and the ground truth is the independent IBTrACS record. The 2023 test year is explicitly separated from the recommended training/validation years: 'We require that the year 2023 be left for testing, given that the models we include as baselines are not trained or tuned on this year' (Sec. 3.3). No metric is defined in terms of a model output, and no reported 'prediction' is the fitted value itself. The only self-citation burden is the postprocessing baselines that carry the intensity and RI claims: they 'follow (Gomez et al., 2025)' (Sec. 5) and use 'the details provided for these architectures and associated hyperparameters provided by Gomez et al. (2025)' (App. D.4), while the Fig. 3 caption concedes 'The post-processing model (dotted; e.g., PANGU_POST_ANN) deviates from our protocol (non-6 h lead grid; trained/validated on different years)'. The training years of these postprocessors are never stated, so the headline 'Only the postprocessed AI models capture RI' rests on a baseline whose 2023 holdout status is not documented. This is a legitimate benchmark-integrity and reproducibility caveat, but it is not demonstrated circularity: evaluating a trained postprocessor on a held-out year is not equivalent, by construction, to its training target unless 2023 was in the training set, which the paper does not claim. No equation or definition in the paper reduces a predicted quantity to its own input, so no circular step rises above the minor-self-citation level.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The benchmark's central claims lean on: (i) tracking-derived intensities from coarse global fields being commensurable with agency best-track values, (ii) persistence-fill as a fair basis for inter-model comparison, (iii) the postprocessing baselines being comparable to protocol-compliant baselines despite a stated protocol deviation with unspecified training years, and (iv) the propriety of a Haversine-based track CRPS. No new physical entities are introduced; the paper's postprocessors are methods drawn from a self-cited preprint.

free parameters (2)
  • TempestExtremes closed-contour/stitch thresholds (MSLP delta 200 Pa; min/max distance 6.5°/5.5°; Δz300,500 = -58.8 m2 s- = Table 2
    Hand-chosen thresholds (largely from Ullrich et al., 2021, with min duration changed to 12 h) determine which model vortices become tracks and thereby which storms each model is scored on; changing them changes the baseline scores.
  • Postprocessor architectures and hyperparameters (MLR/ANN/UNet) = Not stated; deferred to Gomez et al. (2025, arXiv:2508.17903)
    The models that produce the headline intensity/RI skill are adopted from a self-cited preprint; their training/validation years for TCBench are unspecified in this manuscript.
assumptions (6)
  • domain assumption IBTrACS (US-agency 1-min sustained wind, MSLP) is a sufficient ground truth for global TC track and intensity evaluation
    Sec. 3.1; the paper acknowledges agency definition inconsistencies (Schreck III et al., 2014) and basin-dependent reliability (App. C.5), yet uses single-agency values as the evaluation target.
  • domain assumption Tracks and intensities extracted from coarse (0.25-0.5°) global model fields by TempestExtremes are commensurable with best-track values once spurious tracks are filtered by IBTrACS matching
    Sec. 3.1 and Table 2; TC cores are not resolved at these resolutions; matching to IBTrACS removes false positives but not intensity-representation bias.
  • domain assumption Filling missing model forecasts with the persistence forecast yields a fair inter-model comparison
    Sec. 5; this makes scores a mix of skill and coverage, acknowledged by the paper's own Fig. 7, which explains AIFS's deterministic track lead by higher coverage.
  • ad hoc to paper Postprocessor scores obtained on a non-6 h lead grid and with non-protocol training/validation years are comparable to the other baselines
    Fig. 3 caption states PANGU_POST_ANN 'deviates from our protocol (non-6 h lead grid; trained/validated on different years)'; the actual years are never given, so 2023 holdout cannot be verified.
  • ad hoc to paper Track CRPS defined with Haversine distance is a proper scoring rule
    Sec. 4, Eq. 1; CRPS/energy-score propriety requires a negative-definite semi-metric; great-circle distance on the sphere is not negative definite, so this generalization is not justified.
  • domain assumption Postprocessed intensity distribution is Gaussian (mean + std output sampled to 50 members)
    App. D.4; distributional regression assumes Gaussianity of the intensity-change distribution; no calibration check is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale." pith.science (2026). https://pith.science/paper/BEAF6WAG

@misc{pith2026260123268,
  author       = {Pith},
  title        = {Pith review of: TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BEAF6WAG}},
  note         = {Machine review of arXiv:2601.23268}
}
read the original abstract

TCBench is a benchmark for evaluating global, short to medium-range (1-5 days) forecasts of tropical cyclone (TC) track and intensity. To allow a fair and model-agnostic comparison, TCBench builds on the IBTrACS observational dataset and formulates TC forecasting as predicting the time evolution of an existing tropical system conditioned on its initial position and intensity. TCBench includes state-of-the-art physics-based (TIGGE) and Artificial Intelligence Weather Prediction (AIWP) models (AIFS, Pangu-Weather, FourCastNet v2, GenCast, FNV3). If not readily available (e.g., from the NOAA website as is done with TIGGE), TC tracks are consistently derived from model outputs using the TempestExtremes library. TCBench provides deterministic and probabilistic storm-following metrics. On 2023 test cases, AIWP models skillfully forecast TC tracks, while skillful intensity forecasts require additional steps such as post-processing or task-specific training. Designed for accessibility, TCBench helps AI practitioners tackle domain-relevant TC challenges and equips tropical meteorologists with data-driven tools and workflows to improve prediction and TC process understanding. By lowering barriers to reproducible, process-aware evaluation of extreme events, TCBench aims to democratize data-driven TC forecasting.

Figures

Figures reproduced from arXiv: 2601.23268 by the authors.

Figure 1
Figure 1. TCBench defines TC forecasting as predicting time-series of track and intensity knowing the [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (a) 2023 tropical cyclones in the TCBench test year, from IBTrACS. The lines represent the position of each tropical cyclone over time, with the line color representing the storm’s intensity at that position. (b) IBTrACS estimate of tropical cyclone numbers from 2017-2022 (corresponding to TCBench’s training and validation years). Tropical cyclone counts binned into a 1◦ latitude by 1◦ longitude grid. (c) Mean sea l… view at source ↗
Figure 3
Figure 3. FAIR per-lead comparison on TCBench-2023. Deterministic (left) and probabilistic (right) scores from 6–120 h for: (a) DPE, (b) CRPS–track, (c) AE–pressure, (d) CRPS–pressure, (e) AE–wind, (f) CRPS–wind. Means are computed on IBTrACS verification keys (00/12Z inits), with missing entries filled via persistence for fair comparison. Baselines: persistence (black dashed) and MT-LB (cyan dashed; mean tendency by lead & b… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: TCBench deterministic scorecard (2023, baseline = Persistence). Cells show mean error; [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Illustration of track error metrics adapted from (Heming, 2017) [PITH_FULL_IMAGE:figures/full_fig_p023_5.png]
Figure 6
Figure 6. Figure 6: Per–lead verification on TCBench-2023 (6–120 h), non-filled. Rows: track, pressure, wind; columns: AE, R2 , CRPS (track uses DPE, CRPS-track, along-track). Curves use raw model coverage—no persistence filling—so each mean is over the forecast–verification pairs availab…
Figure 7
Figure 7. Figure 7: Coverage on the 2023 test set (% of IBTrACS verification pairs) by lead. An IBTrACS pair is a unique (SID, t0, t0+L) observed on the 6-hour grid with t0 ∈ {00, 12} UTC. A model covers a pair if it outputs any row for that key. Shaded regions indicate the fraction cover…
Figure 8
Figure 8. Figure 8: CRPS scorecards (probabilistic). Track displacement (km), max wind (kt), and min pressure (hPa) on TCBench-2023 for 6–120,h leads. Each cell is the percent difference in CRPS relative to the Persistence baseline at the same lead (lower is better); the Persistence row r…
Figure 9
Figure 9. Figure 9: Rapid Intensification (RI) skill—overall by model. Bars show overall Critical Success Index (CSI; left) and Peirce Skill Score (PSS = TPR−FPR; right) computed against the IBTrACS RI ground truth for 2023. Scores use the common (SID, t0, t0+L) key set on the 6 h IBTrACS…
Figure 10
Figure 10. Figure 10: Rapid Intensification (RI) skill by lead time. CSI (left) and PSS (right) as a function of lead time (6–120 h, 6 h steps), evaluated vs. the IBTrACS 2023 RI ground truth on the 6 h grid. Axes are capped for readability. G POTENTIAL APPLICATIONS Accurate assessments of…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 1 canonical work pages

  1. [1]

    Boris Bonev, Thorsten Kurth, Christian Hundt, Jaideep Pathak, Maximilian Baust, Karthik Kashinath, and Anima Anandkumar

    interannual to interdecadal variability.Journal of Geophysical Research: Atmospheres, 107 (D24):ACL–26, 2002. Boris Bonev, Thorsten Kurth, Christian Hundt, Jaideep Pathak, Maximilian Baust, Karthik Kashinath, and Anima Anandkumar. Spherical Fourier neural operators: Learning stable dynamics on the sphere, 2023. URLhttp://arxiv.org/abs/2306.03838. P. Bouge...

  2. [2]

    Model-specific track outputs (e.g., TIGGE XML files, neural weather model forecasts in netCDF format) are converted into a uniform CSV format that includes the storm identifier (taken from IBTrACS), position, maximum sustained wind, and minimum sea-level pressure

  3. [3]

    11 Tomer Burg and Sam Lillo

    Accessed May 14, 2025. 11 Tomer Burg and Sam Lillo. TroPYcal, 2019. URL https://github.com/tropycal/ tropycal. Github repository. J.P Cangialosi, E. Blake, M. DeMaria, A. Penny, A. Latto, E. Rappaport, and V . Tallapragada. Recent progress in tropical cyclone intensity forecasting at the National Hurricane Center.Weather and Forecasting, 35:1913–1922, 202...

  4. [4]

    Gridded predictors (e.g., neural weather model forecasts) are stored as multidimensional arrays [samples, time, lat, lon, variables] for use in ML-based experi- ments. C.3 DATAPROCESSING Each dataset undergoes preprocessing to ensure comparability: 17 Data Source Description Website Provided Reanalysis ERA5 European Re-Analysis 5https://cds.climate. coper...

  5. [5]

    Forecast and reanalysis fields are subset in space (storm-centered region) and time (forecast lead times) using storm initialization from IBTrACS

  6. [7]

    A TCTrack object is created for each storm, encapsulating observed and predicted values across time steps

  7. [9]

    Define the observed motion vector from the storm position 12 hours prior to VT to the observed position at VT

  8. [10]

    Draw a perpendicular line from the forecast position to this vector

Show all 12 references
  1. [11]

    Find the intersection point of this perpendicular with the observed vector

  2. [12]

    yes" events from

    CTE is the great circle distance between the forecast position and this intersection point. • In theNorthern Hemisphere, a positive CTE indicates the forecast is to therightof the observed path. • In theSouthern Hemisphere, the interpretation is reversed. Note:CTE is undefined...

  3. [2021]

    URL https://www.star.nesdis.noaa.gov/star/documents/meetings/ 2020AI/presentations/202101/20210128_Slocum.pptx. H. Su, L. Wu, J.H. Jiang, R. Pai, A. Liu, A.J. Zhai, P. Tavallali, and M. DeMaria. Applying satellite observations of tropical cyclone internal structures to rapid i...

  4. [2024]

    The New York Times, Published July 29,

    URL https://www.nytimes.com/interactive/2024/07/29/science/ ai-weather-forecast-hurricane.html . The New York Times, Published July 29,

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.