{"id":"a9a2d4b3-afa7-4b14-93ed-c7b16c4c1192","arxiv_id":"2505.01455","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A hybrid ML physics model, NeuralGCM, with persisted SST and sea ice anomalies, yields skillful seasonal forecasts of tropical cyclone frequency in the North Atlantic and East Pacific (r around 0.7 over 1990-2023).","lead":"This study uses Google's NeuralGCM, a hybrid machine-learning physics climate model, to forecast tropical cyclone activity for July through November. The model produces skill comparable to conventional physical models in the North Atlantic and East Pacific, at a fraction of the computing cost.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Most of the evaluation period (1990–2017) overlaps the NeuralGCM training window, so the headline TC skill r≈0.7 may be an in-sample artifact; the out-of-training-period skill is only reported qualitatively as 'some skill'.","rationale":"The reader's weakest_assumption focused on the persisted-SST boundary forcing and the replacement of unstable stochastic runs, and the reader's rationale also noted that the evaluation period overlaps the model training period. I agree that the persistence baseline is an important attribution issue, but I identify the training-period overlap as the single most load-bearing concern because it directly threatens the validity of the headline correlation, not just its interpretation. For an ML model, skill computed on the training distribution is not a sound estimate of predictive skill; here, 82% of the deterministic verification years are in-sample, and the out-of-training period is only six years with no reported correlation or p-value. The paper does include a brief statement that 2018–2023 shows 'some skill,' but this is not quantified, so a reader cannot assess whether the central r≈0.7 is driven by in-sample behavior. I would not reject the paper: the experimental design is reasonable for a feasibility study, and the authors are transparent about many caveats. However, conditional acceptance should require the out-of-training skill breakdown as a condition. Since the reader already recommended CONDITIONAL, my recommendation is UNCHANGED, but with a sharper specific requirement. The concrete test above would settle the concern: if out-of-training correlations are significant and comparable, the claim survives; if not, the central claim must be weakened.","tokens_in":16883,"tokens_out":11468,"duration_ms":125980,"concrete_test":"Recompute the basin-wide TC count and ACE anomaly correlations separately for 1990–2017 (training period) and 2018–2023 (out-of-training period for the deterministic model, which the main text focuses on; for the stochastic model, use 2020–2023 as out-of-training years), using the same detrending and significance testing as the paper. If the out-of-training correlations are not significant at p<0.05 or are substantially lower than the full-period r≈0.7, the reported skill is primarily an in-sample artifact, and the abstract's claim of useful seasonal prediction should be weakened or explicitly restricted to the training period.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is the r≈0.7 correlation between predicted and observed basin-wide TC frequency over 1990–2023 (Figure 4, Section 3.2). Section 2.1 states that the deterministic NeuralGCM was trained on ERA5 data from 1979–2017 and the stochastic version on 1979–2019. Thus, 28 of the 34 deterministic verification years and 24 of the 34 stochastic verification years are inside the training period. Because this is an ML model, evaluating on the training distribution is the classic source of inflated apparent skill. The model was not trained directly to predict seasonal TC counts, so this is not a simple memorization claim, but the learned mapping from initial and boundary states to atmospheric evolution was optimized on exactly these years' data, and the persisted SST anomalies used as boundary forcing are also from the verification period. The Supplementary Materials note that the 2018–2023 seasons are not used to train the deterministic model and 'also show some skill in the North Atlantic and Northeast Pacific,' but the out-of-training correlation coefficient and its significance are not reported. Without a separate out-of-sample verification, the headline 'useful seasonal predictions' claim is not secure. The reader's persistence-baseline concern is related, but the training-period overlap is more fundamental because it questions whether the reported r≈0.7 is even a valid estimate of predictive skill.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports retrospective seasonal predictions of Northern Hemisphere tropical cyclone (TC) activity using the 1.4-degree NeuralGCM hybrid ML-physics atmospheric model. Starting from 1 July initial conditions for 1990–2023, the authors run 20-member ensembles forced with the climatological seasonal cycle of SST and sea ice superimposed with anomalies persisting from the initialization date. They evaluate July–November large-scale fields against ERA5 and TC tracks, counts, and ACE against IBTrACS, using TempestExtreme with adjusted vorticity and duration thresholds. The central claim is that the model reproduces observed TC climatology and achieves statistically significant interannual skill, notably basin-wide TC frequency correlations of about 0.7 in the North Atlantic and Northeast Pacific, comparable to physical GCMs with similar persistence-type boundary forcing. The paper also reports computational costs (~8 minutes per 100 simulation days on one GPU) and compares with SEAS5 in the supplementary material.","tokens_in":17217,"tokens_out":4557,"duration_ms":44730,"significance":"If the central claim holds, this is a useful demonstration that a hybrid ML-physics model can provide a computationally cheap baseline for seasonal TC prediction and motivate coupling such models to ocean and sea-ice components. The paper has several strengths: it evaluates against an external observational dataset (IBTrACS) rather than reanalysis-derived tracks; it provides confidence intervals by resampling; it benchmarks against published physical-model results (Chen and Lin 2013; Zhang et al. 2019) and includes a preliminary SEAS5 comparison; and code and data links are provided. The main caveats are that most evaluation years lie inside the model's training period and that the persistence-SST forcing may be responsible for much of the apparent skill. These issues are addressable with additional quantitative out-of-sample and baseline analyses.","major_comments":[{"comment":"The deterministic NeuralGCM was trained on ERA5 1979–2017 and the stochastic version on 1979–2019 (Section 2.1), while the headline correlation r≈0.7 is computed over 1990–2023. This means 28 of 34 deterministic verification years are inside the training window, so the reported skill is not a clean out-of-sample estimate. The Supplementary Materials state that 2018–2023 'also show some skill' but do not report the correlation coefficient or its significance. Please provide explicit out-of-sample skill metrics, such as the correlation for 2018–2023 alone or a leave-period-out analysis, and compare them with in-sample skill. If out-of-sample skill is much lower, the claim of useful seasonal predictions needs to be softened.","section":"Section 2.1, 2.2, 3.2, Figure 4"},{"comment":"Because the boundary forcing is simply the observed July 1 SST and sea-ice anomaly persisted through the season, much of the TC skill may be inherited from the observed SST state rather than from NeuralGCM's learned dynamics. The paper shows environmental skill relative to a persistence baseline in Supplementary Figure 11, but no equivalent persistence baseline is provided for basin-wide TC counts or ACE. Please add a baseline that correlates observed TC activity with the initial June/July SST anomaly alone, or with a simple statistical persistence forecast, and show whether NeuralGCM adds skill beyond it. Without such a baseline, the attribution of r≈0.7 to the model dynamics is not secure.","section":"Section 2.2 and 3.2"},{"comment":"The tracker's vorticity threshold (4×10^-5 s^-1) and duration threshold (54 h) are explicitly tuned to match IBTrACS climatological counts. This is not a direct fit to interannual skill, but it is an additional tunable choice that could affect the reported correlations. Please report how the interannual correlations vary over a plausible range of these thresholds, and state whether any tuning was performed against the skill metrics themselves. In addition, about 10% of stochastic-physics simulations are flagged as unstable and replaced with climatological fields (Section 2.3, Supplementary Figure 1); please show that the deterministic headline results are insensitive to excluding unstable years, or that the stochastic results are robust to alternative handling of those members.","section":"Section 2.3"}],"minor_comments":[{"comment":"There is a typo: 'NeruralGCM' should be 'NeuralGCM'.","section":"Section 3.2, paragraph 2"},{"comment":"The two NeuralGCM rows are not explicitly labeled as deterministic and stochastic in the table; adding labels would make the comparison much clearer.","section":"Table 1"},{"comment":"The unit 'm-2 s-2' should be 'm^2 s^-2'.","section":"Supplementary Figure 15 caption"},{"comment":"The vorticity threshold '4×10!\"!\" s!\"#' is a rendering error; it should be 4×10^-5 s^-1.","section":"Section 2.3"},{"comment":"The notation '10-3 to 10-5' should be formatted as 10^-3 to 10^-5.","section":"Introduction"},{"comment":"With 20 individual ensemble members plotted as light blue lines, the ensemble mean is difficult to distinguish; please consider plotting a subset of members or a shaded spread.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the training-period overlap: most of the evaluation years are inside the training window of NeuralGCM, and the out-of-sample period (2018–2023) is only described qualitatively. If the authors can provide a credible out-of-sample skill assessment and a persistence baseline for TC metrics, the paper could be acceptable. I would not require a fully independent retraining, but the headline claim needs the out-of-sample number reported explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The genuinely new thing is the first multi-decade seasonal TC hindcast with NeuralGCM: 1990–2023, 20-member ensembles, basin-level skill scores, and a comparison against physical models like HiRAM and FLOR. The cost is striking—about 8 minutes per 100 simulation days on a single GPU—and the paper is honest about its simplifications: persisted SST anomalies, no ocean coupling, known intensity biases. Using IBTrACS rather than reanalysis tracks for verification is a good choice.\n\nThe soft spot is the one the stress-test flags: most of the evaluation period overlaps the training window. The deterministic model was trained on ERA5 1979–2017, so 28 of the 34 verification years are in-sample. The stochastic model was trained through 2019. NeuralGCM wasn't trained to predict seasonal TC counts, so this isn't simple memorization, but the learned dynamics are optimized on exactly these years, and the SST anomalies used as boundary forcing come from the same period. The 2018–2023 out-of-sample years are mentioned only as \"some skill,\" without a correlation coefficient or confidence interval. That leaves the headline r≈0.7 as an estimate whose validity as true predictive skill is unproven.\n\nTwo smaller issues: the tracker thresholds were chosen to match climatological TC counts, which is fine for climatology but could inflate interannual skill if the tuning implicitly fits noise; and unstable stochastic runs are replaced with climatological fields, which assumes failures are random with respect to TC activity. Neither is fatal, but both deserve sensitivity analyses.\n\nOn balance, I think the paper is a solid baseline demonstration, not a resolved claim of operational skill. The central idea—a cheap hybrid model with persisted SST can emulate what physical GCMs do—is plausible and worth engaging. But the abstract overstates the confidence in the quantitative skill. If this goes to peer review, the main requirement should be an explicit out-of-sample evaluation: report the 2018–2023 correlation, its significance, and also the skill on a strict persistence baseline for TC metrics. With that, the paper would be much stronger.\n\nMy recommendation: yes, send it to peer review. It's a serious, well-executed study with a real flaw that revision can address. I'd bring it to a reading group and cite it as a baseline, though I'd be careful about citing the r≈0.7 as evidence of predictive skill.","headline":"NeuralGCM yields a cheap, plausible seasonal TC hindcast, but the headline r≈0.7 is compromised by training-period overlap; out-of-sample skill is only reported qualitatively.","tokens_in":17712,"tokens_out":2344,"would_cite":true,"duration_ms":23779,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid AI-physics climate model, forced only with persisted July 1 ocean anomalies, predicts Northern Hemisphere tropical cyclone activity for July through November with skill comparable to physical climate models.","keywords":["seasonal prediction","tropical cyclone activity","NeuralGCM","hybrid machine-learning physics model","persistent sea-surface temperature anomalies","hindcast skill","accumulated cyclone energy"],"falsifier":"Run the same NeuralGCM hindcast protocol with climatological SST and sea ice (no July 1 anomalies) for 1990–2023. If the North Atlantic and Northeast Pacific TC-frequency correlations with observations remain near r=0.7, the claimed seasonal skill is not coming from the model's response to the persisted boundary anomaly; if they collapse, the skill is genuinely forced by the SST information.","tokens_in":16698,"feed_emoji":"🌀","tokens_out":6294,"duration_ms":57175,"temperature":0.7,"pith_summary":"This paper asks whether a machine-learning weather model can be reconfigured to do climate prediction, not just forecasting. It shows that NeuralGCM, a hybrid model pairing a conventional dynamical core with learned subgrid physics, can produce useful five-month (July–November) hindcasts of tropical cyclone activity when driven by observed July 1 sea-surface temperature and sea-ice anomalies that are persisted through the season. The predicted basin-wide tropical cyclone counts in the North Atlantic and Northeast Pacific correlate with observations at about r=0.7 over 1990–2023, and accumulated cyclone energy is significantly correlated in those basins as well. The authors argue this demonstrates a computationally cheap route to seasonal tropical cyclone prediction and a baseline for future ML climate models.","feed_headline":"AI-physics model forecasts hurricane seasons at climate-GCM skill","feed_subtitle":"NeuralGCM, driven by July 1 ocean conditions, skillfully reproduces July–November cyclone counts from 1990 to 2023.","key_machinery":"The central object is NeuralGCM, a differentiable hydrostatic dynamical core whose unresolved subgrid processes, principally convection, are replaced by a learned single-column neural network shared across grid columns. The experimental mechanism is the boundary-forcing simplification borrowed from physical seasonal predictions: sea surface temperature and sea ice follow the climatological seasonal cycle with the July 1 anomaly held fixed, which captures the ocean's slow thermal evolution without an ocean model. Twenty-member ensembles are initialized on July 1 each year from 1990 to 2023, with perturbations to the encoder's learned correction; a vorticity-based tracker converts model output into cyclone tracks. This combination, rather than any new training or model modification, carries the argument: it lets a weather-trained ML model respond to the SST conditions that govern tropical cyclone activity.","core_discovery":"The central claim is that a hybrid ML-physics atmospheric model, given only persisted boundary anomalies from the initialization date plus the climatological seasonal cycle, can skillfully predict the tropical atmosphere and Northern Hemisphere tropical cyclone activity months ahead. On the paper's terms, this establishes that the model has learned enough of the atmospheric response to boundary forcing that simplified persistence forcing is sufficient, at least in TC-active basins. The supporting evidence is the 1990–2023 hindcast correlation (r≈0.7) for TC frequency in the North Atlantic and Northeast Pacific, significant correlations for sub-basin track density (p<0.1) and basin-wide accumulated cyclone energy (p<0.01) in the North Atlantic and North Pacific, and a realistic simulated TC climatology. The authors also position the skill as comparable to earlier physical GCM seasonal prediction studies, while noting that the simplified forcing makes the reported skill a lower-bound estimate.","pith_inferences":["A direct test not performed in the paper would be to replace the persisted SST anomalies with climatological SST in the same protocol; if TC-frequency correlations stay near r=0.7, much of the apparent skill is inherited from the observed June SST anomaly rather than from the model's learned dynamics.","The paper's assumption that unstable stochastic rollouts are random with respect to TC activity is untested; replacing them with climatology could bias skill upward if instabilities correlate with active seasons.","The skill in the North Atlantic and Northeast Pacific likely traces to the strong SST-TC relationship in those basins; the same protocol applied to the North Indian Ocean or to Atlantic regions like the Caribbean and Gulf of Mexico should show weaker skill, and indeed the paper reports such regional underprediction.","A natural extension is to feed NeuralGCM predicted SST fields from a coupled or statistical ocean model instead of persisting July 1 anomalies; comparing the two would separate model dynamics from boundary information."],"forward_implications":["If the claim holds, a single GPU can produce a 20-member, 5-month seasonal tropical cyclone hindcast in roughly 8 minutes per 100 simulation days, making ensemble seasonal prediction far cheaper than current physical models.","Skillful TC frequency prediction in the North Atlantic and Northeast Pacific becomes achievable without an ocean model or atmosphere-ocean coupling, so the bottleneck shifts to boundary-condition quality.","The same setup could be rerun for other initialization months, other basins, or other TC metrics, because the protocol is simple and modular.","The demonstrated response to persisted SST anomalies gives a benchmark against which future ML climate models that add coupled ocean or land components can be judged.","Because the model runs far faster than physical GCMs, the hindcast record can be extended and ensemble size increased, tightening skill estimates."],"supporting_citations":[{"why":"Supplies the NeuralGCM model, its learned subgrid physics, and the pre-trained 1.4-degree deterministic and stochastic versions used for all hindcasts.","marker":"Kochkov et al. (2024)"},{"why":"Establishes the persistent-SST-anomaly boundary forcing technique that this paper adapts for TC seasonal prediction.","marker":"Zhao et al. (2010)"},{"why":"Provides the comparison physical-model hindcast (HiRAM) and supports the choice of observed initial conditions plus persistent SST anomalies.","marker":"Chen and Lin (2013)"},{"why":"Supplies ERA5 reanalysis data used to train NeuralGCM and as the atmospheric reference for evaluating hindcast skill.","marker":"Hersbach et al 2020"},{"why":"Supplies the IBTrACS best-track dataset used as the observational target for TC frequency and ACE.","marker":"Knapp et al (2010)"},{"why":"Provides the FLOR physical-model comparison, the skill estimation methods, and the predictability analysis that identifies regions where TC activity should be predictable.","marker":"Zhang et al (2019)"},{"why":"Documents the SEAS5 operational seasonal forecast system used as a reference for environmental skill comparisons.","marker":"Johnson et al (2019)"},{"why":"Supplies the TempestExtreme tracking package whose vorticity-based tracker converts NeuralGCM output into TC tracks.","marker":"Ullrich et al (2021)"}],"fun_headline_variants":["Hybrid AI-physics model predicts hurricane seasons with GCM-level skill","NeuralGCM forecasts hurricane activity months ahead at GCM skill","Hybrid model matches GCMs in seasonal hurricane prediction skill","AI-physics model with GCM-level skill for seasonal hurricane forecasts","NeuralGCM simplified forcing yields skillful hurricane season forecasts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that holding the July 1 sea-surface temperature and sea-ice anomalies fixed through the climatological cycle is a sufficient boundary forcing, so the reported tropical cyclone skill could largely be inherited from the observed SST anomaly rather than produced by the model's learned dynamics.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid AI-physics model predicts hurricane seasons with GCM-level skill","NeuralGCM forecasts hurricane activity months ahead at GCM skill","Hybrid model matches GCMs in seasonal hurricane prediction skill","AI-physics model with GCM-level skill for seasonal hurricane forecasts","NeuralGCM simplified forcing yields skillful hurricane season forecasts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000618,"raw_usage":{"total_tokens":2902,"prompt_tokens":1014,"completion_tokens":1888,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":1799}},"tokens_in":630,"tokens_out":1888,"duration_ms":14730,"temperature":1.0,"reasoning_tokens":1799,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:49:44.287081+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same NeuralGCM hindcast protocol with climatological SST and sea ice (no July 1 anomalies) for 1990–2023. If the North Atlantic and Northeast Pacific TC-frequency correlations with observations remain near r=0.7, the claimed seasonal skill is not coming from the model's response to the persisted boundary anomaly; if they collapse, the skill is genuinely forced by the SST information.","supporting_citations":[],"review_version":1}