Pith. sign in

REVIEW 2 major objections 5 minor 20 references

Machine Learning (ML)-Physics Fusion Model Outperforms Both Physics-Only and ML-Only Models in Typhoon Predictions

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A hybrid that spectrally nudges FuXi's large-scale forecasts into the Shanghai Typhoon Model yields lower typhoon track and intensity errors than FuXi, SHTM, or ECMWF HRES alone across every 2024 western North Pacific typhoon.

desk verdict A legitimate season-scale test of an ML-physics hybrid, but the headline skill numbers are not yet statistically backed and the manuscript has a figure/caption swap. read the letter →

arxiv 2504.20852 v2 pith:RPHIYNQT submitted 2025-04-29 physics.ao-ph

classification physics.ao-ph PACS 92.60.Wc
keywords typhoonforecastingspectralnudginghybridML-physicsmodelFuXiShanghaitrackerrorintensitywesternNorthPacific
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a hybrid typhoon forecasting system, FuXi-SHTM, outperforms both its physics-only and ML-only parents when tested on all 26 typhoons of the 2024 western North Pacific season. The fusion works by spectral nudging the large-scale wind and temperature fields of the machine-learning model FuXi into the physics-based Shanghai Typhoon Model during its integration. Across the season, the hybrid reduces track error relative to FuXi by 16.5% at 72 h and 5.2% at 120 h, and reduces intensity error by 59.7% and 47.6%. It also reproduces satellite-observed cloud patterns and 10-m wind structure more faithfully than either parent, and raising the physics model resolution from 9 km to 3 km further improves intensity forecasts. If the result holds, operational forecasting gains ML-level large-scale skill and physics-level mesoscale detail in one system.

What carries the argument

The load-bearing mechanism is spectral nudging, a coupling that continuously relaxes the physics model's planetary-scale wind and temperature toward the ML forecast during integration, with truncation wavenumbers 8 and 7, a wavelength cutoff around 1000 km, and a six-hour e-folding time; relative humidity is deliberately excluded and sea surface temperature is prescribed. This carries the argument because it is what converts FuXi's accurate large-scale flow into SHTM's mesoscale forecasts, preventing the ML model's intensity smoothing and the physics model's long-lead track drift.

What would settle it

Run the same spectral-nudging configuration with the same FuXi version and SHTM physics across the 2018-2023 seasons or over Atlantic and Central Pacific basins and compare 72-h and 120-h track and intensity errors against FuXi; if the 16.5% and 5.2% track reductions and the 59.7% and 47.6% intensity reductions do not approximately reproduce, the claim is seasonal rather than systematic.

Watch

Extended reading notes

Core claim

The discovery is that the hybrid inherits the complementary strengths of its components: FuXi supplies accurate large-scale steering flow that keeps long-lead tracks on course, while SHTM supplies mesoscale dynamics, clouds, and boundary-layer wind structure that ML models smooth away. After nudging only U, V, and T at wavelengths longer than 1000 km with a six-hour relaxation time, and prescribing sea surface temperature from ECMWF HRES analysis, the hybrid reports mean track errors below 200 km through 108 h, intensity errors near 7.5 m/s versus about 15 m/s for FuXi, and realistic cloud and wind structures for Typhoon Yagi. A 3-km version of SHTM inside the same hybrid nudging leaves track skill nearly unchanged but improves intensity forecasts, indicating that the physical model's resolution, not the ML large-scale constraint, currently limits intensity skill.

Load-bearing premise

The entire case for the hybrid rests on one typhoon season (26 storms, with the 120-h sample shrinking to 35 forecasts) and on no significance testing or year-to-year check, so the reported percentage gains could be a 2024-specific artifact rather than a general property of the fusion method.

Editorial extensions

If this is right

  • At 72 h and 120 h, track error against FuXi drops by 16.5% and 5.2%; intensity error drops by 59.7% and 47.6%.
  • The hybrid mean track error stays within 200 km through 108 h, while SHTM and FuXi both degrade at long leads.
  • The hybrid's intensity error, near 7.5 m/s, is comparable to SHTM and far below standalone FuXi's roughly 15 m/s and ECMWF HRES's roughly 10 m/s.
  • Cloud-top brightness temperatures and 10-m wind fields from the hybrid match FY-4B and SAR observations more closely than SHTM or FuXi, including realistic eye size and outer rainbands.
  • Raising SHTM resolution from 9 km to 3 km under the same nudging strengthens intensity forecasts while leaving track forecasts nearly unchanged.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper evaluates only one season, so a natural extension is to reforce 2018-2023 seasons; the same error-reduction pattern would confirm the mechanism is general rather than a 2024 peculiarity.
  • Because relative humidity is excluded from the nudging, a testable corollary is that nudging FuXi's humidity fields would erode intensity skill, a suspicion the paper's own citations support.
  • The hybrid's correction of FuXi's extreme track errors on non-landfalling storms suggests physics constraints could serve as a safety net for ML forecasts in other basins and for sudden-turning storms, which the paper notes ML models handle poorly.
  • The 3-km result implies further gains in hybrid typhoon intensity skill will come mainly from better physics resolution and data assimilation, not from improving the large-scale ML forecast alone.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript presents FuXi–SHTM, a hybrid typhoon forecasting system that spectrally nudges large-scale fields from the ML model FuXi into the physics-based Shanghai Typhoon Model (SHTM). The system is evaluated on all 26 western North Pacific typhoons of 2024 using CMA best tracks, with consistent samples from 171 down to 35 forecasts per lead time. The authors report that FuXi–SHTM reduces mean track errors relative to FuXi by 16.5% at 72 h and 5.2% at 120 h, and reduces intensity errors by 59.7% and 47.6%, while also outperforming SHTM and ECMWF HRES. Additional case studies of Typhoon Yagi (2024) are used to argue for improved cloud and 10-m wind structure, and a 3-km resolution version is shown to improve intensity on that case.

Significance. If the reported improvements are robust, the hybrid approach is a practically valuable interim solution for operational typhoon forecasting, addressing a known weakness of ML models in intensity and structure. The notable strengths of the paper are the use of the complete 2024 typhoon season, verification against independent CMA best-track data, and spectral-nudging configurations adopted from prior published work rather than tuned on the 2024 verification set, which reduces circularity concerns. However, the significance is currently limited by the absence of uncertainty quantification and the reliance on one typhoon for the structural and resolution claims.

major comments (2)
  1. [§3.1, Fig. 2] The central claim that FuXi–SHTM significantly outperforms SHTM, FuXi, and ECMWF-IFS rests on mean track and intensity errors computed over forecast samples that are not independent: each of the 26 storms contributes multiple 12-hourly initializations, so the effective sample size at any lead time is at most 26 (and fewer at 120 h), far below the reported sample sizes (171 down to 35). No confidence intervals, significance tests, or storm-blocked resampling are presented, and the 5.2% track reduction at 120 h and 47.6% intensity reduction could be within sampling noise. Please provide per-storm error tables or a storm-blocked bootstrap/permutation test to quantify the uncertainty of the headline reductions and to support the word 'significantly'.
  2. [§3.2, Figs. 4–5; §3.3, Fig. 5/6] The structural conclusions (cloud structures, eye size, 10-m wind representation) and the resolution-sensitivity conclusion are based on a single typhoon (Yagi, initialized at 0000 UTC 3 September 2024) and are presented qualitatively. Figures 4 and 5 compare one case, and the 3-km experiment in §3.3 also uses only this case. The abstract and discussion generalize these case studies into claims that FuXi–SHTM 'simulates cloud structures more realistically' and that higher resolution 'further enhances intensity forecasts.' These general claims require either multi-case objective verification (e.g., spatial correlation of TB, wind-field pattern statistics) or explicit caveats that they are case-dependent findings.
minor comments (5)
  1. [§3.3] The text refers to 'Figure 6' for the 9-km vs 3-km comparison, but the displayed figure is labeled 'Figure 5'; please correct the cross-reference and renumber figures consistently.
  2. [§3.1] The phrase 'consistent samples' should be defined; clarify that the sample sizes are the numbers of initialization times for which all four models have valid forecasts at each lead time.
  3. [§3.2] The units for effective radii are written as 'um'; use 'μm' or 'micrometers'.
  4. [Introduction; §2] The manuscript cites 'Table 1' but the table is not included in the provided text; ensure it is present and that the terminology 'ECMWF-IFS' and 'ECMWF HRES' is made consistent.
  5. [§3.1] Please describe how the ML model's track position and intensity are derived from FuXi outputs (e.g., vortex tracker, minimum sea-level pressure), since this affects the error comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the FuXi–SHTM evaluation is an independent comparison against CMA best-track data, with spectral nudging settings adopted from prior published work rather than fitted to the 2024 verification set.

full rationale

The paper's central claim is an empirical comparison of forecast errors among FuXi–SHTM, FuXi, SHTM, and ECMWF HRES, verified against CMA best-track observations that are external to all four forecasting systems. The headline reductions (e.g., 16.5% and 5.2% track-error reductions at 72 h and 120 h relative to FuXi) are direct arithmetic summaries of those independent error series, not quantities derived by construction from the model inputs. The spectral nudging configuration (tau = 6 h, truncation 8x7, wavelength > 1000 km, U/V/T only) is stated in Section 2 as adopted from prior studies, including the same research group's earlier work, rather than tuned on the 2024 forecasts; no forecast error information is fed back into the configuration. The only self-citations appear in the introduction as contextual references to the hybrid modeling framework, and one methodological choice (excluding relative humidity) is attributed to an external reference (Husain et al., 2024). These citations are descriptive, not load-bearing reductions of the claimed result. The evaluation's numerical fragility under storm-correlated samples is a statistical robustness concern, not a circularity concern. Therefore no circular step is present and the score is 0.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim depends on several configuration choices (nudging parameters, variable selection, prescribed SST) that are taken from prior literature rather than fitted to 2024 data. The evaluation itself is independent, but the generalizability rests on the representativeness of one season and on the accuracy of the CMA best-track dataset.

free parameters (4)
  • Spectral nudging relaxation time tau = 6 hours
    Set to match FuXi's 6-hourly output interval; chosen by hand, not fitted to 2024 verification data.
  • Spectral truncation numbers = 8 zonal, 7 meridional
    Targets planetary-scale waves >1000 km; adopted from prior studies, not tuned here.
  • Nudged variables = U, V, T only (RH excluded)
    Choice based on Husain et al. (2024); affects the hybrid's intensity behavior.
  • CRTM effective radii = 20, 40, 400, 600, 800 um
    Fixed values for the cloud radiative transfer simulation; affect the TB comparison but not the track/intensity claim.
assumptions (4)
  • domain assumption Spectral nudging of large-scale ML forecasts into a regional model improves TC track and intensity forecasts.
    The entire hybrid design assumes this mechanism works, based on prior studies (Husain et al. 2024; Xu et al. 2025; Niu et al. 2025).
  • domain assumption FuXi's large-scale forecasts are accurate enough to serve as a nudging target.
    The method's benefit depends on FuXi's skill at large scales; no independent verification of FuXi's validity in 2024 is provided beyond the reference to prior benchmarks.
  • domain assumption CMA best-track data are the correct ground truth for verification.
    All verification uses CMA best-track positions and intensities as truth; any biases in this operational dataset would propagate into the error statistics.
  • domain assumption The 2024 typhoon season is representative enough for the conclusions.
    All conclusions are drawn from a single season; no interannual robustness check is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Machine Learning (ML)-Physics Fusion Model Outperforms Both Physics-Only and ML-Only Models in Typhoon Predictions." pith.science (2026). https://pith.science/paper/RPHIYNQT

@misc{pith2026250420852,
  author       = {Pith},
  title        = {Pith review of: Machine Learning (ML)-Physics Fusion Model Outperforms Both Physics-Only and ML-Only Models in Typhoon Predictions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RPHIYNQT}},
  note         = {Machine review of arXiv:2504.20852}
}
read the original abstract

Data-driven machine learning (ML) models, such as FuXi, exhibit notable limitations in forecasting typhoon intensity and structure. This study presents a comprehensive evaluation of FuXi-SHTM, a hybrid ML-physics model, using all 2024 western North Pacific typhoon cases. The FuXi-SHTM hybrid demonstrates clear improvements in both track and intensity forecasts compared to the standalone SHTM, FuXi, and ECMWF HRES models. Compared to FuXi alone, FuXi-SHTM reduces typhoon track forecast errors by 16.5% and 5.2% at lead times of 72 h and 120 h, respectively, and reduces intensity forecast errors by 59.7% and 47.6%. Furthermore, FuXi-SHTM simulates cloud structures more realistically compared to SHTM, and achieves superior representation of the 10-m wind fields in both intensity and spatial structure compared to FuXi and SHTM. Increasing the resolution of FuXi-SHTM from 9 km to 3 km further enhances intensity forecasts, highlighting the critical role of the resolution of the physical model in advancing hybrid forecasting capabilities.

Figures

Figures reproduced from arXiv: 2504.20852 by the authors.

Figure 1
Figure 1. (a) Best tracks from the China Meteorological Administration (CMA) for all 2024 Northwest Pacific typhoons. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. (a) Mean track errors (unit: km) and (b) mean intensity errors (unit: m/s) for forecasts from 0 to 120 hours for [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Scatterplots of mean track errors (km) for (a) landfalling and (b) non-landfalling typhoons during 2024 from [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Spatial distributions of brightness temperature (unit: K) observed by FY-4B AGRI infrared channel 13 (left [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Forecasted (a) track errors (km) and (b) maximum wind speeds (Vmax, m/s) for Typhoon Yagi (2024), [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

20 extracted references · 8 canonical work pages

  1. [1]

    Oskarsson, L

    Adamov, S., J. Oskarsson, L. Denby, T. Landelius, K. Hintz, S. Christiansen, & S. Schemm (2025). Building machine-learning limited-area models: Kilometer-scale weather forecasting in realistic settings. arXiv preprint, arXiv:2504.09340. https://arxiv.org/abs/2504.09340

  2. [2]

    Markou, W

    Allen, A., S. Markou, W. Tebbutt, J. Requeima, W. P. Bruinsma, T. R. Andersson, & R. E. Turner (2025). End-to-end data-driven weather prediction. Nature, 1–3. https://doi.org/10.1038/s41586-025-08897-0

  3. [3]

    Bi, K., L. Xie, H. Zhang, X. Chen, X. Gu, & Q. Tian (2023). Accurate medium-range global weather forecasting with 3D neural networks. Nature, 619(7970), 533–538. https://doi.org/10.1038/s41586-023-06185-3

  4. [4]

    Zhong, F

    Chen, L., X. Zhong, F. Zhang, Y . Cheng, Y . Xu, Y . Qi, & H. Li (2023). FuXi: A cascade machine-learning forecasting system for 15-day global weather forecast. npj Climate and Atmospheric Science, 6(1), 190. https: //doi.org/10.1038/s41612-023-00512-1

  5. [5]

    Zhang, C

    Huang, Y ., Y . Zhang, C. Zhang, B. Zheng, G. Dai, & M. Li (2024). An assessment of model capability on rapid intensification prediction of tropical cyclones in the South China Sea. Dynamics of Atmospheres and Oceans, 106, 101431. https://doi.org/10.1016/j.dynatmoce.2023.101431

  6. [6]

    Husain, S. Z., L. Separovic, J. F. Caron, R. Aider, M. Buehner, S. Chamberland, et al. (2024). Leveraging data- driven weather models for improving numerical weather-prediction skill through large-scale spectral nudging. arXiv preprint, arXiv:2407.06100

  7. [7]

    Sánchez-González, M

    Lam, R., A. Sánchez-González, M. Willson, P. Wirnsberger, M. Fortunato, F. Alet, et al. (2023). Learning skillful medium-range global weather forecasting. Science, 382(6677), 1416–1421. https://doi.org/10.1126/ science.adi2336

  8. [8]

    Alexe, M

    Lang, S., M. Alexe, M. Chantry, J. Dramsch, F. Pinault, B. Raoult,et al. (2024). AIFS—ECMWF’s data-driven forecasting system. arXiv preprint, arXiv:2406.01465. https://doi.org/10.48550/arXiv.2406.01465

Show all 20 references
  1. [9]

    Liu, H. Y ., Z. M. Tan, Y . Wang, J. Tang, M. Satoh, L. Lei,et al. (2024). A hybrid machine-learning/physics-based modeling framework for two-week extended prediction of tropical cyclones. Journal of Geophysical Research: Machine Learning and Computation, 1(3), e2024JH000207. ...

  2. [10]

    Nipen, T. N., H. H. Haugen, M. S. Ingstad, E. M. Nordhagen, A. F. S. Salihi, P. Tedesco,et al. (2024). Regional data-driven weather modeling with a global stretched-grid. arXiv preprint, arXiv:2409.02891

  3. [11]

    Zhang, Y

    Niu, Z., L. Zhang, Y . Yang, Y . Han, H. Li, D. Wang,et al. (2024). Assimilating FY-4B AGRI three water-vapor channels in the operational Shanghai Typhoon Model using a GSI-based 3DVar approach.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing. h...

  4. [12]

    Huang, L

    Niu, Z., W. Huang, L. Zhang, L. Deng, H. Wang, Y . Yang, et al. (2025). Improving typhoon predictions by integrating a data-driven machine-learning model with a physics model via spectral nudging and data assimilation. Earth and Space Science, 12, e2024EA003952. https://doi.or...

  5. [13]

    Messori (2024)

    Olivetti, L., & G. Messori (2024). Do data-driven models beat numerical models in forecasting weather extremes? A comparison of IFS HRES, Pangu-Weather and GraphCast. EGUsphere, 2024, 1–35. https://doi.org/10. 5194/egusphere-2024-1042

  6. [14]

    Hoyer, A

    Rasp, S., S. Hoyer, A. Merose, I. Langmore, P. Battaglia, T. Russell,et al. (2024). WeatherBench 2: A benchmark for the next generation of data-driven global weather models. Journal of Advances in Modeling Earth Systems, 16(6), e2023MS004019. https://doi.org/10.1029/2023MS004019

  7. [15]

    Subich, C., S. Z. Husain, L. Separovic, & J. Yang (2025). Fixing the double penalty in data-driven weather forecasting through a modified spherical-harmonic loss function. arXiv preprint, arXiv:2501.19374

  8. [16]

    Sun, Y . Q., P. Hassanzadeh, M. Zand, A. Chattopadhyay, J. Weare, & D. S. Abbot (2024). Can AI weather models predict out-of-distribution gray-swan tropical cyclones? arXiv preprint, arXiv:2410.14932

  9. [17]

    Xu, H., Y . Zhao, D. Zhao, Y . Duan, & X. Xu (2024). Improvement of disastrous extreme-precipitation forecasting in North China by a Pangu-Weather AI-driven regional WRF model. Environmental Research Letters, 19(5), 054051. https://doi.org/10.1088/1748-9326/ad41f0

  10. [18]

    Xu, H., Y . Zhao, D. Zhao, Y . Duan, & X. Xu (2025). Exploring typhoon-intensity forecasting by integrating AI weather forecasts with a regional numerical-weather model. npj Climate and Atmospheric Science , 8(1), 38. https://doi.org/10.1038/s41612-025-00926-z

  11. [19]

    Xu, P., W. Li, C. Xu, & J. Li (2024). Large-scale road-network partitioning: A deep-learning method based on a convolutional autoencoder model. arXiv preprint, arXiv:2409.05132

  12. [20]

    Zhang, Z., W. Wang, J. D. Doyle, J. Moskaitis, W. A. Komaromi, J. Heming,et al. (2023). A review of recent advances (2018–2021) on tropical-cyclone intensity change from operational perspectives—Part 1: Dynamical-model guidance. Tropical Cyclone Research and Review, 12(1), 30–...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.