{"id":"fd3771ae-de70-49e5-af6b-61ff22fe0060","arxiv_id":"2504.20852","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hybrid model that nudges machine-learning forecasts into a physics-based typhoon model reduces track and intensity errors compared to either approach alone in a full-season evaluation.","lead":"This paper combines the FuXi machine-learning weather model with the Shanghai Typhoon Model, a physics-based model, and reports that the hybrid forecasts typhoon tracks and intensity more accurately than either model alone across all 2024 western Pacific typhoons. A smart generalist should read it because it suggests a practical way to get the best of both AI and physics for operational storm forecasting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline reductions rest on single-season, storm-correlated samples without significance tests; effective sample may be well below the reported n, so the claimed percentage improvements are not yet established.","rationale":"The reader's weakest assumption correctly identifies the single-season, small-sample basis as the main vulnerability. I agree that this is the load-bearing concern, but I would sharpen it: the reported sample sizes overstate the evidence because forecasts from the same storm are not independent, so the effective sample size is at most 26 storms and often much smaller. This makes the headline percentages, especially the 120 h intensity reduction, harder to trust without storm-level significance testing. The paper's internal logic is otherwise coherent; the hybrid construction via spectral nudging is clearly described, the all-2024 evaluation is systematic, and the structural case studies (FY-4B and SAR comparisons) provide independent qualitative support. No machine-checked proof or public code is provided, but the absence of code/data is a reproducibility issue rather than a flaw in the argument itself. Since the reader's conditional verdict already accounts for the statistical weakness, the verdict should remain unchanged: acceptance conditional on demonstrating that the improvements are robust to storm-level resampling and ideally to additional seasons.","tokens_in":7456,"tokens_out":5580,"duration_ms":61217,"concrete_test":"Recompute the headline error metrics as storm-blocked means: for each of the 26 storms, average forecast errors at each lead time across its 12-hourly initializations, then compute paired differences (FuXi–SHTM minus FuXi, and FuXi–SHTM minus SHTM) across storms. Apply a Wilcoxon signed-rank test and a block bootstrap that resamples storms (not individual forecasts) at 72 h and 120 h. If the 95% bootstrap confidence interval for the mean difference excludes zero for the 72 h track and 120 h intensity improvements, the concern is resolved; if it includes zero, the quantitative claims should be downgraded to case-study status.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that FuXi–SHTM lowers track and intensity errors relative to FuXi, SHTM, and ECMWF HRES—rests on mean errors computed over all 2024 forecasts, with reported sample sizes of 171 down to 35 at 0–120 h. Two intertwined problems make these means fragile. First, the individual forecasts are not independent: each of the 26 storms contributes multiple 12-hourly initializations, so the effective sample size at any lead time is at most 26 storms (and fewer at 120 h), far below the reported n. Second, no confidence intervals, significance tests, or storm-blocked resampling are provided; Figure 3 shows per-storm scatter for track only, and no intensity scatter at all. The headline reductions (16.5%/5.2% track at 72/120 h; 59.7%/47.6% intensity) could therefore be driven by a handful of unrepresentative cases, especially the 120 h intensity figure based on n=35. Because the paper does not present storm-level error tables, a reader cannot assess whether the 5.2% track reduction at 120 h is within sampling noise. The hybrid idea is plausible and the case studies are illustrative, but the evidence as presented does not yet quantify the uncertainty around the claimed improvements.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents FuXi–SHTM, a hybrid typhoon forecasting system that spectrally nudges large-scale fields from the ML model FuXi into the physics-based Shanghai Typhoon Model (SHTM). The system is evaluated on all 26 western North Pacific typhoons of 2024 using CMA best tracks, with consistent samples from 171 down to 35 forecasts per lead time. The authors report that FuXi–SHTM reduces mean track errors relative to FuXi by 16.5% at 72 h and 5.2% at 120 h, and reduces intensity errors by 59.7% and 47.6%, while also outperforming SHTM and ECMWF HRES. Additional case studies of Typhoon Yagi (2024) are used to argue for improved cloud and 10-m wind structure, and a 3-km resolution version is shown to improve intensity on that case.","tokens_in":7762,"tokens_out":6298,"duration_ms":60396,"significance":"If the reported improvements are robust, the hybrid approach is a practically valuable interim solution for operational typhoon forecasting, addressing a known weakness of ML models in intensity and structure. The notable strengths of the paper are the use of the complete 2024 typhoon season, verification against independent CMA best-track data, and spectral-nudging configurations adopted from prior published work rather than tuned on the 2024 verification set, which reduces circularity concerns. However, the significance is currently limited by the absence of uncertainty quantification and the reliance on one typhoon for the structural and resolution claims.","major_comments":[{"comment":"The central claim that FuXi–SHTM significantly outperforms SHTM, FuXi, and ECMWF-IFS rests on mean track and intensity errors computed over forecast samples that are not independent: each of the 26 storms contributes multiple 12-hourly initializations, so the effective sample size at any lead time is at most 26 (and fewer at 120 h), far below the reported sample sizes (171 down to 35). No confidence intervals, significance tests, or storm-blocked resampling are presented, and the 5.2% track reduction at 120 h and 47.6% intensity reduction could be within sampling noise. Please provide per-storm error tables or a storm-blocked bootstrap/permutation test to quantify the uncertainty of the headline reductions and to support the word 'significantly'.","section":"§3.1, Fig. 2"},{"comment":"The structural conclusions (cloud structures, eye size, 10-m wind representation) and the resolution-sensitivity conclusion are based on a single typhoon (Yagi, initialized at 0000 UTC 3 September 2024) and are presented qualitatively. Figures 4 and 5 compare one case, and the 3-km experiment in §3.3 also uses only this case. The abstract and discussion generalize these case studies into claims that FuXi–SHTM 'simulates cloud structures more realistically' and that higher resolution 'further enhances intensity forecasts.' These general claims require either multi-case objective verification (e.g., spatial correlation of TB, wind-field pattern statistics) or explicit caveats that they are case-dependent findings.","section":"§3.2, Figs. 4–5; §3.3, Fig. 5/6"}],"minor_comments":[{"comment":"The text refers to 'Figure 6' for the 9-km vs 3-km comparison, but the displayed figure is labeled 'Figure 5'; please correct the cross-reference and renumber figures consistently.","section":"§3.3"},{"comment":"The phrase 'consistent samples' should be defined; clarify that the sample sizes are the numbers of initialization times for which all four models have valid forecasts at each lead time.","section":"§3.1"},{"comment":"The units for effective radii are written as 'um'; use 'μm' or 'micrometers'.","section":"§3.2"},{"comment":"The manuscript cites 'Table 1' but the table is not included in the provided text; ensure it is present and that the terminology 'ECMWF-IFS' and 'ECMWF HRES' is made consistent.","section":"Introduction; §2"},{"comment":"Please describe how the ML model's track position and intensity are derived from FuXi outputs (e.g., vortex tracker, minimum sea-level pressure), since this affects the error comparison.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The main risk is whether the headline improvements survive a better-justified uncertainty analysis. I would encourage the editor to request storm-level error statistics or block-bootstrap confidence intervals as part of the revision, and to ensure the case-study claims are clearly labeled as illustrative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: the paper does what it says — spectral nudging of FuXi into SHTM gives lower mean track errors than either model alone over the 2024 western North Pacific season, and the intensity error is much lower than FuXi's. The catch is that the headline percentages come from one season and storm-correlated samples without significance tests, so I'd call them promising, not yet established.\n\nWhat's new is the full-season evaluation (all 26 typhoons) and the 3-km vs. 9-km resolution test. The hybrid idea itself is not new — it's in the authors' prior work and Xu et al. (2025). The paper does several things well: consistent verification samples from CMA best tracks, comparison against ECMWF HRES and both parents, structural checks with FY-4B brightness temperatures and SAR winds, and a clear description of the spectral nudging settings. I also appreciate that the nudging parameters come from earlier work rather than being tuned on the 2024 data, so there's no obvious circularity.\n\nThe weak spots are real but not fatal. The reported n (171 down to 35) counts forecast initializations, not independent storms; each storm contributes multiple 12-hourly forecasts, so the effective sample at any lead time is at most 26 storms and fewer at 120 h. Without storm-blocked confidence intervals or significance tests, the 16.5% / 5.2% track reductions and especially the 59.7% intensity reduction could be within sampling noise. The figure/caption mismatch is also a genuine problem: the text for Figure 5 describes a 10-m wind spatial map, but the caption for Figure 5 shows track errors and Vmax (which is what Figure 6 should show). The resolution experiment is a single case (Yagi), so it's illustrative, not systematic. The phrase 'greater than the sum of their parts' overstates the data — intensity is roughly SHTM-level, not better than both.\n\nWho benefits? Operational typhoon centers and researchers working on ML-physics coupling will get a useful case study and a clear workflow, even if the statistics are not yet convincing. It deserves a serious referee, not a desk reject.\n\nMy recommendation: send it out, but ask the authors to add storm-blocked uncertainty quantification, fix the captions, and soften the conclusions accordingly.","headline":"A legitimate season-scale test of an ML-physics hybrid, but the headline skill numbers are not yet statistically backed and the manuscript has a figure/caption swap.","tokens_in":8249,"tokens_out":4380,"would_cite":true,"duration_ms":41633,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["92.60.Wc"],"model":"deepseek-v4-flash","headline":"A hybrid that spectrally nudges FuXi's large-scale forecasts into the Shanghai Typhoon Model yields lower typhoon track and intensity errors than FuXi, SHTM, or ECMWF HRES alone across every 2024 western North Pacific typhoon.","keywords":["typhoon forecasting","spectral nudging","hybrid ML-physics model","FuXi","Shanghai Typhoon Model","track error","intensity error","western North Pacific"],"falsifier":"Run the same spectral-nudging configuration with the same FuXi version and SHTM physics across the 2018-2023 seasons or over Atlantic and Central Pacific basins and compare 72-h and 120-h track and intensity errors against FuXi; if the 16.5% and 5.2% track reductions and the 59.7% and 47.6% intensity reductions do not approximately reproduce, the claim is seasonal rather than systematic.","tokens_in":7309,"feed_emoji":"🌀","tokens_out":6026,"duration_ms":57451,"temperature":0.7,"pith_summary":"This paper tries to establish that a hybrid typhoon forecasting system, FuXi-SHTM, outperforms both its physics-only and ML-only parents when tested on all 26 typhoons of the 2024 western North Pacific season. The fusion works by spectral nudging the large-scale wind and temperature fields of the machine-learning model FuXi into the physics-based Shanghai Typhoon Model during its integration. Across the season, the hybrid reduces track error relative to FuXi by 16.5% at 72 h and 5.2% at 120 h, and reduces intensity error by 59.7% and 47.6%. It also reproduces satellite-observed cloud patterns and 10-m wind structure more faithfully than either parent, and raising the physics model resolution from 9 km to 3 km further improves intensity forecasts. If the result holds, operational forecasting gains ML-level large-scale skill and physics-level mesoscale detail in one system.","feed_headline":"Hybrid typhoon model cuts track error 16.5% at 72 hours","feed_subtitle":"Fusing FuXi machine-learning forecasts into a physics model also cuts intensity error by 59.7 percent.","key_machinery":"The load-bearing mechanism is spectral nudging, a coupling that continuously relaxes the physics model's planetary-scale wind and temperature toward the ML forecast during integration, with truncation wavenumbers 8 and 7, a wavelength cutoff around 1000 km, and a six-hour e-folding time; relative humidity is deliberately excluded and sea surface temperature is prescribed. This carries the argument because it is what converts FuXi's accurate large-scale flow into SHTM's mesoscale forecasts, preventing the ML model's intensity smoothing and the physics model's long-lead track drift.","core_discovery":"The discovery is that the hybrid inherits the complementary strengths of its components: FuXi supplies accurate large-scale steering flow that keeps long-lead tracks on course, while SHTM supplies mesoscale dynamics, clouds, and boundary-layer wind structure that ML models smooth away. After nudging only U, V, and T at wavelengths longer than 1000 km with a six-hour relaxation time, and prescribing sea surface temperature from ECMWF HRES analysis, the hybrid reports mean track errors below 200 km through 108 h, intensity errors near 7.5 m/s versus about 15 m/s for FuXi, and realistic cloud and wind structures for Typhoon Yagi. A 3-km version of SHTM inside the same hybrid nudging leaves track skill nearly unchanged but improves intensity forecasts, indicating that the physical model's resolution, not the ML large-scale constraint, currently limits intensity skill.","pith_inferences":["The paper evaluates only one season, so a natural extension is to reforce 2018-2023 seasons; the same error-reduction pattern would confirm the mechanism is general rather than a 2024 peculiarity.","Because relative humidity is excluded from the nudging, a testable corollary is that nudging FuXi's humidity fields would erode intensity skill, a suspicion the paper's own citations support.","The hybrid's correction of FuXi's extreme track errors on non-landfalling storms suggests physics constraints could serve as a safety net for ML forecasts in other basins and for sudden-turning storms, which the paper notes ML models handle poorly.","The 3-km result implies further gains in hybrid typhoon intensity skill will come mainly from better physics resolution and data assimilation, not from improving the large-scale ML forecast alone."],"forward_implications":["At 72 h and 120 h, track error against FuXi drops by 16.5% and 5.2%; intensity error drops by 59.7% and 47.6%.","The hybrid mean track error stays within 200 km through 108 h, while SHTM and FuXi both degrade at long leads.","The hybrid's intensity error, near 7.5 m/s, is comparable to SHTM and far below standalone FuXi's roughly 15 m/s and ECMWF HRES's roughly 10 m/s.","Cloud-top brightness temperatures and 10-m wind fields from the hybrid match FY-4B and SAR observations more closely than SHTM or FuXi, including realistic eye size and outer rainbands.","Raising SHTM resolution from 9 km to 3 km under the same nudging strengthens intensity forecasts while leaving track forecasts nearly unchanged."],"supporting_citations":[{"why":"Defines FuXi, the ML model whose large-scale forecasts are nudged into SHTM, including its ERA5 training and 0.25-degree resolution.","marker":"[4]"},{"why":"Supplies the SHTM physics model and its GSI-based data assimilation system, the mesoscale backbone of the hybrid.","marker":"[11]"},{"why":"Prior demonstration of integrating an ML model with SHTM via spectral nudging and data assimilation; this paper extends that approach to a full season.","marker":"[12]"},{"why":"Establishes large-scale spectral nudging of data-driven models into NWP and motivates excluding relative humidity from the nudged variables.","marker":"[6]"},{"why":"Provides the precedent of an AI-driven regional WRF model via nudging, supporting the hybrid design.","marker":"[17]"},{"why":"Explores AI-plus-regional-NWP for typhoon intensity and documents ML models' weakness on sudden-turning typhoons, framing the paper's robustness argument.","marker":"[18]"},{"why":"Explains the double-penalty effect in ML loss functions that underlies FuXi's intensity underestimation, motivating the need for a physical model.","marker":"[15]"}],"fun_headline_variants":["Hybrid ML-physics typhoon model cuts track error 16.5% and intensity 59.7%","ML-physics fusion improves typhoon forecasts beyond pure ML or physics","Typhoon track and intensity errors shrink with hybrid ML-physics model","Hybrid model beats both pure ML and pure physics for typhoon forecasts","Fusing FuXi into physics model reduces typhoon intensity error by 59.7%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire case for the hybrid rests on one typhoon season (26 storms, with the 120-h sample shrinking to 35 forecasts) and on no significance testing or year-to-year check, so the reported percentage gains could be a 2024-specific artifact rather than a general property of the fusion method.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid ML-physics typhoon model cuts track error 16.5% and intensity 59.7%","ML-physics fusion improves typhoon forecasts beyond pure ML or physics","Typhoon track and intensity errors shrink with hybrid ML-physics model","Hybrid model beats both pure ML and pure physics for typhoon forecasts","Fusing FuXi into physics model reduces typhoon intensity error by 59.7%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000553,"raw_usage":{"total_tokens":2637,"prompt_tokens":946,"completion_tokens":1691,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":1583}},"tokens_in":562,"tokens_out":1691,"duration_ms":13545,"temperature":1.0,"reasoning_tokens":1583,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:16:59.070274+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same spectral-nudging configuration with the same FuXi version and SHTM physics across the 2018-2023 seasons or over Atlantic and Central Pacific basins and compare 72-h and 120-h track and intensity errors against FuXi; if the 16.5% and 5.2% track reductions and the 59.7% and 47.6% intensity reductions do not approximately reproduce, the claim is seasonal rather than systematic.","supporting_citations":[{"cited_title":"Huang, L","cited_arxiv_id":null,"evidence_quote":"Prior demonstration of integrating an ML model with SHTM via spectral nudging and data assimilation; this paper extends that approach to a full season."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the precedent of an AI-driven regional WRF model via nudging, supporting the hybrid design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explores AI-plus-regional-NWP for typhoon intensity and documents ML models' weakness on sudden-turning typhoons, framing the paper's robustness argument."}],"review_version":1}