{"id":"0a0866f3-056e-4611-ba53-73b9f24bd8bf","arxiv_id":"2507.07952","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In free runs, four AI-ML weather models reproduce Kelvin wave composites reasonably well but show incorrect vertical temperature and divergence structure for equatorial Rossby waves.","lead":"This paper tests whether four modern AI weather models, PanguWeather, GraphCast, FourCastNet, and Aurora, can reproduce the large-scale Kelvin and Rossby waves that shape tropical weather. It finds the models capture Kelvin wave structure well but fail on the vertical structure and internal consistency of Rossby waves.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Rossby-wave T–omega sign mismatch may be an artifact of the kinematically derived vertical velocity, not a learned-model deficiency; the paper needs an ERA5 control using its own omega derivation.","rationale":"Stress-testing the central claim requires separating the learned model physics from the diagnostic chain. The strongest, most general statement is the Kelvin success/Rossby failure asymmetry; within that, the Rossby temperature–vertical velocity inconsistency is the load-bearing result. The paper's own Section 2 reveals that vertical velocity is not always a model output: it is derived by kinematic integration. This derivation has known failure modes in free-running and coarse-level models, and the paper provides no validation of the derived omega against model-native or reanalysis-native omega. Thus the central negative result could be manufactured by the analysis pipeline. I do not see a more decisive internal weakness: the spectral arguments are supportive, the composites are plausible, and the model differences described are internally consistent. The concern is therefore conditional, exactly as the reader concluded, but more specific than 'free runs may drift.' An ERA5 control that uses the paper's own omega derivation is the one check that would settle it. If such a control preserves the sign mismatch, the claim is robust; if not, the Rossby-wave failure needs to be reframed as a diagnostic issue.","tokens_in":14063,"tokens_out":5344,"duration_ms":66921,"concrete_test":"Computational test: Apply the paper's exact event selection, spectral filtering, and kinematic omega derivation (continuity integration with omega_sfc=0) to ERA5 for the same 10 initialized dates and four-month segments, then construct the Rossby composite at 10N. Compare the T–omega phase relation from this ERA5 control with (a) the paper's model composites and (b) an ERA5 composite using ERA5's native omega. If the control reproduces the observed warm-with-upward/downward sign only with native omega and flips with the kinematic omega, the reported model deficiency is a diagnostic artifact. If the kinematic ERA5 composite preserves the correct sign relation, the model finding survives. A secondary check is to recompute the GraphCast Rossby composite using GraphCast's native omega if available; agreement with the kinematic version would further isolate the issue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most novel claim—that all four AI-ML models produce a temperature anomaly inconsistent with vertical velocity in Rossby waves—is made on the basis of a pressure-velocity field that is partly derived, not predicted. Section 2 states that, for models that do not output specific humidity and pressure velocity, omega is computed by kinematic integration of the continuity equation from the surface with omega_sfc = 0. No section of the paper reports which models provide native omega, which models use the kinematic estimate, or how the two compare. Because the kinematic integral accumulates errors from divergence biases, missing surface-pressure tendency, and the coarse 13-level vertical grid, the derived omega can be decorrelated from the model's own temperature field even if the model is internally consistent. The same caveat also affects the Kelvin-wave tilt and phase-speed statements, though those claims are less sensitive to the sign relationship. As written, the Rossby conclusion cannot distinguish 'the model has unphysical thermodynamics' from 'the diagnostic omega is inconsistent with model T'. The lack of an identical-methodology ERA5 control means the reference composites from Nakamura and Takayabu (2022) use different filters, event selection, and native omega, so a direct comparison is not apples-to-apples.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines intraseasonal equatorial Kelvin and Rossby waves in four AI-ML weather models (PanguWeather, GraphCast, FourCastNet, Aurora) through free runs of up to four months. It presents wavenumber-frequency spectra, composite horizontal maps and vertical cross-sections, and Hovmöller diagrams for Kelvin and Rossby waves, comparing them with theory and with reanalysis-based structures from the literature. The authors report that all four models capture the basic Kelvin wave structure including convergence-divergence patterns, vertical tilts, and temperature-vertical velocity phase relations, but that all four models fail to represent the vertical structure of Rossby waves, with temperature anomalies inconsistent with the vertical velocity. The paper is a pure diagnostic study with no parameter fitting and treats the models as black boxes.","tokens_in":14241,"tokens_out":4730,"duration_ms":47138,"significance":"If the main conclusions are robust, this is a useful and timely contribution: it provides one of the first systematic comparisons of intraseasonal equatorial wave structure in modern AI-ML weather models, using multiple models and free runs rather than a single architecture. The Kelvin wave results are encouraging and, if confirmed with appropriate uncertainty quantification, would be a meaningful positive result for the field. The reported Rossby wave temperature-vertical velocity inconsistency is potentially important for the evaluation of physical consistency in data-driven models. The study is self-contained in its use of external benchmarks (shallow-water dispersion curves, reanalysis composites) and does not fit parameters. The authors also deserve credit for transparently describing the kinematic derivation of vertical velocity and for noting that some model differences exist, even if the analysis lacks formal statistical support.","major_comments":[{"comment":"The paper's most novel claim—that in all four models Rossby-wave temperature anomalies are inconsistent with the vertical velocity—rests on an omega field whose provenance is not documented. Section 2 states that for models that do not provide pressure velocity, omega is computed kinematically by integrating the continuity equation from the surface with omega_sfc=0, but no section states which models output native omega and which use the derived estimate, nor how the two diagnostics compare. Because the kinematic integral accumulates divergence bias and omits surface-pressure tendency, derived omega can be decorrelated from the model's own temperature field even in an internally consistent model. The abstract's statement 'the temperature anomaly was inconsistent with the nature of the vertical velocity' therefore cannot yet be attributed to a model deficiency. Please report omega provenance per model and add an ERA5 control processed with the identical kinematic omega derivation and identical filtering/compositing, so that the reference comparison is apples-to-apples.","section":"Section 2 (Models and Methodology)"},{"comment":"The wavenumber-frequency diagrams and the accompanying 'equivalent depths in accord with observations' claim are based on one run per model, despite the ten four-month runs described in Section 2. With only one realization there is no measure of run-to-run spread, and the qualitative statement that all four models show Kelvin/Rossby bands cannot be separated from sampling variability. Please show at least the run-to-run spread (e.g., mean spectrum with inter-run envelope) or state explicitly why one run is representative.","section":"Section 3 (Wavenumber Frequency Diagrams)"},{"comment":"The composite comparisons—including the model-difference statements such as weak PanguWeather convergence, GraphCast's upright Kelvin tilt, and FourCastNet's 'cleanest' Rossby humidity—are made without event counts, significance tests, or confidence intervals. The composites may be based on very few events, in which case the visual differences (and even the sign of some anomalies) could be sampling noise. Please include event counts for each model and wave type and add a statistical significance assessment (e.g., bootstrap confidence intervals or field significance on the composites) for the key sign and tilt claims.","section":"Sections 4 and 5 (composites)"},{"comment":"The entire comparison assumes that four-month free runs of forecast-trained models produce physically realistic intraseasonal variability. The paper provides only indirect evidence—red background spectra and decorrelation timescales—and does not quantify whether model climatology, variances, or wave activity levels drift over the four months, or whether any runs become unstable. If a run drifts to an unrealistic mean state, the composites and even the spectral bands could be artifacts. Please add time series or maps of key variables (e.g., equatorial zonal wind, temperature, precipitation) over the runs, and a comparison of model climatology and variance to ERA5 over the same period.","section":"Section 2 (free runs) and Section 6"}],"minor_comments":[{"comment":"The Acknowledgements section spells 'Aurora' as 'Auroa'; please correct the typo.","section":"Acknowledgements"},{"comment":"The text refers to '200 mbar flow' for the Rossby gyres while the composite maps are labeled 250 mbar; please reconcile the level references.","section":"Section 5"},{"comment":"Several reference entries are malformed or incomplete (e.g., the Guo et al. entry lists 'WM Guo, J Waliser' and the Wheeler and Kiladis entry has broken text); please supply a careful reference cleanup.","section":"References"},{"comment":"The order in which Rossby-wave figures are discussed (Figure 8, then Figure 10, then Figure 9, then Figure 11) is confusing; please renumber or reorder the figures so they appear in the order discussed.","section":"Section 5"},{"comment":"The phrase 'as is observed in from real-world data' appears to contain a typo; please revise to 'as is observed in real-world data.'","section":"Section 3"},{"comment":"The phase speeds estimated from the Hovmöller diagrams are quoted without uncertainty estimates; at minimum, please give the range across runs or state that these are single-composite estimates.","section":"Sections 4 and 5 (Hovmöller diagrams)"}],"recommendation":"major_revision","confidential_remarks":"No confidential concerns. The manuscript is within the scope of the journal and the main technical issues are addressable within the manuscript's framework. The most important fix is the ERA5 control with identical omega derivation; without it, the central Rossby-wave claim remains underdetermined."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper does something useful: it runs four major AI weather models freely for four months and applies the standard Wheeler-Kiladis spectral and composite framework to see whether the models have learned the vertical structure of equatorial Kelvin and Rossby waves. The Kelvin results are encouraging, and the model-by-model differences are concrete. That part is a genuine contribution, extending Hakim and Masanam (2024) from a single model and forced response to free-running multi-model diagnostics.\n\nThe main claim to watch is the Rossby-wave failure: the authors say that in all four models the temperature anomaly is inconsistent with the vertical velocity. The stress-test note gets this exactly right. The vertical velocity is not a native output of these models; Section 2 says omega is computed by kinematic integration of the continuity equation from the surface with zero surface omega. The paper never states which models provide native omega (likely none) or validates the derived field. So the T–omega sign mismatch may be an artifact of the derived omega accumulating error from divergence biases, the coarse 13-level grid, and missing surface pressure tendency. An ERA5 control using the identical omega derivation is necessary before trusting that conclusion. The paper compares to reanalysis composites using native omega and different event selection, which is not apples-to-apples.\n\nOther soft spots are minor by comparison but still real: composites lack event counts and significance tests; free-run stability is only indirectly supported via red spectra and decorrelation times, with no check on climatological drift or variance changes over the four months; and the wavenumber-frequency diagrams come from a single run per model.\n\nI want to credit the authors for a clear, honest write-up. They don't overstate. The paper deserves a serious referee, but the referee should push for the ERA5 control and event-count reporting. If the T–omega inconsistency survives the control, the finding is important. If not, the paper remains a useful benchmark, just with a different punchline.\n\nI'd bring this to reading group; it's a good discussion piece about diagnostics for AI models.\n\nRecommendation: accept for peer review, with the derived-omega issue front and center.","headline":"Useful first multi-model benchmark of intraseasonal wave structure in AI weather models, but the headline Rossby-wave finding is likely contaminated by the kinematically derived vertical velocity.","tokens_in":14797,"tokens_out":3038,"would_cite":true,"duration_ms":33573,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Modern AI weather models reproduce the three-dimensional structure of intraseasonal Kelvin waves, but all four fail on the vertical structure of equatorial Rossby waves, where temperature anomalies point opposite to vertical motion.","keywords":["convectively coupled equatorial waves","Kelvin waves","Rossby waves","wavenumber-frequency spectra","AI-ML weather models","intraseasonal variability","tropical predictability","free-running forecasts"],"falsifier":"Run the same four models in a short-lead configuration, for example 15-day forecasts starting from reanalysis states, filter the Rossby band, and composite temperature and vertical velocity; if temperature anomalies align with vertical velocity in any model at short lead, the claim that all four models fail on Rossby vertical structure would be an artifact of free-run drift rather than a learned deficiency.","tokens_in":13829,"feed_emoji":"🌊","tokens_out":7593,"duration_ms":84524,"temperature":0.7,"pith_summary":"This paper asks whether data-driven weather models, trained only for short-range forecasting, have internalized the basic equatorial wave modes of the tropical atmosphere. The authors run PanguWeather, GraphCast, FourCastNet and Aurora freely for four months, filter the output in wavenumber-frequency space, and composite the structures of intraseasonal Kelvin and Rossby waves. They find that all four models produce Kelvin waves whose horizontal convergence patterns, vertical tilts in temperature, humidity and vertical velocity, and temperature-vertical velocity phase relations closely match reanalysis and observations. For Rossby waves, however, the horizontal gyres come out right while the vertical structure does not: temperature anomalies are systematically inconsistent with the sign of vertical velocity in all four models, even though moisture and vertical velocity anomalies are closer to observed. The result matters because these models are candidates for subseasonal prediction, and a correct representation of convectively coupled waves is thought to underpin tropical predictability.","feed_headline":"All four AI weather models nail Kelvin waves, flub Rossby structure","feed_subtitle":"Kelvin wave structure matches reanalysis; Rossby temperature anomalies contradict vertical motion in every model.","key_machinery":"The method is a space-time spectral filter and event-compositing pipeline applied to free runs of the four models. Wavenumber-frequency diagrams of symmetric and antisymmetric zonal wind are compared with theoretical dispersion curves for shallow-water equivalent depths of 12, 25 and 50 m to identify Kelvin and Rossby bands; events are selected where zonal wind variance for Kelvin waves, or geopotential height variance for Rossby waves, exceeds one standard deviation above the climatological mean in specified tropical boxes; composites of divergence, horizontal wind, temperature, humidity and pressure velocity are then examined in horizontal maps and vertical cross-sections. This machinery lets the authors isolate the wave signature from chaotic variability and compare each field's structure against reanalysis-based expectations.","core_discovery":"The paper's central claim is that all four AI-ML models capture the three-dimensional structure of intraseasonal Kelvin waves, including lower-level convergence and upper-level divergence, westward tilt of anomalies with height up to roughly 200 mb and eastward tilt above, and the observed phase relationship between temperature and vertical velocity, while none of them represents the vertical structure of equatorial Rossby waves correctly. In the Rossby composites, the upper-level cyclonic and anticyclonic gyres are present, but the temperature anomaly is of the wrong sign for the sense of vertical motion in all four models, and the divergence field shows an anomalous mid-tropospheric bias and unexpected tilts. The moisture field is much closer to observations, and only GraphCast and FourCastNet show the simultaneous build-up of moisture and deep vertical motion. The authors interpret this as evidence that the models have learned the more divergent Kelvin mode well but have not learned the physical consistency among thermodynamic and dynamical fields required for the more rotational Rossby mode.","pith_inferences":["A testable extension would be to compute the same Rossby composites over the first month of each free run rather than the full four months; if the temperature sign error disappears, the failure is a drift artifact rather than a learned property.","The temperature-vertical velocity inconsistency in the Rossby waves resembles a form of geostrophic or hydrostatic imbalance; if so, adding balance constraints during training or post-processing could improve rotational mode structure.","The same composite pipeline could be applied to the MJO or to mixed Rossby-gravity waves once longer free runs with more events become available.","The contrast between good moisture and bad temperature in Rossby composites suggests the models may be learning moisture-dynamics coupling but not the thermodynamic energy balance that ties temperature to vertical motion."],"forward_implications":["Free-running AI-ML models can serve as testbeds for intraseasonal tropical variability, because their spectra already show Kelvin and Rossby bands with observed equivalent depths.","The Kelvin wave composites, including vertical tilts and phase relationships, provide a benchmark that many traditional GCMs have struggled to reach.","Since all four models share the Rossby temperature-vertical velocity sign error, the flaw is likely common to the way these models represent rotational large-scale flow, not specific to one architecture.","The diagnostic of wavenumber-frequency filtering plus vertical compositing can be reused to evaluate future AI-ML models, including foundation models, before they are used for subseasonal prediction."],"supporting_citations":[{"why":"Supplies the wavenumber-frequency filtering method and dispersion-curve framework used to isolate Kelvin and Rossby bands and to estimate equivalent depths.","marker":"[Wheeler and Kiladis, 1999]"},{"why":"Provides the observed Kelvin wave vertical structure, tilts, and phase relationships against which the model composites are judged.","marker":"[Straub and Kiladis, 2003]"},{"why":"Supplies the reanalysis-based composite structures of Kelvin and Rossby waves, including vertical profiles of divergence, temperature, humidity, and vertical velocity.","marker":"[Nakamura and Takayabu, 2022]"},{"why":"Defines the convectively coupled equatorial wave framework and the expected horizontal and vertical signatures, including Rossby temperature structure.","marker":"[Kiladis et al., 2009]"},{"why":"Earlier experiment showing that PanguWeather produces Kelvin convergence patterns and Rossby gyres in response to an imposed tropical heat source; motivates this study.","marker":"[Hakim and Masanam, 2024]"},{"why":"Reference for the red background spectrum and zonal wind decorrelation timescales used to characterize the free runs.","marker":"[Hendon and Wheeler, 2008]"},{"why":"Theoretical equatorial wave solutions used to interpret the dispersion and expected structure of Kelvin and Rossby waves.","marker":"[Matsuno, 1966]"}],"fun_headline_variants":["AI models nail Kelvin waves, botch Rossby vertical structure","Kelvin waves pass, Rossby waves fail in AI weather models","AI weather models ace Kelvin, trip on Rossby thermodynamics","All four AI models agree: Kelvin good, Rossby flawed","AI weather models: Kelvin waves spot-on, Rossby waves off"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that four-month free runs of these forecasting models produce physically realistic intraseasonal variability that can be meaningfully compared with reanalysis; the paper checks red spectra and decorrelation times but does not directly verify that model climatology, variances, or wave activity levels remain stable through the runs.","fun_headline_variants_meta":{"raw":{"variants":["AI models nail Kelvin waves, botch Rossby vertical structure","Kelvin waves pass, Rossby waves fail in AI weather models","AI weather models ace Kelvin, trip on Rossby thermodynamics","All four AI models agree: Kelvin good, Rossby flawed","AI weather models: Kelvin waves spot-on, Rossby waves off"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000322,"raw_usage":{"total_tokens":1841,"prompt_tokens":1006,"completion_tokens":835,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":748}},"tokens_in":622,"tokens_out":835,"duration_ms":8263,"temperature":1.0,"reasoning_tokens":748,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:28:09.579406+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four models in a short-lead configuration, for example 15-day forecasts starting from reanalysis states, filter the Rossby band, and composite temperature and vertical velocity; if temperature anomalies align with vertical velocity in any model at short lead, the claim that all four models fail on Rossby vertical structure would be an artifact of free-run drift rather than a learned deficiency.","supporting_citations":[{"cited_title":"Convective couplings with equatorial rossby waves and equatorial kelvin waves","cited_arxiv_id":null,"evidence_quote":"Supplies the reanalysis-based composite structures of Kelvin and Rossby waves, including vertical profiles of divergence, temperature, humidity, and vertical velocity."},{"cited_title":"Some space-time spectral analyses of tropical convection and planetary-scale waves","cited_arxiv_id":null,"evidence_quote":"Reference for the red background spectrum and zonal wind decorrelation timescales used to characterize the free runs."}],"review_version":1}