{"id":"9448a57c-f37c-4ff9-bc58-27a933a8cb3b","arxiv_id":"2411.11268","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 450M-parameter AI atmospheric model trained on reanalysis data stably reproduces 80 years of climate variability and trends, but its separate CO2 sensitivity is not physically realistic.","lead":"Researchers built a machine learning model that mimics the atmosphere and reproduced 80 years of weather patterns, storms, and temperature trends at a fraction of the usual computing cost. The model runs thousands of simulated years per day, which could make climate experiments much cheaper, but it still handles some carbon dioxide effects incorrectly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CO2 sensitivity experiment undermines the forced-response claim; the near-surface warming is wrongly attributed to CO2 rather than SST, and the training data cannot disentangle the two.","rationale":"The reader's CONDITIONAL verdict is appropriate. The 2001-2010 held-out test period provides genuine out-of-sample support for climate skill, ENSO response, tropical cyclone statistics, and MJO behavior. The 1000-year stability run with climatological forcing and the physical conservation constraints are strong supporting evidence. The CO2 sensitivity failure in Section 2.4 is the most load-bearing weakness because it directly targets the paper's claim of learning 'forced responses' and the Abstract's phrase 'accurately reproduces... global trends of temperature over the past 80 years.' The key technical point: during 1940-2020, SST and CO2 are highly correlated in the AMIP forcing, so the model cannot cleanly attribute the learned trend to either variable. The authors themselves flag this in Section 2.4 and in the Discussion limitation paragraph, so the critique is not manufactured; it is a disclosed limitation with concrete implications for how the headline claim should be read. The checkpoint-selection issue (validation-based model selection with a manually downweighted variable) is secondary because the held-out test period still gives out-of-sample support. I agree with the reader's identification of the CO2/SST confounding as the weakest assumption. The concrete test I propose would determine whether this is a fixable training-data limitation or an intrinsic model deficiency; either way, the current evidence does not support separate CO2 and SST sensitivity claims, so CONDITIONAL should remain.","tokens_in":24819,"tokens_out":1719,"duration_ms":15600,"concrete_test":"Run paired SHiELD AMIP simulations with (a) historical SST plus fixed 1940 CO2 and (b) fixed SST plus historical CO2, then retrain ACE2-SHiELD including these paired runs as training data and rerun the Section 2.4 sensitivity test. If the retrained model produces near-surface warming under (a) comparable to the paired SHiELD run, the CO2-attribution problem is a training-data limitation that is fixable. If the model still loses most near-surface warming under (a), the learned SST sensitivity is intrinsically too weak and the forced-response claim for near-surface temperature is not reliable even for the historical period.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that ACE2 'accurately reproduces the atmospheric response to El Niño variability and global trends of temperature over the past 80 years' and can emulate forced responses. The load-bearing premise is that the model has learned physically separable forced responses. Section 2.4 directly tests this by holding CO2 fixed at the 1940 value while SST rises. ACE2-SHiELD then loses most of the near-surface warming, which is not physically expected if SST is the dominant driver. The paper states 'this is largely due to lack of warming over high-latitude land... despite evidence that such polar amplification should be driven largely by SST and sea ice coverage forcing (Screen et al., 2012).' This is a within-model inconsistency: the model attributes the historical near-surface warming trend to CO2 rather than SST, whereas the reference SHiELD and physical understanding attribute it largely to SST/sea-ice forcing. Since the 1940-2020 record has strongly correlated SST and CO2, the training data do not identify which forcing drives the learned trend. The model's separate sensitivities are thus not physically correct, and the 'forced response' claim is only established for the combined historical forcing. The reader's weakest_assumption captures this correctly, and Section 3's limitation statement concedes it. The issue is narrower than the headline claim because the ENSO response and combined-forcing trends remain supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ACE2, a 450M-parameter autoregressive machine learning emulator run at 1-degree resolution and 6-hour steps, trained on either ERA5 or an AMIP-style GFDL SHiELD simulation. It reports stable multi-decadal and millennial rollouts, exact dry-air mass and moisture conservation, realistic ENSO regression patterns, tropical cyclone statistics, MJO propagation, polar stratospheric variability, medium-range weather skill, and roughly 1500 simulated years per wall-clock day. The central claim is that ACE2 accurately captures subseasonal-to-decadal atmospheric variability and forced responses over 1940-2020, while the authors acknowledge that separately varying SST and CO2 produces non-realistic sensitivities.","tokens_in":25071,"tokens_out":6057,"duration_ms":59078,"significance":"If the main claims hold, this is a substantial advance for learned climate emulators: it demonstrates stable long autoregressive simulations under historically varying SST and CO2 forcing, with emergent phenomena such as TCs, MJO, and SSWs, and with enforceable conservation properties. The paper is unusually strong on reproducibility: training targets, code, and trained checkpoints are public, and the 10-year held-out test period, the 1000-year stability check, and the ENSO regression comparisons against reference internal variability provide concrete evidence. The main caveat is that the forced-response claim is broader than what the current evidence supports, because the model's separate SST and CO2 sensitivities are shown in Section 2.4 to be physically questionable.","major_comments":[{"comment":"The fixed-CO2 experiment directly tests the learned separate CO2 sensitivity, and it fails: when CO2 is held at 307 ppm while SST rises, ACE2-SHiELD loses most of the near-surface warming, including high-latitude land amplification that the paper itself (citing Screen et al., 2012) expects to be driven mainly by SST and sea-ice forcing. Because SST and CO2 rise together over 1940-2020, the training data cannot identify which forcing produced the learned trend, so the model's separate sensitivities are not physically grounded. This is load-bearing for the title/abstract claim of \"forced responses\": the paper establishes an accurate response to the combined historical SST+CO2 forcing, but not an accurate response to the individual forcing agents, which is what scenario interpolation would require. The Discussion in Section 3 acknowledges the limitation, but the framing should be revised, and/or the suggested SHiELD runs with historical SST/fixed CO2 and vice versa should be performed to test whether training-data augmentation fixes the attribution.","section":"§2.2.1"},{"comment":"The 81-year trend evaluation overlaps substantially with the training data (1940-1995 and 2011-2019) and with the validation/checkpoint-selection period (1996-2000; see Section 4.3). Although ACE2 is trained only on 6-hour transitions, the long-trend R2 values in Figure 1 reflect this overlap and therefore do not by themselves prove out-of-sample generalization of the 80-year trend. The held-out 2001-2010 test period supports the 10-year climate-skill and ENSO-response claims, but it is too short to validate the \"past 80 years\" trend claim. Please either report trend skill on a fully held-out period (for example, 2001-2010 only) or explicitly restate the trend claim as an in-sample/emergent property of the 6-hourly training objective.","section":"§2.1"},{"comment":"The final model is selected by climate skill over twelve 5-year inference runs spanning 1940-2000, with the q0 channel downweighted by a factor of 10, rather than by held-out test skill. This introduces a mild selection bias into the long-run skill statistics in Figures 1 and 13. The 10-year test period and Figure 20 mitigate the concern, but the paper should quantify how much of the reported 81-year skill is robust across the four training seeds and whether the q0 downweighting changes the qualitative conclusions.","section":"§4.3"}],"minor_comments":[{"comment":"The phrase \"forced responses\" should be qualified as \"responses to combined historical SST and CO2 forcing\" unless the separate-forcing experiments are added, because the paper itself reports non-realistic separate SST and CO2 sensitivities.","section":"Abstract and §2.4/§3"},{"comment":"The moisture-correction formula appears mis-typed: it should read ⟨E(t) − (TWP(t) − TWP(t−1))/Δt⟩ rather than ⟨E(t) − TWP(t) − TWP(t−1)/Δt⟩.","section":"§4.3 (Eq. 6)"},{"comment":"There are several typos: \"that ERA5\" (Section 2.2.3), \"of of\" and \"variabilty\" (Section 2.2.4), and \"simmilar\" (Appendix A.5).","section":"Typos"}],"recommendation":"major_revision","confidential_remarks":"This is a strong, reproducible contribution with a clear limitation. The main risk is overclaiming forced-response skill in the title and abstract given the SST-CO2 collinearity and the acknowledged failure in Section 2.4. I would recommend major revision focused on either reframing the forced-response claim or running the proposed factorially designed SHiELD experiments; the variability, stability, and weather-skill results are solid and should not be delayed by overreach."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ACE2 is worth a careful read. It is the first learned atmospheric emulator I know of that combines stable century-scale rollouts, exact dry-air mass and moisture conservation, and a credible ENSO response on a held-out decade. The 1000-year stability check under climatological forcing is a real result, and the tropical cyclone, MJO, and stratospheric vortex sections are honest evaluations rather than cherry-picked metrics. They also shipped the code and data, which is reproducible practice.\n\nThe main caveat is exactly the one the authors disclose: the separated forcing sensitivities are not physically right. Section 2.4 holds CO2 fixed at 1940 while SST rises, and the model loses most of the near-surface warming. That means the learned trend is largely attributed to CO2, which contradicts the SHiELD reference and physical expectation that SST/sea-ice forcing drive polar amplification. Since SST and CO2 are collinear over 1940-2020, the training data cannot identify separate causal contributions. So the 'forced response' claim should be read as 'response to the combined historical forcing,' and the paper's own abstract says the separate sensitivities are not entirely realistic. That is a genuine limitation, but it is stated in the paper, not hidden.\n\nA second, milder concern is checkpoint selection. The best of four seeds is chosen using a validation-period climate metric, with q0 downweighted by a manual factor. That is a form of model selection that can inflate the headline numbers. The authors show all four seeds and the held-out test period, so it is not fatal, but it should be disclosed more prominently.\n\nThe 81-year trend evaluation overlaps the training period, as the paper itself notes. The held-out 2001-2010 test period is the real out-of-sample evidence, and it is convincing for climate mean state and ENSO. The CO2 sensitivity issue means the long-trend skill could partly be memorization of the correlated forcing, but the held-out period still supports the combined-forcing claim.\n\nBottom line: this paper deserves a serious referee. The weaknesses are real but bounded, and the authors are unusually transparent about them. A careful revision should either retrain on SHiELD runs with decoupled SST and CO2, or clearly scope the claim to combined historical forcing. I would bring it to our reading group and would cite it for the stability and conservation results.","headline":"A serious step forward for learned climate emulators, but the forced-response claim is only established for combined historical forcing, not separable SST and CO2 responses.","tokens_in":25672,"tokens_out":2141,"would_cite":true,"duration_ms":20176,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ACE2, a 450-million-parameter learned atmospheric model, can be stepped forward stably for arbitrarily many steps and reproduces atmospheric variability from days to decades, including the response to El Niño and 80-year temperature trends.","keywords":["machine learning emulator","climate variability","forced response","El Niño-Southern Oscillation","tropical cyclones","Madden-Julian Oscillation","subseasonal-to-decadal prediction","mass and moisture conservation"],"falsifier":"The decisive test is to generate paired SHiELD simulations with historical SST plus fixed CO2 and fixed SST plus historical CO2, then run ACE2 with the same factorial forcings; if ACE2's separated responses do not match the physics model's, the forced-response claim holds only for the combined historical forcing. The paper's Figure 14 already hints at this, since fixing CO2 removes most near-surface warming.","tokens_in":24567,"feed_emoji":"🌍","tokens_out":14885,"duration_ms":127322,"temperature":0.7,"pith_summary":"This paper tries to establish that a single learned atmospheric model can do what has been reserved for physics-based climate models: run stably for decades while responding to changing sea surface temperature and CO2. If true, a 450-million-parameter emulator could simulate climate variability and forced response at about 1,500 simulated years per wall-clock day, making century-scale ensembles and rare-event studies cheap. ACE2 is autoregressive, runs at 1° resolution with eight vertical layers and 6-hour steps, and is trained separately on the ERA5 reanalysis and on an AMIP-style SHiELD simulation. In 81-year rollouts it reproduces the reference datasets' temperature and moisture trends, the ENSO-driven precipitation pattern, tropical cyclone frequency, the Madden-Julian Oscillation, and sudden stratospheric warmings, while exactly conserving global dry air mass and moisture. The paper states plainly that its sensitivities to separately changing SST and CO2 are not entirely realistic, so the forced-response result is established for the combined historical forcing.","feed_headline":"Learned weather model reproduces 80 years of climate variability","feed_subtitle":"ACE2 runs 1,500 simulated years per day and captures El Niño responses plus decade-scale temperature trends.","key_machinery":"The central object is ACE2 itself: an autoregressive Spherical Fourier Neural Operator that maps a 6-hourly atmospheric state plus forcing variables (SST, CO2, solar radiation, surface fractions) to the next state, with a physical-corrector module appended as part of the architecture. The corrector enforces exact global dry-air-mass conservation and a closed global moisture budget by adjusting surface pressure and precipitation and deriving the advective moisture tendency as a residual. The other load-bearing mechanisms are the use of CO2 as an input feature, training on two 80-year datasets with historical SST variability, and a checkpoint-selection criterion based on time-mean climate skill rather than short-term loss. Together these let ACE2 roll out for centuries under changing boundary conditions instead of drifting to a fixed climatology.","core_discovery":"ACE2's central claim is that a model trained only to predict two 6-hour steps ahead can be integrated autoregressively over 81 years and beyond without instability, while tracking the observed atmospheric response to changing boundary conditions. The paper reports that ACE2-ERA5 matches the global-annual mean 2-meter temperature of ERA5 with an $R^2$ of 0.93, that the ENSO-regressed precipitation map is as close to the reference as the reference's own internal variability, and that a 1000-year run under climatological forcing shows no drift in total water path. It generates tropical cyclones, the Madden-Julian Oscillation, and sudden stratospheric warmings as emergent behavior. The authors state the model 'can be stepped forward stably for arbitrarily many steps' and 'accurately reproduces the atmospheric response to El Niño variability and global trends of temperature over the past 80 years.' The same experiments show the separation of CO2 and SST forcing is incomplete: fixing CO2 at its 1940 value removes most near-surface warming and all stratospheric cooling, which is not physically expected.","pith_inferences":["Editorial extension: because SST and CO2 rise together in the 1940–2020 training record, the fixed-CO2 experiment suggests the two forcings are entangled in what the model learns; a factorial training set with historical SST at fixed CO2 and vice versa, which the authors note could be generated from SHiELD, is the natural test of whether separable sensitivities are learnable at all.","Editorial extension: the model is differentiable and cheap, so it invites use in data assimilation or parameter-estimation loops, where many forward integrations are needed.","Editorial extension: the same architecture and conservation constraints could extend to ocean or coupled emulation, since ACE2 itself is atmosphere-only with prescribed SST and sea ice."],"forward_implications":["Century-scale simulations become cheap enough for large ensembles: ACE2 runs about 1,500 simulated years per wall-clock day on one GPU, so separating forced response from internal variability no longer requires thousands of node-hours.","Because ACE2-ERA5 reproduces Madden-Julian Oscillation propagation, tropical cyclone statistics, and polar stratospheric vortex variability, it is a plausible fast platform for subseasonal-to-seasonal predictability studies.","The 4-degree version retains most of the 1-degree model's climate skill at a fraction of the cost, which would make paleoclimate and biogeochemistry applications tractable with a learned emulator.","The hard dry-air-mass and moisture constraints eliminate the long-term drift seen in the earlier ACE model; a 1000-year simulation forced by climatological 1990–2020 boundary conditions shows no drift in total water path.","Weather forecast skill is a separate axis: ACE2-ERA5 sits behind the IFS and GraphCast in medium-range RMSE, so climate fidelity does not automatically buy forecast skill."],"supporting_citations":[{"why":"Provides the ERA5 reanalysis used to train and evaluate ACE2-ERA5.","marker":"Hersbach et al. (2020)"},{"why":"Provides the SHiELD model and the AMIP-style simulation used to train and evaluate ACE2-SHiELD.","marker":"Harris et al. (2020)"},{"why":"Describes the predecessor ACE model that ACE2 extends and that serves as the main baseline.","marker":"Watt-Meyer et al. (2023)"},{"why":"Supplies the Spherical Fourier Neural Operator architecture that ACE2 is built on.","marker":"Bonev et al. (2023)"},{"why":"Provides NeuralGCM as the current comparison for long-run stability and total-water-path error.","marker":"Kochkov et al. (2024)"},{"why":"Supplies the CMIP6/AMIP experimental framework and historical forcing data used for the SHiELD runs.","marker":"Eyring et al. (2016)"},{"why":"Supplies the TempestExtremes tracking package used to detect and evaluate tropical cyclone statistics.","marker":"Ullrich et al. (2021)"},{"why":"Provides the lag-correlation diagnostic used to show eastward-propagating Madden-Julian Oscillation variability.","marker":"Waliser et al. (2009)"}],"fun_headline_variants":["AI climate emulator reproduces 80 years of variability in a day","AI emulator: 1500 years/day, 80-year climate skill","AI runs 80-year climate with accurate ENSO and trends","Stable AI emulator simulates 80 years of climate variability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 1940–2020 record, in which sea surface temperature and CO2 rise together, provides enough signal for the model to learn physically correct separate responses to each forcing, a premise the paper's fixed-CO2 test only partially confirms.","fun_headline_variants_meta":{"raw":{"variants":["AI climate emulator reproduces 80 years of variability in a day","AI emulator: 1500 years/day, 80-year climate skill","AI runs 80-year climate with accurate ENSO and trends","Stable AI emulator simulates 80 years of climate variability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001099,"raw_usage":{"total_tokens":4591,"prompt_tokens":956,"completion_tokens":3635,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":3558}},"tokens_in":572,"tokens_out":3635,"duration_ms":27487,"temperature":1.0,"reasoning_tokens":3558,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:44:24.745343+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The decisive test is to generate paired SHiELD simulations with historical SST plus fixed CO2 and fixed SST plus historical CO2, then run ACE2 with the same factorial forcings; if ACE2's separated responses do not match the physics model's, the forced-response claim holds only for the combined historical forcing. The paper's Figure 14 already hints at this, since fixing CO2 removes most near-surface warming.","supporting_citations":[],"review_version":1}