{"id":"35a25639-a6d2-42d7-a7c1-1a0a3ed2914d","arxiv_id":"2509.02061","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A lightweight 3D climate emulator trained on 30 years of reanalysis reproduces CO2-driven surface warming and stratospheric cooling with long-term stability.","lead":"This paper introduces LUCIE-3D, a compact neural network that emulates the three-dimensional atmosphere and is trained on 30 years of ERA5 data. It reproduces surface warming and stratospheric cooling when CO2 rises, aiming to make fast climate experiments and large ensembles accessible.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training/evaluation split for ERA5 is unspecified; forced-response claim may be in-sample.","rationale":"The reader's weakest_assumption—that the training/evaluation split is unspecified and may render the forced response in-sample—is the single most load-bearing concern. I agree with that assessment. The paper's central claim depends on demonstrating that the model generalizes to CO2 forcing, and the only evidence for that is the reproduction of observed trends and the stationary-CO2 counterfactual. If the evaluation period is inside the training period, the trend reproduction is trivial memorization and the counterfactual is a within-distribution interpolation. The paper does not explicitly report the training years, and Section 4.1 evaluates climatology over 1981–2020, which strongly suggests the 40-year period may overlap the 30-year training window. The stationary-CO2 experiment does not fully mitigate this because it only shows that fixing the input forcing produces a stable trajectory, not that the response is causal. The Zenodo repository may contain the exact data split, but the manuscript itself is incomplete on this critical methodological point. Therefore the reader's CONDITIONAL verdict is appropriate: the claim is plausible and partially supported, but requires a clear out-of-sample demonstration or explicit confirmation of the split. I see no other concern that is more load-bearing than this one; other limitations (stratospheric biases, SST extrapolation failure) are explicitly acknowledged and do not directly undermine the central claim as much as data leakage would. My recommendation is UNCHANGED, meaning I would keep the reader's CONDITIONAL verdict.","tokens_in":13721,"tokens_out":3125,"duration_ms":36377,"concrete_test":"Inspect the Zenodo repository (linked in Code and data availability) to identify the exact years of ERA5 used in the training split. If the training period overlaps any part of 1981–2020, retrain LUCIE-3D on a strictly disjoint period (e.g., train on 1991–2010, test on 1981–1990 and 2011–2020) and recompute the surface temperature trend and stratospheric temperature trend under observed CO2. If the out-of-sample trends do not reproduce the observed +0.20 K/decade surface warming and −0.47 K/decade stratospheric cooling (within a reasonable tolerance), the claim of a learned CO2-forcing relationship is not supported. If the authors can confirm that the training window excludes the evaluation decades, the current claim stands without modification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 4.2) that LUCIE-3D 'has not memorized the effect of CO2 but has learned the relationship between the dynamics and the forcing' rests on comparing climatological differences between 1981–1990 and 2010–2020 (Fig. 2) and on 40-year forced trends (Fig. 3). However, the paper never states the exact years used for training. Section 2 says only '30 years of ERA5 reanalysis data' without a date range, while the evaluation windows span 1981–2020. If the training window overlaps either the baseline decade (1981–1990) or the later decade (2010–2020), the reported climate-change response is partially or entirely an in-sample fit. The stationary-CO2 experiment is a useful internal control, but it does not resolve this leakage: a model trained on data with a strong secular CO2 trend could produce near-zero trend under fixed CO2 simply because its learned baseline state is stable, without having learned a physically generalizable forcing–response mapping. The paper's own Discussion acknowledges that validation on future climates is impossible, but it does not address the more immediate need to demonstrate that the evaluation period is out-of-sample relative to training. Without a clear statement of the data split, the paper's strongest claim is not established as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"LUCIE-3D is a three-dimensional climate emulator built on a Spherical Fourier Neural Operator, trained on 30 years of ERA5 data at eight sigma levels, with atmospheric CO2 and optionally SST as forcing inputs. The paper reports that the model reproduces ERA5 climatology, variability (Kelvin waves, MJO, annular modes), long-term forced responses (surface warming and stratospheric cooling under increasing CO2), and remains stable over 40-year integrations. It also documents spin-up from climatological and 'zero-atmosphere' initial conditions, and a stationary-CO2 control that yields near-zero trends. The central claim (Section 4.2) is that the model has not memorized the CO2 effect but has learned a physically consistent dynamics-forcing relationship.","tokens_in":14005,"tokens_out":3913,"duration_ms":47809,"significance":"If the central claim is supported, LUCIE-3D would be a valuable lightweight tool for rapid climate experimentation, ablation studies, and exploratory paleoclimate/future-scenario work, complementing larger emulators like ACE2 and CAMulator. The paper has notable strengths: the code and trained models are released, the model demonstrably remains stable over 40-year simulations, spin-up from strongly out-of-distribution initial states is examined, and diagnostics such as the Wheeler–Kiladis diagram, annular modes, and PDF tails provide a broad evaluation. The main weakness is that the evidence for the 'learned relationship' claim depends on an unspecified training/evaluation split and on an interpretation of the stationary-CO2 experiment that is not fully warranted. These issues are addressable and do not invalidate the engineering contribution, but they are load-bearing for the paper's headline claim.","major_comments":[{"comment":"The training period is never specified: Section 2 says only '30 years of ERA5 reanalysis data', while Figure 1 and Figure 2 evaluate the 1981–2020 period, and the climate-change response is computed as a difference between 1981–1990 and 2010–2020. If the 30 training years lie inside this 40-year window, the reported forced response is partly or entirely an in-sample fit. This matters because the central claim that LUCIE-3D 'has not memorized the effect of CO2' (Section 4.2) rests on this response being an out-of-sample generalization. Please state the exact training years and, if any overlap exists, provide a non-overlapping train/test evaluation (e.g., train on 1990–2020 and validate on 1981–1989) or otherwise demonstrate that the reported trends are not fitted values.","section":"Sections 2 and 4.2"},{"comment":"The stationary-CO2 experiment is a useful control, but it does not by itself establish a physically generalizable forcing–response mapping. A model trained on a period with a strong secular CO2 increase could learn to map any constant CO2 input to the training-era climatology, thereby producing near-zero trends under fixed CO2, without having learned a physical relationship. The paper's own Discussion (Section 5) acknowledges that 'validation on future climates is not possible' but does not address this non-identifiability. A concrete additional test, such as a response to CO2 values outside the historical training range or a Green's-function-style perturbation, is needed before the phrase 'has learned the relationship between the dynamics and the forcing' can be accepted as stated.","section":"Section 4.2 and Section 5"}],"minor_comments":[{"comment":"Typo in affiliation: 'Allen Insitute' should be 'Allen Institute'.","section":"Title page / affiliations"},{"comment":"Caption contains 'Souther Hemisphere Annualr Mode' — should be 'Southern Hemisphere Annular Mode'.","section":"Figure 7 caption"},{"comment":"Typo: 'olar amplification' should be 'polar amplification'; also '2+' and '4+K' in the text should be '+2 K' and '+4 K' for consistency.","section":"Section 4.2"},{"comment":"Missing space in 'att + ∆t'; also 'full-field precipitation' could be clarified as total precipitation (TP) at the next step.","section":"Section 3.1"},{"comment":"The SSW example says 'inference initialized in 1980' but the training interval is unspecified; please clarify whether the initialization year lies outside the training record, which is directly relevant to the out-of-sample discussion.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":"The data-split ambiguity is the key gate: if the authors can show that the evaluation decades are genuinely out-of-sample (or supply an additional non-overlapping validation), the paper's central claim becomes credible. I would request that the exact training period and train/test split be made explicit in both the main text and the figure captions, and I would ask for a response to the non-identifiability concern about the stationary-CO2 control."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: LUCIE-3D is a credible, incremental extension of LUCIE-2D to full 3D with CO2 forcing. The model trains fast, runs stable, and the authors are transparent about its failures (QBO absent, SSW timing off, moisture extremes underpredicted). The genuinely new bits: Kelvin waves show up in the Wheeler–Kiladis spectrum, and the stationary-CO2 control produces near-flat trends while observed CO2 gives surface warming and stratospheric cooling. The paper also ships code, trained models, and data, which is real evidence and deserves credit.\n\nThe main soft spot is the data split. The paper never states which 30 years of ERA5 go into training, while the climate-change evaluation compares 1981–1990 against 2010–2020. If those decades overlap the training window, the claimed forced response is a fitted value, not a prediction. The stationary-CO2 experiment is a useful internal control, but it sits inside the training distribution and does not resolve leakage. So the sentence in Section 4.2 claiming the model \"has not memorized the effect of CO2 but has learned the relationship between the dynamics and the forcing\" is stronger than the evidence. The authors do note in Section 5 that validating on future climate is impossible; what's missing is showing the evaluation period is out-of-sample relative to training. That is a fixable revision, not a fatal flaw.\n\nMinor gripes: the trends come without error bars or ensemble spread, and the +2/+4K SST perturbation responses are underpredicted, which the authors flag honestly. The stratospheric biases are discussed as openly as one could expect.\n\nOverall, the paper is a solid systems contribution and the authors behave well empirically, but the central claim needs a clearly stated data split and ideally a train-then-validate-on-holdout-decades experiment. I'd send this to peer review; the paper deserves referee time, and the requested changes are straightforward.","headline":"Solid, honest 3D emulator work; the missing train/eval split makes the headline CO2-forcing claim unproven as written.","tokens_in":14534,"tokens_out":2839,"would_cite":true,"duration_ms":34400,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a lightweight 3D climate emulator trained on 30 years of reanalysis data, with CO2 as an input, reproduces observed surface warming and stratospheric cooling under rising CO2 and stays flat when CO2 is fixed.","keywords":["climate emulator","Spherical Fourier Neural Operator","ERA5 reanalysis","CO2 forcing","forced response","stratospheric cooling","long-term stability","Madden-Julian Oscillation"],"falsifier":"Inspect the released code and data record to determine the exact 30-year training interval; if it includes 1981-1990 or 2010-2020, the reported trend reproduction is an in-sample fit and the central claim is not established. Alternatively, train the same model only on data before 1990 and test its CO2-driven trend on 1991-2020: if the trend disappears, the learned-mapping claim fails.","tokens_in":13608,"feed_emoji":"🌡️","tokens_out":4645,"duration_ms":46155,"temperature":0.7,"pith_summary":"LUCIE-3D is a three-dimensional machine-learning climate emulator that takes atmospheric CO2 as a forcing variable. The paper's central claim is that the model has learned a physically consistent mapping between CO2 forcing and atmospheric response, rather than memorizing a fixed trend. Evidence includes surface warming of +0.20 K per decade under observed CO2 (matching ERA5's +0.205), stratospheric cooling, and near-zero temperature trends when CO2 is held constant at 1981 levels. The claim matters because it suggests a cheap, reanalysis-trained emulator could be used for scenario experiments, paleoclimate studies, and fast coupling tests.","feed_headline":"AI reproduces CO2-driven warming from 30 years of weather data","feed_subtitle":"LUCIE-3D matches ERA5's +0.20 K/decade surface trend and shows flat trends when CO2 is fixed.","key_machinery":"The central mechanism is a Spherical Fourier Neural Operator (SFNO) backbone combined with an Euler-integration tendency constraint: the model ingests current prognostic fields plus forcing variables (including monthly CO2 interpolated to six-hourly values) and outputs the fields at the next time step. The architecture is extended to twelve layers with a latent dimension of 256 and trained on eight sigma levels spanning the troposphere and stratosphere, which the paper argues is essential for capturing equatorial Kelvin waves and the vertical structure of the forced response. A two-phase training scheme with validation-loss-scaled weighting and a spectral regularizer is used to maintain long","core_discovery":"The paper argues that LUCIE-3D reproduces the climate system's forced response to increasing CO2: global-mean surface temperature warms at +0.20 K per decade under observed CO2, almost identical to ERA5's +0.205 K per decade, while stratospheric temperature cools at -0.66 K per decade (ERA5: -0.47). When CO2 is held fixed at 1981 values, the surface trend drops to -0.015 K per decade and the stratosphere shows only a weak positive drift. The same separation holds in a variant with prescribed SST forcing. The authors conclude from this contrast that the model 'has not memorized the effect of CO2 but has learned the relationship between the dynamics and the forcing.' They further show the mode","pith_inferences":["A natural testable extension is a strict temporal holdout: train only on years before 1990 and evaluate the CO2 response on 1991-2020. The paper does not report such a split, and the central 'learned mapping' conclusion would be much stronger if the trend persists out-of-sample.","If the learned mapping is real, the emulator could be probed with single-forcing Green's function perturbations to recover its linear response kernel, connecting to the GFMIP-style protocol the paper cites as future work.","The stationary-CO2 experiment is a within-distribution counterfactual, so it demonstrates sensitivity to the forcing variable but does not by itself establish extrapolation to CO2 levels far outside the training range.","The underwhelming response to +2 K and +4 K SST perturbations suggests that despite the CO2 result, the model's extrapolation to strong boundary forcings remains limited, which the paper acknowledges."],"forward_implications":["If the claim holds, a model trained on only 30 years of reanalysis data can produce credible decadal trends in surface temperature and stratospheric temperature under realistic CO2 forcing.","The near-flat response under stationary CO2 suggests the emulator could be used to isolate the forced component of climate change in a way that is transparent and cheap to run.","The ability to spin up from arbitrary initial states, including a 'zero atmosphere,' points toward use in paleoclimate and idealized dynamical experiments where initial conditions are uncertain.","The model's AMIP-style SST-forced variant provides a testbed for coupled ocean-atmosphere emulation, though the paper notes two-way coupling is still missing.","The reported training cost (under five hours on four GPUs) makes systematic ablation studies of architecture and loss design feasible for a wider community."],"supporting_citations":[{"why":"Supplies the original LUCIE-2D architecture, loss formulation, and spectral fine-tuning approach that LUCIE-3D extends.","marker":"(Guan et al., 2024)"},{"why":"Provides the Spherical Fourier Neural Operator backbone used for stable dynamics on the sphere.","marker":"(Bonev et al., 2023)"},{"why":"Defines the T30 Gaussian grid and eight-sigma-level vertical interpolation applied to the ERA5 data.","marker":"(Arcomano et al., 2022)"},{"why":"Supplies the validation-loss-scaled weighting scheme used in the training loss.","marker":"(Ocampo et al., 2024)"},{"why":"ACE2 is the closest reanalysis-trained emulator that also demonstrates forced responses; used as a comparison and a benchmark for stratospheric behavior.","marker":"(Watt-Meyer et al., 2025)"},{"why":"CAMulator provides a climate-model-trained comparison point and motivates the paper's distinction between reanalysis-trained and simulation-trained emulators.","marker":"(Chapman et al., 2025)"},{"why":"Neural GCM is compared for long-term stability and is also noted to underrespond to +2 K and +4 K SST perturbations.","marker":"(Kochkov et al., 2024)"},{"why":"Proposes the Green's function experiment protocol that the paper recommends as a future test of out-of-distribution forcing response.","marker":"(Bloch-Johnson et al., 2024)"},{"why":"Supplies the analysis of long-term instabilities in deep learning climate models and the spectral regularizer used to stabilize LUCIE-3D.","marker":"(Chattopadhyay and Hassanzadeh, 2023)"}],"fun_headline_variants":["AI emulator ties warming to CO2, not memorized","3D climate AI: CO2 drives warming, freezing it stops the trend","Climate emulator: CO2 forcing learned, not just weather patterns","Warming in this AI climate model disappears when CO2 is frozen","AI emulator shows CO2 alone causes warming trend"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The 30 years of ERA5 training data must not overlap the 1981-1990 and 2010-2020 decades used to evaluate the warming trend; the paper does not state the exact training years.","fun_headline_variants_meta":{"raw":{"variants":["AI emulator ties warming to CO2, not memorized","3D climate AI: CO2 drives warming, freezing it stops the trend","Climate emulator: CO2 forcing learned, not just weather patterns","Warming in this AI climate model disappears when CO2 is frozen","AI emulator shows CO2 alone causes warming trend"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000591,"raw_usage":{"total_tokens":2634,"prompt_tokens":799,"completion_tokens":1835,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":1746}},"tokens_in":543,"tokens_out":1835,"duration_ms":13067,"temperature":1.0,"reasoning_tokens":1746,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:55:49.484157+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the released code and data record to determine the exact 30-year training interval; if it includes 1981-1990 or 2010-2020, the reported trend reproduction is an in-sample fit and the central claim is not established. Alternatively, train the same model only on data before 1990 and test its CO2-driven trend on 1991-2020: if the trend disappears, the learned-mapping claim fails.","supporting_citations":[],"review_version":1}