{"id":"2bdbdbbd-da9f-4a1d-802f-4f5fc5338ecb","arxiv_id":"2608.10277","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A stochastic AI atmosphere-ocean emulator trained on 105 years of E3SMv3 reproduces its mean climate and internal variability over an independent 400-year segment, but underestimates rare tropical precipitation extremes.","lead":"This paper builds a machine-learned twin of the E3SMv3 climate model, coupling a fast AI atmosphere to a fast AI ocean. Running roughly 40 times faster than the original on one GPU, it reproduces the average climate and much of the natural variability, but misses the most extreme tropical rainfall.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The variability claim hinges on atmosphere-driven stochasticity being sufficient for ocean and sea-ice internal variability; the paper's own persistent deficits in Gulf Stream SST and MIZ sea-ice variance suggest a missing ocean-internal noise source that the architecture cannot supply.","rationale":"The reader's weakest_assumption identifies the same issue; I agree. I considered the two-seed sampling and the piControl-versus-late-20th-century observation comparison as alternative concerns, but they affect confidence bands and the framing of the mean-state comparison rather than the core mechanism of the paper's novelty. The ocean-internal noise question is load-bearing because it determines whether the architecture can, in principle, achieve the variability it claims to target. The paper is otherwise strong: open data/code, independent 400-year evaluation, and explicit admission of remaining gaps and training-data inconsistencies. A revision that either adds internal ocean noise and shows the gaps close, or analyzes the forced-versus-internal decomposition of E3SMv3's variance and shows the residual is small, would substantially strengthen the central claim. Until then, the current CONDITIONAL verdict is appropriate; I do not recommend changing it.","tokens_in":17597,"tokens_out":5464,"duration_ms":57001,"concrete_test":"Retrain the stochastic configuration with a stochastic noise process injected internal to Samudra (e.g., red-noise perturbations to the ocean/sea-ice state representing unresolved scales, tuned so the uncoupled ocean matches E3SMv3's residual variability), holding all other training settings fixed, and compare 400-year Gulf Stream SST anomaly standard deviation and MIZ sea-ice anomaly ratios against the current atmosphere-only stochastic run. If these metrics move substantially toward the E3SMv3 targets (1.27 K; ratios near 1.0), the atmosphere-only stochasticity is insufficient and the headline variability claim needs qualification; if they remain near 0.74 K and 0.70/0.76, the gap has a different cause and the current attribution is more secure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central novelty—that stochastic training preserves internal variability better than a deterministic baseline—rests on the assumption that the stochastic atmosphere (ACE2S) is the only needed noise source for the coupled ocean. In Section 2.2.4, no noise is injected into Samudra; both ensemble members share identical ocean dynamics and differ only through ACE2S's 5-day mean surface forcing. The authors explicitly flag in Section 4 that adding ocean-internal noise would be physically reasonable and that its importance 'remains be explored.' This is not a peripheral caveat: the remaining, honestly-reported deficits are concentrated exactly where ocean and sea-ice internal dynamics are expected to matter. The Gulf Stream Extension per-gridpoint SST anomaly standard deviation is 0.74 K in the stochastic run versus 1.27 K in E3SMv3 (Section 3.4), and MIZ sea-ice anomaly ratios reach only 0.70 (NH) and 0.76 (SH) of target. If a substantial fraction of this missing variance is internally generated in E3SMv3's ocean/sea-ice, then no amount of atmosphere stochasticity can close the gap, and the claim that 'stochastic coupled emulators can reproduce long-timescale variability with high fidelity' overstates what this architecture demonstrates. The improvement over the deterministic baseline is real and supported, but the causal attribution to atmospheric stochasticity, and the promise of reaching target variability, is unproven. The mean-state claims are not affected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"SamudrACE-E3SMv3 couples the stochastic ACE2S atmosphere emulator with the Samudra ocean/sea-ice emulator, both pretrained on 105 years of E3SMv3 preindustrial control output and then fine-tuned jointly with a probabilistic (CRPS + spectral energy score) objective. The coupled emulator is evaluated on a 400-year independent segment of the same E3SMv3 control run. The paper reports climatological RMSBs of 0.62 K in surface temperature and 0.19 mm/day in precipitation, well below E3SMv3's own biases against observations, and shows that the emulator reproduces the daily precipitation distribution out to the 99.99th percentile except for rare tropical extremes. Relative to a deterministic baseline, the stochastic emulator better sustains ENSO spectral power, Gulf Stream SST variance, and marginal-ice-zone sea-ice variability. Code, trained weights, and data are publicly archived.","tokens_in":17928,"tokens_out":8250,"duration_ms":72877,"significance":"If the main claims hold, this is an important step toward fast, fully coupled climate emulation: a roughly 40x speedup on a single GPU with century-scale stability would enable large ensembles and rapid model iteration. The paper's evaluation design is a strength: a held-out 400-year stationary segment, block-averaged spectral uncertainties, and comparisons against E3SMv3's own 40-year block spread. The authors are also commendably honest about limitations (Nordic Seas bias, tropical precipitation tail, residual SST and sea-ice variance deficits, and the open question of ocean-internal noise). The open-data and open-code practices are exemplary. However, the central claim that stochastic training—rather than the concurrent change to a probabilistic loss or the particular random seed—is responsible for the variability improvements is not fully established by the two-seed comparison, and the remaining variability gaps are in fields where ocean-internal noise would be expected to matter. These issues are correctable with additional experiments or careful hedging, so the paper warrants revision rather than rejection.","major_comments":[{"comment":"The claim that stochastic training is more reliable than deterministic training for ENSO variability rests on only two random seeds per configuration. One deterministic seed produces a plausible spectrum while the other collapses onto an overly regular oscillation, and both stochastic seeds are plausible. With n=2, the collapsed deterministic run could be an unlucky seed rather than evidence of a systematic property of deterministic training, and §3.3's statement that 'We attribute the more realistic ENSO variability ... to the use of a stochastic rather than deterministic atmospheric emulator' goes beyond what these data establish. I recommend either training additional seeds (even three per condition) or reframing the claim as a case study and explicitly stating that the seed count is small. This is load-bearing because the headline 'stochastic training maintains internal variability' is the paper's central novelty and the only direct evidence for it is this comparison.","section":"§3.3, Fig S2"},{"comment":"The stochastic-versus-deterministic comparison changes more than the presence of stochasticity: the stochastic configuration uses ACE2S with a CRPS + energy-score loss and randomly sampled loss windows, while the deterministic baseline uses ACE2 with an MSE loss and fixed four-ocean-step / two-atmosphere-step windows. The attribution in §3.4 that 'Adding stochasticity to the emulator recovers much of this missing variance' therefore conflates stochasticity with the probabilistic objective and the rollout-sampling scheme. A cleaner attribution would require a deterministic run trained with the same probabilistic loss (or a stochastic run trained with MSE), or at least an explicit caveat that the improvement may reflect the combination of these changes. This matters because §4's mechanistic statement that 'the deterministic ocean inheriting its variability from the stochastic atmosphere' presupposes that the improvement is due to the stochastic source rather than the loss function.","section":"§2.2.4, §3.4"},{"comment":"The paper's own reported deficits are concentrated where internally generated ocean and sea-ice variability should be important: Gulf Stream per-gridpoint SST anomaly standard deviation is 0.74 K versus 1.27 K in E3SMv3, and marginal-ice-zone sea-ice anomaly ratios reach only 0.70 (NH) and 0.76 (SH) of target. Because no noise is injected into Samudra and all stochasticity enters through ACE2S's 5-day mean surface forcing, a substantial fraction of the missing variability may be irreducible with the current architecture if E3SMv3's ocean/sea-ice internal variability is partly generated within those components. The authors acknowledge this in §4 ('the importance of such ocean-internal noise ... remains be explored'), but the Abstract's closing claim that 'stochastic coupled emulators can reproduce long-timescale variability with high fidelity' overstates the evidence. I recommend softening that sentence to describe improvement relative to the deterministic baseline rather than high-fidelity reproduction, and explicitly noting that the demonstrated variability is atmosphere-forced only.","section":"§2.2.4, §3.4, §4, Abstract"}],"minor_comments":[{"comment":"There is a typo in the Conclusions: 'variaiblity' should be 'variability', and 'remains be explored' should be 'remains to be explored'.","section":"§4"},{"comment":"The sentence 'further displace the ice edge' should read 'further displaces the ice edge'.","section":"§3.1"},{"comment":"The total energy budget correction imposes a fitted constant residual of 0.09 W/m2; consider clarifying that the model enforces energy balance only up to this imposed residual, not exact conservation.","section":"§2.2.2"},{"comment":"The observation datasets cover different periods (1870-1900 for HadISST and 1985-2014 for GPCP); the text states this, but the caption should also restate the periods or note that they are defined in the text.","section":"Fig 2 caption"},{"comment":"The spectral land-fill diagnostic parameters (n=4, k=5, sigma=1) are presented as defaults; a brief sensitivity check would help show that the reported spectral metrics are robust to these choices.","section":"Text S2"}],"recommendation":"major_revision","confidential_remarks":"This is a strong, honest paper with exemplary open science practices. The main revisions I would ask for are (1) a more careful statement of the stochasticity attribution given the confounded training comparison and the small seed count, and (2) tempering the abstract's 'high fidelity' claim to match the paper's own documented deficits. I do not see any fatal flaw; the empirical benchmarks are well designed and the code/data availability is a model for the field."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of arXiv:2608.10277. This is a solid, useful continuation of the SamudrACE line: they couple ACE2S, the stochastic atmosphere, to the Samudra ocean emulator, fine-tune the coupled system with a probabilistic loss, and evaluate against a 400-year held-out segment of E3SMv3 piControl. The assembly is new, the evaluation is clean, and the data/code are available. Mean climate biases vs E3SMv3 (0.62 K, 0.19 mm/day) are correctly benchmarked against E3SMv3's own observation biases, and the stochastic version clearly beats the deterministic baseline at sustaining ENSO and eddy-region SST variance. The precipitation tail discussion is honest: faithful to the 99.99th percentile, underestimates rare tropical extremes.\n\nThe main soft spot is the one the authors themselves flag: all stochasticity enters through the atmosphere. No noise is injected into Samudra. The residual deficits they report—Gulf Stream gridpoint SST std 0.74 K vs 1.27 K target, MIZ sea-ice ratios 0.70/0.76—sit exactly where ocean-internal variability should matter. So the abstract's claim that stochastic coupled emulators can reproduce long-timescale variability \"with high fidelity\" overstates what is demonstrated. What is demonstrated is that atmosphere-driven stochasticity recovers a substantial fraction of the missing variance, not that the architecture can reach target variability without ocean-internal noise. The authors are upfront about this in Section 4, but it deserves emphasis in revision.\n\nTwo smaller issues. The stochastic-vs-deterministic comparison rests on two seeds per configuration; one deterministic seed collapses, which is striking, but there are no uncertainty bars on the variability metrics. And the model-to-observation bias comparison uses the pre-industrial control against GPCP 1985-2014 without quantifying the forced contribution. Minor for an emulation fidelity study, worth a sentence.\n\nOverall, this is a worthwhile paper that should go to review. The central claim—stochastic training preserves internal variability better than a deterministic baseline, and this transfers to a coupled ocean—holds up. The gap between \"improves\" and \"fully reproduces\" needs to be kept honest in the abstract and conclusions.\n\nRecommend: send to peer review; ask the authors to clarify the atmosphere-only stochasticity limitation and add uncertainty estimates to the seed comparison.","headline":"A strong, honest coupled-emulator study whose central claim holds up, but whose abstract overstates the variability result given that only the atmosphere carries stochasticity.","tokens_in":18524,"tokens_out":1722,"would_cite":true,"duration_ms":17117,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A stochastic coupled emulator reproduces a reference climate model's mean state and centuries-long internal variability, with biases smaller than the model's own errors against observations.","keywords":["stochastic climate emulation","coupled atmosphere-ocean emulator","E3SMv3","internal variability","ENSO","precipitation extremes","sea ice variability","machine learning climate models"],"falsifier":"Add stochastic noise inside the Samudra ocean emulator and retrain under the same protocol; if the Gulf Stream SST anomaly standard deviation rises from 0.74 K toward the target 1.27 K and the marginal-ice-zone sea ice variance ratios rise from 0.70 and 0.76 toward 1, the atmosphere-only assumption is falsified. Conversely, finding a deterministic training seed that yields a realistic Niño 3.4 spectrum would weaken the claim that stochasticity is necessary.","tokens_in":17395,"feed_emoji":"🌊","tokens_out":8139,"duration_ms":69297,"temperature":0.7,"pith_summary":"The paper sets out to show that a fully coupled AI emulator of a climate model can reproduce not only the time-mean climate but also the internal variability that matters for long-term behavior, if the atmosphere component is stochastic rather than deterministic. The authors train the coupled system on 105 years of a pre-industrial control run of a reference earth system model and test it on 400 independent years. On the mean state, the emulator's root-mean-square biases against the reference model (0.62 K surface temperature, 0.19 mm/day precipitation) are much smaller than the reference model's own biases against observations (1.13 K, 1.04 mm/day). With stochastic training the emulator keeps ENSO spectral power, Gulf Stream SST variance, and marginal-ice-zone sea ice variability closer to the reference, while deterministic training tends to damp or collapse that variability. If correct, this makes cheap, fast stochastic emulators a credible tool for generating the long ensembles needed to sample internal climate variability.","feed_headline":"Stochastic AI climate twin tracks variability for 400 years","feed_subtitle":"Stochastic training beats deterministic baselines on mean state, ENSO, and sea ice while running about 40x faster.","key_machinery":"The load-bearing mechanism is ACE2S, a stochastic version of an atmospheric emulator that outputs an ensemble of equally likely states rather than one deterministic forecast, trained with a probabilistic loss combining the fair continuous ranked probability score and a spectral energy score. During coupled fine-tuning, its 5-day mean surface fluxes and wind stress drive two parallel Samudra ocean rollouts, and the ocean returns SST and sea ice fraction. No noise is injected inside Samudra; all stochasticity enters through the atmosphere. This mechanism carries the argument because it is what distinguishes the stochastic runs from the deterministic baseline and what recovers ENSO and eddy variability.","core_discovery":"The central discovery is that replacing a deterministic atmospheric emulator with a stochastic one, and fine-tuning the coupled atmosphere–ocean system with a probabilistic scoring rule, turns the atmosphere into a source of internal variability for the ocean. The resulting system, SamudrACE-E3SMv3 (the Samudra full-depth ocean emulator coupled to the ACE2S stochastic atmosphere emulator), produces a 400-year free-running simulation with no drift in its precipitation distribution, reproduces the Niño 3.4 power spectrum within the spread of the reference model's own 40-year blocks, and recovers about half the Gulf Stream SST variance that the deterministic baseline loses (0.74 K versus 1.27 K anomaly standard deviation in the target). Biases against the reference are smaller than the reference's biases against observations. The paper also documents what is not captured: the emulator underestimates the rarest tropical daily precipitation extremes above roughly 150 mm/day.","pith_inferences":["An implication not drawn by the paper: the failure mode of one deterministic seed, a collapsed overly regular oscillation with normal mean-state skill, suggests deterministic MSE training has multiple solutions with equal mean fidelity, so variability diagnostics should be part of checkpoint and seed selection.","A natural testable extension: injecting noise inside the ocean emulator and comparing the Gulf Stream and sea-ice variance gaps would isolate how much missing variance is due to absent ocean-internal stochasticity rather than to atmosphere forcing.","The tropical precipitation tail deficit, present in both training and evaluation, points to a concrete training fix, such as tail-weighted or up-sampled losses, that could be tested on the same 400-year evaluation protocol.","If the stochastic coupling recipe transfers to another climate model without re-architecting, it may become a general method for producing cheap internal-variability ensembles; the transfer to a second GCM here is a first step toward that conclusion."],"forward_implications":["A 400-year free-running stochastic rollout shows no precipitation drift, so the emulator can stand in for the reference control climate for centennial variability studies.","Because emulator-to-reference mean biases (0.62 K, 0.19 mm/day) are smaller than reference-to-observation biases (1.13 K, 1.04 mm/day), the emulator is a valid substitute for the model in model-vs-observation comparisons.","Stochastic training gives ENSO block-to-block spectral spread comparable to the reference, meaning the emulator reproduces not just the mean spectral peak but the intrinsic randomness of ENSO.","The deterministic baseline's variability collapse is invisible in mean-state metrics, so variability-aware diagnostics are needed to select among emulator training seeds.","Daily precipitation is reliable to the 99.99th percentile globally and into the far tail over CONUS; only tropical extremes above about 150 mm/day are under-represented."],"supporting_citations":[{"why":"Supplies the SamudrACE coupled emulator framework and its GFDL-CM4 baseline, which this paper extends with stochastic training.","marker":"(Duncan et al., 2026)"},{"why":"Provides the ACE2S stochastic atmosphere emulator with the probabilistic training objective.","marker":"(Perkins et al., 2026)"},{"why":"Provides the Samudra full-depth ocean emulator and its pretraining procedure.","marker":"(Dheeshjith et al., 2025)"},{"why":"Describes E3SMv3, the reference GCM whose control simulation supplies all training and evaluation data.","marker":"(Golaz et al., 2026)"},{"why":"Provides the deterministic ACE2 atmosphere emulator used in the deterministic baseline and as the basis for ACE2S.","marker":"(Watt-Meyer et al., 2025)"},{"why":"Supplies the total energy budget correction applied during stochastic fine-tuning.","marker":"(Clark et al., 2026)"},{"why":"Provides the wider Samudra configuration with increased channel widths used for ocean pretraining.","marker":"(Yuan et al., 2026)"},{"why":"Cited as the approach for tail-weighted losses or up-sampling that would address the tropical extreme precipitation shortfall.","marker":"(Sun et al., 2025)"},{"why":"Supplies HadISST surface temperature observations used to quantify the reference model's own bias.","marker":"(Hurrell et al., 2008)"},{"why":"Supplies GPCP v2.3 precipitation observations used for the same bias comparison.","marker":"(Adler et al., 2018)"}],"fun_headline_variants":["Stochastic AI twin mimics 400-year climate variability","Stochastic emulator hits 400-year E3SMv3 run, 40x faster","Stochastic coupling preserves ENSO and sea ice for centuries","AI climate emulator sustains variability, misses rarest rain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that atmosphere-driven stochasticity alone is sufficient to generate the ocean's internal variability; if ocean-internal noise also matters, the remaining variance gaps may not close.","fun_headline_variants_meta":{"raw":{"variants":["Stochastic AI twin mimics 400-year climate variability","Stochastic emulator hits 400-year E3SMv3 run, 40x faster","Stochastic coupling preserves ENSO and sea ice for centuries","AI climate emulator sustains variability, misses rarest rain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000445,"raw_usage":{"total_tokens":2251,"prompt_tokens":946,"completion_tokens":1305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":1230}},"tokens_in":562,"tokens_out":1305,"duration_ms":11545,"temperature":1.0,"reasoning_tokens":1230,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:11:02.252286+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Add stochastic noise inside the Samudra ocean emulator and retrain under the same protocol; if the Gulf Stream SST anomaly standard deviation rises from 0.74 K toward the target 1.27 K and the marginal-ice-zone sea ice variance ratios rise from 0.70 and 0.76 toward 1, the atmosphere-only assumption is falsified. Conversely, finding a deterministic training seed that yields a realistic Niño 3.4 spectrum would weaken the claim that stochasticity is necessary.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SamudrACE coupled emulator framework and its GFDL-CM4 baseline, which this paper extends with stochastic training."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ACE2S stochastic atmosphere emulator with the probabilistic training objective."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the total energy budget correction applied during stochastic fine-tuning."}],"review_version":1}