{"id":"1cd85aac-016d-4d17-b81d-844fc84931c4","arxiv_id":"2412.03795","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A ConvNeXt UNet trained on OM4 ocean model output reproduces full-depth ocean climatology and variability for centuries, while under-responding to climate-change forcing.","lead":"Samudra is an AI emulator trained on a high-resolution ocean climate model that predicts full-depth ocean temperature, salinity, currents, and sea level. It stays stable for simulated centuries and runs about 150 times faster than the parent model, but it underestimates climate-forcing trends.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Century-long 'no drift / stable' result is demonstrated only for repeated near-zero-heat-flux forcing; the paper's own forced runs show weak trends and unstable rollouts, so the headline claim is broader than the evidence.","rationale":"Reading the paper in good faith, the advance is real: Samudra reproduces full-depth ocean structure and interannual variability, is robust to seeds and initial conditions, and the code, weights, and data are publicly released. The central claim, however, is that the emulator is stable for centuries under realistic time-dependent forcing with no drift. The load-bearing condition is that stability under the repeated 1990-2000 near-zero-heat-flux forcing cycle transfers to time-varying or climate-change forcing. That condition is the least secure part of the argument. The 100-year and 400-year control runs in Section 3.2 are driven by a periodic forcing cycle chosen to minimize drift, and the comparison is to the same 10-year repeated OM4 segment, not to a century-long free-running OM4 integration or a transient-forcing integration. The paper's own Section 4 and Supporting Information show that under forced heat-flux increases the emulator's response is too weak, that stronger forcing leads to unstable rollouts, and that alternative configurations that improve trend capture become unstable after decades. These are disclosed limitations, which supports a conditional rather than a reject verdict, but they directly bound the scope of the headline claim. My concern matches the reader's weakest assumption, and I recommend keeping the verdict unchanged: the paper should be accepted only with the stability claim scoped to the repeated near-zero-forcing control configuration, not as a general statement of century-scale stability under time-dependent climate forcing.","tokens_in":15390,"tokens_out":4102,"duration_ms":43300,"concrete_test":"Download the released Samudra weights and force Fthermo with a repeated 10-year cycle drawn from 1975-1985 (a period with non-negligible global heat flux, unlike 1990-2000) for 100 years; compute the global-mean theta-O trend and compare with OM4's 1975-1985 segment. If the emulator drifts, underestimates the trend by more than the 20-50% seen in the 8-year test, or destabilizes, then the century-stability/no-drift claim is specific to the chosen near-zero-forcing cycle and does not generalize to time-dependent forcing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Introduction claim Samudra is 'the first ocean emulator capable of reproducing the full-depth ocean temperature structure and its variability, while running for multiple centuries in a realistic configuration with time-dependent forcing' and that it 'exhibits no drift relative to the truth.' The evidence for century-scale stability (Section 3.2) consists of rollouts forced with a repeated 10-year cycle from 1990-2000, selected specifically because it has near-zero global heat flux (Section 2.5); the 400-year run uses the same cycle. The comparison target is the same repeated 10-year OM4 segment, so this tests equilibration to a periodic zero-drift forcing, not stability or accuracy under transient forcing. The paper itself supplies the counterexample: Section 4 and Figures S16 and S24 show that under increasing heat flux (0.25-1 W/m2/yr) the emulator's trend response is too weak and stronger forcing produces unstable rollouts; the cumulative-forcing and tendency-learning variants improved trends but were unstable after roughly 50-80 years (Figures S23, S27). Thus the load-bearing assumption that control-cycle stability transfers to time-varying climate forcing is not supported and is contradicted by the paper's own forced experiments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Samudra, a global ocean emulator built on a modified ConvNeXt UNet architecture, trained on 1975–2014 output from the OM4 ocean general circulation model. Samudra autoregressively predicts full-depth potential temperature, salinity, sea surface height, and horizontal velocities on a 1-degree grid with a 5-day timestep, using atmospheric forcing as boundary conditions. The authors evaluate the emulator on an 8-year held-out period (2014–2022), compare it against OM4 and, in the Supporting Information, against the GODAS reanalysis, and perform 100-year and 400-year control rollouts forced by a repeated 1990–2000 atmospheric cycle chosen for near-zero global heat flux. They report accurate climatologies and interannual variability, century-scale stability, a roughly 150x speedup relative to OM4, and robustness across training seeds and initial conditions. The paper also reports a central limitation: under increasing surface heat flux, the emulator underestimates warming trends and some configurations become unstable after roughly 50–80 years.","tokens_in":15623,"tokens_out":5432,"duration_ms":52189,"significance":"If the central claims hold, Samudra would be a valuable computational tool for large-ensemble studies, data assimilation, and model development for the contemporary ocean, and the released code and weights would make it reproducible. The manuscript includes several commendable practices: an 8-year held-out test, multiple training seeds, a GODAS comparison, and a clear disclosure of the trend-response deficiency in Section 4. The main significance, however, hinges on the scope of the stability claim. The demonstration of century-scale stability is confined to a repeated near-zero-heat-flux forcing cycle, which tests equilibration to a periodic control forcing rather than behavior under transient, climate-change-like forcing. The paper's own forced experiments show that the emulator cannot simultaneously capture trends and remain stable, so the abstract's and introduction's unqualified statements about stability 'for centuries' in 'climate-change simulations' overstate the evidence. With appropriate qualification, the paper would be a solid contribution to the growing literature on learned climate emulators.","major_comments":[{"comment":"The introduction states that Samudra 'can retain skill and remain stable for centuries for experiments equivalent to both control and climate-change simulations,' and the abstract claims the emulator 'is stable for centuries.' The evidence in Section 3.2 consists exclusively of 100-year and 400-year rollouts forced with a repeated 10-year cycle from 1990–2000, a period chosen specifically for its near-zero global heat flux (Section 2.5). This tests stability under a periodic zero-drift forcing, not under climate-change forcing. Section 4 and Supporting Information Figures S23, S24, and S27 show that forced runs with increasing heat flux produce too-weak trends, that the cumulative-forcing variant becomes unstable after about 80 years, and that a tendency-learning variant shows instabilities during short rollouts. The claims in the abstract and introduction should be revised to state that stability is demonstrated for control forcing only, and the phrase 'no drift relative to the truth' should be qualified accordingly.","section":"Introduction and Section 3.2"},{"comment":"The abstract says Samudra 'exhibits no drift relative to the truth.' In the 8-year test, the emulator underestimates the global-mean potential temperature trends by 20–50% at most depths (Section 3.1, Figures S1 and S3), and the GODAS comparison in the Supporting Information (Figures S21 and S22) shows that the emulator loses track of the observed warming trend and exhibits larger errors than OM4. In the 100-year control run, the 'truth' is the repeated 10-year OM4 segment, not an independent century-scale integration. The 'no drift' claim should be limited to the repeated near-zero-heat-flux control experiment; otherwise, it is contradicted by the paper's own results.","section":"Abstract and Section 3.1"},{"comment":"The reported 150x speedup compares 8 days of OM4 integration on 4,671 CPU cores with 1.3 hours for Samudra on a single 40GB A100 GPU. Because the two are run on different hardware types and the emulator uses a much larger timestep and coarser grid, the speedup is not a like-for-like comparison. The manuscript should either provide an estimate on comparable hardware or explicitly state that the speedup is hardware-dependent. In addition, Section 2.1 says OM4 used a 20-minute timestep, while Section 4 says '5 day time step (vs. 15 minutes in OM4)'; these statements are inconsistent and should be reconciled.","section":"Section 4: speedup comparison"},{"comment":"The paper claims Samudra is 'the first ocean emulator capable of reproducing the full-depth (from the surface down to the ocean floor) ocean temperature structure and its variability, while running for multiple centuries in a realistic configuration with time-dependent forcing.' The qualifier 'realistic configuration with time-dependent forcing' is potentially misleading because the multi-century runs use a repeated 10-year atmospheric forcing cycle, and the primary configuration struggles with transient forcing. The sentence should be rephrased to accurately reflect the experiment design, e.g., 'with prescribed atmospheric forcing from a reanalysis product, over control-style repeated forcing.'","section":"Introduction, 'first' claim"}],"minor_comments":[{"comment":"The loss function is garbled in the text: 'Lt = P NX n=1 1 C Y X ...' is not a properly typeset equation. The symbol 'P' is not defined in the equation, and the normalization factors are unclear. Please rewrite the equation in standard mathematical notation and define P and N explicitly.","section":"Section 2.4, Eq. (2)"},{"comment":"The sentence 'The global mean temperatures are 3.225 ◦C/yr for Fthermo and 3.215 ◦C/yr for Fthermo+dynamic' uses incorrect units; the values are temperatures, not rates, so the units should be degrees Celsius (◦C), not ◦C/yr.","section":"Section 3.2"},{"comment":"The caption for Figure S24 reads 'OHC trends (same caption as S23),' but Figure S23 is an 8-year test-set comparison while S24 shows 100-year forced runs; the cross-reference is misleading and should be corrected.","section":"Supporting Information, Figure S24 caption"},{"comment":"The phrase 'conservatively remap onto 19 fixed-depth levels' and 'conservatively coarsen the data in time' would benefit from a one-sentence clarification of what 'conservatively' means in this context (e.g., conservation of volume-weighted integrals or fluxes), since the term is used for both spatial and temporal coarsening.","section":"Section 2.1"},{"comment":"The robustness claim 'The emulators’ skill is unchanged when using different seeds and start dates' is supported by standard deviations of 0.0033 and 0.00225 in RMSE, but the absolute RMSE values are not reported in the main text, making it hard for the reader to gauge whether these standard deviations are small relative to the skill. Please report the mean RMSE alongside the standard deviation.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is appropriate for GRL in scope, and the authors are to be commended for releasing code, weights, data, and for including multiple seeds and an external reanalysis comparison. My major concern is the mismatch between the headline stability/trend claims and the actual experimental design: the multi-century runs only use repeated near-zero heat-flux forcing, and the paper's own forced experiments show instability and weak trend response. This is not a fatal flaw, because the limitations are disclosed in Section 4, but the abstract and introduction need substantial rephrasing before the paper can be accepted. I would also encourage the authors to provide a more careful speedup comparison and to fix the typesetting of Eq. (2)."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Samudra is a real step forward for ocean emulation, but you should read the abstract as a claim about control-run stability, not about climate-change response. The architecture is an extension of the group's surface emulator to 19 depth levels, which matters: full-depth temperature/salinity structure, ENSO phase, and thermocline depth are reproduced at useful accuracy in the 8-year held-out test, and the multi-seed robustness plus the GODAS comparison in the SI give credible evidence of contemporary-ocean skill. The code, data, and weights are released, which makes the work reproducible. No formal proofs here, but that is not that kind of paper; the empirical protocol is the content.\n\nThe main soft spot is exactly where the stress-test note lands. The century-scale stability results come from repeatedly forcing the model with a 1990-2000 10-year cycle chosen for near-zero net heat flux. That is a valid control experiment, and the emulator equilibrates rather than drifting. But 'no drift relative to the truth' and 'multiple centuries in a realistic configuration with time-dependent forcing' overstate it. The paper's own forced experiments show the emulator underestimates warming trends by 20-50% on the test period, and stronger heat-flux forcing produces unstable rollouts. So the stability claim does not generalize to time-varying climate forcing, and the authors basically admit this in the discussion. That limits the practical value for climate-change studies, though not for ensembles, spin-up, or data assimilation where contemporary forcing is the target.\n\nI do not think this is a reject. The limitations are disclosed rather than hidden, the GODAS comparison grounds the skill externally, and the trend weakness is a known hard problem shared with other learned climate emulators. The citation pattern looks normal, with self-citations only to prior surface-emulator work that this extends. What the paper needs is a revision that scopes the headline: century stability is demonstrated for repeated control forcing; trend response under imposed warming is an open problem. If the authors fix that framing, the work is publishable.\n\nWould I send to peer review? Yes. It deserves serious referees. The data/code release and careful multi-seed experiments put it above the usual threshold. I would cite it for the full-depth emulator architecture and the evaluation protocol, not for climate-projection ability.","headline":"A genuine full-depth ocean emulator with credible control-run stability, but the abstract oversells it: century-scale 'no drift' is shown only under repeated near-zero-forcing conditions, and the paper's own forced runs expose weak trends and instabilities.","tokens_in":16195,"tokens_out":2122,"would_cite":true,"duration_ms":23014,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the neural-network emulator Samudra reproduces the full-depth ocean temperature structure and variability of the OM4 model, remains stable for centuries, and runs about 150 times faster, though it under-responds to…","keywords":["ocean emulator","machine learning","climate model emulation","full-depth ocean","ConvNeXt UNet","OM4","ENSO variability","long-term stability"],"falsifier":"Force Samudra for 100 years with a sustained 1 W/m2 per year increase in surface heat flux and compare the global ocean heat content trend against OM4 run under the same forcing; the central claim fails if the emulator's trend is systematically more than 20-50% too weak or if the rollout becomes unstable within a few decades.","tokens_in":15157,"feed_emoji":"🌊","tokens_out":7827,"duration_ms":71963,"temperature":0.7,"pith_summary":"This paper claims that a single neural network can stand in for the ocean component of a global climate model across the full ocean depth and over centennial timescales. The network, named Samudra, is trained on output of the OM4 ocean model and predicts temperature, salinity, sea-surface height, and horizontal currents on nineteen depth levels from the surface to the seafloor. On an eight-year held-out test, the emulator reproduces the climatological depth structure and ENSO variability with little drift, and under a repeating ten-year atmospheric forcing cycle it remains stable for 400 years while running about 150 times faster than OM4. If these claims hold, ocean control runs and large ensembles become dramatically cheaper. The paper is candid that the emulator under-reproduces forced warming trends and that adding trend sensitivity tends to destabilize long rollouts.","feed_headline":"AI ocean emulator runs stable for centuries at 150x speed","feed_subtitle":"It reproduces full-depth ocean temperature, salinity, and El Niño variability — but climate trends remain too weak.","key_machinery":"The load-bearing object is the modified ConvNeXt UNet: a U-shaped convolutional network in which each block uses GeLU activations, dilated 3x3 convolutions, batch normalization, and inverted channel bottlenecks, with average-pooling downsampling and bilinear upsampling, periodic padding in longitude, and zero padding at the poles. The model is autoregressive in a two-input/two-output configuration: two previous 5-day ocean states plus the current atmospheric forcing produce the next two states, which are then fed back in. Each channel encodes one variable at one depth level, giving 158 input and 154 output channels for the full model; a separate thermodynamic-only version predicts temperature, salinity, and sea-surface height. The depth-varying land mask keeps land cells at zero. This learned map is what must simultaneously reproduce the model's climatology, its interannual variability, and its long-term equilibrium under repeated forcing.","core_discovery":"The contribution is a global, autoregressive, machine-learning emulator of a full-depth ocean model. The authors train Samudra on 65 years of output from OM4, the ocean component of the CM4 climate model, conservatively remapped to 19 vertical levels and a 1 degree horizontal grid, with a 5-day time step. It predicts potential temperature, salinity, sea-surface height, and the two horizontal velocity components, and is driven by atmospheric wind stress, downward heat flux, and its anomaly. On a held-out 8-year period, the emulator reproduces the depth-latitude structure of temperature and salinity, the upper-ocean response to atmospheric forcing, and the phase and structure of ENSO events; the authors also report that training is robust to random seeds and initial conditions. Under a repeated 10-year atmospheric forcing cycle from 1990-2000, chosen for its near-zero global heat flux, both variants of the emulator stay stable for 100 years, and a 400-year run shows no drift, with century rollouts taking about 1.3 hours on one GPU versus roughly 8 days for OM4 on thousands of CPU cores. The paper's central claim is that this is the first ocean emulator to reproduce the full-depth ocean temperature structure and its variability over multiple centuries in a realistic, time-dependent forced configuration. The same experiments show the main limitation: when heat flux increases steadily, the emulator's warming trend is too weak, and some stronger-forced rollouts become unstable.","pith_inferences":["The near-zero-heat-flux forcing cycle used for the stability tests removes the sustained drift that defines climate change, so centennial stability under this cycle is a weaker result than it appears for forced applications.","The weak trend response may stem from the training data itself: the model still carries initialization adjustment and the atmospheric forcing already reflects the coupled ocean state, so the emulator learns a damped forcing-to-trend map; a promising testable fix is predicting tendencies rather than states and adding a conservation penalty.","A practical test would be to fine-tune Samudra on only the most recent decade of OM4 output and measure whether the trend bias shrinks, since the earlier adjustment period may be contaminating the learned response.","The thermodynamic-only emulator shows more aperiodic variability and less noise than the full model, suggesting that treating fast velocities separately from slow thermodynamics is a workable route toward stable forced runs."],"forward_implications":["Samudra can replace OM4 for control-type ocean simulations and for large ensemble studies, since a century-long run takes about 1.3 hours on a single GPU rather than 8 days on thousands of CPU cores.","Because the emulator reproduces the full-depth temperature structure and interannual variability including ENSO, it can be used to study contemporary ocean variability and extreme events at much lower cost.","The demonstrated stability over 400 years under repeated forcing means the emulator can accelerate spin-up integrations and support perturbed-parameter experiments for model calibration.","Coupling Samudra with an atmospheric emulator would make fast, full coupled-climate surrogates feasible, following the role the paper proposes for emulating the coupled model.","The demonstrated limitation restricts these uses to contemporary-ocean and control settings; forced climate-change projections are not yet supported by this emulator."],"supporting_citations":[{"why":"Supplies the OM4 ocean model whose output is the training data and the ground truth for all evaluations.","marker":"Adcroft et al. (2019)"},{"why":"Provides the surface ocean emulator and the ConvNeXt UNet architecture that Samudra extends to multiple depth levels.","marker":"Dheeshjith et al. (2024)"},{"why":"Establishes the prior ocean climate emulator approach that Samudra extends and identifies the timescale-separation problem that motivates the thermodynamic-only variant.","marker":"Subel & Zanna (2024)"},{"why":"The atmospheric climate emulator ACE, used as the comparison for weak generalization under climate change and as a candidate coupling partner.","marker":"Watt-Meyer et al. (2023)"},{"why":"Defines the OMIP-2 protocol and the JRA reanalysis forcing that generated the OM4 dataset used for training.","marker":"Tsujino et al. (2020)"},{"why":"Describes the CM4 coupled climate model of which OM4 is the ocean component, defining the target for eventual coupled emulation.","marker":"Held et al. (2019)"},{"why":"Provides ACE2-SOM, which shows that coupling an atmospheric emulator to a slab ocean improved trend response, serving as a comparison for the trend-stability trade-off.","marker":"Clark et al. (2024)"}],"fun_headline_variants":["Ocean emulator: 150x faster, stable for centuries, no drift","AI ocean model runs centuries without drift, 150x speed","Full-depth ocean emulator stable for centuries, but trend weak","Samudra: AI ocean emulator with 150x speed and stability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The centuries-long stability is demonstrated only under a repeating 10-year atmospheric forcing cycle from 1990-2000 chosen for its near-zero global heat flux, and the paper's own forced runs show that under steadily increasing heat flux the emulator's trend is too weak and some rollouts become unstable.","fun_headline_variants_meta":{"raw":{"variants":["Ocean emulator: 150x faster, stable for centuries, no drift","AI ocean model runs centuries without drift, 150x speed","Full-depth ocean emulator stable for centuries, but trend weak","Samudra: AI ocean emulator with 150x speed and stability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000429,"raw_usage":{"total_tokens":2224,"prompt_tokens":1007,"completion_tokens":1217,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":1140}},"tokens_in":623,"tokens_out":1217,"duration_ms":11487,"temperature":1.0,"reasoning_tokens":1140,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:05:34.644865+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Force Samudra for 100 years with a sustained 1 W/m2 per year increase in surface heat flux and compare the global ocean heat content trend against OM4 run under the same forcing; the central claim fails if the emulator's trend is systematically more than 20-50% too weak or if the rollout becomes unstable within a few decades.","supporting_citations":[{"cited_title":", Urakawa, L S","cited_arxiv_id":null,"evidence_quote":"Defines the OMIP-2 protocol and the JRA reanalysis forcing that generated the OM4 dataset used for training."},{"cited_title":"ACE2-SOM: Coupling an ML atmospheric emulator to a slab ocean and learning the sensitivity of climate to changed CO$_2$","cited_arxiv_id":"2412.04418","evidence_quote":"Provides ACE2-SOM, which shows that coupling an atmospheric emulator to a slab ocean improved trend response, serving as a comparison for the trend-stability trade-off."}],"review_version":1}