{"id":"9afa5078-a91f-4ad6-b78e-76555b63d358","arxiv_id":"2608.11545","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DLESyM-Ocean is a probabilistic deep learning ocean-sea ice emulator trained with a patch energy score loss that stays stable over multi-year simulations.","lead":"This paper introduces DLESyM-Ocean, a deep learning model that simulates global upper ocean and sea ice from atmospheric conditions, generating many possible future states at low cost. It is a step toward fast probabilistic Earth system models that could forecast marine heatwaves, El Nino, and sea ice extremes months ahead.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Skill claim rests on sufficiency of three atmospheric forcings; 2023 El Niño underprediction shows this assumption can fail exactly when the system is most active.","rationale":"The reader's weakest_assumption identifies the sufficiency and stationarity of the three prescribed atmospheric forcing fields as the load-bearing condition. My analysis converges on the same point, sharpened by the paper's own 2023 El Niño underprediction, which is direct evidence that the assumption can fail. This is more consequential than the calibration overstatement or the lack of code/weights, because it targets the model's predictive power under the exact conditions it is designed for. The concrete test I propose would isolate whether the failure is due to the OLR forcing itself or to internal model drift. The reader's verdict is already CONDITIONAL, and this concern reinforces rather than overturns it; no verdict change is warranted, but the condition should explicitly include verification of forcing sufficiency via the proposed experiment.","tokens_in":27082,"tokens_out":8913,"duration_ms":96776,"concrete_test":"Run the 2023 50-member ensemble with the OLR input replaced by a statistically consistent value (e.g., OLR predicted from observed SST using the 1994-2018 linear regression), or add a perturbation to OLR reflecting the observed SST warming. If the ensemble-mean Niño3.4 anomaly reaches the observed ~2 K by December 2023, the three-field forcing set is insufficient because OLR is the bottleneck; if the underprediction persists, the cause is internal drift or a missing dynamical process, which would require a different fix.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that DLESyM-Ocean is a skillful, well-calibrated ensemble emulator depends on the assumption that the three atmospheric forcing inputs (1000-hPa geopotential height, 10-m wind speed, OLR) together with the oceanic state are sufficient to determine future upper-ocean and sea-ice evolution, and that the learned mapping is stationary. The paper's own strongest failure, the 2023 El Niño underprediction (Fig. 5D, where no member reaches the observed amplitude), is explicitly attributed in Section 4 to a possible mismatch between observed OLR and surface warming or to internal drift. This is not a peripheral error: the 2023 event is the most energetic interannual ocean anomaly in the test set, and the model misses its magnitude. If the forcing set is incomplete (e.g., OLR alone does not capture the ocean-atmosphere energy flux responsible for ENSO growth), then the good 90-day skill reported for 2022 may be specific to the atmospheric forcing of that year and not representative of general skill. The conclusion that the model is stable and skillful is thus conditional on the unverified sufficiency of the three fields; the paper's own evidence indicates this assumption can fail precisely when the system is most active.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"DLESyM-Ocean is a global, probabilistic deep-learning emulator for upper-ocean and sea-ice fields, built on a ConvNeXt U-Net with HEALPix tiling and an almost-fair patch energy score (afPES) loss. The model ingests current and previous ocean/ice states plus three atmospheric forcing fields (1000-hPa height, 10-m wind speed, and OLR) and outputs 4- and 8-day residuals, with ensemble spread generated by conditioning-layer norm noise. Training uses ERA5 and UFS-Replay data from 1994-2018; validation is 2019-2021 and the test period 2022-2023. The paper reports 90-day forecast skill against persistence and climatology, a 5-year stable climatological integration, case studies of the 2019 Blob, the 2023 El Nino transition, and the 2023 Antarctic sea-ice minimum, and a 29-year ENSO integration. The central claims are that the model is stable, skillful, and well-calibrated, with minimal bias relative to reanalyses.","tokens_in":27393,"tokens_out":7037,"duration_ms":66307,"significance":"If the claims are upheld, the paper would demonstrate a fast (3.1M-parameter), stable, probabilistic ocean/sea-ice emulator that can be driven by atmospheric forcing and potentially coupled to AI atmosphere models. The afPES loss and the careful treatment of fair-score degeneracies are a useful contribution to probabilistic ML for Earth systems. The multi-year stability and the diverse ensemble case studies are encouraging. However, the calibration claim is contradicted by the paper's own diagnostics, and the 2023 El Nino underprediction raises a substantial question about the sufficiency of the chosen forcing set. These issues currently limit the strength of the central contribution.","major_comments":[{"comment":"The abstract claims DLESyM-Ocean 'produces a well-calibrated, spatially coherent, and skillful ensemble,' but the paper's own calibration diagnostics show otherwise: Section 3.1 reports that the spread-skill ratio is 'modestly underdispersive for most variables and lead times' (Figures 1I-P, S5), and the rank histograms in Figure S6 are U-shaped, which the caption identifies as underdispersion. Underdispersion means the ensemble is not well calibrated in the standard probabilistic sense; the spatial correlation between spread and RMSE (r=0.85, Figure 2) does not establish calibration. This is a load-bearing claim because the abstract's headline is a well-calibrated ensemble. Please either revise the abstract and conclusions to say 'slightly underdispersive but spatially coherent,' or provide a different calibration metric that supports the original wording. The distinction matters for users who will interpret ensemble spread as predictive uncertainty.","section":"Abstract; Section 3.1, Figures 1I-P, S5, S6"},{"comment":"The 2023 El Nino is the most energetic interannual event in the test set, and no ensemble member reaches the observed Nino-3.4 amplitude by year-end (Figure 5D). The paper attributes this in Section 4 to a 'possible mismatch between observed OLR and surface warming' or 'some drift,' but offers no quantitative test of either explanation. This matters because the central skill claim rests on the sufficiency of the three atmospheric forcings (z1000, w10, OLR) chosen in Section 2.4; if OLR does not capture the air-sea fluxes that drive ENSO growth, the 90-day skill shown for 2022 may not generalize. Please either add a test (e.g., recompute the 2023 case with additional or alternative forcing fields, or compare the model's implied surface heat fluxes against reanalysis) or explicitly narrow the skill claim to the variables and periods where the forcing assumption is supported.","section":"Section 4 (Conclusion) and Section 3.3 (Figure 5D, 6)"},{"comment":"The 29-year ENSO evaluation (Figure 5A) includes the training period (1994-2018), and the paper acknowledges this. The reproduction of the 1997/98 and 2015/16 El Nino events is therefore not an out-of-sample test. The only out-of-sample interannual event, 2023, is underpredicted. To support the claim that the model 'reproduces ENSO variability,' please separate training-period from validation/test-period skill (e.g., show skill only for 2019-2023 for the 2019-initialized run) or discuss the training-period results explicitly as a consistency check rather than predictive evidence. As it stands, the evidence for interannual predictive skill outside the training distribution is limited to the underpredicted 2023 event.","section":"Section 3.3 (Figure 5A)"}],"minor_comments":[{"comment":"The text states channel depths D1=136, D2=64, D3=34, but Figure S1 labels D2=68. Please reconcile the two values.","section":"Section 2.1 and Figure S1"},{"comment":"The main text describes 50-member forecasts initialized weekly from January 2022 through December 2022, while the SI captions refer to 25-member forecasts initialized weekly from 2021-01 through 2023-12. Please clarify which configuration underlies each figure and ensure the calibration discussion is consistent with the ensemble size actually used.","section":"Section 3.1 vs SI Figures S2-S6"},{"comment":"The variables Z in MLP_gamma(Z) and MLP_beta(Z) are not defined in the text; presumably they should be the conditioning noise vector nv. Please clarify.","section":"Equation (2)"},{"comment":"The phrase '600 epochs using 4-10 AR steps (100 epoch per AR step)' should read '100 epochs per AR-step length' for consistency with Table 2.","section":"Section 2.3"},{"comment":"The caption refers to '(K-T) Hovmoller diagrams,' but only panels K and L are shown in the figure; please correct the panel range.","section":"Figure 8 caption"},{"comment":"The paper states that all data are publicly available, but no code or model checkpoints are listed. Given the novelty of the afPES loss and the architecture details, releasing code or weights would substantially strengthen reproducibility.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The abstract's calibration claim is directly contradicted by the body's own diagnostics, so the paper needs a substantive revision before acceptance. I also note the SI/main-text discrepancy in ensemble size and period (25 vs 50 members; 2021-2023 vs 2022), which suggests that some figures may not correspond to the described experiments. The absence of code or model checkpoints is a scope concern for a methods paper, though it could be addressed by adding a code-availability statement during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this one is mostly what it claims, with one important caveat in the abstract. The real new content is the sea ice and subsurface extension of DLESyM, the almost-fair patch energy score loss with a land-sea mask, and the conditional layer norm noise for ensemble generation. It also fixes two known HEALPix artifacts (checkerboarding and seam padding), which is a useful practical contribution. The architecture is incremental, no shame in that, and the evaluation is genuinely careful: held-out 2022-23 forecasts, multiple baselines, and honest disclosure that the ensemble is underdispersive and that every 2023 El Niño member misses the observed amplitude.\n\nThe headline skill claim—that 90-day RMSE beats persistence and climatology under concurrent ERA5 forcing—is supported. So is the multi-year stability claim: 5-year and 29-year rollouts stay climatologically reasonable without blowing up. The 2019 marine heatwave and Antarctic sea ice case studies are convincing demonstrations that the model produces plausible, diverse trajectories under common forcing.\n\nThe soft spots are real but not fatal. First, the abstract says \"well-calibrated\" while the paper's own figures show SSR mostly below 1 and U-shaped rank histograms. That's an overstatement, and it should be fixed before publication. Second, the 2023 El Niño underprediction is the most energetic event in the test set, and no member gets close. The paper attributes this to insufficient OLR information or drift; that is plausible, but it also tells you that the three-forcing-field assumption has limits exactly when the system is most active. Worth stating more explicitly as a limitation rather than as a side comment. Third, no code or weights are released, which limits reproducibility. The 29-year ENSO run includes training-period forcing, though the authors disclose this and point to the 80-day autoregressive training horizon.\n\nThe stress-test concern about forcing sufficiency is fair, but I don't think it sinks the paper. The central claim is about emulation under perfect atmospheric forcing, not operational forecast skill. The model's own admitted failures are where that forcing set is insufficient. A serious referee should push on this and on the calibration wording, but the core is solid.\n\nWho is this for? Anyone working on AI ocean emulators, coupled AI Earth system models, or S2S forecasting. It is a meaningful step forward for the field, and it deserves a real peer review. I'd send it out with a request to fix the abstract, discuss the forcing-sufficiency limitation head-on, and ideally release code and weights.","headline":"Solid within-subfield advance that overstates calibration in the abstract and misses the 2023 El Niño amplitude, but the core multi-year stability and skill results are credible and worth refereeing.","tokens_in":27926,"tokens_out":1345,"would_cite":true,"duration_ms":17161,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A compact neural network trained on reanalysis can emulate the global upper ocean and sea ice probabilistically over multi-year autoregressive simulations, reproducing recent extremes and remaining stable.","keywords":["deep learning","Earth system model","sea ice","upper ocean","probabilistic ensemble forecasting","patch energy score","HEALPix","autoregressive stability"],"falsifier":"Take the 2023 El Niño case and rerun the 50-member ensemble with the outgoing longwave radiation forcing replaced by its 1994–2018 climatology, leaving winds and geopotential unchanged. If the underprediction of the Niño3.4 anomaly persists, the paper's proposed explanation (an OLR anomaly that is too weak or model drift) is falsified; if the forecast degrades further, OLR is carrying part of the signal.","tokens_in":26904,"feed_emoji":"🌊","tokens_out":8832,"duration_ms":85584,"temperature":0.7,"pith_summary":"DLESyM-Ocean is a compact neural network that aims to act as a probabilistic emulator for the global upper ocean and sea ice: given the current and previous ocean state and three atmospheric forcing fields, it predicts states four and eight days ahead, and can be looped autoregressively for years. The paper's central claim is that this trained model produces well-calibrated, spatially coherent ensembles with less error than persistence or ocean climatology at 90-day leads, and that it reproduces the 2023 El Niño transition, the 2019 northeast Pacific marine heatwave, and the 2023 Antarctic sea ice record while remaining stable over five- and twenty-nine-year integrations. If true, the model would supply a fast, probabilistic ocean–sea ice component that could be coupled to AI atmosphere models for subseasonal-to-seasonal forecasting and for generating counterfactual ocean states.","feed_headline":"Sea ice, heatwaves, El Niño: one small AI model reproduces them all","feed_subtitle":"A 3.1M-parameter network forecasts sea ice, heatwaves, and El Niño from just three atmospheric fields.","key_machinery":"The load-bearing mechanism is the almost-fair patch energy score (afPES), a proper scoring rule that generalizes CRPS to multivariate joint distributions by evaluating an energy score over localized 3×3 spatial patches on a HEALPix sphere. The patch restriction avoids the high-dimensional distance concentration that makes global energy scores unstable, while the 'almost-fair' parameterization interpolates between the biased empirical estimator and the exactly fair estimator, whose zero-gradient extremes would otherwise detach the most extreme ensemble members from training. Ensemble members are produced by injecting a 32-dimensional Gaussian noise vector into conditional layer norms (CLNs) that apply channel-wise scale and shift; the network learns residuals on top of a global skip connection, with hard clipping to keep sea ice concentration and thickness in physical ranges. Custom isolatitude padding and a smoother-based upsampling scheme remove the face-seam and checkerboard artifacts that the authors identify in prior HEALPix U-Nets.","core_discovery":"The paper claims that a single U-Net with roughly 3.1 million trainable parameters, trained on ERA5 and UFS-Replay data, learns sufficient autoregressive ocean and sea ice dynamics to simulate present-day upper-ocean and sea ice states for multiple years when driven by the observed atmospheric state. The central methodological choice is training with an almost-fair patch energy score (afPES) instead of a marginal loss like CRPS, which the authors argue is what lets individual ensemble members remain spatially coherent. The evidence offered includes 90-day ensemble forecasts that beat persistence and climatology baselines at nearly all variables and leads; a 5-year, 50-member integration whose climatology and interannual variability track reanalysis; and 29-year integrations that capture the ENSO cycle. The authors also show that the ensembles bracket observed extremes, including the 2023 Antarctic sea ice minimum and the 2019 Blob 2.0 heatwave, while noting a consistent underprediction of the 2023 El Niño amplitude that they attribute to the OLR forcing or model drift.","pith_inferences":["The 2023 El Niño amplitude error is the most informative failure mode: if it stems from the OLR forcing, then the model's three-variable forcing set is insufficient for capturing the energy balance of rapid ENSO growth, and adding surface flux or momentum predictors should be tested.","Because the model is trained on reanalysis, its 'internal variability' is really the statistical spread of learned ocean-atmosphere correlations; whether that spread generalizes to non-stationary climates (for example, a warming Arctic with thinner ice) can only be tested with out-of-sample decades.","One could extend the same architecture to predict additional ocean variables (e.g., biogeochemical tracers) by adding channels and reweighting the afPES, since the loss is agnostic to variable type.","If coupled to an AI atmosphere model, the 4-day and 8-day lead structure could enable online data assimilation, since the residual formulation gives a natural prior for state updates."],"forward_implications":["If DLESyM-Ocean's skill holds when coupled to forecast atmospheric fields rather than perfect reanalysis, it would give AI Earth system models a multi-year-stable ocean and ice component for the first time.","The afPES objective provides a template for training other spatially coherent probabilistic emulators without the Fourier spectral losses used by atmospheric models, which are hard to apply across ocean coastlines.","The ensemble spread scaling with RMSE in eddy-active regions (r=0.85 for SST) implies the model has learned a useful uncertainty map for upper-ocean predictability, not just a mean climatology.","The model's ability to produce diverse subsurface trajectories from identical atmospheric forcing suggests it can generate counterfactual ocean heat-content states for marine heatwave studies, though the authors stop short of claiming those are dynamically equivalent to NWP ensemble members."],"supporting_citations":[{"why":"Supplies the original DLESyM architecture and the ocean-module design that DLESyM-Ocean adapts, including the U-Net structure and residual learning.","marker":"Cresswell-Clay et al. [2025]"},{"why":"Introduces the HEALPix mesh formulation for deep learning weather models; its padding and upsampling schemes are modified here.","marker":"Karlbauer et al. [2024]"},{"why":"Provides the ERA5 reanalysis used for training, forcing, and verification of surface and sea ice fields.","marker":"Hersbach et al. [2020]"},{"why":"Documents the UFS-Replay coupled reanalysis that supplies subsurface temperature, salinity, sea surface height, and sea ice thickness targets.","marker":"Orbe et al. [2017]"},{"why":"Defines the CRPS baseline that the patch energy score generalizes and replaces.","marker":"Hersbach [2000]"},{"why":"Defines fair scores for finite ensembles, the exact estimator whose degeneracies motivate the almost-fair interpolation.","marker":"Ferro [2014]"},{"why":"Introduces the almost-fair CRPS parameterization that the paper adapts to the multivariate patch energy score.","marker":"Lang et al. [2026]"},{"why":"Provides the CESM2 wind-nudging experiment used as a physical baseline for the 2023 Antarctic sea ice case study.","marker":"Espinosa et al. [2024]"}],"fun_headline_variants":["3.1M-param AI simulates ocean and sea ice for years","Tiny AI model recreates ocean extremes from atmospheric data","Deep learning ocean model: multi-year stability, real extremes","AI ocean ensemble: sea ice, heatwaves, El Nino from 3 inputs","Compact neural net predicts ocean and ice from atmospheric forcing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model's skill rests on the idea that three weather variables—the height of the 1000-hPa pressure surface, wind speed 10 meters above the surface, and outgoing heat radiation—plus the initial ocean state are enough to predict how the upper ocean and sea ice will evolve; if the real ocean needs more information from the atmosphere, or if the relationships learned from 1994–2018 change over time, the skill claim breaks.","fun_headline_variants_meta":{"raw":{"variants":["3.1M-param AI simulates ocean and sea ice for years","Tiny AI model recreates ocean extremes from atmospheric data","Deep learning ocean model: multi-year stability, real extremes","AI ocean ensemble: sea ice, heatwaves, El Nino from 3 inputs","Compact neural net predicts ocean and ice from atmospheric forcing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000664,"raw_usage":{"total_tokens":3071,"prompt_tokens":1025,"completion_tokens":2046,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":1956}},"tokens_in":641,"tokens_out":2046,"duration_ms":17038,"temperature":1.0,"reasoning_tokens":1956,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:37:10.901504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 2023 El Niño case and rerun the 50-member ensemble with the outgoing longwave radiation forcing replaced by its 1994–2018 climatology, leaving winds and geopotential unchanged. If the underprediction of the Niño3.4 anomaly persists, the paper's proposed explanation (an OLR anomaly that is too weak or model drift) is falsified; if the forecast degrades further, OLR is carrying part of the signal.","supporting_citations":[],"review_version":1}