{"id":"342a92aa-f282-47e6-93fe-fe1a7e453396","arxiv_id":"2412.10656","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A multiple-input residual neural network that combines satellite surface fields with sparse subsurface profiles reconstructs filtered subsurface eddy kinetic energy in the upper 2000 m better than the tested baselines.","lead":"This paper trains neural networks to estimate monthly subsurface ocean eddy kinetic energy from satellite sea surface fields plus sparse subsurface profiles, and finds that a multiple-input residual network matches reanalysis best. The result matters because deep eddy energy is hard to observe globally and is needed to improve ocean climate model parameterizations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MI models receive filtered subsurface velocities and gradients at the exact depths where EKE is predicted, so reported deep-ocean skill may reflect learning a local closure on the target field rather than reconstructing unobserved EKE.","rationale":"I agree with the reader's weakest_assumption: the chief risk is that the deep-ocean skill of the MI models is inflated by feeding them filtered subsurface velocity fields and gradients at the same depths and months as the EKE target. This concern is not about the architecture or the reproducibility of the training, which are credibly documented with deposited code and a clear train/test split. Rather, it is about interpretation: the abstract and conclusions frame the result as reconstruction of unobserved subsurface EKE, whereas the input set already contains variables that Eq. 5 and Eq. 6 tie almost directly to the target. An ablation or direct-computation check can settle whether the MI-ResNet is learning a physically interesting closure or effectively reading off the target through a learned proxy. I do not think this requires rejection, because the paper is transparent about including subsurface variables and a re-framed claim about learning a subgrid-scale EKE closure could be valid. However, as written, the central claim overstates what is demonstrated, so the reader's CONDITIONAL verdict is appropriate and I would keep it unchanged pending the proposed check.","tokens_in":28035,"tokens_out":8875,"duration_ms":92053,"concrete_test":"Run the deposited code to recompute test-period (2017–2020) R2 with the subsurface branch reduced to density profiles only, removing u_g, v_g and their gradients. If deep-ocean R2 collapses toward the surface-only ResNet values, the near-target velocity/gradient channel is the load-bearing input. As a complementary check, compute EKE directly from Eq. 5 using the exact u_g, v_g fields supplied as inputs; if that direct field achieves R2 comparable to or above the model's value at every depth, the network is effectively approximating a known identity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim (Abstract; §5) is that MI-ResNet reconstructs subsurface EKE in the upper 2000 m. The load-bearing assumption is that the subsurface branch inputs in §3.2.2 and Table 3 are independent of the target. They are not independent in the relevant sense: the model receives filtered velocities u_g, v_g and their gradients at exactly the five depths where EKE is output, and EKE is defined in Eq. 5 as a quadratic functional of those same velocity fields. Eq. 6 is then invoked to argue that EKE is locally approximated by filtered velocity gradients, so the subsurface channel already carries a near-deterministic proxy of the target. The reported deep skill (e.g., R2 = 0.78 at 2000 m globally) therefore measures how well the network learns a local mapping from filtered velocity/gradient fields to EKE, not how well unobserved subsurface EKE can be reconstructed from surface observables and sparse independent profiles. This also weakens the comparison with FCNN/ResNet, which are deliberately denied that near-target channel. The abstract's phrase 'subsurface climatological variables' does not resolve the issue: the actual inputs are contemporaneous monthly filtered velocities, not climatology. If the intended contribution is a subgrid-scale closure for EKE, that is a legitimate but different claim and should be stated as such.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes multiple-input neural network models (MI-FCNN and MI-ResNet) to estimate spatially filtered subsurface eddy kinetic energy (EKE) from sea surface variables together with vertical profiles of subsurface density and geostrophic velocities, motivated by a Taylor-series expansion of EKE. The models are trained on five regions of a global eddy-resolving reanalysis (GLORYS12V1) over 2001-2016, tested on 2017-2020 and on non-training regions, and transferred to observational inputs. The authors report that MI-ResNet outperforms surface-input FCNN/ResNet and physics-based first-mode models, with test-period R2 around 0.86 at the surface and 0.78 at 2000 m globally.","tokens_in":28318,"tokens_out":5741,"duration_ms":52320,"significance":"If the central claim were valid, the paper would provide a practical tool for mapping filtered subsurface EKE from satellite and gridded Argo data, with useful implications for eddy parameterizations. The study has clear strengths: a held-out temporal test period, evaluation on non-training regions, comparison with several baselines, transfer-learning experiments, and deposit of the main code at Zenodo. However, the validity of the central claim is undermined by the construction of the subsurface input branch, which supplies contemporaneous filtered velocities and gradients at exactly the depths where EKE is predicted; because EKE is defined from those same velocities (Eq. 5) and locally approximated by products of the gradients (Eq. 6), the reported deep-ocean skill largely reflects fitting a local target-proxy mapping rather than reconstructing unobserved EKE from independent observables.","major_comments":[{"comment":"The subsurface branch supplies ρ, u_g, v_g, and their horizontal gradients at exactly the five depths where EKE is output (0, -500, -1000, -1500, -2000 m). Since Eq. (5) defines EKE as a quadratic functional of the filtered velocities and Eq. (6) states that ab - \\bar a \\bar b ≈ α (∂a/∂x_k)(∂b/∂x_k), the filtered velocity gradients at depth are an approximate deterministic proxy for the target EKE at the same depth. The network can therefore learn a fitted local closure from near-target fields, and the deep-ocean R² values in Fig. 12 (e.g., 0.775 at 2000 m globally) do not measure the ability to reconstruct unobserved EKE from independent surface and Argo data. This also invalidates the comparison against the FCNN/ResNet models, which are deliberately denied this near-target channel, and undermines the headline claim in the Abstract and §5. Please remove the contemporaneous subsurface velocity/gradient inputs, use genuine climatological (long-term mean) fields, or explicitly reframe the contribution as a subgrid-scale EKE closure; a control experiment with the subsurface branch withheld at test time is needed to quantify the leakage.","section":"§3.2.2, Table 3, Eqs. (5)–(6)"},{"comment":"The abstract and Table 3 describe the subsurface inputs as 'climatological variables,' but §3.2.3 states that the monthly mean subsurface filtered velocities and gradients are calculated at 1/2° resolution from the same reanalysis period as the target, and the evaluation uses monthly fields contemporaneous with each EKE month. These are not climatological fields. The distinction matters because the interpretation of the transfer-learning results depends on whether the subsurface branch carries time-varying target information. Please correct the terminology and clearly specify which fields (contemporaneous monthly versus long-term climatology) are used in each experiment.","section":"Abstract; §3.2.3, Table 3"},{"comment":"The observational transfer-learning experiments do not describe how the subsurface velocities u_g, v_g and their gradients in Table 3 are obtained from the observational data (SSH, SST, ISAS20 temperature and salinity). Argo measures temperature and salinity, not velocity; the only description in §2.1 is that vertical profiles of velocities can be derived from the thermal wind relation using surface geostrophic velocities as references, but no equation or implementation detail is given for the observational branch. Without this information, the MI-ResNet (Obs.) results in Figures 4-9 and Figure 12 are not reproducible. Please specify the derivation, including the choice of reference velocity and how the gradients are computed after interpolation to the common grid.","section":"§4, MI-ResNet (Obs.) transfer-learning experiments"}],"minor_comments":[{"comment":"The text calls the statistic 'volume-weighted' R², but the displayed definition uses a spatial average ⟨•⟩ over the region; please clarify whether vertical weighting is included or whether the metric is computed separately at each depth.","section":"Eq. (14)"},{"comment":"The term 'Strurm-Liouville problem' is a typo and should be 'Sturm-Liouville problem.'","section":"§3.1"},{"comment":"The caption refers to 'EKEResNet GLORYS,' which should presumably be 'EKEMI-ResNet GLORYS'; the PDF panels also appear to be missing axis labels and units.","section":"Figure 11 caption"},{"comment":"The sentence 'Although a fixed filter kernel is used at all locations, Only the water area is considered...' has an awkward capitalization and word order; please rephrase for clarity.","section":"§2.2"},{"comment":"The legend labels 'GLORYS' and 'Observation' are not defined in the caption; please state explicitly that these denote the input-data source for the MI-ResNet model.","section":"Figure 12 caption"}],"recommendation":"major_revision","confidential_remarks":"The circular-input issue is severe and should be the central focus of the revision. If the authors can reframe the paper as a subgrid-scale EKE closure or retrain the models without the near-target subsurface velocity/gradient inputs, the work may be publishable; in its current form, the reported deep-ocean generalization metrics do not support the stated reconstruction claim. I would also ask the editor to verify that the deposited code implements the architecture and data pipelines described in the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a competent ML study with deposited code and a sensible held-out test protocol, but the central claim needs to be narrowed. Actually new here is the two-branch network that feeds both surface fields and subsurface velocity/gradient profiles to estimate 1-degree spatially filtered EKE from 0 to 2000 m, plus a transfer-learning application to Argo-plus-altimetry observations. That is a worthwhile target, and the authors were careful to test on unseen years (2017-2020) and on regions outside the training boxes.\n\nThe big soft spot is the subsurface branch. The inputs to MI-FCNN and MI-ResNet are the filtered geostrophic velocities and their horizontal gradients at exactly the five depths where EKE is the target. Since EKE is defined from those same velocities (Eq. 5) and the paper's own Taylor expansion says it is locally approximated by products of those gradients (Eq. 6), the network is essentially fitting a local closure on the target field. The reported deep-ocean R2 of ~0.78 at 2000 m mostly tells us how well a network can learn that local mapping, not how well unobserved subsurface EKE can be reconstructed from independent surface observables. The abstract says 'subsurface climatological variables,' but these are contemporaneous monthly filtered velocities, not climatology. That undercuts the comparison with FCNN/ResNet, which are deliberately denied this nearly-deterministic channel.\n\nThere are smaller issues worth noting: the physics baselines (BC1/SM1) are given the true surface EKE, so the NN-vs-physics comparison is not apples-to-apples; the 'global' product excludes the tropics and latitudes beyond 60; and there is no uncertainty quantification. None of these are fatal on their own, but they compound the framing problem.\n\nIf the authors reframe the contribution as a learned subgrid-scale closure for spatially filtered EKE, or if they retrain with subsurface inputs at depths other than the prediction depths (or only climatological stratification), the results would be convincing and useful for eddy parameterization work. As written, the headline claim overstates what the experiment demonstrates.\n\nThis deserves a serious referee—the issue is real but addressable, and the code and data make it easy to verify. I wouldn't cite the headline product, but the closed-form interpretation and the transfer-learning setup are worth a look. Bring it to reading group if you want a discussion of circularity in ML-based reconstruction.","headline":"Capable ML study with a circularity problem in the deep-ocean skill claim; deserves review but needs a reframe.","tokens_in":28844,"tokens_out":2682,"would_cite":false,"duration_ms":24385,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-branch residual neural network reconstructs monthly mean, 1-degree-filtered eddy kinetic energy down to 2000 m across most of the global ocean, outperforming earlier networks and mode-based physical reconstructions.","keywords":["eddy kinetic energy","mesoscale eddies","residual neural network","top-hat spatial filter","satellite altimetry","subsurface reconstruction","transfer learning","ocean reanalysis"],"falsifier":"Retrain MI-ResNet holding out, from both inputs and outputs, all filtered velocity and gradient data at the deepest target level (for example, 2000 m), so the network must predict that level from surface fields and shallower profiles; if R² at 2000 m falls toward the surface-only ResNet's value rather than staying near 0.78, the claimed deep skill traces to the network reading the target itself.","tokens_in":27829,"feed_emoji":"🌊","tokens_out":13568,"duration_ms":112160,"temperature":0.7,"pith_summary":"Mesoscale eddy kinetic energy (EKE) is the standard measure of eddy intensity and a key ingredient in ocean-climate eddy parameterizations, yet its subsurface form has been hard to measure because Argo-style observations are sparse. This paper claims that a multiple-input residual neural network (MI-ResNet) can reconstruct monthly mean, 1-degree-filtered EKE in the upper 2000 m over the 10°S-60°S and 10°N-60°N bands from satellite sea surface fields plus sparse subsurface climatological profiles, and that it does so more accurately than surface-only networks and physics-based baroclinic/surface-mode reconstructions. On the 2017-2020 test period, the paper reports global R² of about 0.86 at the surface and 0.78 at 2000 m using reanalysis inputs, with comparable skill after transfer learning to observational inputs. If the claim holds, it offers a practical way to generate global subsurface EKE fields for eddy parameterization without waiting for dense in-situ coverage.","feed_headline":"Neural network maps hidden ocean eddy energy to 2 km","feed_subtitle":"Subsurface eddy-energy maps come from satellite fields plus sparse Argo profiles, sharpening ocean climate models.","key_machinery":"The load-bearing object is the multiple-input residual neural network (MI-ResNet): a two-branch architecture whose surface branch processes a 17×17×10 tensor of satellite-derived variables around the target column and whose subsurface branch processes a vertical profile of density, geostrophic velocities, and velocity gradients at five depths. The theoretical thread that ties the inputs to the target is the Taylor-series expansion of the filtered product, $\\overline{ab}-\\bar{a}\\bar{b}\\approx \\alpha \\frac{\\partial a}{\\partial x_k}\\frac{\\partial b}{\\partial x_k}$, which the paper uses to motivate why filtered velocity gradients should carry most of the information needed to predict EKE. In the multiple-input models, those gradients are supplied at exactly the depths where EKE is predicted, so the network is effectively learning a local regression from the pieces of the target itself plus surface information.","core_discovery":"The paper's central claim is that spatially filtered EKE, defined by a top-hat filter at 1 degree as $\\mathrm{EKE}=\\frac{1}{2}(\\overline{uu}+\\overline{vv}-\\bar{u}\\bar{u}-\\bar{v}\\bar{v})$, is learnable as a function of sea surface variables (SSH, SST, bathymetry, surface geostrophic velocities and their gradients) combined with sparse vertical profiles of density, geostrophic velocity, and velocity gradients. Using the Taylor-series relation $\\overline{ab}-\\bar{a}\\bar{b}\\approx \\alpha \\frac{\\partial a}{\\partial x_k}\\frac{\\partial b}{\\partial x_k}$, the authors argue that filtered velocity gradients are the natural bridge between the inputs and the target. The proposed MI-ResNet integrates a surface branch that reads a 17×17 spatial tensor around each location with a subsurface branch that reads vertical profiles at five depths, and it is trained on five high-EKE regions of the GLORYS12V1 reanalysis. The reported result is that this model outperforms the surface-only FCNN and ResNet, the multiple-input FCNN, and the first baroclinic mode (BC1) and first surface mode (SM1) reconstructions in every tested region, with global test-period R² of 0.859 at the surface and 0.775 at 2000 m, and that transfer learning preserves most of this skill when the inputs are switched to satellite and ISAS20 observational fields.","pith_inferences":["Because the subsurface branch supplies velocity gradients at the target depths, a natural next test, not reported here, is to retrain with those gradients withheld and see how much deep skill remains; that experiment would separate genuine surface-driven reconstruction from local regression of the target.","The same architecture should transfer to other filtered quadratic quantities, since any product of filtered variables has the same Taylor-series structure; filtered enstrophy or tracer variance would be plausible candidates.","For practical deployment, the pointwise R² and relative-error numbers should be accompanied by uncertainty maps, because eddy parameterizations respond nonlinearly to EKE magnitude and confidence intervals matter for climate use."],"forward_implications":["If the MI-ResNet result is right, monthly maps of 1-degree-filtered EKE down to 2000 m can be produced for the global ocean outside the tropics from data that already exist: satellite SSH and SST plus gridded Argo climatology.","The reported improvement over BC1 and SM1 means that mode-decomposition reconstructions, which lose skill in the deep ocean, can be replaced by a data-driven vertical structure that keeps R² above roughly 0.6-0.8 at 2000 m.","Transfer learning from the reanalysis-trained network to observational inputs would let the method be applied without retraining from scratch in regions where only altimetry and Argo products are available.","Because the same inputs are global, the model can supply EKE fields to GM-style and backscatter eddy parameterizations, including estimates of the vertical structure of eddy mixing.","The paper's own discussion extends the framework to other subsurface variables such as temperature, salinity, currents, subgrid stress, and biochemical tracers."],"supporting_citations":[{"why":"Supplies the GLORYS12V1 eddy-resolving reanalysis, which provides the training inputs, target EKE, and ground truth for validation.","marker":"Jean-Michel et al. 2021"},{"why":"Establishes the coarse-graining/top-hat spatial filter formulation the paper uses to define scale-dependent EKE.","marker":"Aluie et al. 2018"},{"why":"Derives the Taylor-series expansion of filtered products that motivates using filtered velocity gradients as predictors.","marker":"Jakhar et al. 2024"},{"why":"Provides the first-surface-mode vertical-structure method and the 1-degree filtering convention used as a comparison baseline.","marker":"Stanley et al. 2020"},{"why":"Represents the physics-based surface-mode reconstruction of subsurface EKE that the MI models are benchmarked against.","marker":"Groeskamp et al. 2020"},{"why":"Supplies the residual-network approach for inferring eddy quantities and the R²/relative-error evaluation metrics.","marker":"George et al. 2021"},{"why":"Shows residual CNNs can infer subsurface flow structure from sea surface height, a direct architectural precedent.","marker":"Manucharyan et al. 2021"},{"why":"Provides the transfer-learning framework used to fine-tune the reanalysis-trained model on observational inputs.","marker":"Pan & Yang 2010"},{"why":"Provides the ISAS20 gridded Argo-based temperature/salinity fields used as observational subsurface input.","marker":"Fabienne et al. 2016"}],"fun_headline_variants":["Multiple-input ResNet maps eddy energy down to 2 km","Deep learning maps hidden ocean eddy energy to 2 km","Fusing surface and profile data maps eddy energy to 2 km","Global subsurface eddy energy mapped by deep learning","ResNet unlocks ocean eddy energy hidden below 2 km"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that it is legitimate to feed the model filtered subsurface velocities and their gradients at exactly the depths where EKE is predicted; since EKE is defined from those velocities, the deep inputs are near-deterministic proxies of the target, and if the intended task is reconstruction from independent surface observables this premise fails.","fun_headline_variants_meta":{"raw":{"variants":["Multiple-input ResNet maps eddy energy down to 2 km","Deep learning maps hidden ocean eddy energy to 2 km","Fusing surface and profile data maps eddy energy to 2 km","Global subsurface eddy energy mapped by deep learning","ResNet unlocks ocean eddy energy hidden below 2 km"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001244,"raw_usage":{"total_tokens":5211,"prompt_tokens":1157,"completion_tokens":4054,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":773,"completion_tokens_details":{"reasoning_tokens":3968}},"tokens_in":773,"tokens_out":4054,"duration_ms":23730,"temperature":1.0,"reasoning_tokens":3968,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:44:26.524250+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain MI-ResNet holding out, from both inputs and outputs, all filtered velocity and gradient data at the deepest target level (for example, 2000 m), so the network must predict that level from surface fields and shallower profiles; if R² at 2000 m falls toward the surface-only ResNet's value rather than staying near 0.78, the claimed deep skill traces to the network reading the target itself.","supporting_citations":[{"cited_title":", Bathman, S D","cited_arxiv_id":null,"evidence_quote":"Provides the first-surface-mode vertical-structure method and the 1-degree filtering convention used as a comparison baseline."},{"cited_title":", Siegelman, L","cited_arxiv_id":null,"evidence_quote":"Shows residual CNNs can infer subsurface flow structure from sea surface height, a direct architectural precedent."},{"cited_title":"\\ Yang, Q","cited_arxiv_id":null,"evidence_quote":"Provides the transfer-learning framework used to fine-tune the reanalysis-trained model on observational inputs."}],"review_version":1}