{"id":"20dd2e25-3e81-4deb-91fd-30dedc0eef05","arxiv_id":"2411.14106","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Back-propagating through the fluid model during training (online learning) makes CNN subgrid parameterizations of two-layer quasi-geostrophic turbulence more skillful and stable than offline training.","lead":"Researchers compared three ways to train a machine-learned replacement for missing ocean turbulence effects: offline training on snapshots, online training that runs the fluid model during training, and a cheaper approximation. Online training produced more accurate and stable model runs, with the full adjoint version doing best.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Online stability advantage rests mainly on unshown no-constraint runs; the removal policy can bias exactly those results, so per-seed no-constraint survival must be reported.","rationale":"The reader's weakest assumption correctly identifies the ensemble-removal policy as a potential source of bias, but in the main constrained experiments the paper states that no seed crashed, so the removal policy does not confound the reported skill and stability scores there. The sharper and more load-bearing version of the concern sits in the no-constraint stability claims of Sec. 3.3 and Sec. 5, which are the primary evidence that online training is inherently stabilizing rather than merely benefiting from the hard constraint. Those claims are made qualitatively and marked '(not shown)', and the removal policy could selectively exclude offline and approximate-online failures from the qualitative summary. The proposed check would settle the issue by forcing a complete per-seed accounting of failures and successes in the no-constraint setting. The paper remains plausible and the conditional verdict stands; the concern is an evidence gap rather than a demonstrated error, so no verdict change is needed.","tokens_in":23180,"tokens_out":6788,"duration_ms":69848,"concrete_test":"Run the Sec. 3.3 no-constraint variants for offline, full online, and approximate online with the same 10 random seeds and architecture used in the constrained ensembles, and report each seed's outcome separately: training convergence, 10-year prognostic completion, and time-to-blow-up if any. Compute the ensemble KE spectra and long-term similarity scores (Eqs. 11-13) with all members included, and report the number of survivors per method. If full online shows 10/10 survival and no grid-scale energy kick-back while at least one non-online member fails on every seed set, the stability claim is confirmed; if online members also fail, or if offline and approximate-online members survive and show no kick-back on rerun seeds, the headline advantage is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest version of the central claim—that full online training confers inherent numerical stability, not just skill—depends on the no-hard-constraint experiments described in Sec. 3.3 ('in the absence of the imposed hard constraint ... (not shown)') and summarized in Sec. 5 ('offline model and the approximately online model do [depend on the hard constraint] (not shown, but see descriptions in body text)'). In the main constrained experiments, the paper explicitly states that none of the chosen seeds crashed (Sec. 3), so the Sec. 2.3 policy of removing failed ensemble members does not bias those headline scores. But for the no-constraint comparison, the failure behavior is the effect being measured, and the text gives no per-seed counts, no failure statistics, and no quantitative q-accumulation values. If failed members are discarded before 'when the CNNs converge' is assessed, offline and approximate-online failure rates could be understated relative to online. Thus the central stability advantage is currently supported mainly by an unreported, potentially selection-prone comparison. The jet-regime transfer (Sec. 5, 'not shown') is a second omitted support, but the no-constraint test is the more load-bearing because it underpins the mechanistic claim that adjoint-based online training removes the need for the PV hard constraint.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"An empirical comparison of offline, full adjoint-based online, and approximately online training for a CNN subgrid parameterization in a two-layer quasi-geostrophic model, following the Ross et al. (2023) benchmark in the eddy-dominated regime. Ten-member ensembles are trained using the same architecture and zero-mean PV hard constraint, then evaluated a priori and a posteriori via Q-Q distribution plots, kinetic energy spectra, spectral energy budgets, and similarity scores. The paper reports that full online training removes the grid-scale energy kick-back, gives better distribution agreement and spectral fluxes, remains stable without the hard constraint, and that the approximate online method inherits some but not all of these benefits.","tokens_in":23446,"tokens_out":5446,"duration_ms":53653,"significance":"The comparison is carefully constructed and largely non-circular: skill is measured on held-out a posteriori statistics (KE spectra, energy fluxes, similarity scores) that are not training objectives. The use of ten-member ensembles, a common architecture, and the public pyqg-JAX/Equinox stack strengthens reproducibility, and the paper provides code and data access. If the stability claims are confirmed, the paper offers concrete evidence that adjoint-based online training is a viable route for ocean eddy closures and quantifies what the approximate online method gives up. The main risk is that the strongest mechanistic conclusion depends on omitted no-constraint experiments, so the empirical core is sound but the headline requires additional reporting.","major_comments":[{"comment":"The central stability claim—that full online training removes the need for the zero-mean PV hard constraint of Eq. (9)—is asserted with '(not shown)' in Sec. 3.3 and repeated in Sec. 5. Section 2.3 states that any ensemble member that fails is removed from score calculations, so in the no-constraint comparison the failure rate is precisely the outcome being measured. The text gives no per-seed survival counts, no blow-up times, and no quantitative q-accumulation values for offline, approximate online, and full online under the no-constraint condition. Without these numbers, the reader cannot rule out selection bias in the reported stability advantage. Please report the per-seed outcomes and the q-accumulation statistics, or explicitly qualify the conclusion as a preliminary observation.","section":"Sec. 3.3, Sec. 5, Sec. 2.3"},{"comment":"Analogous investigations in the jet regime are summarized only as '(not shown)' with claims of 'moderately positive similarity scores' and a resolution dependence. Since the conclusion states online learning is preferable 'over a wider range of conditions,' the jet-regime transfer is part of the paper's scope. Please include at least a compact summary (e.g., a table of similarity scores and stability counts for the jet regime, with and without the hard constraint), or restrict the conclusion to the eddy-dominated regime that is actually documented.","section":"Sec. 5"}],"minor_comments":[{"comment":"'Others details' should be 'Other details'; the key-points bullet 'with the not requiring a differentiable model' is missing a noun, e.g., 'with the approximate approach not requiring a differentiable model.'","section":"Abstract and key points"},{"comment":"Typos: 'a posteori' should be 'a posteriori' (Sec. 1); 'geostrohic' should be 'geostrophic' (Sec. 2.2); 'at least two orders of magnitude layer than' should be 'larger than' (Sec. 2.3).","section":"Sec. 1, Sec. 2.2, Sec. 2.3"},{"comment":"The ensemble standard deviations overlap substantially for many metrics, so the statement that 'it is generally the case that the online models further improve' would be strengthened by a paired significance test or at least a statement of how many of the ten members improve for each metric.","section":"Fig. 5"},{"comment":"The mixed-loss weight α̃ is introduced but no values or normalization are reported for the mixed-loss experiments; please state the values used or note that they were scanned.","section":"Eq. (14)"},{"comment":"The text says Fig. 6 is shown with a linear scale, but the axis labels are in scientific notation with uneven spacing; please clarify the axis scaling or revise the wording.","section":"Sec. 4.1 and Fig. 6"},{"comment":"The phrase 'the choice of seeds is uniform across all three ensembles' should be accompanied by the actual seed values or a repository pointer to the exact seeds, since the stability conclusions depend on these seeds.","section":"Sec. 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical comparison and would be publishable once the missing no-constraint and jet-regime results are supplied; the pattern of relying on '(not shown)' for load-bearing claims is the main weakness. I would not require redoing the main experiments, only reporting what was already run. The removal policy in Sec. 2.3 should be confronted directly in the no-constraint analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper extends online (adjoint-based and approximate) training of neural subgrid closures from one-layer QG to a two-layer baroclinic system, using the Ross et al. (2023) benchmark and pyqg-JAX. That extension is real and useful: it is the first like-for-like comparison of offline, full online, and the List et al. (2024) truncated online method in a setting with baroclinic instability. The main shown results are credible: ten-member ensembles, KE spectra, energy budgets, similarity scores, and Q-Q plots are all reported. The online models clearly remove the grid-scale kinetic-energy kick-back that offline models show (Fig. 3), do better on distribution similarity for Sq (Fig. 2), and hold up reasonably when explicit small-scale dissipation is switched off (Fig. 7). I believe the skill claims.\n\nThe soft spot is the stability claim without the hard constraint. The paper states that offline and approximate online models accumulate q and blow up without the zero-mean constraint, while the full online model is stable, but the quantitative support is '(not shown)' (Sec. 3.3) and repeated in the conclusion. Since Sec. 2.3 says unstable ensemble members are removed from scores, the no-constraint comparison could be biased exactly where the claim matters. The paper says no seeds crashed in the constrained runs, so the removal policy does not affect those headline results, but for the unconstrained runs we need per-seed counts and the actual blow-up statistics. The jet-regime comments in Sec. 5 are also unshown, but that is a minor omission compared to the no-constraint claim. The loss-function explorations (q- and psi-based) are preliminary but honestly reported.\n\nThe reader's take largely matches mine, with one correction: the removal-policy concern is a real issue for the no-constraint experiments only, not for the main constrained comparison, since none of the chosen seeds crashed there. The stress-test headline is fair: the online stability advantage currently rests on an unreported, potentially selection-prone comparison.\n\nThis paper deserves a serious referee. The empirical comparison is an asset, the code and data are available, and the questions it opens (window size, loss design, approximate online) are directly useful for researchers building differentiable ESMs. A referee should ask for the no-constraint per-seed results and the jet-regime numbers before acceptance; the skill claims can stand on what is shown.","headline":"Plausible and mostly well-supported extension of online learning to two-layer QG, but the headline stability-without-constraint claim rests on unshown runs and a removal policy that could bias them.","tokens_in":24000,"tokens_out":3871,"would_cite":true,"duration_ms":34765,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Online training, with the ocean model inside the learning loop, produces more accurate and stable ML subgrid closures than offline regression; the full adjoint-based version eliminates grid-scale energy accumulation and stays stable…","keywords":["online learning","adjoint","subgrid parameterization","baroclinic turbulence","quasi-geostrophic model","machine learning","ocean eddy parameterization","differentiable model"],"falsifier":"Re-run the eddy-regime experiment with a larger number of seeds (e.g., 30) per training strategy, explicitly counting failed training runs and integrating them into the averaged scores; if offline or approximately online models do not show higher failure rates, or if the full online model crashes for some seeds, the claimed stability advantage is not robust. A sharper test: train the full online model without the zero-mean PV hard constraint at longer windows (e.g., K=20, 40); if the grid-scale kick-back reappears or the model becomes unstable, the attribution of stability to the online procedure rather than to window size or the constraint would be undercut.","tokens_in":22992,"feed_emoji":"🌊","tokens_out":6019,"duration_ms":56718,"temperature":0.7,"pith_summary":"This paper argues that machine-learned subgrid parameterizations for ocean turbulence should be trained with the fluid model inside the training loop, rather than as a static regression on precomputed data. In a two-layer quasi-geostrophic baroclinic turbulence test, the full adjoint-based online procedure produces CNN closures that are more accurate and more stable than offline-trained ones. These online-trained closures match the high-resolution kinetic energy spectrum without the grid-scale energy accumulation, show better distribution agreement for the predicted subgrid forcing, and remain stable even when the zero-mean potential-vorticity hard constraint is removed. An approximately online variant that truncates the back-propagation inherits some but not all of these benefits and does not require a differentiable model. If the claims hold, online training offers a reliable and computationally modest route to building ML closures for ocean and climate models, since a training window of only ten time steps suffices.","feed_headline":"Online learning fixes the energy blow-up offline ML closures leave","feed_subtitle":"Full adjoint-based training removes grid-scale kinetic energy kick-back and stays stable without hard potential-vorticity constraints.","key_machinery":"The argument turns on a hybrid dynamical model: a low-resolution two-layer quasi-geostrophic solver that supports algorithmic differentiation, coupled to convolutional neural networks that predict the sub-grid forcing $S_q$, defined as the difference between the coarse-grained advective potential-vorticity tendency and the advective tendency computed from the coarse-grained fields. Training minimizes, over a rolling window of $K$ time steps, the squared error between predicted and diagnosed $S_q$ with separate losses per layer, and the online gradient is obtained by back-propagating through the differentiable solver via the adjoint. All networks are trained with a zero-mean (zero net PV tendency) constraint; the full online model is shown to be the one that does not need this constraint for stability.","core_discovery":"The paper's central claim is that the training procedure itself, not the network architecture or the data, determines the prognostic skill and stability of a learned subgrid closure. In the two-layer quasi-geostrophic 'eddy' benchmark, the full adjoint-based online approach (loss defined over a trajectory of the hybrid fluid–ML model, back-propagated end-to-end through a differentiable solver) removes the grid-scale kinetic energy kick-back that offline models robustly display, makes the predicted subgrid potential-vorticity forcing closer in distribution to the high-resolution truth, and stays stable when the zero-mean potential-vorticity constraint is dropped—while offline and approximately online models accumulate top-layer potential vorticity and blow up. The full online model also retains some skill out-of-sample when explicit small-scale dissipation is switched off. These benefits are achieved with a training window of just ten time steps, so the additional cost over offline training is small.","pith_inferences":["One testable extension is whether online training also removes the need for other constraints (e.g., energy conservation) in more realistic primitive-equation models, since the stability appears to arise from the trajectory-based loss rather than from any specific constraint term.","The paper's comparison could be extended to derivative-free online methods, such as ensemble Kalman inversion; if such methods match the full-online stability, the requirement of a differentiable model would become less binding.","The observation that the approximate online model needs the zero-mean constraint while the full model does not suggests that truncating back-propagation removes precisely the gradient information that prevents top-layer potential-vorticity accumulation; this could be diagnosed by comparing learned sensitivity maps.","The finding that offline training cannot tune away the kick-back while online training removes it hints that the stability property is structural to trajectory-based training rather than a hyper-parameter effect; testing across other closure architectures in the same benchmark would clarify this."],"forward_implications":["If the result carries to operational ocean models, ML subgrid closures can be trained to be stable in prognostic use without ad hoc constraint terms, removing a major obstacle to deployment.","The full online approach removes the grid-scale kinetic energy kick-back that marks unresolved energy accumulation, implying longer and more faithful integrations at coarse resolution.","A training window of ten time steps is enough to obtain these benefits, so the computational overhead of online training is modest and feasible for larger models.","The approximately online variant offers a practical route for existing non-differentiable model codes, inheriting improved energy spectra and fluxes though not the stability without the hard constraint.","Because the full online model keeps skill when explicit small-scale dissipation is removed, it learns a more complete representation of sub-grid processes, requiring less dissipative tuning afterward."],"supporting_citations":[{"why":"Supplies the benchmark framework, CNN architecture, and the zero-mean PV hard constraint that all trained models in this paper use.","marker":"Ross et al. (2023)"},{"why":"Showed a posteriori (online) learning in a one-layer quasi-geostrophic model and introduced the curriculum/window-growing strategy adopted here.","marker":"Frezat et al. (2022)"},{"why":"Introduced online calibration of ML sub-models in a one-layer QG setting and discussed the approximate gradient alternative.","marker":"Ouala et al. (2023)"},{"why":"Proposed the truncated back-propagation 'approximately online' approach that this paper adopts for its second online variant.","marker":"List et al. (2024)"},{"why":"Demonstrated large-scale adjoint-based online training in an atmospheric general circulation model, motivating the differentiable-model requirement.","marker":"Kochkov et al. (2024)"},{"why":"Provides pyqg-JAX, the differentiable two-layer quasi-geostrophic solver that enables end-to-end adjoint back-propagation.","marker":"Otness et al. (2023)"},{"why":"Provides the spectral energy budget diagnostics used to attribute backscatter and dissipation effects in the hybrid models.","marker":"Jansen & Held (2014)"},{"why":"Supplies the adjoint-optimization framing that the online loss and its gradient computation are built on.","marker":"Gunzburger (2003)"}],"fun_headline_variants":["Adjoint-based online learning beats offline for ocean turbulence closures","Online training prevents energy blow-up in learned ocean eddy closures","End-to-end adjoint learning keeps baroclinic turbulence models stable","Training with the fluid model beats offline ML for eddy parameterization","Adjoint online learning fixes grid-scale energy kickback in closures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The stability comparison assumes that removing failed ensemble members from the score calculations does not bias the results; if offline or approximately online models fail more often and are excluded, the reported stability advantage of full online training could be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Adjoint-based online learning beats offline for ocean turbulence closures","Online training prevents energy blow-up in learned ocean eddy closures","End-to-end adjoint learning keeps baroclinic turbulence models stable","Training with the fluid model beats offline ML for eddy parameterization","Adjoint online learning fixes grid-scale energy kickback in closures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2627,"prompt_tokens":963,"completion_tokens":1664,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":1586}},"tokens_in":579,"tokens_out":1664,"duration_ms":11486,"temperature":1.0,"reasoning_tokens":1586,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:31:37.205463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the eddy-regime experiment with a larger number of seeds (e.g., 30) per training strategy, explicitly counting failed training runs and integrating them into the averaged scores; if offline or approximately online models do not show higher failure rates, or if the full online model crashes for some seeds, the claimed stability advantage is not robust. A sharper test: train the full online model without the zero-mean PV hard constraint at longer windows (e.g., K=20, 40); if the grid-scale kick-back reappears or the model becomes unstable, the attribution of stability to the online procedure rather than to window size or the constraint would be undercut.","supporting_citations":[{"cited_title":"Online Calibration of Deep Learning Sub-Models for Hybrid Numerical Modeling Systems","cited_arxiv_id":"2311.10665","evidence_quote":"Introduced online calibration of ML sub-models in a one-layer QG setting and discussed the approximate gradient alternative."},{"cited_title":", Zanna, L","cited_arxiv_id":null,"evidence_quote":"Provides pyqg-JAX, the differentiable two-layer quasi-geostrophic solver that enables end-to-end adjoint back-propagation."},{"cited_title":"\\ Held, I M","cited_arxiv_id":null,"evidence_quote":"Provides the spectral energy budget diagnostics used to attribute backscatter and dissipation effects in the hybrid models."},{"cited_title":"APACrefauthors \\ 2003","cited_arxiv_id":null,"evidence_quote":"Supplies the adjoint-optimization framing that the online loss and its gradient computation are built on."}],"review_version":1}