{"id":"02140d5a-d6d9-4187-a1ce-c82f19ff13ce","arxiv_id":"2505.10191","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LanTu, a non-autoregressive regional AI ocean forecasting model with a dynamics-enhanced multiscale loss, outperforms operational numerical and global AI ocean forecasts for the Northwestern Pacific for lead times up to 30 days.","lead":"This paper introduces LanTu, a regional AI ocean forecasting model that beats several operational numerical systems for temperature, salinity, currents and sea level up to 10 days ahead. It adds a physics-inspired training constraint that keeps small ocean eddies sharp, a problem that blurs other AI ocean forecasts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central novelty claim—that the dynamics-enhanced multiscale constraint drives LanTu's skill—is untested because no ablation isolates the loss terms; comparisons to LTAR and XiHe change more than one factor at once.","rationale":"The reader's identification of the missing ablation is the most load-bearing concern. The headline performance numbers against IV-TT are credible independent evidence that LanTu as a system forecasts well; no issue arises there except that the 'more than 10 days' phrasing goes beyond the benchmark horizon. But the paper's stated contribution is not merely a new regional AI forecast system—it is a dynamics-enhanced constraint (Eq. 4) that 'effectively improves' performance. The only comparisons designed to probe this are LTAR, which changes autoregressive versus non-autoregressive strategy and adds atmospheric outputs, and XiHe, which changes architecture and training domain. Both violate the ceteris paribus condition. Without an ablation, a plausible alternative explanation—regional capacity, extra input variables, lead-time-specific training, or fitting GLORYS statistics—cannot be excluded. The circularity of using GLORYS as ground truth for eddy dynamics compounds the issue: closeness to GLORYS is expected for any model trained on GLORYS and is not itself proof of dynamical forecasting skill. Nonetheless, because the core engineering result—competitive operational-level forecasts at 1–10 days against independent IV-TT observations—is credible and important, the correct verdict is conditional, not rejection. A focused ablation would settle whether the physics-informed loss earns its central role, and the paper's claims should be adjusted if the ablation shows no effect.","tokens_in":13481,"tokens_out":4975,"duration_ms":51512,"concrete_test":"Retrain LanTu from scratch exactly as described in Section 4.2.4 but set lambda_D = 0 (loss = Eq. (3) only), keeping architecture, input variables, training data, lead-time-specific models, optimizer, epochs, and random seed fixed. Compare this no-dynamics variant against the full LanTu on the IV-TT Class 4 metrics (RMSE for temperature, salinity, SLA, and currents at 1–10 days) and on the DUACS SLA-based eddy metrics used in Figs. 3–5, using 2021–2023. If the RMSE and eddy-skill differences are small relative to baseline variability (e.g., bootstrap 95% confidence intervals overlap), then the increment-correlation term is not the driver and the paper's novelty claim needs revision; if the full model is clearly better, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that dynamics-enhanced learning is what makes LanTu superior rests entirely on Eqs. (2)–(4), but the experiments never vary the loss while holding everything else fixed. LTAR (Section 2.1, Figs. S2–S3) changes the forecasting strategy from non-autoregressive to autoregressive and adds atmospheric outputs; XiHe (Section 2.2, Figs. 3–5) changes architecture, training domain, and possibly training details. Neither isolates the increment-correlation term D_loss. In addition, the eddy-dynamics evidence in Figs. 3–5 uses GLORYS as ground truth, which is also the training target; a model trained on GLORYS can be closer to GLORYS than XiHe is without being a better dynamical forecaster. The independent DUACS SLA comparison in Fig. 1A and Fig. S4 supports overall SLA pattern skill, but it does not separate the proposed constraint from the regional architecture and training procedure. Therefore, the specific causal claim—that incorporating eddy dynamics through Eq. (4) yields the reported improvements—is not established by the present experiments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents LanTu, a regional eddy-resolving ocean forecasting system for the Northwestern Pacific, based on a Vision Transformer / Adaptive Fourier Neural Operator architecture with a non-autoregressive forecasting strategy. The model is trained on GLORYS reanalysis and ERA5 atmospheric forcing. The claimed novelty is a 'dynamics-enhanced multiscale constraint' (Eqs. 2–4) that adds a Pearson-correlation loss on forecast increments to the standard MSE loss. The paper reports that LanTu outperforms operational numerical ocean forecasts (IV-TT Class 4 systems) in temperature, salinity, sea level anomaly, and currents at lead times up to 10 days, with RMSE reductions of 46.78% for temperature and 20.08% for salinity at 6 days. Beyond 10 days, LanTu is compared against persistence and an autoregressive variant (LTAR), and its eddy dynamics are compared against the global AI model XiHe using GLORYS as ground truth. The authors argue that the dynamics-enhanced constraint mitigates the blurring effect and improves eddy splitting/merging forecasts and 3D eddy structure.","tokens_in":13821,"tokens_out":2911,"duration_ms":29783,"significance":"If the results hold, LanTu would be a useful demonstration that regionally trained AI ocean forecasting systems can exceed global operational numerical forecasts in a limited domain, with substantial computational savings. The use of independent IV-TT observations and DUACS altimetry for the headline skill metrics is a genuine strength, as is the evaluation against an established operational intercomparison framework. The central causal claim, however—that the dynamics-enhanced multiscale constraint is what drives the skill improvement—is not isolated by any ablation experiment. Because the comparisons to LTAR and XiHe change multiple factors simultaneously, the paper currently does not establish that the proposed loss term is the source of the improvement. The scientific significance of the methodological contribution therefore remains unproven, even though the empirical forecast skill is credible.","major_comments":[{"comment":"The paper's central claim—that the dynamics-enhanced multiscale constraint improves forecast skill—is not tested. No experiment trains the same architecture with and without the D_loss term while holding all other factors fixed. The LTAR baseline changes the forecasting strategy (autoregressive vs. non-autoregressive) and adds atmospheric outputs, and the XiHe comparison changes architecture and training domain. Thus the reported improvements cannot be causally attributed to the increment-correlation loss. An ablation that removes D_loss (or varies lambda_D) is required to support the abstract's and Section 3's attribution.","section":"Section 4.2.4, Eqs. (2)–(4)"},{"comment":"The eddy-dynamics comparisons use GLORYS reanalysis as ground truth, and GLORYS is also the training target for both LanTu and XiHe. A model trained on a particular reanalysis can be closer to that reanalysis than another model without being a better predictor of the real ocean. While this does not favor LanTu over XiHe in an obvious one-sided way, it does mean that the claimed superiority in capturing eddy splitting, merging, and 3D structure is only a measure of closeness to the training product, not to independent ocean state. An independent eddy census (e.g., from altimetry or Argo-based eddy detection) would strengthen the eddy-dynamics claim. As it stands, the only fully independent evidence for eddy skill is the DUACS SLA comparison in Fig. 1A and Fig. S4, which is limited to surface patterns.","section":"Section 2.2 and 2.3, Figs. 3–5"},{"comment":"The abstract states that LanTu outperforms NOFS and AI-OFS 'with a lead time of more than 10 days.' The direct comparison to IV-TT numerical forecasts is limited to lead times of 1–10 days (Fig. 1, Fig. 2, Fig. S1). Beyond 10 days, the only benchmarks are persistence and the autoregressive LTAR variant (Figs. S2–S3); no same-lead-time head-to-head with an operational NOFS is shown at 11–30 days. The comparison in Fig. S1 of LanTu-30 against IV-TT-10 is a cross-lead-time comparison and does not support the phrasing that the outperformance over NOFS extends beyond 10 days. The claim should be reworded to distinguish 'skillful beyond 10 days' from 'outperforming NOFS beyond 10 days.'","section":"Abstract and Section 2.1"}],"minor_comments":[{"comment":"The term 'Hovmöller' is misspelled as 'Hofmöller' in the caption of Fig. S4 and in Table S2; the spelling should be made consistent.","section":"Supplementary Materials, Tables and Figs."},{"comment":"Equation (1) appears garbled in the manuscript text; the mathematical notation for the LanTu mapping is not rendered correctly and should be fixed.","section":"Section 4.2.3, Eq. (1)"},{"comment":"The caption states 'Forecast lead times of 1-7 days' but panels A–D and E–H appear to cover 1–7 days; the text should clarify whether the 10-day data shown in the figure or only in Fig. S1, as the caption is ambiguous.","section":"Section 2.1, Fig. 2 caption"},{"comment":"The manuscript states that inference code and weights for XiHe are available, but no code or weights are provided for LanTu. Given the paper's emphasis on a new model, releasing code/weights would improve reproducibility.","section":"Materials and Methods, Data availability"},{"comment":"The sentence 'The PSS of LanTu is consistently greater than 0 and shows a positive trend over the 10-30 days lead time' would benefit from a precise definition of the averaging period used to compute each PSS point, since the figure shows a curve over lead time.","section":"Section 2.1, 'PSS' paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a credible empirical system with independent validation for the core skill metrics. The main deficiency is the absence of an ablation isolating the proposed loss term, which is the paper's stated novelty. I would encourage the editor to request the ablation as a condition for publication, or, failing that, to require the authors to substantially soften the causal claims. The comparison to XiHe is useful but confounded; an independent eddy-validation product would materially raise confidence. The overstatement in the abstract regarding lead times beyond 10 days should also be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"LanTu is a regional eddy-resolving ocean forecast model trained on GLORYS/ERA5. The headline result is a 46.8% RMSE reduction in temperature and 20.1% in salinity at 6 days against the IV-TT operational NOFS benchmark, measured on independent Argo and drifter data. That's a solid, useful result. The model also beats persistence out to 30 days, which is what you'd want to see from a non-autoregressive forecaster. The dynamic constraint—adding a Pearson correlation on forecast increments to MSE—is a simple, reasonable idea that deserves attention.\n\nThe problem is that the paper attributes most of its eddy-forecasting improvement to that constraint, but never actually tests it. The LTAR baseline differs by being autoregressive, and XiHe differs by architecture and domain. Neither isolates the loss term. So the specific contribution of Eq. (4) is untested. An ablation with the dynamic term switched off and everything else fixed would settle this, and its absence is a real gap.\n\nThe eddy-dynamics comparisons in Figs. 3-5 use GLORYS as ground truth, the same reanalysis used for training. That's a mild circularity, and it's partially mitigated by the independent DUACS SLA comparisons, which do show LanTu capturing observed eddy patterns. But it still can't support the causal claim about the loss.\n\nThe abstract overstates slightly: claiming outperformance 'with a lead time of more than 10 days' against NOFS is only supported up to 10 days; beyond that, the comparison is against persistence and LTAR. The paper is transparent about this in the text, but the abstract is easy to misread.\n\nOn the positive side, the evaluation framework is appropriate, the verification against IV-TT is credible, and the model description is detailed enough to reproduce. No code or weights are released, which limits direct verification.\n\nWho is this for? Anyone working on AI-based ocean forecasting will want to see this. The result is a new trained model with convincing skill, and the increment-correlation loss is worth trying in other settings. But a serious referee should require an ablation before the novelty claim is accepted.","headline":"LanTu shows convincing operational gains against independent observations, but the paper never isolates its claimed novelty, so the causal story isn't established.","tokens_in":14224,"tokens_out":3733,"would_cite":true,"duration_ms":32836,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A regional AI ocean model trained with an eddy-dynamics loss reports lower forecast errors than numerical and AI operational systems for lead times beyond 10 days.","keywords":["eddy-resolving ocean forecasting","deep learning","mesoscale eddies","dynamics-enhanced multiscale constraint","non-autoregressive forecasting","sea level anomaly","regional ocean modeling"],"falsifier":"Run LanTu's training pipeline with $L_D$ removed (static MSE loss only), keeping architecture, data region, resolution, and the one-model-per-lead-time scheme identical; if the eddy metrics and RMSE advantages largely persist, the paper's central mechanism is not responsible for them.","tokens_in":13271,"feed_emoji":"🌊","tokens_out":6999,"duration_ms":67037,"temperature":0.7,"pith_summary":"LanTu is a regional AI ocean forecasting system for the upper ocean (0–643 m) that predicts temperature, salinity, currents, and sea level anomaly at 1/12° resolution over the Northwestern Pacific. The paper aims to establish that adding a dynamics-informed training constraint—one that rewards matching the spatial pattern of change between the initial state and the forecast—lets an AI model preserve mesoscale eddies instead of smoothing them away. The reported result is that LanTu beats operational numerical forecast systems and a global AI forecast system on all four variable types for lead times of at least 10 days, with RMSE reductions of 46.78% for temperature and 20.08% for salinity at 6 days. If this holds, it matters because eddy-resolving forecasts are costly for numerical models, and a lightweight AI system that keeps eddies sharp for two to four weeks would be a practical tool for fisheries, navigation, and regional ocean services.","feed_headline":"AI ocean forecaster beats numerical models at 10+ days","feed_subtitle":"A pattern-of-change loss keeps mesoscale eddies sharp for weeks of skillful forecasts.","key_machinery":"The load-bearing mechanism is the dynamics-enhanced multiscale constraint, a two-part loss:\n$$\\mathrm{Loss} = \\lambda_S L_S + \\lambda_D L_D, \\quad \\lambda_S = \\lambda_D = 0.5,$$\nwith $L_S$ the mean squared error between forecast and target state and $L_D = 1 - \\mathrm{Corr}(\\Delta \\hat{O}, \\Delta O)$, where $\\Delta \\hat{O}$ and $\\Delta O$ are the forecast and ground-truth increments relative to the initial field. Minimizing $L_D$ maximizes the Pearson correlation between predicted and true change, so the model is forced to get the spatial pattern of evolution right, not just the magnitude. This is paired with a non-autoregressive strategy—one trained model per lead time—and an architecture that mixes Fourier spatial tokens with cross-variable linear layers so atmospheric driver fields and ocean variables interact. Together they counteract the smoothing from norm-based losses that otherwise erases mesoscale eddy signals.","core_discovery":"On its own terms, the paper's central claim is that LanTu, a non-autoregressive vision-transformer model trained separately for each forecast lead time, outperforms both the operational numerical ocean forecast systems represented in the Class 4 verification benchmark and the global AI forecast system on temperature, salinity, sea level anomaly, and current forecasts in its regional domain, with the advantage persisting beyond 10 days and up to 30 days. The quantitative headline is a 46.78% lower RMSE for temperature and 20.08% lower for salinity at 6-day lead time versus the numerical benchmark, and vertical profile errors at 10 days that are 35.93% and 14.81% lower. The paper further claims that LanTu captures mesoscale eddy splitting and merging events and 3D eddy structure that the global AI system blurs or misses, and that its non-autoregressive strategy avoids the cumulative error that makes the autoregressive variant unskillful beyond 10 days. On the paper's terms, these results establish dynamics-enhanced deep learning as a viable regional eddy-resolving forecasting paradigm.","pith_inferences":["A testable extension the paper leaves implicit: an ablation of the dynamic loss would isolate how much of the gain comes from the increment-correlation term versus from regional training, lead-time-specific models, or larger effective domain.","Because the same smoothing problem afflicts other learned spatiotemporal forecasters, the increment-correlation loss could transfer to coastal circulation, sea-ice, or weather nowcasting tasks where conserving small-scale features matters.","The paper's non-autoregressive advantage is partly relative to an unfine-tuned autoregressive baseline; error-correcting or fine-tuned autoregressive models might narrow that gap, so the claim should be read as comparing strategies under equal training effort.","If the loss's effect is confirmed, it suggests that RMSE alone understates forecast quality; pattern-of-change correlation is what preserves dynamically important eddy features."],"forward_implications":["At lead times of 10–30 days, LanTu reports forecasts of temperature, salinity, and sea level anomaly with anomaly correlation above 0.6 and persistence skill scores that stay positive, meaning the model remains informative for two to four weeks.","The paper's eddy case studies imply that a properly constrained AI model can reproduce eddy splitting and merging as discrete events rather than smoothed averages, which is what regional forecasting applications need.","Because LanTu is a single regional model trained once per lead time, it offers a computationally cheap alternative to eddy-resolving numerical systems for regions with limited supercomputing capacity.","The comparison with the autoregressive variant indicates that non-autoregressive, lead-time-specific training is what prevents error accumulation in this setup; autoregressive rollouts lose skill beyond 10 days."],"supporting_citations":[{"why":"Supplies the operational numerical forecasts and observing datasets used for the 1–10 day skill comparison.","marker":"(28)"},{"why":"Provides the Fourier token-mixing architecture adapted as the model backbone for high-resolution inputs.","marker":"(37)"},{"why":"Defines the global AI forecasting baseline whose smoothing and eddy-evolution behavior are compared against.","marker":"(25)"},{"why":"Supplies the global eddy-resolving ocean reanalysis used as training target and forecast ground truth.","marker":"(35)"},{"why":"Provides the atmospheric driver fields used as cross-scale input variables.","marker":"(36)"},{"why":"Explains why norm-based losses smooth forecasts toward statistical averages, motivating the added dynamic constraint.","marker":"(31)"},{"why":"Supplies independent satellite sea level observations used to validate sea level anomaly forecasts.","marker":"(38)"}],"fun_headline_variants":["Dynamics-enhanced AI outscores ocean models past 10 days","AI model with eddy dynamics beats ocean forecasters","Deep learning sees ocean eddies clearly for weeks","LanTu predicts eddies better than numerical and AI","Eddy-resolving AI forecast leads 10 days"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed advantage of the new training loss over all alternative explanations is assumed, because the paper never runs the same model with that loss removed.","fun_headline_variants_meta":{"raw":{"variants":["Dynamics-enhanced AI outscores ocean models past 10 days","AI model with eddy dynamics beats ocean forecasters","Deep learning sees ocean eddies clearly for weeks","LanTu predicts eddies better than numerical and AI","Eddy-resolving AI forecast leads 10 days"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2890,"prompt_tokens":997,"completion_tokens":1893,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":1816}},"tokens_in":613,"tokens_out":1893,"duration_ms":15220,"temperature":1.0,"reasoning_tokens":1816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:13:57.211174+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run LanTu's training pipeline with $L_D$ removed (static MSE loss only), keeping architecture, data region, resolution, and the one-model-per-lead-time scheme identical; if the eddy metrics and RMSE advantages largely persist, the paper's central mechanism is not responsible for them.","supporting_citations":[],"review_version":1}