{"id":"c2f6e714-2281-405d-9b2f-fba21f378a92","arxiv_id":"2608.08954","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A diffusion model for coastal sea level has skillful single-station forecasts but joint spatial structure worse than a climatological draw, and the gap persists with training volume in Lorenz-96 experiments.","lead":"This paper trains an AI diffusion model to forecast coastal sea levels and finds the forecasts are locally accurate but wrong about how locations move together. This matters because AI forecast ensembles may look skillful while misrepresenting spatial uncertainty, which could mislead decisions that depend on correlations between stations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The L96 'structural inadequacy' test conditions on exact initial states, so the true conditional distribution is a point mass; the persistent CRPS–VS gap may reflect irreducible generative stochasticity rather than a general failure of learned joint distributions.","rationale":"Good-faith reading: the paper's empirical core is solid—a DDPM trained on coastal SSH has positive CRPS skill and negative variogram skill, and the shuffle decomposition supports the claim that standard marginal metrics can miss joint miscalibration. Block bootstrap and public code are pluses. The concern is not with the real-data result but with the generalization from L96. In L96, conditioning on the exact state makes the true predictive distribution degenerate. The fact that DDPM, MVN, and deterministic emulator all show a VS/CRPS gap then does not isolate a failure to learn joint dependence; it may simply quantify stochastic spread around a deterministic map. The MVN baseline is also not a neutral control: a constant residual covariance cannot represent the state-dependent, eventually zero conditional covariance of a nonlinear deterministic system. The reader's weakest_assumption (L96 transferability) points to the same region; I sharpen it by specifying the degenerate-target mechanism and proposing a stochastic L96 control. This does not overturn the case-study finding, so I agree with CONDITIONAL: the manuscript should either add the stochastic control or explicitly restrict the structural claim to the tested deterministic setting.","tokens_in":11324,"tokens_out":11612,"duration_ms":124130,"concrete_test":"Run the L96 training-volume experiment with the same DDPM architecture and data sizes, but add stochastic forcing to the L96 equations (e.g., additive Gaussian noise with known amplitude) so the true conditional distribution has nonzero, state-dependent spread; then compare CRPS and VS skill across training set sizes. If the CRPS-VS gap persists up to 249k training steps in the stochastic system, the structural-inadequacy claim is supported. If the gap shrinks or closes, the deterministic L96 result is specific to a point-mass target and the paper should soften 'structural inadequacy' to a property of the tested deterministic setup.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest load-bearing premise is the inference in Sections 3.2.1–3.2.2 that the CRPS–VS gap is structural. In the L96 experiments the DDPM and MVN baseline are conditioned on the exact initial state of a deterministic ODE system (Section 2.3), so the correct conditional forecast is a point mass. Any probabilistic generative model must add variance the target does not have; the MVN baseline, whose residual covariance is constant and cannot shrink to zero, is misspecified for the nonlinear L96 map. The observed plateau of VS skill with training volume may therefore be a property of stochastic model classes on a degenerate target, not evidence that learned emulators cannot represent realistic joint conditional distributions with nonzero, state-dependent spread. This matters because the real coastal sea level application is one case; the generalization in the abstract and conclusions ('structural inadequacy', conclusions 2 and 4 in Section 4) leans on this idealized experiment. Without a stochastic-control experiment, the persistence of the gap is not sufficient to establish that more data cannot help in a nondegenerate setting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper trains a denoising diffusion probabilistic model (DDPM) to predict weekly coastal sea-level anomalies at eight US East Coast tide gauges, conditioned on ocean-atmosphere predictors, and evaluates 500-member ensembles on a locked 2022-2025 test set against a climatological reference. Using CRPS, energy score, and variogram score, the authors report positive marginal CRPS skill at all lead times but negative variogram-score skill, and a shuffle-based decomposition indicates that the energy score is largely insensitive to the joint miscalibration. To test whether the gap closes with more data, they run Lorenz-96 experiments with training volumes equivalent to 0.7-170 years and find the CRPS-versus-variogram gap persists for both the DDPM and a multivariate-normal baseline. A dynamical ensemble shows a VS/CRPS ratio near unity, while a deterministic emulator reproduces the gap. The paper concludes that the joint-calibration failure is a structural inadequacy of learned emulators that more data cannot remedy.","tokens_in":11483,"tokens_out":6693,"duration_ms":68873,"significance":"The empirical finding that a probabilistic AI ensemble can be marginally skillful yet jointly miscalibrated is important and timely, and the evaluation design using proper scoring rules, block-bootstrap confidence intervals, and a locked test set is a strength. The pair-level attribution of the variogram-score deficit is also informative. If the structural conclusion survives revision with a nondegenerate idealized target, this would be a valuable caution for the growing use of generative AI forecast ensembles. The main weakness is that the Lorenz-96 experiments, which carry the load of the 'structural' claim, condition on exact initial states of a deterministic system and therefore do not by themselves support the generalization to realistic, nondegenerate conditional forecast problems.","major_comments":[{"comment":"The Lorenz-96 experiments condition on the exact initial state of a deterministic ODE (Section 2.3: the system is integrated with a fourth-order Runge-Kutta scheme with no stochastic forcing, and the DDPM is 'conditioned on the initial state'). The true conditional forecast distribution is therefore a point mass. A probabilistic generative model must add variance that the target does not have, and the MVN baseline's residual covariance is constant and cannot shrink to zero with more data. The persistence and plateau of the CRPS-versus-VS gap with training volume (Figures 3a-3d) is thus an expected property of any stochastic model class on a degenerate target, not evidence that learned emulators cannot represent realistic joint conditional distributions with nonzero, state-dependent spread. This undermines the abstract and Conclusions 2 and 4, which claim a structural inadequacy that is insensitive to data volume. I recommend adding stochastic forcing to the L96 system, conditioning on noisy or partial observations so the true conditional distribution is nondegenerate, or explicitly restricting the claim to degenerate targets.","section":"2.3, 3.2.1"},{"comment":"The abstract states 'positive skill at every station and lead time marginally,' but the only displayed CRPS result (Figure 1a) is the station-mean CRPS skill. No per-station CRPS values are shown in the main text. If per-station results are in the supplement, cite them explicitly; otherwise the 'every station' claim is unsupported and the text should be qualified to the station-mean result.","section":"3.1, Figure 1"},{"comment":"The text states that variogram-score skill is 'negative at all lead times' and uses this to conclude that the joint spatial structure is 'worse than a climatological draw,' but it does not state whether the negative skill is statistically significant at each lead. Because block-bootstrap confidence intervals are shown, the authors should state explicitly which leads have 95% CIs excluding zero. The Figure 1 caption also says the DDPM achieves positive ES skill at all leads, which contradicts the text reporting negative (though not significant) ES skill at 8 and 12 weeks; the caption should be corrected.","section":"3.1, Figure 1c"}],"minor_comments":[{"comment":"There are numerous typographical and encoding issues, including 'submitted toGeophysical Research Letters' (missing space), 'Ni˜ no 3.4' (broken tilde), 'variogramscore' in Section 3.2.1 (missing space), and several formatting artifacts in the reference list; these should be cleaned up before resubmission.","section":"Throughout"},{"comment":"Figure 3 uses two x-axes with duplicated labels ('1 yr 3 yr 7 yr10 yr...' appears twice); simplifying the axis labeling would improve readability.","section":"Figure 3"},{"comment":"Section 2.3 reports 'approximately 1.7 Lyapunov times' at 1 TU, but the Lyapunov time in model time units is not defined; please specify the Lyapunov time so the reader can interpret the predictability regimes.","section":"2.3"},{"comment":"The climatological reference is drawn from the training period (1993-2018) and evaluated on 2022-2025; any trend or mean-state change between periods will affect the absolute skill scores. The relative decoupling of CRPS and VS is not affected by this choice, but the phrase 'worse than climatological draws' should be qualified as 'worse than a training-period climatological draw.'","section":"2.2, 3.1"}],"recommendation":"major_revision","confidential_remarks":"No editor-only concerns beyond the technical issues in the major comments. The degenerate-target problem in the Lorenz-96 experiments is addressable within revision, and the central empirical observation about the coastal sea-level DDPM is potentially valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a serious referee. The thing that is actually new is the empirical, real-world demonstration: a DDPM trained for subseasonal coastal sea level forecasts at eight US East Coast stations gets positive CRPS skill at every lead time while its variogram score skill is negative at every lead time, i.e., the joint spatial structure of the ensemble is worse than a climatological draw. The shuffle-based decomposition is a clean diagnostic—it shows why energy score misses the failure (it is dominated by marginal terms) while variogram score detects it. That combination of a concrete phenomenon and a reusable tool is a genuine contribution.\n\nThe soft spot is the Lorenz-96 'structural inadequacy' inference. The L96 experiments condition on the exact initial state of a deterministic ODE, so the true conditional distribution is a point mass. Both the DDPM and the multivariate-normal baseline necessarily add variance that the target does not have, and the MVN's residual covariance is constant by construction and cannot shrink to zero. The persistent CRPS–VS gap and its insensitivity to training volume are exactly what you would expect from a model class that cannot represent a sharp conditional distribution. That does not tell us whether a learned emulator can capture a realistic conditional distribution with nonzero, state-dependent spread. The claim that the failure is 'structural' and that more data 'cannot be assumed to remedy' it is too strong without a control experiment in a nondegenerate setting (e.g., observation noise on the initial state, or a stochastic version of L96). This is not a fatal flaw in the sea level result, but it should be fixed before publication.\n\nMinor issues: the abstract says 'positive skill at every station' but the displayed evidence is station-mean; the climatological reference pool comes from the training period and may not represent the test period under trend or regime shift; and the architecture details are in a supplement not included in the PDF. The citation pattern is fair, including the relevant prior work on covariance errors in AI models.\n\nBottom line: the empirical finding and the diagnostic are solid, and the paper deserves peer review. I would recommend major revision, mainly to rework the Lorenz-96 analysis and moderate the 'structural inadequacy' wording.","headline":"A genuinely useful empirical finding about marginal-joint decoupling in AI forecast ensembles, but the 'structural inadequacy' conclusion overreaches because the Lorenz-96 control conditions on a deterministic target.","tokens_in":12044,"tokens_out":4506,"would_cite":true,"duration_ms":46683,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI forecast ensembles can pass every marginal test while getting the joint spatial structure wrong—worse than climatology—and more training data will not fix it.","keywords":["ensemble-forecasting","diffusion-models","conditional-distribution","variogram-score","CRPS","joint-calibration","coastal-sea-level","Lorenz-96"],"falsifier":"Retrain the Lorenz-96 DDPM with an additional variogram-score-based penalty or with a covariance-matching postprocessor; if the CRPS-variogram gap closes without reducing CRPS, the paper's claim that the joint failure is a structural property of learned emulators, rather than a property of the training objective, would be refuted.","tokens_in":11071,"feed_emoji":"🌊","tokens_out":7280,"duration_ms":71066,"temperature":0.7,"pith_summary":"Ensemble forecasting is only honest if its members sample the full conditional distribution of outcomes, including how locations vary together. This paper tests that joint requirement for an AI forecast model and finds it fails in a specific way: a diffusion model trained on reanalysis to predict subseasonal coastal sea level is skillful at each of eight tide gauge stations individually, but its spatial correlations are worse than drawing randomly from the historical record. The paper shows this marginal-vs-joint decoupling is invisible to the energy score but exposed by the variogram score, and that it persists in Lorenz-96 experiments from 0.7 to 170 equivalent years of training data. Because a linear baseline reproduces the gap and a dynamical ensemble does not, the authors conclude the joint miscalibration is a structural property of learned emulators rather than a data shortage or a quirk of diffusion models. This matters because standard verification and training metrics would certify such forecasts as good while their spatial risk information is actively misleading.","feed_headline":"AI forecasts pass station tests, fail the spatial test","feed_subtitle":"Skillful at every station, yet its spatial correlations are worse than historical draws—extra data won't fix it.","key_machinery":"The load-bearing object is the denoising diffusion probabilistic model (DDPM)—a generative model that learns to reverse a noise process in order to draw samples from a learned conditional distribution—used here to produce 500-member ensembles of sea surface height anomalies at eight stations simultaneously, conditioned on a 131-dimensional vector of ocean and atmosphere predictors. The evaluation machinery is the trio of proper scoring rules: CRPS for marginal skill, energy score for multivariate skill (shown to be insensitive to joint structure), and variogram score with power $p=0.5$, which is specifically sensitive to the dependence structure and is the detector of the failure. The shuffle-based permutation decomposition separates each score's joint contribution by permuting ensemble members across stations, preserving marginals while destroying spatial correlation. The Lorenz-96 system, its linear multivariate normal baseline, and the dynamical-versus-deterministic emulator comparisons are the controls that convert a single-model observation into the claim of structural inadequacy.","core_discovery":"The paper's central claim is that a trained probabilistic AI forecast ensemble can have positive marginal skill and negative joint skill at the same time. For the eight-station coastal sea level diffusion model, CRPS skill is positive at all leads of 2–12 weeks while variogram score skill is negative at all leads, meaning the ensemble's cross-station spatial structure is worse than a climatological draw. A shuffle-based permutation decomposition shows why standard metrics miss this: the energy score is dominated by marginal contributions, so the DDPM's joint structure barely differs from climatology under ES, whereas the variogram score detects a large, significantly non-climatological joint contribution that is still miscalibrated. Lorenz-96 experiments demonstrate that the CRPS-variogram gap does not close with training data out to 170 equivalent years and is shared by a linear multivariate normal baseline. Comparing ensemble types at matched spread, a dynamical ensemble shows no gap while a deterministic AI emulator does, which the authors interpret as evidence that the failure is intrinsic to data-driven emulation rather than to ensemble forecasting in chaotic systems.","pith_inferences":["Inference: the shuffle-decomposition diagnostic suggests a concrete fix the paper does not test—training with a variogram-score auxiliary loss or post-processing the ensemble with a copula or correlation reordering. If that closes the gap without eroding CRPS, the 'structural' conclusion would need to be softened to 'structural under current training objectives.'","Inference: the same decoupling likely affects AI weather emulators used operationally, where users extract spatial products such as storm surge, wind energy, or hydrology; the paper notes consistency with data-assimilation covariance errors, but the verification implication extends beyond sea level.","Inference: because the linear MVN baseline shows the same insensitivity to data volume, the bottleneck may be the conditional-mean or regression flavor of the learned mapping rather than neural-network capacity—a testable hypothesis distinguishing diffusion, flow-based, and GAN ensembles of comparable marginal skill.","Inference: the per-pair variogram attribution implies a targeted postprocessing route: instead of global calibration, correct only the over-correlated pairs (for example Atlantic City–Battery); the heterogeneity the paper documents suggests a low-rank correlation repair could recover most joint skill."],"forward_implications":["Standard marginal verification is insufficient: a forecast can post positive CRPS skill at every station and positive short-lead energy score while its joint spatial structure is worse than climatology.","Adding training data is not a guaranteed remedy for joint miscalibration in learned emulators; the CRPS-variogram gap persists through 170 equivalent years in the idealized experiments.","Deterministic AI emulators used with perturbed initial conditions inherit the joint calibration deficit, so downstream users should not assume ensemble spread encodes correct spatial uncertainty.","Variogram-score-style dependence checks should be part of routine evaluation and training of AI forecast ensembles, since CRPS and energy score can certify the wrong distribution.","Dynamical ensembles retain a role as a complement or reference for AI forecasts, because they do not show the same joint failure mode."],"supporting_citations":[{"why":"Supplies the denoising diffusion probabilistic model architecture used to generate the forecast ensembles.","marker":"Ho et al. (2020)"},{"why":"Supplies the variogram score and the recommendation to use power parameter p=0.5; this is the score that detects the joint failure.","marker":"Scheuerer & Hamill (2015)"},{"why":"Supplies CRPS and the proper scoring rule framework used to measure marginal forecast skill.","marker":"Gneiting & Raftery (2007)"},{"why":"Supplies the energy score and the framing of ensemble forecasting as sampling the conditional distribution.","marker":"Gneiting & Katzfuss (2014)"},{"why":"Supplies the Lorenz-96 chaotic system used to test whether the skill gap depends on training data volume.","marker":"Lorenz (2005)"},{"why":"Supplies the GLORYS reanalysis SSH and SST data used to train and evaluate the coastal sea level forecast.","marker":"Lellouche et al. (2021)"},{"why":"Establishes the ensemble forecasting objective—sampling the conditional distribution—that the paper tests.","marker":"Leutbecher & Palmer (2008)"},{"why":"Provides an independent prediction that AI models exaggerate sensitivities and distort covariances, cited as supporting the relevance of the joint failure.","marker":"Tian et al. (2026)"},{"why":"Provides empirical confirmation of covariance amplification in AI model data assimilation, cited as consistent with the paper's joint miscalibration finding.","marker":"Slivinski et al. (2025)"}],"fun_headline_variants":["Skillful at every port, worse than chance for the whole coast","AI forecast: right per station, wrong per pattern","Smart at each tide gauge, dumb at the big picture","Station-level wins, spatial losses: AI ensemble's split","AI gets each port right but the coast wrong"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 'structural, won't improve with data' conclusion assumes the idealized Lorenz-96 system reproduces the same joint-calibration failure as the real coastal sea level forecast, and that the climatological reference pool from the training period correctly defines what 'worse than climatology' means in the test period.","fun_headline_variants_meta":{"raw":{"variants":["Skillful at every port, worse than chance for the whole coast","AI forecast: right per station, wrong per pattern","Smart at each tide gauge, dumb at the big picture","Station-level wins, spatial losses: AI ensemble's split","AI gets each port right but the coast wrong"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000603,"raw_usage":{"total_tokens":2797,"prompt_tokens":908,"completion_tokens":1889,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":1809}},"tokens_in":524,"tokens_out":1889,"duration_ms":13623,"temperature":1.0,"reasoning_tokens":1809,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:19:48.349872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the Lorenz-96 DDPM with an additional variogram-score-based penalty or with a covariance-matching postprocessor; if the CRPS-variogram gap closes without reducing CRPS, the paper's claim that the joint failure is a structural property of learned emulators, rather than a property of the training objective, would be refuted.","supporting_citations":[],"review_version":1}