{"id":"ca359b48-e7fb-4d91-8928-64a1a647fe77","arxiv_id":"2412.04475","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Global monthly marine heatwave forecasts are produced by combining GraphSAGE, imbalanced regression losses, and temporal diffusion, with a new public SSTA graph dataset.","lead":"This paper combines graph neural networks, imbalanced regression losses, and temporal diffusion to forecast monthly marine heatwaves globally from sea surface temperature anomalies. The authors report improved forecasts versus a numerical model baseline in several ocean regions and claim skillful predictions up to six months ahead, though the evaluation has methodological flaws.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported skill is inflated because the test set is used for early stopping (Appendix), so the headline SEDI/CSI values and the comparison to numerical models are selected on the same data used for evaluation.","rationale":"The reader's weakest_assumption correctly identifies the use of the test set for early stopping as the central methodological flaw, and I agree that it invalidates the reported performance claims. This is not a matter of external consensus or stylistic disagreement; it is a direct violation of the requirement that test data remain untouched until final evaluation. The Appendix confirms the practice in explicit language, and the paper's own results section further undermines the abstract by noting that six-month forecasts are essentially unpredictable. The comparison to numerical models is also compromised because the test-selected models are being compared on the same data used for selection. A remedy is straightforward—introduce a validation split and re-estimate all reported metrics—but until then, no quantitative claim in the paper can be accepted as evidence. I therefore see no basis to change the reader's REJECT verdict: the contribution may be salvageable after major revision, but the current evaluation cannot support the stated conclusions.","tokens_in":19987,"tokens_out":2137,"duration_ms":27014,"concrete_test":"Re-run the one-month and six-month forecast experiments with a strict three-way temporal split: train on 1940–2004, validate on 2005–2009, and test only on 2010–2022. Use validation SEDI for early stopping and evaluate the test set exactly once. Compare the resulting average SEDI and CSI to the reported values (~0.68 SEDI at lead 1; ~0.14 SEDI and ~0.01 CSI at lead 6). If the test SEDI drops materially or the regional SEDI advantage over Jacox et al. (2022) disappears, the headline claims are artifacts of test-set selection. Additionally, report the fraction of nodes with undefined SEDI instead of averaging only over defined nodes, since excluding undefined nodes can bias the average upward.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claims—that the integrated DL framework outperforms numerical models in specific ocean regions and that forecasts remain useful up to six months—rest entirely on quantitative skill estimates. The Appendix states that 'the model configuration with the largest overall SEDI over the test data was saved' and that SEDI was used as the early-stopping metric on the test data. This makes the reported test-set metrics in-sample selection maxima, not unbiased estimates of generalization. With patience 40 and up to 200 epochs, repeatedly choosing the epoch with the best test SEDI can inflate scores substantially, especially for a rare-event metric like SEDI computed on a relatively short test period (2010–2022, 156 months). The regional superiority over Jacox et al. (2022) and the diffusion-related improvements are therefore not trustworthy: the models may simply be tuned to noise in the same test data against which they are scored. A separate validation split is absent; no honest held-out evaluation is reported. The body's admission that six-month CSI is 'almost zero' already contradicts the abstract's 'up to six months' claim, but even that pessimistic reading is based on test-selected models, so the true skill could be lower still. This flaw is load-bearing because every quantitative contribution depends on the validity of these test-set numbers, and removing the bias requires re-running the entire model-selection protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a global monthly marine heatwave (MHW) forecasting framework that combines GraphSAGE graph representation with imbalanced regression losses (BMSE, WMSE) and a temporal-diffusion training process. The authors introduce a sorted Kendall-correlation graph construction that guarantees no isolated nodes, release a new public SSTA graph dataset, and evaluate one- to six-month-ahead forecasts using precision, recall, CSI, and SEDI, comparing qualitatively with the numerical-model forecasts of Jacox et al. (2022). The central claims are that the integrated DL approach outperforms numerical models in several ocean regions and that temporal diffusion enables useful forecasts up to six months ahead.","tokens_in":20276,"tokens_out":6673,"duration_ms":63566,"significance":"If the evaluation were unbiased, the paper would make a useful contribution: a novel graph-construction method, a public dataset, and a comparison of imbalanced losses and diffusion for a high-impact climate extreme. The framework is original in combining these three threads, and the authors provide code and data for reproducibility. However, the quantitative evidence for these claims is currently undermined by the test-set-based early stopping and by internal inconsistencies in the long-lead results, so the significance cannot be assessed from the reported numbers.","major_comments":[{"comment":"The early stopping protocol selects the model checkpoint with the largest SEDI on the test data: 'After each training epoch, the model configuration with the largest overall SEDI over the test data was saved.' Because the same test data (2010-2022, 156 months) is then used for all reported SEDI/CSI values in Tables 2-4 and Figure 2, the scores are selection maxima over up to 200 training epochs rather than unbiased estimates of generalization. With rare-event metrics and patience 40, this can substantially inflate skill, so the comparison with Jacox et al. (2022) and the six-month claim are not supported. A separate validation split must be used for early stopping and hyperparameter selection, with the test set used only once for final evaluation.","section":"Appendix: Experiment Details"},{"comment":"The abstract claims 'achieving improved prediction up to six months in advance,' but the body states that six-month-ahead forecasts had average SEDIs around 0.14 and average CSIs 'almost zero, implying the unpredictability for six-month-ahead MHW forecasts so far.' This is a direct contradiction. The abstract should be revised to match the actual results, and the authors should clarify what 'improved prediction' means when the CSI is near zero.","section":"Experiments: Prediction for Longer Terms (vs. Abstract)"},{"comment":"The loss hyperparameters alpha=2, w=2, and sigma=0.02 are stated to be selected via a Friedman test in 'earlier experiments (not included in this manuscript).' This makes the loss-function comparison non-reproducible and the claim of 'optimized' losses unverifiable. The evidence for these choices should be included in the paper or the hyperparameters should be treated as exploratory, with a sensitivity analysis reported.","section":"Methodology: Imbalanced loss functions"},{"comment":"The evidence that temporal diffusion improves long-lead forecasts is mixed and often within one standard deviation. For example, in Table 3 at lead 6, the mean CSI with diffusion (0.0102) is lower than without (0.0424), while in Table 4 at lead 6 diffusion increases CSI (0.14 vs 0.091) but decreases SEDI (0.1555 vs 0.1681). The claim that diffusion 'achieves improved prediction up to six months' is therefore not clearly supported by the reported metrics. The authors should provide a consistent, statistically grounded comparison (e.g., confidence intervals or significance tests) and temper the conclusion accordingly.","section":"Experiments: Prediction for Longer Terms, Tables 3 and 4"}],"minor_comments":[{"comment":"The second author's name appears as 'Varvara V etrova' with a stray space; it should be 'Varvara Vetrova.'","section":"Author line"},{"comment":"In Algorithm 1, line 10 uses the variable tau_ij in 'sort(correlations, by -|tau_ij|)' without defining it in the pseudocode; tau_ij should be defined as the Kendall rank correlation coefficient computed in line 7.","section":"Algorithm 1"},{"comment":"The SEDI formula is written as a single fraction without parentheses around the numerator and denominator; adding parentheses would remove ambiguity about the order of operations.","section":"Equation (3)"},{"comment":"The caption states 'All used them = 25graph construction method,' which has a missing space and should read 'the m = 25 graph construction method.'","section":"Figure 2 caption"},{"comment":"The 12 hotspot locations are identified only by abbreviations in the appendix; providing coordinates or a reference map in the main text would improve reproducibility.","section":"Appendix: Additional Experiment Results"}],"recommendation":"major_revision","confidential_remarks":"The test-set early stopping issue is the central obstacle: every quantitative claim in the paper rests on scores that were selected on the evaluation data. This is fixable by re-running the experiments with a proper validation split, but the revision will require substantial new computations. If the authors cannot provide unbiased held-out results, the paper should not be published. The abstract's 'up to six months' phrasing overstates the body's own conclusion and needs correction regardless."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper is the first to combine GraphSAGE, imbalanced regression losses (BMSE/WMSE), and DYffusion-style temporal diffusion for marine heatwave forecasting, and it ships a new graph-construction rule (sorted Kendall correlations with a minimum degree) plus a public SSTA graph dataset and code. Those are real contributions. But the evaluation protocol undercuts the headline numbers: the appendix states that the model configuration with the largest overall SEDI over the test data was saved at each epoch, i.e., the test set was used for early stopping. Every reported SEDI and CSI is then a selection maximum over the test data, not an unbiased estimate of generalization. The comparison to Jacox et al. (2022) and the six-month skill claim rest on those numbers.\n\nWhat the paper does well: the ablations over graph sizes and loss functions are useful; the empirical finding that imbalanced losses boost MHW detection at one-month lead is plausible; and the new graph dataset is a practical contribution. The authors are transparent about the early-stopping protocol, which is at least auditable.\n\nThe soft spots are mostly the one big one. Test-set early stopping inflates rare-event metrics, and for a 156-month test period the inflation can be substantial. The body's own admission that six-month CSI is 'almost zero' contradicts the abstract's 'improved prediction up to six months in advance'; even that pessimistic reading is optimistic because the six-month models were selected on the test set. The loss hyperparameters alpha=2, w=2, sigma=0.02 are fixed without showing the 'earlier experiments'; that is a reproducibility gap. The regional comparison to the numerical model ensemble is verbal; there is no error bar or significance test on the differences.\n\nThe framework itself is not circular or incoherent; it is an empirical forecasting paper with a broken evaluation protocol. The right fix is a proper train/validation/test split with early stopping on validation, then re-report skill. If the one-month results survive that re-analysis, the paper is publishable. The six-month claim needs the same treatment.\n\nWho gets value: readers working on GNNs for climate extremes and on rare-event forecasting; the dataset may be useful even before the numbers are re-run.\n\nRecommendation: I would not cite it in its current form, but I would send it to peer review—not desk reject—because the contributions are genuinely new and the flaw is fixable with re-analysis. A serious referee should ask for the re-run before publication.","headline":"Test-set early stopping bakes the headline scores, so the six-month claim is unreliable; the framework and dataset are still worth a re-run.","tokens_in":20793,"tokens_out":2764,"would_cite":false,"duration_ms":30188,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An integrated deep-learning framework combining graph representation, imbalanced regression, and temporal diffusion can forecast marine heatwaves up to six months ahead and outperform numerical models in several ocean regions.","keywords":["marine heatwaves","deep learning","graph neural network","temporal diffusion","imbalanced regression","sea surface temperature anomaly","climate forecasting","extremal dependence index"],"falsifier":"Re-run the same experiments on the same ERA5-derived data with the $m=25$ graph, the same losses, and the same diffusion procedure, but choose all hyperparameters and the early-stopping epoch using a validation period disjoint from the test period (for example, validate on 2000–2012 and test on 2013–2022). If the one-month SEDI falls below roughly 0.6 or the six-month critical success index drops to zero, the paper's claimed skill and its comparison with numerical models would not stand.","tokens_in":19801,"feed_emoji":"🌊","tokens_out":7644,"duration_ms":71353,"temperature":0.7,"pith_summary":"This paper claims that an integrated deep-learning system can forecast marine heatwaves on a global grid from one to six months ahead, and that it outperforms physics-based numerical models in several regions, including the middle south Pacific, the equatorial Atlantic near Africa, the south Atlantic, and the high-latitude Indian Ocean. The system combines graph representation of sea surface temperature anomalies, imbalanced-regression losses that emphasize warm extremes, and a temporal diffusion process that refines multi-step forecasts without a long sliding window. If the claims hold, a data-driven model could complement or partially replace computationally expensive dynamical forecast systems and provide earlier practical warning for marine ecosystems and fisheries. The paper also introduces a graph-construction method that guarantees no isolated nodes and releases a new public SSTA graph dataset.","feed_headline":"Graph AI forecasts marine heatwaves six months ahead","feed_subtitle":"Deep-learning model matches and beats numerical forecasts in several ocean regions.","key_machinery":"The central object is a graph whose nodes are grid locations and whose edges keep the top $m$ Kendall rank correlations between each location's sea surface temperature anomaly time series; ranking the correlations instead of thresholding them guarantees every node has at least $m$ edges and eliminates isolated nodes. The predictor is a two-layer GraphSAGE network (with mean, pooling, or LSTM aggregation) trained by standard MSE, balanced MSE, or a custom weighted MSE that up-weights errors above the 90th percentile. For long leads, the paper adapts temporal diffusion: a forecaster predicts the target month, interpolator networks reconstruct intermediate months, and the forecaster is then refined using the interpolated fields, so that the model can advance one step at a time without a long input window.","core_discovery":"The paper's central claim is that a graph neural network (GraphSAGE) trained on rank-correlation graphs of monthly sea surface temperature anomalies, with balanced or weighted mean-squared-error losses and temporal diffusion, produces marine heatwave forecasts whose average symmetric extremal dependence index (SEDI) is about 0.68 at one-month leads and remains above chance at six-month leads (SEDI around 0.14, CSI around 0.14 when a single input time step is used). Spatially, the model shares the strongest skill regions with numerical models—equatorial Pacific, northwest Pacific near North America, south Pacific near South America, and equatorial Atlantic near South America—but additionally shows higher SEDI than the numerical ensemble in the middle south Pacific, equatorial Atlantic near Africa, south Atlantic, and high-latitude Indian Ocean. The authors further claim that temporal diffusion makes a conventional 12-month sliding window unnecessary, reducing input requirements while improving the critical success index for long leads.","pith_inferences":["If the six-month skill survives a clean validation protocol, a global data-driven early-warning system could run on a single GPU in minutes, making seasonal marine heatwave outlooks accessible to regions without operational dynamical forecast centers.","The guarantee that every node has at least $m$ edges is likely transferable to other gridded geophysical prediction tasks where correlation-based graphs with fixed thresholds suffer from isolated nodes.","Because the reported models were selected using the test data's SEDI as the early-stopping criterion, the numerical-model comparison should be re-run with a held-out validation set before operational use; that re-run is a direct test the paper leaves implicit."],"forward_implications":["A purely data-driven model can match or beat dynamical model skill in specific ocean regions at three-to-four-month leads, suggesting machine learning is a viable alternative for global marine heatwave outlooks.","Temporal diffusion plus a single time step input produces similar or better long-lead skill than the conventional 12-month sliding window, cutting input data requirements.","Imbalanced regression losses (BMSE and WMSE) improve detection of marine heatwave events at the cost of some precision, giving a useful lever for forecast users who prioritize recall.","The minimum-degree graph construction removes isolated nodes and yields a reusable public SSTA graph dataset for other climate forecasting problems.","Skill degrades with lead time and is marginal at six months; forecasts beyond six months are not usable, so the practical horizon of this system is about half a year."],"supporting_citations":[{"why":"Provides the global numerical model MHW forecast baseline and the SEDI evaluation framework that the paper compares against.","marker":"(Jacox et al. 2022)"},{"why":"Supplies the graph representation workflow, the Kendall-correlation graph construction, and the GraphSAGE baseline this study extends.","marker":"(Ning et al. 2024)"},{"why":"Introduces the DYffusion temporal diffusion mechanism that the paper adapts for long-lead forecasting.","marker":"(Cachay et al. 2023)"},{"why":"Defines marine heatwaves using the 90th-percentile threshold on sea surface temperature anomalies.","marker":"(Hobday et al. 2016)"},{"why":"Defines the extremal dependence index SEDI used as the primary evaluation and early-stopping metric.","marker":"(Ferro and Stephenson 2011)"},{"why":"Identifies the 12 historical MHW hotspot locations used for regional evaluation.","marker":"(Oliver et al. 2021)"},{"why":"Provides the balanced MSE (BMSE) loss for imbalanced regression.","marker":"(Ren et al. 2022)"},{"why":"Introduces GraphSAGE, the base graph neural network architecture.","marker":"(Hamilton, Ying, and Leskovec 2017)"},{"why":"Provides the data preprocessing (gridded ERA5 extraction, 2° latitude by 2.8125° longitude, monthly) and the sliding-window convention.","marker":"(Taylor and Feng 2022)"},{"why":"Supplies the ERA5 reanalysis dataset used to build the SSTA graphs.","marker":"(Hersbach et al. 2020)"}],"fun_headline_variants":["Graph AI beats numerical models for marine heatwaves","Deep learning forecasts ocean heatwaves six months ahead","New graph dataset enhances marine heatwave predictions","Integrated AI model improves heatwave forecasts globally","Graph neural nets predict marine heatwaves with less data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported skill depends on treating the test set as untouched: the paper says 'the model configuration with the largest overall SEDI over the test data was saved,' meaning the test data guided model selection; if that guidance is not honest, the headline SEDI values and the six-month skill claim are inflated.","fun_headline_variants_meta":{"raw":{"variants":["Graph AI beats numerical models for marine heatwaves","Deep learning forecasts ocean heatwaves six months ahead","New graph dataset enhances marine heatwave predictions","Integrated AI model improves heatwave forecasts globally","Graph neural nets predict marine heatwaves with less data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00061,"raw_usage":{"total_tokens":2844,"prompt_tokens":957,"completion_tokens":1887,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":1817}},"tokens_in":573,"tokens_out":1887,"duration_ms":16589,"temperature":1.0,"reasoning_tokens":1817,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:44:21.582638+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same experiments on the same ERA5-derived data with the $m=25$ graph, the same losses, and the same diffusion procedure, but choose all hyperparameters and the early-stopping epoch using a validation period disjoint from the test period (for example, validate on 2000–2012 and test on 2013–2022). If the one-month SEDI falls below roughly 0.6 or the six-month critical success index drops to zero, the paper's claimed skill and its comparison with numerical models would not stand.","supporting_citations":[],"review_version":1}