{"id":"352ddf8c-8b97-44e7-b74f-dbfc1c61070c","arxiv_id":"2602.16579","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An LSTM pre-trained on ERA5-Land and fine-tuned on IFS forecasts reaches median KGE' 0.66 / NSE 0.53 for global daily streamflow over 2021-2024.","lead":"AIFL is a global streamflow-forecasting LSTM trained first on 40 years of ERA5-Land reanalysis and then fine-tuned on IFS operational forecasts; on 2,003 gauged basins in 2021–2024 it reports median KGE' 0.66 and NSE 0.53. The paper is worth reading as a simple, transparent operational baseline and a direct comparison to Google's global flood model, but its key two-stage-training claim lacks the promised ablation results and the code is not yet public.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's central claim—two-stage pretraining/fine-tuning outperforms IFS-only and mixed single-stage baselines—is never actually reported; no ablation table, figure, or comparison appears in §3 or §4. The paper's main contribution is therefore unsupported as written.","rationale":"I read the paper as an empirical claim about a global LSTM streamflow model plus a methodological claim about a two-stage reanalysis-to-forecast training strategy. The empirical core—LT1-based median KGE' 0.66 on 2,003 temporally held-out basins—is plausible and partially supported, and the benchmarking against Google is honest about the small median gap. The methodological claim, however, is the paper's stated primary contribution and is asserted in the abstract but absent from the results. The reader's formal weakest assumption is the LT1-to-LT9 transfer, which is a genuine limitation and is explicitly acknowledged in §3.2 as future work. I consider the missing ablation more load-bearing because it concerns the central claim that the paper's novel strategy outperforms alternatives; if that claim is false or unsupported, the paper's contribution reduces to 'a standard LSTM with reasonable results,' while the lead-time issue only weakens the operational generalization beyond LT1. The concrete test I propose is a direct reproducibility check: re-run the missing baselines. This is feasible with the described data pipeline and would settle whether the central claim is factual. I therefore keep the reader's CONDITIONAL verdict unchanged, pending either the ablation evidence or a revised claim that does not depend on it.","tokens_in":15816,"tokens_out":8264,"duration_ms":81667,"concrete_test":"Run the two missing ablation configurations on the same 2,003-basin 2021–2024 test set with identical architecture and hyperparameters: (a) an IFS-only model trained from scratch on 2016–2019 IFS LT1 forcings, and (b) a single-stage model trained on concatenated ERA5-Land and IFS LT1 data. Report median/mean KGE' and NSE, plus a lower-tail metric (e.g., 10th percentile KGE'), for each configuration. If either baseline matches or beats the two-stage model on the headline metrics, the abstract's central claim is not supported. If the two-stage model is worse on medians but better on the tail, the paper should be reframed to claim a targeted robustness improvement rather than overall superiority.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central methodological claim is stated in the abstract: 'Ablation experiments confirm that this two-stage approach outperforms both a naive IFS-only baseline and a mixed-forcing single-stage alternative.' No such ablation experiments appear anywhere in §3 or §4. §3.2 reports only a pre-trained-versus-fine-tuned comparison on the test basins (median ΔKGE' = −0.013; mean KGE' 0.21→0.44; mean NSE −11.40→−3.26). That comparison is not the claimed IFS-only or mixed-forcing control. This is load-bearing because the entire novelty of AIFL is attributed to the two-stage strategy. Without the two missing baselines, the reported test skill could in principle be achieved by pre-training alone, by training directly on IFS, or by a single-stage mixed-forcing model. The omission is especially consequential because fine-tuning slightly lowers median skill while improving the mean, so the strategy's benefit is concentrated in a lower tail that the paper never decomposes against the claimed control conditions. The lead-time transfer limitation raised by the reader is real but secondary: it is explicitly acknowledged as future work and does not invalidate the LT1 skill claim. The unsupported ablation claim is the more load-bearing gap: it is the paper's stated primary contribution and it is currently unverifiable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AIFL, a single-layer deterministic LSTM for global daily streamflow forecasting, pre-trained on 40 years of ERA5-Land reanalysis over 18,588 Caravan basins and then fine-tuned on IFS control forecasts at LT1 (2016–2019). The model is evaluated on a temporally independent test period (2021–2024) at 2,003 gauged basins, reporting median KGE' 0.66 and NSE 0.53, and is benchmarked against the Google global flood model on 1,218 shared stations. The central methodological claim is that this two-stage reanalysis-to-forecast transfer strategy outperforms a naive IFS-only baseline and a mixed-forcing single-stage alternative, and that the resulting model is a competitive, operationally robust baseline. Flood-event verification and a case study are also presented.","tokens_in":16118,"tokens_out":5273,"duration_ms":53407,"significance":"If the central claims hold, AIFL would be a useful transparent baseline for global operational streamflow forecasting, with a simple LSTM architecture, a clear temporal split, and an open data pipeline (Caravan/MultiMet). The independent temporal test design is a genuine strength: fine-tuning on 2016–2019 and test on 2021–2024 avoids trivial circularity. The comparison with the Google model over shared stations is informative, and the Hokitika/Storm Henk examples illustrate the intended behavior. However, the paper's stated primary contribution—the superiority of the two-stage training strategy over IFS-only and mixed single-stage alternatives—is asserted in the abstract but never demonstrated in Sections 3 or 4. The only reported comparison is pre-trained versus fine-tuned, which does not establish the claimed ablation. Until that evidence is supplied, the methodological novelty of AIFL is unverified, and the operational claims are also limited by the lack of systematic lead-time evaluation.","major_comments":[{"comment":"The abstract states: 'Ablation experiments confirm that this two-stage approach outperforms both a naive IFS-only baseline and a mixed-forcing single-stage alternative.' I could not find these ablation experiments anywhere in the manuscript. §3.2 reports only the pre-trained versus fine-tuned comparison on the 2,003 test basins (median ΔKGE' = −0.013; mean KGE' 0.21→0.44; mean NSE −11.40→−3.26). That comparison is not an IFS-only baseline and is not a mixed-forcing single-stage alternative. Because the two-stage strategy is the paper's stated primary contribution, this omission is load-bearing: the reported test skill could in principle be achieved by pre-training alone, by direct IFS training, or by a single-stage mixed-forcing model. Please add the missing ablation results with a full description of the training setups and a metric table, or remove the claim.","section":"Abstract; §3.2"},{"comment":"The operational system is claimed to provide global daily streamflow forecasts over a 10-day horizon, but all systematic results are for LT1 forcings. §3.2 states that lead-time-specific drift is 'reserved for future analyses,' and §4.1 reports skill only for IFS LT1. The only multi-lead evidence is the single Storm Henk case study (Fig. 8). The title and conclusion promise a 10-day forecasting model; without a lead-time-resolved evaluation (e.g., median KGE'/NSE by LT1–LT9), or an explicit statement that skill claims apply only to LT1, the operational claim is unsupported. Please add a systematic multi-lead evaluation or substantially qualify the claims.","section":"§3.2, §4.1"},{"comment":"The reported global precision of 1.0 and zero false alarms for all return periods under exact day matching are extraordinary and require more careful definition. Hits/misses/false alarms depend on the threshold construction, which is asymmetric: forecast event thresholds are derived from a 45-year ERA5-Land simulation, while observed event thresholds come from the historical observational record. The text does not define the counting unit (station-days, event episodes, or individual exceedance days) or how sustained exceedances are handled. As written, the zero-false-alarm result may be an artifact of the threshold/event definition rather than a meaningful property of the model. Please report precision and recall with a time tolerance, explicit event definitions, and base rates, or restrict the claims accordingly.","section":"§4.2, Table 4"}],"minor_comments":[{"comment":"Figure 3 legend labels the test period as '2021–2023' while the text and table use '2021–2024'. Please correct the inconsistency.","section":"§2.4, Fig. 3"},{"comment":"The text says code and weights are 'intended to be hosted' and that 'the repository is still under preparation and not yet publicly accessible.' This is inconsistent with the paper's repeated emphasis on transparency and reproducibility. Please provide a URL or DOI, or state clearly that the artifacts are not yet available.","section":"Code and Data Availability"},{"comment":"The Google model comparison is based on publicly released outputs, but the manuscript does not state whether the Google model's training period overlaps the 2021–2024 test period. A sentence clarifying the benchmark setup and any potential temporal overlap would strengthen the comparison.","section":"§4.3"},{"comment":"The notation for the variability ratio in Figure 6 uses α, while the text in §4.3 refers to γ. Please standardize the notation for the KGE' decomposition.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The missing ablation is the gating issue: the paper's central methodological contribution is not verifiable from the current text. If the authors can supply the IFS-only and mixed single-stage ablation results, the paper may become publishable. I would also like to see some form of lead-time-resolved skill evaluation before endorsing the operational framing. These are fixable within the scope of a revision, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a useful empirical contribution: it gives the global hydrology community a transparent, reproducible LSTM baseline trained entirely inside the Caravan ecosystem, with a sensible two-stage recipe—pretrain on ERA5-Land, fine-tune on IFS—and a clean temporal test on 2021–2024 at 2,003 basins. The reported median KGE' of 0.66 and NSE of 0.53 are plausible, and the Google benchmark is honest: AIFL is slightly worse on medians but competitive, with the failure modes mostly in large basins where Google's multi-source forcing helps. The fine-tuning story also has some internal support: mean KGE' jumps from 0.21 to 0.44 while the median barely moves, consistent with correcting a lower tail of bad basins rather than a uniform shift. That is a genuinely informative result even if the median gain is negative.\n\nThe problem is the abstract. It says \"ablation experiments confirm that this two-stage approach outperforms both a naive IFS-only baseline and a mixed-forcing single-stage alternative.\" I read sections 3 and 4 carefully. No such experiments appear. There is no IFS-only baseline, no mixed-forcing single-stage model, no ablation table or figure. The only transfer comparison shown is pretrained versus fine-tuned on IFS LT1 forcings, which does not test the claimed controls. This is load-bearing because the two-stage strategy is the paper's stated primary novelty. Without those baselines, the reader cannot tell whether the reported skill comes from pretraining, from direct IFS training, or from the two-stage combination. That is a fixable gap—the authors clearly know what baselines they mean—but as written the central claim is unverifiable.\n\nTwo smaller concerns. First, flood-event precision of exactly 1.0 across all return periods, with zero false alarms out of millions of hits, is suspicious enough that the threshold definitions and the zero-day matching rule need a careful audit before anyone trusts the \"zero false alarms\" headline. Second, the operational claim extends to 10-day leads, but fine-tuning and most evaluation happen at LT1; the paper acknowledges this as future work, which is fair, but it does limit what the title implies.\n\nOn the positive side, the data curation is thoughtful, the temporal split is not circular, the paper cites prior transfer-learning work (Ryd & Nearing, Konold et al.) rather than pretending the idea is new, and the writing is clear. Code and weights are not yet public, which is another reason to treat the results as provisional.\n\nWho is this for? Operational hydrologists and ML-for-Earth-systems researchers who want a baseline for global forecasting. It deserves serious peer review, but only if the authors supply the missing ablations and clean up the flood-event analysis.","headline":"AIFL is a solid global streamflow baseline with a real two-stage transfer idea, but the abstract advertises ablation results that simply are not in the manuscript.","tokens_in":16704,"tokens_out":1194,"would_cite":false,"duration_ms":12358,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Global daily streamflow forecasting with a deterministic LSTM reaches a median KGE' of 0.66 when the model is pre-trained on reanalysis and fine-tuned on operational forecast forcings.","keywords":["global streamflow forecasting","LSTM","transfer learning","reanalysis-to-forecast domain shift","ERA5-Land","IFS forecasts","Caravan dataset","flood detection"],"falsifier":"Run the fine-tuned model on operational forecasts at lead times 2 through 9 over the same 2021–2024 test basins and compare skill to the 1-day-lead numbers; if KGE' or NSE decays substantially with lead time, the operational ten-day claim collapses. Also obtain the pre-trained-only model and the single-stage/mixed baselines and reproduce the ablation; if the two-stage gain disappears, the central methodological claim fails.","tokens_in":15667,"feed_emoji":"🌊","tokens_out":5779,"duration_ms":51798,"temperature":0.7,"pith_summary":"This paper aims to show that a deliberately simple, deterministic LSTM can be made operationally competitive for global daily streamflow forecasting despite the mismatch between the reanalysis data hydrologists train on and the numerical weather forecasts available in real time. The proposed fix is a two-stage training strategy: first learn general rainfall–runoff behavior from four decades of reanalysis across 18,588 basins, then fine-tune all weights on operational control forecasts to absorb forecast-specific biases. On an independent 2021–2024 test at 2,003 gauged basins driven by 1-day lead-time operational forcing, the model reports median KGE' 0.66, median NSE 0.53, and a median bias ratio of 1.00. The paper further claims that fine-tuning improves the worst-performing basins enough to lift mean global skill, and that the resulting model matches or exceeds a leading existing global flood model on a large subset of shared stations while producing zero false alarms in flood-event detection.","feed_headline":"Two-stage LSTM hits median KGE' 0.66 on global streamflow","feed_subtitle":"A deliberately simple LSTM matches a more complex global flood model at 43% of shared stations and reports zero false alarms.","key_machinery":"The central mechanism is the two-stage transfer-learning schedule. A single-layer LSTM (1024 hidden units) reads 180-day sequences of five meteorological variables through a dynamic embedding MLP and 203 static basin attributes through a static embedding MLP, and outputs 10-day streamflow sequences. Pre-training on reanalysis learns generic rainfall–runoff response; fine-tuning on operational control forecasts adapts the same weights to the systematic wet bias and other error structures of the numerical weather prediction stream. The model is deliberately simple, with no routing graph, no probabilistic loss, and a single deterministic forecast.","core_discovery":"AIFL is a global deterministic LSTM for daily streamflow forecasting that solves the reanalysis-to-forecast shift by explicit two-stage transfer learning: pre-training on 40 years of ERA5-Land reanalysis across 18,588 curated basins, then fine-tuning all weights on IFS control forecasts from 2016–2019. On a 2021–2024 temporal test with IFS 1-day lead-time forcing, the model achieves median KGE' 0.66, median NSE 0.53, median correlation 0.81, and median bias ratio 1.00. The paper argues that fine-tuning acts as an operational stabilizer, correcting severe forecast-induced errors in previously low-performing basins at only modest cost to well-calibrated ones, and that the model's flood-event d","pith_inferences":["The paper leaves the multi-lead-time case largely untested: fine-tuning used only 1-day lead-time control forecasts while operational inference extends to day 10. If forecast error structure changes with lead time, 1-day-derived skill may not transfer; a systematic lead-time evaluation is the natural next experiment.","The claimed advantage of the two-stage strategy over single-stage or forecast-only training is stated but not documented in the displayed results. Until the ablation is published, the causal claim that reanalysis pretraining is what drives the gain should be treated as provisional.","Because the test basins are skewed toward larger catchments (due to station record availability after 2015), the headline median skill may overstate performance for small headwater basins, where the paper itself reports mixed results relative to the benchmark.","The zero-false-alarm result follows from a strict same-day matching rule; relaxing the timing tolerance would likely reveal a trade-off between missed events and false alarms and is a cheap way to stress-test the claim."],"forward_implications":["Global operational streamflow forecasting can be done with a plain deterministic LSTM rather than a physically routed or probabilistic system, provided the reanalysis-to-forecast shift is explicitly bridged.","Fine-tuning on operational forecasts acts as an operational stabilizer: it corrects severe forecast-induced errors in low-performing basins while only modestly degrading well-calibrated ones.","With a median bias ratio of 1.00, the model conserves volume across the global test network, a prerequisite for trustworthy water-resource accounting.","The zero-false-alarm flood detection profile, with recall between 0.32 and 0.54, implies the model can be used in early-warning settings where user trust depends on eliminating spurious alerts.","Because the model matches a larger, more complex global flood model at roughly 43% of shared stations and performs more consistently across basin sizes, it offers a transparent baseline for community experimentation."],"fun_headline_variants":["Global LSTM streamflow: two-stage training hits KGE' 0.66","AIFL: first Caravan-trained global flood LSTM, KGE' 0.66","Two-stage transfer LSTM tops IFS-only for global streamflow","New global streamflow LSTM: median KGE' 0.66 on test set"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The model assumes that correcting forecast bias at a one-day horizon is enough to correct the full ten-day forecast horizon, because the fine-tuning and all headline results use only one-day lead-time forcing.","fun_headline_variants_meta":{"raw":{"variants":["Global LSTM streamflow: two-stage training hits KGE' 0.66","AIFL: first Caravan-trained global flood LSTM, KGE' 0.66","Two-stage transfer LSTM tops IFS-only for global streamflow","New global streamflow LSTM: median KGE' 0.66 on test set"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1066,"prompt_tokens":856,"completion_tokens":210,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":119}},"tokens_in":600,"tokens_out":210,"duration_ms":2971,"temperature":1.0,"reasoning_tokens":119,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T22:28:22.689948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the fine-tuned model on operational forecasts at lead times 2 through 9 over the same 2021–2024 test basins and compare skill to the 1-day-lead numbers; if KGE' or NSE decays substantially with lead time, the operational ten-day claim collapses. Also obtain the pre-trained-only model and the single-stage/mixed baselines and reproduce the ablation; if the two-stage gain disappears, the central methodological claim fails.","supporting_citations":[],"review_version":1}