{"id":"fd8ea6eb-4e32-481b-afdb-e3ba644f0702","arxiv_id":"1908.02781","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of flood prediction studies that classifies machine learning methods by lead time and model structure, then points to hybrid and decomposition methods as the most promising.","lead":"This paper reviews machine learning models used to forecast floods, sorting them by lead time (short-term versus long-term) and by whether they use single or hybrid techniques. It scans thousands of studies and suggests which methods look most promising for each type of flood forecast.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'most promising methods' ranking rests on pooling RMSE/R2 values that are not commensurable across studies; the figures therefore cannot support the comparative claim.","rationale":"The reader identified exactly the load-bearing weakness: the paper averages RMSE and R2 values from heterogeneous studies and then ranks methods on that basis. I agree with that assessment. My read of the full text confirms that the quantitative synthesis is not a valid meta-analysis: the paper acknowledges in Section 4.3 that RMSE values differ across studies but only attempts to equalize units, which does not make them commensurable. The figures and the accuracy ratings in Tables 2 and 5 depend on this problematic pooling, so the central claim about 'most promising' methods is unsupported. I also note secondary internal inconsistencies (Figure 10 is labeled short-term hybrid in one place and long-term hybrid in another; Figure 3 is cited as the source of accuracy values when it is a publication-count chart), which further weaken the reliability of the reported quantitative basis. The conditional verdict is appropriate: the paper remains useful as a qualitative catalogue of ML methods and observed trends, but its comparative rankings should not be taken as established. No change to the reader's verdict is needed, though the final claim should be softened or explicitly caveated in any revision.","tokens_in":38019,"tokens_out":4307,"duration_ms":49060,"concrete_test":"Reconstruct the data behind Figure 8 (hybrid, short-term) from the cited papers. For each study, record the target variable, unit, catchment, lead time, and the RMSE/R2 of each method. Then: (1) separate studies into groups by target variable and lead time; (2) compute average RMSE/R2 within each group; (3) compare the within-group method ranking to the pooled ranking in Figure 8. If the top method changes between pooled and within-group analyses, the pooled ranking is an artifact of mixing non-commensurable studies. As a complementary binary test, for every study that compares methods A and B, record which method won on RMSE/R2; check whether the pooled average ranking agrees with the majority of within-study pairwise comparisons. The ranking is only trustworthy if both analyses agree.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion—that the paper 'introduces the most promising prediction methods' for short- and long-term floods—is supported by Figures 7–10 and by the accuracy columns of Tables 2 and 5, which rank ML methods by averaged RMSE and R2 values taken from different studies. The paper states in Section 4.3: 'we made sure that the unit of RMSE was the same, and, for the multiple RMSEs, the average was calculated.' Equal units are not sufficient for comparability. RMSE is scale-dependent: a water-level RMSE of 0.3 m is excellent for a 6 m tidal range but poor for a 0.5 m range; streamflow RMSE in m3/s depends on catchment size; rainfall RMSE in mm depends on climate and event intensity. R2, although dimensionless, depends on the variance of the observed series and the test period, so it too is study-specific. The figures mix flood-resource variables (streamflow, water level, rainfall, runoff), lead times (1-h, 3-h, 24-h, weekly, monthly), and catchments. No effect size, normalization, or within-study paired comparison is used. Consequently, the pooled averages and the resulting 'most promising' ranking are not a valid quantitative synthesis. Tables 2 and 5 are partly built from the same flawed comparison (the text cites 'accuracy analysis of Figure 3,' though Figure 3 is a publication-count chart, presumably meaning Figures 7–10), making the qualitative rankings dependent on the invalid quantitative pooling. The second part of the claim—that hybridization, decomposition, ensemble, and optimization are the most effective improvement strategies—is a defensible qualitative trend statement, but the specific 'most promising methods' conclusion is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a literature review of machine learning methods for flood prediction. The authors classify the surveyed studies according to prediction lead time (short-term versus long-term) and according to whether the method is single or hybrid. They compile 180 studies and provide narrative summaries of individual applications, supplemented by comparative performance analyses that pool reported RMSE and R2 values across studies. On this basis, the paper claims to identify the most promising prediction methods for short- and long-term floods, and it concludes that hybridization, data decomposition, algorithm ensembling, and model optimization are the most effective strategies for improving ML-based flood prediction.","tokens_in":38328,"tokens_out":6704,"duration_ms":65962,"significance":"If the comparative claims were valid, this review would be a useful synthesis of a fragmented and fast-growing literature. The paper's taxonomy (single versus hybrid, short-term versus long-term) is reasonable, and the compilation of 180 studies, together with the descriptive statistics on publication trends, is a service to the community. The narrative summaries are broadly plausible and the authors are transparent about their inclusion criteria (journal-level quality metrics, comparative content). However, the central quantitative comparison is not methodologically sound: it aggregates RMSE and R2 values from heterogeneous studies without any normalization or meta-analytic correction, so the resulting rankings of methods do not support the paper's headline claims. The paper's value is therefore primarily as a qualitative survey, not as an evidence-based ranking of methods. With a reworked comparison, or with the quantitative ranking removed, the review could be publishable.","major_comments":[{"comment":"The central quantitative comparison that supports the paper's main claim is not methodologically valid. The paper states in Section 4.3 that 'we made sure that the unit of RMSE was the same, and, for the multiple RMSEs, the average was calculated,' but equal units are not sufficient for comparability. RMSE is scale-dependent across flood resource variables (water level in metres, streamflow in m3/s, rainfall in mm) and across catchments of different sizes; R2 depends on the variance of the observed series and on the test period. The studies pooled in Figures 7–10 differ in lead time, region, data period, and evaluation protocol, and no normalization, effect size, or within-study paired comparison is applied. Therefore the averaged RMSE and R2 values, and the resulting rankings in Figures 7–10 and the accuracy ratings in Tables 2 and 5, cannot support the abstract's claim that the paper 'introduces the most promising prediction methods' for short- and long-term floods. The authors should either present the comparison as purely qualitative or conduct a formal meta-analysis with appropriate standardization and a clearly defined, reproducible study-selection protocol.","section":"4.3, Figures 7–10, Tables 2 and 5"},{"comment":"The paper's own taxonomy is internally inconsistent. Section 3.8 defines long-term prediction as lead time greater than one week, yet Table 1, which is labelled 'Short-term predictions using single machine learning methods,' includes the row 'MLP vs. Kohonen NN [154] Flood frequency analysis Long-term China.' This directly contradicts the stated definition and indicates a classification error. Because the division into short-term and long-term is the organizing principle of the entire survey, this inconsistency must be resolved by correcting the table row or revising the definition.","section":"3.8, 4.1, Table 1"},{"comment":"Figure 10 is captioned 'Comparative performance analysis of hybrid methods of ML for short-term prediction,' and the text in Section 6 repeats that 'Figure 10 represents the comparative performance analysis of hybrid methods of ML for short-term prediction,' even though the surrounding paragraph is discussing long-term prediction and Figure 9 is described as covering single methods for long-term prediction. One of the two (text or figure label) is wrong, and the error directly affects the interpretation of the long-term hybrid results, which are central to the paper's conclusions. Please correct the mislabelling and verify that the figure contents match the intended lead-time category.","section":"6, Figures 9–10"},{"comment":"The qualitative ratings in Tables 2 and 5 (e.g., 'Fair', 'High', 'Fairly high') are presented as comparative analyses, but the method for assigning these ratings is not described. The text states that the tables were created 'based on the revisions that were made on the articles of Table 1 and also the accuracy analysis of Figure 3,' yet Figure 3 is a chart of the number of articles per method, not an accuracy analysis. This internal inconsistency and the lack of a reproducible rating protocol make the rankings in these tables unsupported. The authors should specify the rubric used, or remove the ratings and rely on the narrative discussion.","section":"4.1, 4.2, 5.1, 5.2; Tables 2 and 5"}],"minor_comments":[{"comment":"The sentence 'To mimic the complex mathematical expressions of physical processes of floods, during the past two decades, machine learning (ML) methods contributed highly in the advancement of prediction systems providing better performance and cost-effective solutions' is a run-on and should be split or rewritten for clarity.","section":"Abstract and Section 1"},{"comment":"The sentence 'Furthermore, if the prediction leading time to flood is three days longer than the confluence time, the prediction is considered to be long-term [37,58]' is unclear; 'confluence time' is not defined, and the threshold seems to conflict with the later definition of 'greater than a week.'","section":"3.8"},{"comment":"The statement 'generally R2 > 0.8 is considered as an acceptable prediction' is a heuristic that needs a citation or a more nuanced discussion, since acceptable R2 values depend on the variable and context.","section":"4.3"},{"comment":"The sentence 'References [224,226] compared the performances of ANFIS, ANNs, and SVM for the monthly prediction of floods' appears to cite the wrong references: [224] is a review of a multiobjective optimization package and [226] is about solar radiation prediction, not flood prediction. Please correct the citations.","section":"5.2"},{"comment":"The conclusions section is numbered '5' but appears after Section 6; the numbering should be sequential (for example, Section 7).","section":"Section 5 (Conclusions) numbering"},{"comment":"'Reference year: 2008 (source: Scopus)' is ambiguous; the figures appear to plot data from 2008 to 2017, so the caption should say 'Data source: Scopus, 2008–2017' or similar.","section":"Figures 3 and 4 captions"},{"comment":"The manuscript contains numerous typos and grammatical errors (for example, 'reduc tion', 'minimiz ation', 'the results of [149] provides similar conclusions', and 'SVM was demonstrated as a potential candidate'), and a careful language edit is needed.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The reference list contains an unusually high number of self-citations by the first author, which is partly understandable given the author's prior work in this area, but the paper should make clear that its search and selection criteria are independent of authorship. The editorial board may also wish to consider whether the quantitative claims, if not reworked along meta-analytic lines, would meet the journal's standards for review articles. The paper has value as a narrative review, so a major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful catalogue and a reasonable organizing scheme, but the central comparative claim does not survive contact with its own figures. The authors pool RMSE and R2 values from studies with different flood variables, catchments, lead times, and periods, then rank methods. Same units do not make RMSE commensurable—a 0.3 m water-level error means different things in different tidal ranges—and R2 depends on the variance of each observed series. So Figures 7–10 and Tables 2 and 5 cannot support the 'most promising methods' conclusion. The stress-test note is right.\n\nWhat is genuinely useful: the short-term/long-term and single/hybrid split is a sensible way to organize a messy literature; the search procedure is transparent (6596 articles down to 180); the narrative summaries of individual studies are broadly plausible. The qualitative claim that hybridization, decomposition, ensembling, and optimization are common improvement strategies is well supported and probably true. A practitioner entering flood ML would get a decent map of method families and named hybrid systems.\n\nSoft spots beyond the invalid pooling: internal inconsistencies are real—Table 2 references 'accuracy analysis of Figure 3,' which is a publication-count chart, and the two figures labelled 10 appear to be mislabelled in Section 6. Self-citations are heavy, especially to Mosavi's own surveys, though that alone is not disqualifying. The abstract's phrase 'introduces the most promising prediction methods' overstates what a narrative review can establish without a common benchmark.\n\nWho is this for? Someone who wants a broad list of studies and method names, not someone who needs a reliable ranking. With major revision—fix labels, reframe the conclusions as qualitative trends, and either remove the pooled RMSE/R2 ranking or replace it with an explicitly exploratory comparison—it could become a passable survey. As it stands, I would still send it to peer review because the coverage is substantial and the taxonomy deserves scrutiny, but I would expect reviewers to insist that the quantitative ranking be stripped or redone.","headline":"A useful catalogue and a sensible single/hybrid taxonomy, but the central 'most promising methods' ranking is built on pooling RMSE and R2 values that are not commensurable across studies.","tokens_in":38827,"tokens_out":1773,"would_cite":false,"duration_ms":22701,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 180 comparative studies claims hybrid and ensemble machine-learning models—especially ANFIS, wavelet neural networks, and decomposition-based hybrids—are the most promising for flood prediction, with hybridization…","keywords":["flood prediction","machine learning","literature review","hybrid models","ensemble prediction systems","wavelet neural networks","ANFIS","support vector machines"],"falsifier":"A single controlled benchmark run on one or more public streamflow and rainfall datasets with identical train/test splits, showing that plain ANNs or SVMs match or beat the recommended hybrids such as ANFIS, WNN, or EEMD-ANN on both short and long lead times, would undermine the paper's central comparative conclusion.","tokens_in":1564,"feed_emoji":"🌊","tokens_out":2154,"duration_ms":64077,"temperature":0.7,"pith_summary":"This paper is a systematic literature review of machine-learning (ML) methods for flood prediction, screening 6,596 articles and selecting 180 comparative studies. It tries to establish which ML methods are most promising for short-term and long-term flood prediction, and which improvement strategies actually work. The central claim is that hybrid and ensemble methods—particularly ANFIS, wavelet neural networks, and decomposition-based hybrids—tend to outperform single models, and that hybridization, data decomposition, algorithm ensemble, and model optimization are the four effective strategies driving progress. The paper would matter because flood prediction directly affects evacuation planning, risk reduction, and policy, and a clearer map of which method class works for a given lead time could help hydrologists choose models.","feed_headline":"Hybrid ML models win flood prediction review","feed_subtitle":"After 6,596 papers screened, survey names decomposition and ensembles as the levers that improve forecasts.","key_machinery":"The central object of the review is a classification-and-comparison taxonomy rather than a new algorithm. Every selected study is sorted by prediction lead time (short-term versus long-term, with one week as the working boundary) and by model architecture (single versus hybrid), then evaluated on reported R² and RMSE values plus qualitative ratings of complexity, ease of use, speed, accuracy, and input dataset. This apparatus lets the authors aggregate performance signals from heterogeneous case studies into comparative tables and figures, from which they read the trend that hybrids, decomposition, ensembles, and optimization are the recurring levers of improvement.","core_discovery":"On its own terms, the paper claims that no single ML method dominates all flood-prediction tasks, but the comparative evidence organized by lead time points to distinct winners. For short-term prediction (lead times up to about a week), ANN variants, SVM/SVR, ANFIS, and decision-tree models are reported as the most promising single methods, while hybrid models such as ANFIS and wavelet-based networks give better accuracy beyond a two-hour lead time. For long-term prediction (weekly to annual), the paper reports that data-decomposition hybrids—WNN, WARM, EEMD-ANN, modified EMD-SVM, and similar—outperform undecomposed approaches, and that ensemble prediction systems reduce uncertainty. The paper further claims that the field's progress is driven by four strategies: hybridizing ML with other ML, soft-computing, or physical models; decomposing input time series; ensembling predictors; and adding optimizer algorithms for architectural or parameter tuning.","pith_inferences":["The paper's rankings treat RMSE and R² values reported in different studies, catchments, lead times, and data periods as comparable; normalizing these metrics by catchment runoff variance could shift the reported ordering of methods.","The four identified strategies are not fully independent: decomposition and ensemble overlap heavily in models like EEMD-ANN, so the effective number of distinct levers may be smaller than four.","A controlled benchmark on a single large hydrometeorological dataset, with identical train/test splits, would be a natural test of whether the recommended hybrids truly beat plain ANNs and SVMs.","For operational flood warning, the review underweights the trade-off between accuracy and lead time: a model with slightly lower R² but several extra hours of warning could be more valuable than the top-ranked method."],"forward_implications":["Hydrologists building short-term flood warnings should consider ANFIS or wavelet-hybrid ANN/SVR models over plain ANNs, especially for one-to-three-hour lead times.","Longer-lead forecasts, from weekly to annual, are best served by decomposition-based hybrids such as wavelet neural networks, WARM, EEMD-ANN, and modified EMD-SVM.","Ensemble prediction systems built from ANNs, MLP, SVM, or random forests can reduce forecast uncertainty and improve robustness.","Decomposing the input time series before training appears to be a broadly transferable accuracy boost across methods.","Optimization algorithms that tune network architecture and parameters are expected to yield further gains in both short- and long-term flood prediction."],"supporting_citations":[{"why":"Supplies the ARMA, ARIMA, and autoregressive ANN comparison used to support long-term hybrid claims.","marker":"[26]"},{"why":"Provides the long-term rainfall forecasting ANN baseline the review uses for defining long-term prediction tasks.","marker":"[37]"},{"why":"Reports SVR versus ANN results for regional flood frequency analysis, supporting the paper's SVR accuracy claim.","marker":"[48]"},{"why":"Describes wavelet neural network ensembles for streamflow, underpinning the decomposition-ensemble conclusion.","marker":"[50]"},{"why":"Introduces wavelet linear genetic programming for monthly streamflow, a key long-term hybrid example.","marker":"[51]"},{"why":"Compares random forests and SVM for radar rainfall forecasting, supporting the short-term method comparison.","marker":"[69]"},{"why":"Presents the wavelet-bootstrap-ANN hybrid, a central example of decomposition improving short-term flood prediction.","marker":"[105]"},{"why":"Compares ANFIS and ANN for flash floods, supporting the claim that ANFIS is superior for real-time estimation.","marker":"[174]"},{"why":"Demonstrates an ensemble prediction system of ANNs, supporting the ensemble-based uncertainty reduction claim.","marker":"[190]"},{"why":"Compares ANFIS, ANN, and SVM for monthly discharge, a load-bearing comparison for long-term prediction rankings.","marker":"[213]"}],"fun_headline_variants":["Hybrids and decomposition lead flood ML survey","No single ML model rules flood prediction","Short-term vs long-term: distinct ML winners in floods","Decomposition and ensembles boost flood prediction"],"cache_read_input_tokens":40960,"weakest_assumption_plain":"The survey's rankings treat RMSE and R² values reported in different studies, catchments, regions, lead times, and data periods as directly comparable, even though those numbers depend heavily on basin scale, flood magnitude, and dataset length.","fun_headline_variants_meta":{"raw":{"variants":["Hybrids and decomposition lead flood ML survey","No single ML model rules flood prediction","Short-term vs long-term: distinct ML winners in floods","Decomposition and ensembles boost flood prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000202,"raw_usage":{"total_tokens":1408,"prompt_tokens":999,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":352}},"tokens_in":615,"tokens_out":409,"duration_ms":4341,"temperature":1.0,"reasoning_tokens":352,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:33:49.405672+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single controlled benchmark run on one or more public streamflow and rainfall datasets with identical train/test splits, showing that plain ANNs or SVMs match or beat the recommended hybrids such as ANFIS, WNN, or EEMD-ANN on both short and long lead times, would undermine the paper's central comparative conclusion.","supporting_citations":[{"cited_title":"Estimation of instantaneous peak flow using machine -learning models and empirical formula in peninsular Spain","cited_arxiv_id":null,"evidence_quote":"Compares ANFIS and ANN for flash floods, supporting the claim that ANFIS is superior for real-time estimation."},{"cited_title":"Development and operational testing of a super-ensemble artificial intelligence flood-forecast model for a pacific northwest river","cited_arxiv_id":null,"evidence_quote":"Demonstrates an ensemble prediction system of ANNs, supporting the ensemble-based uncertainty reduction claim."},{"cited_title":"A comparison of performance of several artificial intelligence methods for forecasting monthly discharge time series","cited_arxiv_id":null,"evidence_quote":"Compares ANFIS, ANN, and SVM for monthly discharge, a load-bearing comparison for long-term prediction rankings."}],"review_version":1}