{"id":"d769d467-ca4c-40c1-b8cd-b7ade91fd447","arxiv_id":"2507.05761","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A hybrid fuzzy-feature wind speed predictor, CGT-BF, is claimed to improve short-term point and interval forecasts on Penglai wind farm data.","lead":"This paper proposes a wind speed forecasting system that combines fuzzy feature extraction, four machine learning models, and a multi-objective optimization algorithm to tune model weights. The authors report accuracy gains on data from a Chinese wind farm, but the method is incompletely described and the comparison is not fully controlled.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 21.6% improvement is not testable as reported: §2.1–§2.2 never specify whether FIC-MG windows are causal or whether granulation/clustering is fit on training data only; a look-ahead protocol would exactly produce the reported gains.","rationale":"The paper's contribution depends on FIC-MG being a valid, non-leaking preprocessing method. If leakage exists, all downstream point and interval comparisons are invalid; if leakage does not exist, the method is still underspecified but the headline number may be real. The reader's weakest assumption correctly identifies this as the decisive issue. I agree with that identification. The interval-prediction framework has no derivation (no equations for interval construction or for PICP/PINAW/AIS), and some baselines in Table 2 (GRU, XGBOOST) are not described in the method sections, but those are reproducibility problems that independently support REJECT; they do not replace the leakage check as the most direct test of the 21.6% claim. The proposed causal-reconstruction test is a single, concrete ablation that separates a genuine methodological improvement from an artifact of the preprocessing protocol. Since the reader already reached REJECT and my concern supports that verdict, no verdict change is needed.","tokens_in":19147,"tokens_out":5371,"duration_ms":58069,"concrete_test":"Obtain (or reconstruct) the three Penglai series and re-run CGT-BF under exactly the Table 2/3 split, varying only the FIC-MG protocol: (A) causal sliding window ending at t (features for target t+1 use only t−35..t) with FCM and granulation parameters estimated on training+validation only; (B) centered window [t−17..t+18] with full-data fitting; (C) causal window with full-data fitting. If protocol A reproduces the reported MAPE values (e.g., ~3.7–4.2 across Data1–3 for the proposed model), the claim survives; if reported values are only reproduced under B or C, or if A's MAPE rises by more than about 1 percentage point, the 21.6% improvement is a leakage artifact rather than a property of CGT-BF. The same comparison should also be run for the DM-test statistics in Table 7.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—average 21.6% point-prediction improvement and superior interval scores—rests on the FIC-MG preprocessing stage. Section 2.1 states only that \"a window size of 36 was used with the FIC-MG feature extraction method to obtain wind speed feature data for each wind speed point\"; it does not state whether the 36-point window for the feature at forecast time t is [t−35, t], [t−17, t+18], or something else. Section 2.2's cluster-center update (Eqs. 4–6) is written over \"the dataset T\" with no sentence restricting the fuzzy C-means fit and granulation parameter estimation to the 60% training segment. Two concrete leakage paths therefore exist: (i) future-inclusive windows put the target period's observed wind speeds into the regressors; (ii) fitting FCM/granulation on the full 21-series dataset, including the final 20% test segment, transfers distributional information about the test set into the features. Both would inflate MAPE/MSE improvements and DM-test significance. Because no code or data accompanies the paper, the reported 21.6% cannot be distinguished from the artifact of a non-causal protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CGT-BF, a short-term wind speed prediction system that combines a fuzzy information granulation and fuzzy rough C-means feature extraction module (FIC-MG) with four machine-learning predictors (BiLSTM, CNN–GRU, LSTM–XGB, and RF), whose ensemble weights are optimized by a multi-objective sunflower algorithm (IMOSFO). The system is evaluated on three 10-minute wind speed series from the Penglai wind farm under a 60/20/20 train/validation/test split. The authors report point-forecast metrics (MAPE, MSE, R²), interval-forecast metrics (PICP, PINAW, AIS), Diebold–Mariano tests, five-fold cross-validation, and optimization benchmarks on ZDT functions, and they claim an average point-prediction improvement of 21.6% over eight comparison models.","tokens_in":19415,"tokens_out":6699,"duration_ms":74933,"significance":"If the reported protocol is sound, the idea of extracting granular fuzzy features before feeding them to a weighted ensemble of neural and tree models is a plausible contribution to short-term wind speed forecasting, and the paper supplies a fairly broad empirical comparison: three datasets, four single and four ensemble baselines, five-fold cross-validation, DM tests, and benchmark test functions. The explicit holdout split for point forecasting and the use of multiple error metrics are strengths, as is the attempt to provide both point and interval forecasts. However, the contribution's significance hinges on the temporal causality of the FIC-MG preprocessing and on the reproducibility of the interval framework, neither of which is currently documented. The absence of code or data makes these omissions especially consequential: the central numerical claims cannot be independently checked.","major_comments":[{"comment":"The temporal protocol of FIC-MG is not specified. The manuscript states only that \"a window size of 36 was used\" (§2.1) and that the granulation and fuzzy C-means operations are applied to \"the dataset T\" (§2.2, Eqs. 4–6), but it never states whether the 36-point window used for the feature at forecast time t is [t−35, t], centered, or future-inclusive, nor whether the granulation parameters, the rough-set thresholds r1/r2 = 0.3/0.7, and the cluster centers are estimated on the 60% training segment only. If the windows include future observations or the preprocessing is fitted on the full dataset including the final 20% test segment, the MAPE/MSE gains and the DM-test significance reported in Tables 2, 3, and 7 would be inflated by information leakage. The authors must specify a causal, training-only protocol and re-run the experiments if any leakage is found.","section":"§2.1–§2.2"},{"comment":"The FIC-MG feature construction is incomplete. The method produces per-window Low, R, and Up parameters and updated cluster centers, but the manuscript does not define how these objects are assembled into a predictor vector for a given forecast target, how many fuzzy features enter each learner, or whether the same feature matrix is used for all four learners. Without this specification, neither the baselines' input features nor the proposed model's inputs can be reconstructed, and the comparison in §3.1 is not reproducible.","section":"§2.2"},{"comment":"The interval prediction framework is never described. Tables 2 and 3 report PICP, PINAW, and AIS for 95% and 85% confidence intervals, but Section 2 contains no equations or algorithmic steps for constructing the upper and lower bounds, and no statement about whether the intervals are calibrated on the validation set or the test set. The paper's 'dual-frame' claim (Introduction, item 2) therefore cannot be evaluated.","section":"§3.1 and §3.2"},{"comment":"The reported DM statistic is not a standard Diebold–Mariano statistic. The formula as written is a raw sum of squared-error differences with no normalization by the standard deviation of the loss differential or by sample size; as a result the values 8–13 in Table 7 cannot be interpreted as standard-normal test statistics. Please provide the correct DM formula, the loss differential series, and p-values.","section":"§4.2, Eq. (28)"},{"comment":"The claimed 'average point prediction improvement of 21.6%' is not supported by Table 8. Averaging the eight IRI entries per dataset gives approximately 15.1%, 19.7%, and 12.7% for Data1–Data3, with an overall average of roughly 15.9%; the 21.6% figure is not reproduced by any straightforward reading of the table. The conclusion should either report the exact computation or be corrected.","section":"§5 vs. Table 8"},{"comment":"The baseline comparison is not fully defined because the inputs to the benchmark models are not specified. It is unclear whether BILSTM, CNN–GRU, LSTM–XGB, and RF are fed the raw wind speed lags, the FIC-MG features, or the granulated Low/R/Up values; the text only gives hyperparameters. This affects the fairness and interpretability of every point- and interval-forecast comparison in Tables 2 and 3.","section":"§3.1"}],"minor_comments":[{"comment":"Numerous equations are garbled or mis-rendered (e.g., Eqs. 3, 5, 6, 20–21, 24), with symbols missing or misplaced; the manuscript needs a careful typesetting pass before it can be assessed accurately.","section":"Throughout"},{"comment":"In the Data2, CNN–GRU row, the PINAW values of 0.623 and 0.477 are implausible compared with all other entries (roughly 0.04–0.07) and appear to be typographical errors for 0.062 and 0.047.","section":"Table 2"},{"comment":"The text says the baselines are BILSTM, CNN–GRU, LSTM–XGB, and RF, but Table 2 also lists GRU and XGBOOST; please clarify whether these are separate baselines or duplicates.","section":"§3.1 and Table 2"},{"comment":"Terminology is inconsistent: 'Flod' in Table 4 should be 'Fold', and parameter names such as 'a sizeB', 'rateLearn', 'DepthMax', and 'trainMax' in §3.1 are not defined in Table 6.","section":"§3.1 and Table 6"},{"comment":"The reference list contains multiple 'Wang et al. (2023)' entries; the in-text citations should be disambiguated by author initials or reference numbers.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has enough novel components to merit a revision, but the current exposition is too incomplete to support the headline claims. The most serious issue is the unspecified temporal protocol for the 36-point granulation window: if the windows are future-inclusive or the fuzzy clustering is fit on the full dataset, the reported gains would be an artifact. The DM test formula and the 21.6% claim also need correction. I would ask for a revised manuscript that specifies the full algorithm, reports the exact IRI calculation, and either provides code/data or includes a detailed reproducibility appendix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a routine hybrid forecasting paper, and the new preprocessing block (FIC-MG) is an incremental arrangement of known pieces. The paper isn't empty—three real datasets, a 60/20/20 split, five-fold CV, eight comparisons, and a clear lineage from the corresponding author's earlier load-forecasting work. I believe the point-forecast results could be roughly right if the preprocessing is done honestly. But the manuscript as written does not let you verify that, and there are several internal inconsistencies that push me to reject.\n\nThe biggest problem is exactly what the stress-test flagged. Section 2.1 gives a window size of 36 but never says whether the window that produces the feature for time t is [t-35, t] or something centered. Section 2.2 estimates cluster centers and the granulation parameters \"over the dataset T\" with no statement that this is restricted to the training 60%. So the 21.6% average improvement—which I note is not reproducible from Table 8; the IRI entries average to about 15.9%—could be an artifact of look-ahead. This is a load-bearing ambiguity, not a style complaint.\n\nThere's more. The interval-prediction contribution has no derivation; there is no equation or algorithm in §2.5 that explains how PICP/PINAW/AIS are produced, yet it's claimed as a novel dual framework. The baselines' input features are never specified—do BiLSTM/CNN-GRU/etc. use raw wind speed or the same granulated features? If the former, the comparison conflates preprocessing and model. And the DM test section says all values are positive, but Table 7 contains negative DM statistics for MODA, NSMFO, and MOGWO on MSE in multiple datasets; negative values mean the proposed model is worse, contradicting the text. Table 2 also has at least one physically absurd PINAW entry (Data3 BILSTM: 0.399 at 85% confidence vs 0.054 at 95%), which looks like a data-entry error.\n\nWhat does the paper do well? The intention to combine fuzzy granulation with rough-set clustering and an adaptive ensemble is coherent, and the five-fold CV in Experiment III is a reasonable stability check. If the authors supply code/data and clarify the windowing protocol, this could become a modest, citable contribution in the wind-speed-prediction literature. But as it stands, the central empirical claim is not independently verifiable, and the internal contradictions undermine confidence.\n\nWho should read it: anyone who works on hybrid fuzzy/rough-set forecast preprocessing and wants to see the current state of that thread. I would not cite it in its present form. My recommendation to an editor: reject or, at most, send it back for major revision with an explicit request for a causal-windowing proof, a description of interval construction, and a correction of the DM/IRI reporting. The authors deserve a chance to fix it, but the burden of demonstration is entirely on them.","headline":"A plausible hybrid wind-speed forecasting system, but the windowing and interval construction are under-specified, the headline 21.6% gain does not match the paper's own table, and the reported statistics contain internal contradictions.","tokens_in":19948,"tokens_out":6461,"would_cite":false,"duration_ms":67571,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M20","62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"CGT-BF, built on fuzzy-set feature extraction and a weight-optimized ensemble, improves short-term wind speed point forecasts by an average of 21.6% and outperforms eight comparison models on most point and interval metrics.","keywords":["wind speed prediction","fuzzy information granulation","fuzzy rough set","feature extraction","ensemble learning","multi-objective optimization","interval prediction","short-term forecasting"],"falsifier":"Re-run the Penglai wind farm comparison under walk-forward validation, recomputing the 36-point granulation window and all FIC-MG clustering parameters using only past observations at each prediction step; if CGT-BF’s MAPE advantage over the eight baselines disappears or drops below the reported 21.6%, the headline improvement is an artifact of information leakage rather than a real forecasting gain.","tokens_in":18938,"feed_emoji":"🌬️","tokens_out":8274,"duration_ms":79994,"temperature":0.7,"pith_summary":"This paper is trying to establish that a wind speed forecasting system called CGT-BF can make point and interval forecasts that are substantially more accurate than eight alternatives, with an average point-prediction improvement of 21.6% on MAPE. The system replaces raw wind speed inputs with fuzzy-set features extracted from 36-point windows, then combines four machine-learning predictors whose weights are tuned by a multi-objective optimizer. Three datasets from the Penglai wind farm, five-fold cross-validation, DM significance tests, and interval-coverage metrics are used to support the claim. If the results hold, short-term wind power dispatch and turbine control could be planned from more reliable forecasts without collecting additional meteorological variables.","feed_headline":"Fuzzy granulation features lift wind forecasts by 21.6%","feed_subtitle":"Turning raw wind speed into fuzzy features improves short-term forecasts by over a fifth.","key_machinery":"The load-bearing mechanism is the FIC-MG feature extractor. A sliding window of 36 ten-minute observations is fuzzified with a triangular membership function that produces three granule parameters, $Low$, $R$, and $Up$, which serve as initial cluster centers. Fuzzy rough C-means clustering then assigns points to upper approximation sets, lower approximation sets, or boundary regions using thresholds $r_1=0.3$ and $r_2=0.7$, and the cluster centers are updated from the membership matrix to yield the final features. The second mechanism is the T-BF ensemble: BiLSTM, CNN-GRU, LSTM-XGB, and random forest each predict from the fuzzy features, and IMOSFO, a sunflower optimizer initialized with tent mapping and an adaptive t-distribution, sets the ensemble weights by minimizing MAPE and MSE together. Interval forecasts come from the same system, with confidence intervals evaluated by PICP, PINAW, and AIS.","core_discovery":"The central claim is that CGT-BF, an integrated multiframe system, achieves better short-term wind speed prediction than single models (BiLSTM, CNN-GRU, LSTM-XGB, random forest) and combined models (MODA, MSSA, NSMFO, MOGWO). The gain comes from the FIC-MG preprocessing module, which fuzzifies each 36-point granulation window into Low, R, and Up parameters, performs fuzzy rough C-means clustering with upper and lower approximation sets, and updates cluster centers to produce optimal feature values; the T-BF ensemble then predicts those features with four learners whose weights are optimized by IMOSFO under a dual MAPE/MSE objective. The paper reports that the proposed system outperforms the baselines in most point and interval metrics, with an average point-prediction improvement of 21.6%, and that the point-forecast differences are positive in Diebold-Mariano tests.","pith_inferences":["As an extension beyond the paper, the same FIC-MG front end should transfer to other noisy renewable series, such as solar irradiance or wave height, by retuning only the window size and the two approximation thresholds.","The dual MAPE/MSE objective implies an operator could tune the system toward smoothness or peak accuracy; the fixed weights selected here are only one point on that trade-off frontier.","The reported gains are one-step 10-minute forecasts from one wind farm, so they should not be extrapolated to multistep or multi-site horizons until tested.","A stricter check of the temporal alignment of the 36-point windows would determine whether the 21.6% figure survives when granulation is recomputed on past data only; the paper does not state this alignment explicitly."],"forward_implications":["On short-term 10-minute-ahead forecasts, CGT-BF reduces MAPE by an average of 21.6% relative to the eight baselines, which would translate into smaller prediction errors for wind-farm scheduling.","At 95% and 85% confidence levels, the interval predictions improve on PICP and AIS on most datasets, giving grid operators a narrower band that still covers observed wind speeds.","Because five-fold cross-validation keeps MAPE near 4.1–4.8 and $R^2$ near 0.98–0.99 across folds, the accuracy gain is not an artifact of one train/test split.","Positive DM test statistics against all eight comparison models on MAPE indicate the point-forecast differences are reported as significant at the 5% level.","IMOSFO’s Pareto fronts on the ZDT1–ZDT3 benchmark functions are claimed to dominate those of MSSA and MODA, supporting the weight-update module used inside CGT-BF."],"supporting_citations":[{"why":"Supplies the fuzzy information granular structure and triangular granule parameters ($Low$, $R$, $Up$) used to seed the clustering centers in FIC-MG.","marker":"(Xie et al., 2019)"},{"why":"The rough-set, information-granule, multi-objective load-forecasting design that FIC-MG adapts for wind speed feature extraction.","marker":"(Wang et al., 2023)"},{"why":"Defines the multi-objective sunflower optimizer that IMOSFO extends with tent-map initialization and adaptive t-distributions.","marker":"(Pereira & Gomes, 2023)"},{"why":"Provides the XGBoost boosting learner with regularization that the LSTM-XGB branch of the ensemble uses.","marker":"(Kim et al., 2023)"},{"why":"Supplies the random forest regression algorithm and averaging rule used as one of the four base predictors.","marker":"(Khan et al., 2020)"},{"why":"A fuzzy-soft-cluster cascade ensemble for extreme wind speeds, the nearest comparable approach and a motivation for the fuzzy feature front end.","marker":"(Pelaez-Rodriguez et al., 2023)"}],"fun_headline_variants":["Fuzzy wind speed forecasts gain 21.6% accuracy boost","Short-term wind prediction improved 21.6% with fuzzy sets","Fuzzy extraction reduces wind speed forecast error by 21.6%","Integrated fuzzy-system forecasts wind speeds 21.6% better","Wind speed prediction: fuzzy features lift accuracy 21.6%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 36-point fuzzy granulation window for a predicted time step must contain only observations at or before that step; if the window is centered or fitted on the full dataset including the test period, the reported 21.6% improvement could be inflated by information leakage.","fun_headline_variants_meta":{"raw":{"variants":["Fuzzy wind speed forecasts gain 21.6% accuracy boost","Short-term wind prediction improved 21.6% with fuzzy sets","Fuzzy extraction reduces wind speed forecast error by 21.6%","Integrated fuzzy-system forecasts wind speeds 21.6% better","Wind speed prediction: fuzzy features lift accuracy 21.6%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1472,"prompt_tokens":967,"completion_tokens":505,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":414}},"tokens_in":583,"tokens_out":505,"duration_ms":5406,"temperature":1.0,"reasoning_tokens":414,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:18:54.181437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Penglai wind farm comparison under walk-forward validation, recomputing the 36-point granulation window and all FIC-MG clustering parameters using only past observations at each prediction step; if CGT-BF’s MAPE advantage over the eight baselines disappears or drops below the reported 21.6%, the headline improvement is an artifact of information leakage rather than a real forecasting gain.","supporting_citations":[],"review_version":1}