{"id":"b16f1474-212e-47e6-9ca6-c85858447a14","arxiv_id":"2506.10536","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LightGBM outperforms LSTM, XGBoost, and CatBoost for day-ahead electricity price forecasting on short training windows of 45 to 90 days across three European markets.","lead":"Across Greece, Belgium, and Ireland, this paper tests how well four machine learning models forecast next-day electricity prices when trained on only 7 to 90 days of history. It reports that LightGBM, a gradient boosting model, is the most accurate and robust, especially with medium-length training windows.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's impossible RMSE<MAE entry (Greece, 60-day XGBoost) and the starred-best conflict with the abstract's 45–60 day claim mean the central result is not currently supported by the paper's own numbers.","rationale":"The paper's contribution is a comparative benchmark; its central claim is an empirical ranking. The only quantitative table behind that ranking is Table 2, and that table fails a basic inequality (RMSE >= MAE) at a starred entry. This is not a matter of external consensus or hyperparameter choice; it is an internal inconsistency in the primary evidence. If the entry is wrong, the reader cannot trust the rest of the table without re-running. The conflict between the starred entries and the abstract's 45-60 day claim is a second internal inconsistency: the table's own 'optimal' marks in Greece and Belgium point to 90 days, while the text claims 45-60. Together, these mean the headline result is not currently established by the paper's own data.\n\nThe reader identified a generalization risk (one year, fixed hyperparameters) as the weakest assumption; that is real but secondary. The internal table inconsistency is more load-bearing because it undermines even the single-year, single-configuration comparison. A corrected table may or may not preserve the LightGBM advantage; until the error is fixed and the window-level ranking is re-reported honestly, the central claim should be treated as unverified.\n\nMy verdict remains conditional (revision required), consistent with the reader's conditional recommendation rather than a change to it, so I mark the verdict adjustment as UNCHANGED.","tokens_in":17174,"tokens_out":6247,"duration_ms":65864,"concrete_test":"Recompute Table 2, '60 days, Greece': derive MAE and RMSE from the same stored prediction error vector. If MAE=18.303, RMSE must be >=18.303; any value below that disproves the aggregation. Then re-tabulate for each country the window length minimizing LightGBM MAE/RMSE and maximizing FSI. If the argmin is 90 days for Greece and Belgium, revise the abstract's 'particularly 45-60 days' sentence to match the data. Also report repeated-seed mean±std for at least 5 seeds for the best two models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest support for the abstract's claim ('LightGBM consistently achieves the highest forecasting accuracy ... particularly with 45-60 day training windows') is Table 2. That table contains an impossible entry: for Greece, 60-day XGBoost, MAE = 18.303 but RMSE = 12.399 (starred). Since RMSE = sqrt(mean(e_t^2)) >= mean(|e_t|) = MAE for any error vector, this cannot be correct. The entry can only be a transcription or aggregation error, but it invalidates that row as evidence and raises doubt about the rest of the table.\n\nThe star markers also contradict the 45-60-day narrative. For LightGBM in Greece all four metrics are starred at the 90-day window (MAE 11.899, R2 0.817, FSI 0.482; RMSE 17.891 is the LightGBM best at 90 days though not starred because XGBoost's impossible 12.399 undercuts it). In Belgium, LightGBM's starred MAE and FSI are at 90 days. Only Ireland's LightGBM stars sit at 60 days. The paper's own starred 'optimal for all windows' entries do not support 'particularly with 45-60 day windows'; they support 90 days in two of three markets. Section IV-C and IV-D pivot to 'some cases' and to seasonal/peak figures, but Figures 9-10 provide no numeric tables or significance tests, so they cannot carry the 45-60-day conclusion.\n\nAdditionally, the single-year (2023) evaluation and fixed hyperparameters without repeated seeds mean even a corrected table would only show a one-year, one-configuration comparison. The immediate blocker, however, is internal: the quantitative backbone of the central claim is inconsistent.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks four machine learning models (LSTM with FFEC, XGBoost, LightGBM, CatBoost) and a naive persistence baseline for day-ahead electricity price forecasting in Greece, Belgium, and Ireland, using only forecasted features from ENTSO-E. The central claim is that LightGBM consistently achieves the highest accuracy and robustness, especially with 45- and 60-day training windows, which are said to balance temporal relevance and learning depth. The evaluation covers six training-window lengths (7-90 days) over a single year (2023), with per-month 80/20 splits and metrics MAE, RMSE, R2, and FSI. The paper also presents a seasonal and peak-price analysis (Figures 9-10) supporting the 45-60 day optimum. The core result is an out-of-sample benchmark, but the evidence table contains internally inconsistent entries and the summary claims are not fully aligned with the tabulated results.","tokens_in":17509,"tokens_out":4200,"duration_ms":52240,"significance":"If the central claims were fully supported, the paper would make a useful practical contribution: it addresses an operationally relevant question (how little training data suffices for DAM price forecasting) and uses realistically available forecasted features rather than ex-post actuals. The systematic sweep over window sizes and the inclusion of three distinct European markets add empirical breadth, and the FSI relative to persistence is a constructive evaluation choice. However, the current manuscript does not establish the central claim: Table 2 contains an impossible RMSE < MAE entry, the starred 'optimal' entries contradict the 45-60-day narrative, and the 'consistent' and 'robust' language is unsupported by any statistical uncertainty quantification. The paper's usefulness depends on correcting these internal inconsistencies and either adding statistical support or tempering the claims.","major_comments":[{"comment":"Table 2, Greece 60-day XGBoost row reports MAE = 18.303 and RMSE = 12.399. Since RMSE = sqrt(mean(e^2)) >= mean(|e|) = MAE for any fixed error vector, this entry is impossible. The RMSE is also starred as the best among all windows for Greece, which further propagates the error into the summary. Please correct the underlying computation or transcription and re-evaluate the rankings; if the corrected values change the reported best models or windows, revise the abstract and Section IV-C accordingly.","section":"IV-B, Table 2"},{"comment":"The abstract and conclusion claim that LightGBM performs 'particularly with 45-60 day training windows,' but Table 2's own star markers (described in IV-C as 'optimal for all training windows') show LightGBM's best MAE, R2, and FSI in Greece at 90 days, and best MAE and FSI in Belgium at 90 days. Only Ireland's LightGBM stars are at 60 days. Section IV-C itself states that 'longer windows, particularly 90 days, result in more reliable forecasts' and that 60-day beats 90-day only 'in some cases.' The 45-60-day claim is the paper's headline result, yet it is not supported by the table's summary statistics; the authors must reconcile the narrative with the data or provide other quantitative evidence.","section":"Abstract, IV-C, Table 2"},{"comment":"The data-splitting and windowing protocol is underspecified. The text says an 80%/20% training/test split 'applies to each month of the dataset' and that the earliest training date depends on the window size, but it does not clarify whether the split is chronological or random, nor how the time-step-shifting mechanism (Section III-A, Figure 2) treats test samples whose previous-24-hour input sequences overlap with the training period. If the 20% test points are selected randomly within each month and the input features for those points include hourly values from earlier in the same month, the test set may contain training-period information, invalidating the reported generalization. Please specify the exact split rule and confirm that no test sample's lagged inputs come from the training partition.","section":"III-A, IV-A"},{"comment":"The paper's claims of 'consistent' superiority and of optimal 45-60-day windows for seasonal and peak forecasting are made without any uncertainty quantification. All models use one fixed hyperparameter configuration (Table 1) across all markets and windows, and the evaluation covers a single calendar year (2023). No error bars, repeated-seed runs for the LSTM, bootstrap intervals, or significance tests are reported. Figures 9-10 provide only graphical seasonal/peak MAE trends, yet Section IV-D makes concrete quantitative assertions (e.g., 'the 45-day and 60-day training windows tend to outperform the 90-day window in these seasons' and 'the 30 and 60-day window achieves the lowest peak error of all configurations' for Greece) without tabulated numbers. To support the central conclusions, please add statistical comparisons (e.g., paired tests or confidence intervals) and, at minimum, report the numeric values underlying Figures 9-10.","section":"IV-A, IV-D, Table 1, Figures 9-10"},{"comment":"The sentence 'LightGBM consistently outperforms the other models in nearly every metric and training window' is too strong given that Table 2 itself shows instances where XGBoost or CatBoost match or beat LightGBM (e.g., Greece 7-day MAE and RMSE; Belgium 7-day RMSE; Ireland 90-day MAE and FSI). The paper later qualifies this with 'nearly every,' but the abstract and conclusion repeat the unqualified 'consistently highest accuracy.' Please either moderate the claim to match the quantitative evidence or provide statistical tests showing the differences are meaningful across markets and windows.","section":"IV-C, conclusion"}],"minor_comments":[{"comment":"The same work by Tschora et al. (2022) appears as references [9] and [17]; please deduplicate and renumber.","section":"II-B2, references"},{"comment":"MAPE is defined in Eq. (17) but never reported in Table 2 or analyzed in Section IV; please either report the MAPE values or remove the definition to avoid a dangling metric.","section":"IV-B, Table 2"},{"comment":"Table 1 misspells 'Value' as 'V alue' in the column header; also, the text in IV-A says 'respected hyperparameters' and should be 'respective hyperparameters.'","section":"Table 1"},{"comment":"There are several typographical and grammatical issues, e.g., 'it's' for 'its' in Section II-B1, 'Y orat' in the references, 'concluding, accurate forecasting' in Section V, and inconsistent use of 'Naive' vs. 'Naïve' in Table 2; a full proofread is needed.","section":"Throughout"},{"comment":"Figure 2 is referenced as showing the time-step shifting mechanism, but the caption and surrounding text do not define the notation n = 24 beyond 'previous 24 hourly time steps.' Please clarify whether the input sequence for each target hour includes the 24 past hours of all features, and how this interacts with the per-month split described in Section III-A.","section":"III-A, Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's core experimental result cannot be accepted in its current form because Table 2 contains a mathematically impossible entry and the abstract's 45-60-day optimal-window claim contradicts the table's own star markers. The paper also lacks statistical support for 'consistent' and 'robust' claims. These are substantial but potentially fixable issues if the authors can correct the table and rerun the analysis, add uncertainty quantification, and clarify the data-splitting protocol to rule out leakage. The paper is within the scope of the journal's applied energy/ML interests, but the referencing quality (duplicate reference, multiple typos) suggests a need for careful revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the paper asks a genuinely practical question—how much training history do you need for day-ahead price forecasting?—and the short-window boosting answer (especially LightGBM) is plausible. But the quantitative support is currently broken: Table 2 has an impossible entry, and the star markers point to 90 days, not the 45–60 days the abstract advertises. Fix that, add error bars, and there is a useful benchmark here.\n\nWhat is actually new: the systematic comparison of 7–90 day training windows with per-month independent retraining across Greece, Belgium, and Ireland, using only forecasted features available at bid time. That is a legitimate scenario-based extension of the forecasting literature, and it directly addresses a real operational constraint in data-scarce markets. The related-work section is thorough, and the choice to evaluate seasonal and peak behavior is sensible.\n\nThe soft spots are real and they sit in the load-bearing part of the paper. The Greece 60-day XGBoost row lists MAE 18.303 and RMSE 12.399, which is mathematically impossible since RMSE is always at least MAE. That is not a cosmetic typo—it is in the exact table that supports the main claim. The star markers also contradict the abstract: for Greece and Belgium, LightGBM's starred entries are at 90 days, not 45–60, and Ireland's stars are at 60 days. So the paper's own optimal-window table supports 90 days more than the headline. Figures 9–10 are presented as evidence for the 45–60 day advantage, but without numeric tables or significance tests they cannot carry that conclusion.\n\nThe evaluation is also limited to a single market-year (2023) with fixed hyperparameters, no repeated-seed analysis, and no error bars. That means even a corrected table would only show a one-year, one-configuration comparison. No code or data are provided, which makes it hard to verify the aggregation.\n\nWhat the paper does well: the methodology is clearly described, the use of forecasted features is a realistic and commendable choice, and the three-market comparison is informative. The authors are upfront about the scope and cite relevant work.\n\nBottom line: I would send this to peer review, but only with a strong expectation of major revision. The question is worth answering, the benchmark design is sound, and the core finding—boosting works well with short windows—is believable. But the authors need to fix the table, add uncertainty quantification, soften the 45–60 day claim to match their own data, and release the code or at least the exact split protocol. As it stands, the central result is plausible but not supported by the paper's own numbers.","headline":"Useful benchmark, but a table error and overstated abstract undermine the central claim; short-window boosting is plausible but needs a corrected, uncertainty-aware revision.","tokens_in":18101,"tokens_out":3554,"would_cite":false,"duration_ms":40775,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LightGBM, trained on a 45-60 day rolling window, gives the most accurate day-ahead price forecasts across Greece, Belgium, and Ireland in 2023.","keywords":["day-ahead market forecasting","electricity price forecasting","LightGBM","gradient boosting","short training windows","seasonal variations","price spike detection","machine learning"],"falsifier":"Run the exact same 7-90 day benchmark on 2022 and 2024 price data for Greece, Belgium, and Ireland with the same fixed hyperparameters; if LightGBM does not lead in MAE or FSI, or if the best window moves outside 45-60 days, the paper's central claim fails.","tokens_in":16977,"feed_emoji":"⚡","tokens_out":9827,"duration_ms":97470,"temperature":0.7,"pith_summary":"This paper tests whether machine-learning models can forecast day-ahead electricity prices when trained on only a short slice of recent history, from 7 to 90 days. Across the 2023 markets of Greece, Belgium, and Ireland, it finds that LightGBM, a gradient-boosting tree method, yields the lowest error and highest robustness in nearly every window, with the best results at 45 and 60 days. The paper argues that these medium windows capture the recent market regime while avoiding stale trends, and that LightGBM detects seasonal swings and price peaks better than an LSTM with feed-forward error correction, XGBoost, CatBoost, or a naive persistence baseline. If true, this makes short-window boosting a practical recipe for volatile or data-scarce power markets.","feed_headline":"LightGBM beats rivals on short-window power price forecasts","feed_subtitle":"A gradient-boosting model trained on 1-2 months of data wins on accuracy and spike detection across three EU markets.","key_machinery":"The load-bearing mechanism is a rolling look-back sweep: for each calendar month of 2023, an 80/20 train/test split is applied and each model is trained on the preceding 7, 14, 30, 45, 60, or 90 days of data. LightGBM's leaf-wise gradient boosting, which grows trees by the highest loss-reduction gain and downsamples small-gradient samples, is the method that exploits these shallow windows best. The pipeline also relies on 24-hour time-step shifting of the input series and on using only public forecast features for demand, renewable generation, total generation, and net flows, which would be available to a bidder at decision time.","core_discovery":"The paper's central discovery is that LightGBM, a gradient-boosting tree method, delivers the most accurate day-ahead electricity price forecasts when trained on a rolling window of roughly 45 to 60 days of recent data, across all three markets studied (Greece, Belgium, Ireland) in 2023. This result holds both for aggregate error metrics (MAE, RMSE, R2, Forecast Skill Index) and for the harder tasks of seasonal fluctuation and price-spike detection. The authors interpret the medium window as balancing temporal relevance against learning depth: shorter windows (7-30 days) give the model too little structure, while a 90-day window can drag in outdated market regimes. They further show that the other boosting models (XGBoost, CatBoost) and an LSTM with feed-forward error correction trail LightGBM in almost every window and market.","pith_inferences":["The fixed hyperparameters per model may favor LightGBM; retuning the LSTM or running repeated seeds could narrow the reported gap, so the headline ranking is best read as a configuration-level comparison rather than a pure architectural one.","A multi-year extension (for example, 2020-2024) would reveal whether 45-60 days remains the sweet spot across volatility regimes, or whether optimal window length tracks market turbulence.","The peak and valley analysis suggests a LightGBM component could strengthen hybrid forecasters that currently rely on deep sequence models for spike detection.","Because the medium windows win in the most volatile seasons (summer and fall), an adaptive window length that is shortened in calm periods and lengthened in stable ones could squeeze out further gains."],"forward_implications":["An operator with only a month or two of recent market data can obtain competitive day-ahead price forecasts from a LightGBM model instead of a long-history deep network.","The 45-60 day window should be treated as a default starting point for short-window day-ahead market forecasting studies, ahead of 90-day or multi-year histories.","Boosting tree models, not just recurrent networks, deserve a standard place in electricity price forecasting benchmarks for volatile or data-scarce markets.","Because only forecasted features are used, the reported accuracy is attainable in real bidding workflows without look-ahead information.","The underperformance of the 90-day window relative to 45-60 days implies that adding older data can actively hurt forecast skill in non-stationary markets."],"supporting_citations":[{"why":"Supplies the forecast features (demand, renewable generation, total generation, net flows) used to train and test all models.","marker":"[22]"},{"why":"Introduces the LSTM with feed-forward error correction architecture that is the deep-learning baseline the comparison must beat.","marker":"[29]"},{"why":"Provides the comprehensive review of day-ahead forecasting methods and benchmark practices that frame the study's evaluation design.","marker":"[10]"},{"why":"Motivates the use of the Forecast Skill Index alongside MAE, RMSE, and R2 for cross-market model comparison.","marker":"[37]"},{"why":"Demonstrates the value of enriched forecast features and scenario-based training that this paper adapts to short windows.","marker":"[17]"}],"fun_headline_variants":["45-60 day data: LightGBM beats deep learning on EU prices","LightGBM best for price spikes with 45-60 day windows","Short windows: LightGBM leads EU power forecasts","LightGBM wins on 45-60 day price forecasts in EU","Gradient boosting beats LSTM on short-window power prices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings depend on one specific year (2023) and on a single fixed hyperparameter configuration per model, so the results could change with a different market year or if the LSTM and boosting models were each tuned separately.","fun_headline_variants_meta":{"raw":{"variants":["45-60 day data: LightGBM beats deep learning on EU prices","LightGBM best for price spikes with 45-60 day windows","Short windows: LightGBM leads EU power forecasts","LightGBM wins on 45-60 day price forecasts in EU","Gradient boosting beats LSTM on short-window power prices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001826,"raw_usage":{"total_tokens":7158,"prompt_tokens":897,"completion_tokens":6261,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":6170}},"tokens_in":513,"tokens_out":6261,"duration_ms":46435,"temperature":1.0,"reasoning_tokens":6170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:23:58.888763+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the exact same 7-90 day benchmark on 2022 and 2024 price data for Greece, Belgium, and Ireland with the same fixed hyperparameters; if LightGBM does not lead in MAE or FSI, or if the best window moves outside 45-60 days, the paper's central claim fails.","supporting_citations":[{"cited_title":"Transparency platform, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the forecast features (demand, renewable generation, total generation, net flows) used to train and test all models."},{"cited_title":"Georgilakis","cited_arxiv_id":null,"evidence_quote":"Introduces the LSTM with feed-forward error correction architecture that is the deep-learning baseline the comparison must beat."},{"cited_title":"Forecasting day-ahead electricity prices: A review of state-of-the-art al- gorithms, best practices and an open-access benchmark","cited_arxiv_id":null,"evidence_quote":"Provides the comprehensive review of day-ahead forecasting methods and benchmark practices that frame the study's evaluation design."},{"cited_title":"Unsupervised domain adaptation methods for photovoltaic power forecasting","cited_arxiv_id":null,"evidence_quote":"Motivates the use of the Forecast Skill Index alongside MAE, RMSE, and R2 for cross-market model comparison."},{"cited_title":"Elec- tricity price forecasting on the day-ahead market using machine learning","cited_arxiv_id":null,"evidence_quote":"Demonstrates the value of enriched forecast features and scenario-based training that this paper adapts to short windows."}],"review_version":1}