{"id":"e2dc0a19-238a-4252-b82a-8d01aaa69f81","arxiv_id":"2507.15001","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A GDP-based annual forecast combined with stable seasonality indices from the baseline year yields 4-7% MAPE for multi-year-ahead hourly electricity load in Singapore, Belgium, and Bulgaria.","lead":"This paper introduces a parsimonious framework that turns GDP-based annual electricity forecasts into hourly predictions using historically stable load shapes. The method is tested on Singapore, Belgium, and Bulgaria, and the authors report long-term hourly errors near 6%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Horizon label in the central claim is unsupported: the experiments are baseline+4 (four-year-ahead) forecasts, not five- or six-year-ahead.","rationale":"The reader identified the stability of seasonality indices as the weakest assumption, which is indeed central to the method's generalizability. However, that concern does not directly falsify the reported MAPE values: the out-of-sample evaluation already embeds any SI instability for the tested years. The horizon mislabel, by contrast, is a direct factual inconsistency in the central claim itself. Section 3.1 states 'baseline year + 4' while Section 3 calls the same forecasts '5-year ahead' and the abstract says 'six-year-ahead'; the paired baseline years 2013–2018 and forecast years 2017–2022 all differ by 4. This makes the headline claim internally unsupported and should be corrected before acceptance. The reader's rationale did note an 'off-by-one horizon labeling error', so there is partial agreement, but the reader did not elevate it to the primary load-bearing concern. The short-term experiment's look-ahead bias and the nonstandard t-test are real weaknesses, but the horizon mislabel is the most concrete threat to the central claim. If the horizon is corrected to four years, the method may still be credible, so the verdict should remain CONDITIONAL rather than shifting to reject or accept.","tokens_in":24219,"tokens_out":10421,"duration_ms":107217,"concrete_test":"For each of the six experiments in Table 6, compute (forecast year - baseline year) using the baseline years stated in Section 3 ('2013 to 2018, respectively'). If every difference equals 4, then the labels 'five-year-ahead' (Section 3) and 'six-year-ahead' (abstract) are incorrect, and the abstract should be revised to 'four-year-ahead' (or the experimental design should be changed to cover truly longer horizons). Additionally, reconcile Figure 5's 'baseline year is 2014' for a 2019 sample week with the stated baseline+4 scheme.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's headline claim depends on 'six-year-ahead forecasts', and Section 3 says '5-year ahead forecasts', but Section 3.1 explicitly states the models predict 'total yearly load in baseline year + 4'. The stated baseline years (2013–2018) for forecast years (2017–2022) give differences of exactly 4 years, not 5 or 6. Table 6 and Figure 5 add further inconsistency: Figure 5 uses baseline 2014 for a 2019 sample week, a 5-year gap, contradicting the paired baseline+4 scheme. If the actual horizon is four years, the central claim overstates the forecasting difficulty and the MAPE numbers, while potentially valid, do not demonstrate six-year-ahead accuracy. Because the horizon is intrinsic to the claim, this mislabeling is load-bearing and must be corrected before the result can be fully assessed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a parsimonious, top-down framework for long-term hourly electricity demand forecasting. Seasonality indices for month, day of week, and hour are computed from historical load data; pairwise t-tests with an epsilon-adjustment are used to argue that these indices are stable over time. Annual load is regressed on GDP via simple linear regression, and the resulting annual forecast is distributed to individual hours using the baseline year's seasonality indices. The method is applied to Singapore (2004-2022), Belgium, and Bulgaria, with claimed maximum MAPEs of 6.87%, 6.81%, and 5.64%. The stability indices are also plugged into an exponential smoothing model for short-term forecasting and compared with several machine-learning benchmarks.","tokens_in":24389,"tokens_out":3260,"duration_ms":36064,"significance":"If the results hold, the paper offers a notable counterpoint to complex bottom-up and machine-learning approaches: a transparent, low-parameter method for long-term hourly load forecasting, validated in three countries with different economic characteristics. The use of publicly available data, the explicit focus on statistical verification of stability, and the cross-country extension are strengths. However, the headline claim is weakened by an inconsistent and apparently overstated forecast-horizon label, and the stability test is presented without the statistical detail needed to assess its validity. The short-term extension also contains a look-ahead bias. The core idea is plausible and potentially useful, but the evidence as presented is not yet fully convincing.","major_comments":[{"comment":"The forecast horizon is mislabeled in a way that affects the central claim. Section 3.1 states that the models predict 'total yearly load in baseline year + 4', and Table 6 lists forecast years 2017-2022 with baseline years 2013-2018, which is a four-year gap in each case. Yet the abstract claims 'six-year-ahead forecasts' and Section 3 states '5-year ahead forecasts'. Figure 5 further uses baseline year 2014 for a 2019 sample week, a five-year gap. The stated maximum MAPE of 6.87% is therefore for four-year-ahead forecasts, not six-year-ahead. Because the horizon is an intrinsic part of the contribution's claimed difficulty, this mislabeling must be corrected in the abstract, Section 3, the conclusion, and Figure 5, or the experiments must actually be run at the stated horizons.","section":"Abstract and Section 3.1"},{"comment":"The t-test procedure for stability is not fully specified and, as described, does not appear to be a standard statistical test. The text does not report the significance level (alpha) used for the two one-tailed t-tests. The procedure increments epsilon by 0.001 until both null hypotheses are rejected; the final epsilon is then interpreted as the 'maximum difference'. With 19 years, there are 171 pairwise year comparisons for each of 168 hour-by-day combinations, so tens of thousands of tests are performed, yet no multiple-testing correction is mentioned. The reported quantities are therefore data-dependent, and their statistical meaning is unclear. Please specify the significance level, describe how multiple testing is handled (the error rate is otherwise uncontrollable), or present the analysis explicitly as a descriptive stability measure rather than as a formal hypothesis test.","section":"Section 2.2.4"},{"comment":"The short-term exponential smoothing extension uses the full-sample seasonality indices mu_{d,h} from 2004-2022 (Section 5 states 'we make use of the overall seasonality indices ... from (9)') and then evaluates the model on 2018-2022. This means the seasonality components are estimated on data that include the evaluation period, a look-ahead bias. The comparison with machine-learning methods that are fit only on the training period is therefore favorable to the proposed ES model. The seasonality indices should be estimated recursively using only data up to each forecast origin, or the limitation should be explicitly acknowledged and the comparison revisited.","section":"Section 5 and Table 7"}],"minor_comments":[{"comment":"The abstract refers to both 'five-year-ahead total yearly forecasts' and 'six-year-ahead forecasts' for the same experiments; these labels are inconsistent and need to be reconciled with the actual horizons.","section":"Abstract"},{"comment":"The half-hourly to hourly transformation is described with care, but the statement that summing two half-hourly MW values 'is equivalent to 5700MWh' could be clearer: the sum in MW equals the energy in MWh only for a one-hour integration period; consider a more explicit unit conversion.","section":"Section 2.1"},{"comment":"All p-values are reported as '0.00'; they should be reported as '< 0.001' or similar to reflect that they are rounded.","section":"Tables 4 and 5"},{"comment":"The column headers 'Percentage Error (Total Year)' and 'Percentage Error (Hourly)' do not indicate whether these are absolute percentage errors; MAPE implies absolute values elsewhere, so consider renaming for consistency.","section":"Table 6"},{"comment":"The statement that the model was 'trained using the 2006-2017 dataset' is imprecise: the only fitted component is the linear regression on GDP, and the data are used for both regression and seasonality indices; please specify which parts are estimated on which periods.","section":"Section 4"},{"comment":"The description of the right-tailed test contains a typo: the null hypothesis should likely be 'mu_y - mu_y' <= epsilon' rather than 'mu_y - mu_y' <= -epsilon', as written the two tests appear asymmetric.","section":"Section 2.2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is readable and the proposed framework is simple, but the horizon mislabeling is serious and affects the abstract's headline number. The t-test section would benefit from a statistician's review, and the short-term comparison is currently not apples-to-apples. The authors should be asked to correct or clarify these issues before the manuscript is considered for publication. It might also be worth checking whether the supplementary material they reference is actually available, since the reproducibility claim depends on it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper's core idea—verify that hourly seasonality indices are stable over multi-year windows, then use GDP-based annual forecasts with those indices to get hourly forecasts—is sensible and the Singapore/Belgium/Bulgaria results are plausible. But the headline horizon is wrong. The experiments are baseline+4 (four-year-ahead) forecasts, not five- or six-year-ahead as the abstract and Section 3 claim. That matters because the whole selling point is long-horizon accuracy; 4-year errors do not support a 6-year claim.\n\nWhat's genuinely new: combining a formal-ish stability check (t-tests on pairwise year differences of seasonality indices) with a top-down GDP regression to produce hourly long-term forecasts. The cross-country replication is a real plus, and the data and scripts are public, so the numbers are checkable. The MAPEs of 5-7% for hourly forecasts at a 4-year horizon are credible for this kind of parsimonious method.\n\nSoft spots, in order of severity:\n\n1. Horizon mislabeling (load-bearing). Text says 5-year, abstract says 6-year, actual is baseline+4. Also Figure 5 uses a 5-year gap for 2019 (baseline 2014) which contradicts the paired scheme. This needs to be corrected; the reported accuracy is for 4-year-ahead, and the claim should be restated accordingly.\n\n2. The t-test procedure in Section 2.2.4 is nonstandard and under-specified. It doesn't state the significance level, applies repeated pairwise tests without multiplicity adjustment, and the epsilon-increasing rule is a descriptive 'minimum detectable difference' rather than a test of stability. It would be fine as an exploratory descriptive measure, but the paper sells it as statistical verification. Should be reframed or strengthened.\n\n3. The short-term ES comparison has look-ahead bias: the seasonality indices are estimated on the full sample (2004-2022) and then used to forecast 2018-2022, which gives the ES model future information that the ML baselines don't have. The comparison is unfair and the short-term section should either be corrected or dropped.\n\nThe GDP regression itself looks fine—it uses WEO GDP forecasts available at baseline, trained on data up to baseline—and the long-term forecasts are not circular.\n\nWho this is for: energy analysts and utilities that need hourly load shapes for grid planning and don't want to build bottom-up sectoral models. A serious referee should engage with it; the method is potentially useful and the errors are plausible. But the horizon correction is mandatory, and the t-test and short-term sections need to be addressed. I'd recommend conditional acceptance after major revision.","headline":"A sensible and reproducible long-term hourly load forecasting method, but the headline 'six-year-ahead' accuracy is actually a four-year-ahead result and the paper needs to correct that and several statistical issues.","tokens_in":24898,"tokens_out":3376,"would_cite":false,"duration_ms":33968,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","62F03","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Hourly electricity demand can be forecast five years ahead with under seven percent error by combining a GDP regression with the observation that the hourly shape of demand is stable across years.","keywords":["electricity demand forecasting","seasonality indices","load stability","GDP-based forecasting","long-term hourly forecasting","exponential smoothing","t-tests","Singapore electricity demand"],"falsifier":"Recompute the hourly seasonality indices on data from a later window, for example 2023-2025 for Singapore, and compare them with the 2004-2022 means: if the average absolute deviation exceeds the measured 4.24 percent, or the per-hour values exceed the epsilons in the paper's Table 2, the stability premise fails and hourly errors should exceed the reported MAPE bound. A weaker but still decisive test is to run the full procedure on a country with rapid rooftop-solar or EV growth and check whether hourly MAPE over a four-year horizon grows while annual MAPE stays small.","tokens_in":2051,"feed_emoji":"⚡","tokens_out":2864,"duration_ms":77926,"temperature":0.7,"pith_summary":"The paper claims that long-term hourly electricity demand can be forecast accurately from two simple, testable ingredients: the relative shape of hourly demand is statistically stable over many years, and annual demand tracks GDP. On Singapore data from 2004 to 2022, it measures that stability as a 4.24 percent average deviation in hourly seasonality indices, then regresses total yearly demand on GDP and distributes that total across hours using seasonality indices from four to five years earlier. The resulting five-year-ahead hourly forecasts stay within 6.87 percent MAPE, with similar results in Belgium (6.81 percent) and Bulgaria (5.64 percent). The same stability estimates also give a constant-seasonality exponential smoothing model short-term accuracy comparable to machine learning baselines. The payoff is a parsimonious alternative to bottom-up sectoral forecasting, which needs many hard-to-verify assumptions.","feed_headline":"Five-year hourly power forecasts land within 7 percent error","feed_subtitle":"GDP regression sets the yearly total; stable daily demand shapes split it into hours, no heavy models needed.","key_machinery":"The load-bearing object is the seasonality index (SI), a multiplicative factor measuring how much demand in a given month, day of the week, or hour deviates from the yearly average, computed in three nested levels: $\\mathrm{SI}^1_{y,m}$ for months, $\\mathrm{SI}^2_{y,m,d}$ for days of the week, and $\\mathrm{SI}^3_{y,m,d,h}$ for hours. Stability is defined operationally: for every hour and day, a series of two-sample one-tailed t-tests between all pairwise years finds the smallest $\\epsilon$ such that the null hypothesis of a difference at least $\\epsilon$ is rejected, yielding the largest average shift of each index. The forecast then uses a baseline year's indices as a fixed distribution: the GDP-regressed yearly total is deseasonalized to an average hour, and that average is multiplied by the stored $\\mathrm{SI}^1$, $\\mathrm{SI}^2$, and $\\mathrm{SI}^3$ values to rebuild the hourly curve.","core_discovery":"The central discovery the paper argues for is that hourly load shapes are stable enough over multi-year horizons to serve directly as a forecasting device. Multiplicatively decomposing Singapore's half-hourly demand into monthly, day-of-week, and hour-of-day seasonality indices, the paper finds that the largest pairwise year-to-year differences in the hourly indices average just 4.24 percent of the mean index (4.65 percent on weekdays, 3.21 percent on weekends), as verified through paired t-tests. Because annual electricity consumption correlates with GDP (Pearson coefficient 0.9947 for Singapore), a simple linear regression on GDP supplies the yearly total, and the stable indices from the baseline year convert that total into an hourly curve. The reported errors, a maximum 6.87 percent MAPE across six forecast years for Singapore and 6.81 and 5.64 percent for Belgium and Bulgaria, are the paper's evidence that the stability-plus-GDP recipe carries the entire forecasting burden.","pith_inferences":["If hourly shape stability is the real engine, accuracy should degrade precisely where the shape changes, through large rooftop-solar fleets, EV charging ramps, or work-from-home shifts, even when annual GDP correlation holds; a direct test would recompute the seasonality indices for 2023-2025 and compare them with the 2004-2022 baseline.","The paper's worst errors occur in the COVID-recovery years, which suggests the method inherits whatever error the GDP regression makes on the yearly total, so macroeconomic shocks dominate the error budget rather than the stability assumption.","The paper itself flags that GDP decoupling in OECD economies and rapid electrification in non-OECD economies can weaken the predictor, so the sensible deployment is as a baseline curve corrected for known structural changes rather than a blind extrapolation.","The epsilon-search t-test effectively measures the largest average pairwise drift of each index; a complementary check would compare the seasonality-index distribution from a training window directly against a hold-out window, which the baseline-year design already makes possible."],"forward_implications":["A grid planner can produce a defensible hourly load curve for a five-year horizon using only historical load data and GDP, without modeling electrification drivers such as EV or heat-pump uptake.","Hourly error stays below 7 percent even when the yearly-total forecast misses by up to 6.12 percent, so in these cases the seasonality-index distribution is not the dominant error source.","The stability diagnosis transfers across economies: an OECD country (Belgium, 8.41 percent SI deviation) and a non-OECD country (Bulgaria, 13.24 percent) both yield two-year-ahead hourly MAPEs under 7 percent.","Constancy of the seasonality indices lets a Holt-Winters exponential smoothing model fix its seasonal components and optimize only two smoothing constants, matching or beating machine-learning benchmarks in most years.","Any improved yearly-total forecast can be plugged into the same stability machinery to obtain an hourly curve, so the seasonality-index framework is separable from the GDP regression."],"supporting_citations":[{"why":"Supplies the half-hourly system-wide demand data for Singapore from 2004 to 2022, the core test case for the stability analysis and forecasts.","marker":"EMA, 2023"},{"why":"Provides the GDP series used as the predictor in the linear regression for Singapore's annual demand forecasts.","marker":"IMF, 2023"},{"why":"Supplies GDP and electricity consumption data for Belgium and Bulgaria, enabling the cross-country validation.","marker":"OWID, 2023"},{"why":"Provides the historical hourly load data for Belgium and Bulgaria used to compute their seasonality indices and forecasts.","marker":"ENTSOE, 2023"},{"why":"The four-fold seasonality Holt-Winters model that the short-term forecasting extension adopts and simplifies with constant seasonality indices.","marker":"Huang et al., 2017"},{"why":"Establishes the GDP-electricity demand relationship that justifies GDP as a long-term predictor, while also noting OECD decoupling.","marker":"Steinbuks et al., 2017"},{"why":"Empirical support for a causal link between economic growth and electricity consumption in Singapore, grounding the GDP predictor for the main case study.","marker":"Sharif et al., 2017"},{"why":"Evidence that simple linear regression is effective for long-term load forecasting, motivating the chosen annual forecast model.","marker":"Hammad et al., 2020"},{"why":"Quantitative panel evidence that a 1 percent GDP increase raises electricity demand by about 1.5 percent, supporting the elasticity assumed by the GDP regression.","marker":"Bazán Navarro et al., 2024"}],"fun_headline_variants":["Stable hourly load shapes make long-term forecasts simple","GDP plus steady hourly curves forecasts power within 7%","Hourly demand stability enables easy multi-year forecasts","Simple method: GDP totals, stable hourly splits for years","Stable daily patterns let GDP forecast hourly power for years"],"cache_read_input_tokens":27136,"weakest_assumption_plain":"The forecast inherits the shape of the daily demand curve from a baseline year, so if the relative hourly pattern shifts materially over the four-to-five-year horizon, through new technology, changed working habits, or a different generation mix, the hourly forecast remains wrong even when the annual total is right.","fun_headline_variants_meta":{"raw":{"variants":["Stable hourly load shapes make long-term forecasts simple","GDP plus steady hourly curves forecasts power within 7%","Hourly demand stability enables easy multi-year forecasts","Simple method: GDP totals, stable hourly splits for years","Stable daily patterns let GDP forecast hourly power for years"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1404,"prompt_tokens":1011,"completion_tokens":393,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":627,"tokens_out":393,"duration_ms":4979,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:43:02.415593+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the hourly seasonality indices on data from a later window, for example 2023-2025 for Singapore, and compare them with the 2004-2022 means: if the average absolute deviation exceeds the measured 4.24 percent, or the per-hour values exceed the epsilons in the paper's Table 2, the stability premise fails and hourly errors should exceed the reported MAPE bound. A weaker but still decisive test is to run the full procedure on a country with rapid rooftop-solar or EV growth and check whether hourly MAPE over a four-year horizon grows while annual MAPE stays small.","supporting_citations":[{"cited_title":"Half-hourly System Demand Data","cited_arxiv_id":null,"evidence_quote":"Supplies the half-hourly system-wide demand data for Singapore from 2004 to 2022, the core test case for the stability analysis and forecasts."},{"cited_title":"World Economic Outlook Database","cited_arxiv_id":null,"evidence_quote":"Provides the GDP series used as the predictor in the linear regression for Singapore's annual demand forecasts."},{"cited_title":"correlation","cited_arxiv_id":null,"evidence_quote":"Supplies GDP and electricity consumption data for Belgium and Bulgaria, enabling the cross-country validation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the historical hourly load data for Belgium and Bulgaria used to compute their seasonality indices and forecasts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The four-fold seasonality Holt-Winters model that the short-term forecasting extension adopts and simplifies with constant seasonality indices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the GDP-electricity demand relationship that justifies GDP as a long-term predictor, while also noting OECD decoupling."},{"cited_title":"A., and Shahzad, S","cited_arxiv_id":null,"evidence_quote":"Empirical support for a causal link between economic growth and electricity consumption in Singapore, grounding the GDP predictor for the main case study."},{"cited_title":"A., Jereb, B., Rosi, B., and Dragan, D","cited_arxiv_id":null,"evidence_quote":"Evidence that simple linear regression is effective for long-term load forecasting, motivating the chosen annual forecast model."},{"cited_title":"E., Morocho Ruiz , J","cited_arxiv_id":null,"evidence_quote":"Quantitative panel evidence that a 1 percent GDP increase raises electricity demand by about 1.5 percent, supporting the elasticity assumed by the GDP regression."}],"review_version":1}