{"id":"745c5e9f-379d-4eeb-b0a7-dedcf1d6a482","arxiv_id":"2503.04757","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LSTM-type models forecast day-ahead residential demand better than standard profiles today and in simulated 2037 grid states, but all tested methods degrade as rooftop solar and batteries spread.","lead":"The authors simulated how Germany's power grid might look in 2037 by adding rooftop solar and home batteries to real smart meter data from 3,511 households, then tested whether machine learning can forecast daily electricity demand. They found ML models beat simple standard load profiles, but every forecasting method gets worse as solar penetration rises.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 68.5% gain is a current-state comparison against SLP; in future states ML beats the day-before benchmark by only ~4–7% RMSE and loses on MAPE, so 'especially in future grid states' is not supported by Table II.","rationale":"The reader's weakest_assumption focuses on the external validity of the digital twin. That is a legitimate concern: all future-state data come from a self-cited, not independently reproduced simulation, so the ranking in S1/S2 may not transfer to real 2037 conditions. However, the more load-bearing issue for the paper's central claim as written is internal: the abstract says ML outperforms benchmarks \"especially in future grid states,\" but Table II shows the opposite pattern. The 68.5% improvement is a current-state comparison against SLP, which is absent in future states; in S1/S2, ML beats the day-before benchmark by only single-digit percentages on RMSE and loses on MAPE. This means the headline claim is not supported even if the digital twin is perfectly realistic. The reader's rationale does note the abstract overclaims the 68.5% gain, but their weakest_assumption singles out the digital twin; I see the metric/claim mismatch as more central to the paper's stated contribution. Since both issues are fixable with reporting changes and additional validation, the verdict should remain CONDITIONAL rather than moving to ACCEPT or REJECT. The paper's core empirical pattern (ML improves over SLP in the current state, all methods degrade in high-PV scenarios) is plausible and worth reporting, but the abstract and conclusion need correction before the central claim can be taken at face value.","tokens_in":9324,"tokens_out":4159,"duration_ms":36729,"concrete_test":"Recompute the relative improvements from Table II: for each scenario, calculate (RMSE_benchmark - RMSE_ML)/RMSE_benchmark for BM1, BM2, and ARIMA, and compare MAPE rankings across ML and benchmarks. If the S1/S2 RMSE improvements remain below ~10% and MAPE favors BM1, then the abstract's \"especially in future grid states\" should be revised to state that the largest gain is in the current state against SLP, with modest RMSE gains and mixed MAPE results in future states. Also verify whether any SLP result is reported for S1/S2; if not, the claim that ML outperforms SLP in future states cannot be tested from the presented data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on Table II showing that ML approaches \"outperform SLPs as well as simple benchmark estimators ... especially in future grid states.\" The table contradicts the qualifier. The 68.5% reduction is the CS CNN-LSTM vs. SLP comparison (182.4 vs. 579.9), and SLP is not evaluated in S1/S2. In future states, the RMSE advantage of the best ML model over the day-before benchmark BM1 is about 5.7% in S1 (352.8 vs. 374.3 for LSTM) and 7.4% in S2 (255.5 vs. 276.0 for CNN-LSTM). By MAPE, BM1 beats both ML models in S1 (23.2% vs. 26.2%/27.0%) and in S2 (10.6% vs. 11.6%/12.1%). Thus the empirical basis for \"especially in future grid states\" is absent; the strongest ML-vs-benchmark margin occurs in the current state and is against a synthetic SLP. The second part of the claim (all methods perform worse in future states) is supported by the RMSE increases, but the first part is overstated. Since this overstatement appears in the abstract and conclusion, it is load-bearing for the paper's headline message, not merely a stylistic caveat. Fixing it does not require new experiments, only accurate reporting of the paper's own numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies day-ahead forecasting of aggregate residential electricity demand in a current grid state (CS) and two simulated future grid states (S1, S2) of a German town. The future states are generated by a digital twin that adds rooftop PV and battery storage to individual households according to regionalized national expansion targets for 2037. The authors compare LSTM and CNN-LSTM models against practical and naive benchmarks (SLP, day-before, day-one-week-ago, ARIMA) using RMSE and MAPE on an out-of-sample test year. The main reported findings are that LSTM-based models beat SLPs and simple benchmarks with up to 68.5% lower RMSE, and that all methods perform worse in future states with higher PV penetration.","tokens_in":9626,"tokens_out":2922,"duration_ms":25921,"significance":"The study combines a realistic digital twin of a local energy system with deep-learning load forecasting, which is a genuinely useful direction for utility practice. Strengths include the use of real smart meter data from 3,511 households, a clear out-of-sample evaluation scheme, and a relevant set of simple benchmark predictors. If the quantitative claims were accurately reported, the paper would provide a timely warning that high PV penetration degrades the accuracy of current forecasting methods. However, the headline claim in the abstract and conclusion is not supported by the paper's own Table II: the 68.5% improvement is a current-state CNN-LSTM versus SLP comparison, while in future states the ML advantage over the day-before benchmark is small (around 6-7% RMSE) and the benchmark wins on MAPE. The credibility of the future-state conclusions also depends on the validation status of the digital twin, which is cited to prior work but not described or reproduced here.","major_comments":[{"comment":"The claim that LSTM approaches outperform SLPs and simple benchmarks 'especially in future grid states' is contradicted by Table II. The 68.5% RMSE reduction is the current-state comparison between CNN-LSTM (182.4) and SLP (579.9); SLP is not evaluated in S1 or S2. In future states, the best ML RMSE advantage over the one-day-ago benchmark BM1 is only about 5.7% in S1 (LSTM 352.8 vs. 374.3) and 7.4% in S2 (CNN-LSTM 255.5 vs. 276.0), and BM1 has lower MAPE than both ML models in both future states (23.2% vs. 26.2%/27.0% in S1; 10.6% vs. 11.6%/12.1% in S2). The abstract and Section VI should be reworded to report the numbers accurately rather than stating a qualitative advantage in future grid states.","section":"Abstract and Section VI"},{"comment":"The digital twin is described as 'already existing and validated' with a citation to the authors' prior work [10], but the current manuscript provides no information about what that validation consisted of, what error metrics were used, or how well the twin reproduces observed PV generation, battery behavior, and grid demand. Because S1 and S2 are entirely generated by this twin, the transferability of the future-state results to real 2037 conditions rests on this validation. The paper should either summarize the validation evidence or explicitly frame the future-state results as scenario-illustrative rather than predictive, with the associated caveat in the abstract and conclusion.","section":"Section III-A"},{"comment":"The performance comparisons are presented without uncertainty estimates. The simulation is said to be 'executed multiple times' to ensure representativeness, but no details are given on how many runs were performed, whether the reported RMSE/MAPE values are averages across runs, or what the run-to-run variability is. The only statistical test reported is a single t-test for one comparison in S2, and that test is not adjusted for the serial correlation of hourly forecast errors. Without confidence intervals or variance information, the claim that 'the CNN-LSTM approach outperforms all benchmark estimators across all scenarios' is stronger than the evidence supports. Reporting standard deviations or bootstrap intervals over the simulation runs is needed.","section":"Section IV, Table II and Section III-D"}],"minor_comments":[{"comment":"The text states that the standard deviation increases by a factor of 3.8 for S1 and 6.0 for S2, even though S1 has a higher PV expansion than S2. This ordering is counterintuitive and should be explained or checked, since it may reflect differences in battery control or scenario construction rather than PV penetration alone.","section":"Section III-D"},{"comment":"The comparison with [33] ('the LSTM has a MAPE of around 8.4%, which is slightly higher than the MAPE of our results') is not a meaningful state-of-the-art comparison because the dataset, aggregation level, forecasting horizon, and evaluation period differ. It should be removed or clearly qualified.","section":"Section IV"},{"comment":"There is a typo: 'training theLSTM requires 200 epochs' should read 'training the LSTM requires 200 epochs'. Also, the runtime comparison would be more informative if the reported seconds included the same hardware and software environment for both models.","section":"Section IV, footnote 2"},{"comment":"The caption says 'Left: Violinplot of one-year electric demand on grid level caused by residential buildings' and 'Right: Same plot for electric load', but the distinction between 'demand' and 'load' is not clearly defined in the caption. It would help to state explicitly that the left panel shows the sum of residential demand (excluding feed-in) and the right panel shows the net grid load including feed-in.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's core setup is promising and the data are real and relevant, but the abstract and conclusion overstate the evidence in a way that affects the headline contribution. The missing validation details of the digital twin and the absence of uncertainty estimates are also load-bearing for the future-state claims. All of these can be addressed within the manuscript's scope by rephrasing the claims, adding twin validation facts or explicit limitations, and reporting variability across simulation runs. I do not see a fatal flaw, but the revision needs to be substantive rather than purely cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good to know about this one. The headline number is misleading: the 68.5% RMSE gain is CNN-LSTM over SLP in the current state, not in future grid states. Against the day-before benchmark in S1 and S2, the best ML model wins by only 5–7% RMSE and loses on MAPE. So the abstract's 'especially in future grid states' is not backed by their own Table II. The paper's real finding is different and still useful: ML beats the standard SLP profile today by a wide margin, and all methods—ML included—get worse as PV penetration rises in the simulated 2037 scenarios.\n\nWhat's actually new: combining LSTM-style forecasts with digital-twin-generated future grid data. The related work review seems fair; I don't know of similar studies that simulate future grid states at community scale and test forecasting on the simulated demand. The dataset is substantial: 3,511 real households, 34 months of hourly smart meter data, and a clean out-of-sample split (664 training days, 365 test days). The benchmarks are simple and appropriate: SLP, one-day-ago, one-week-ago, ARIMA. That part is defensible.\n\nSoft spots, in order of importance. First, the abstract and conclusion overstate the future-state advantage. That's not a stylistic nit; it changes the practical message. The paper should either reframe as 'ML beats SLP today; future-state advantages are small and may disappear on MAPE' or add experiments that actually show a bigger future-state gap. Second, there are no confidence intervals or simulation-run variability. They say the simulation is run multiple times, but Table II reports a single number. Third, the digital twin that generates all future-state data is cited to the authors' own prior work and not validated here. That doesn't make the paper circular—the forecasting evaluation is on out-of-sample data with external benchmarks—but it does mean the central input is a black box. A summary of the twin's validation would go a long way.\n\nBottom line: worth a serious referee. It's a competent, honest study with a fixable reporting flaw. If the authors fix the abstract and add variance/validation detail, it becomes a solid conference paper. I'd bring it to a reading group as a case study in how simulation-based evaluations can be misreported, but I wouldn't cite the headline number.","headline":"Useful simulation study with a load-bearing overclaim: the 68.5% gain is current-state vs SLP, not a future-state win.","tokens_in":10168,"tokens_out":3408,"would_cite":true,"duration_ms":27278,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that LSTM-based neural networks forecast day-ahead residential electricity demand more accurately than the synthesized load profiles and simple benchmarks that small utilities rely on, but that every tested method —…","keywords":["digital twin","electricity demand forecasting","LSTM","smart meter data","future grid states","photovoltaic expansion","residential load forecasting","neural networks"],"falsifier":"On real hourly data from a distribution grid that already has PV and battery penetration comparable to S1, run the same LSTM and CNN-LSTM day-ahead forecasts; if their RMSE is not higher than on low-PV current-state data (or if the ML models no longer beat the day-before benchmark), the paper's degradation claim is falsified.","tokens_in":9102,"feed_emoji":"⚡","tokens_out":10142,"duration_ms":81800,"temperature":0.7,"pith_summary":"Small and medium-sized utilities often forecast residential electricity demand with standardized load profiles that treat all households alike and ignore rooftop solar, batteries, heat pumps, and electric vehicles. This paper asks whether machine-learning forecasters, specifically LSTM-based neural networks, do better in today's grid and in simulated 2037 grids with much more distributed generation and storage. Using real hourly smart meter data from 3,511 households and a digital twin of a local energy system, the paper finds that LSTM models cut day-ahead RMSE by up to 68.5% compared to the standardized profiles in the current grid state, and that they also beat naive benchmarks and ARIMA in the future scenarios. The paper's second finding is that every method's error grows in the high-PV future states, with RMSE rising from the 182–247 range in the current state to the 255–435 range in the simulated futures. In short, the paper makes the case for adopting ML forecasting now while cautioning that current models are not yet ready for high-renewable grid states.","feed_headline":"ML forecasts beat old load profiles today, but high PV breaks them too","feed_subtitle":"A digital-twin experiment shows why utilities should switch to ML — and why current models still need work.","key_machinery":"The central object is the digital twin of a local energy system: a building-by-building simulation of a town's electricity system built from real smart meter data, geospatial roof data, and a rule-based battery controller that maximizes local PV self-consumption. Two scenarios for 2037, S1 and S2, are derived by regionalizing Germany's official PV and battery expansion targets (S1) or linearly extrapolating current growth (S2), and the twin generates hourly demand time series for each. These simulated series become the training and test data for a vanilla LSTM and a CNN-LSTM encoder-decoder, with SLP, two naive persistence benchmarks, and ARIMA as comparators. The forecasting experiment is a day-ahead, 24-step, univariate prediction of the sum of residential demand, and the paper scores it by RMSE and MAPE.","core_discovery":"The central claim is that LSTM-based approaches are better suited than today's common practice for day-ahead residential demand forecasting in both current and future grid states, yet even they suffer as PV penetration rises. In the current state, the best model (CNN-LSTM) reaches an RMSE of 182.4 kW compared with 579.9 kW for the synthesized load profile, a 68.5% reduction; the ML models also beat the naive “day before” and “day one week ago” benchmarks and ARIMA in the simulated 2037 scenarios. At the same time, the LSTM's RMSE climbs from 190.4 kW in the current state to 352.8 kW in the high-PV scenario S1 and 259.3 kW in the moderate scenario S2, and all other methods degrade similarly. The paper attributes this degradation to the unpredictability of cloud cover on summer days, which makes the residual demand harder to forecast as the share of rooftop PV grows. It concludes that utilities should adopt ML-based forecasting, but that current methods still need to be adapted for future grid states.","pith_inferences":["A testable extension the authors do not run: adding day-ahead solar-irradiance forecasts as exogenous inputs should lower RMSE more in S1 and S2 than in the current state, because the paper's own explanation is that cloudy-summer PV output drives the extra error.","A validation gap: models are trained and tested on the same digital twin's simulated data, so the absolute error levels in 2037 scenarios are untested against reality; comparing against a real high-PV grid would show whether the degradation is as steep.","A reading of the table: the 68.5% improvement is against SLP in the current state; in S1 and S2 the ML models' RMSE is only roughly 6–25% lower than the naive benchmarks, so the practical gain in future states is smaller than the headline suggests."],"forward_implications":["Utilities that still rely on synthesized load profiles should expect their day-ahead forecasts to lose accuracy as rooftop PV and home batteries spread; the measured LSTM advantage argues for a switch to ML-based forecasting.","Forecast error is likely to rise with PV penetration regardless of method, so error estimates from today's models are optimistic for future high-renewable grid states.","Digital-twin simulation provides a way to test forecasting models against plausible future grid states before those states arrive, using current smart meter data and expansion targets as inputs.","MAPE is an unreliable metric in future grid states because residential demand frequently approaches zero; RMSE gives a more stable comparison."],"supporting_citations":[{"why":"This reference supplies the digital twin of the local energy system that generates the simulated future demand profiles used to train and test all forecasters.","marker":"[10]"},{"why":"This reference provides Germany's national 2037 PV and battery expansion targets, which the paper regionalizes to build scenario S1.","marker":"[27]"},{"why":"This reference gives the CNN-LSTM encoder-decoder architecture that yields the best RMSE in the current state and in scenario S2.","marker":"[18]"},{"why":"This reference justifies the battery size (7 kW / 10.5 kWh) and the self-consumption control rule used in the digital-twin simulation.","marker":"[26]"},{"why":"This reference supplies the review of low-voltage load forecasting methods and the standard choice of RMSE and MAPE metrics, as well as LSTM as a state-of-the-art architecture.","marker":"[5]"},{"why":"This reference provides the vanilla LSTM setup that the paper adapts as one of its two ML forecasters.","marker":"[28]"},{"why":"This reference offers the published LSTM MAPE benchmark against which the paper compares its own error levels.","marker":"[33]"}],"fun_headline_variants":["LSTM forecasts beat old load profiles, but PV still stumps them","ML bests legacy demand forecasts; high PV degrades even ML","Digital twin shows ML wins today, but PV breaks forecasts","Even best ML forecast worsens as rooftop PV rises","ML beats old load profiles, yet future grids challenge even ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the digital twin's simulated future demand data faithfully represents how real households with rooftop PV and batteries will draw power from the grid; the paper cites the twin as validated but does not reproduce that validation.","fun_headline_variants_meta":{"raw":{"variants":["LSTM forecasts beat old load profiles, but PV still stumps them","ML bests legacy demand forecasts; high PV degrades even ML","Digital twin shows ML wins today, but PV breaks forecasts","Even best ML forecast worsens as rooftop PV rises","ML beats old load profiles, yet future grids challenge even ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000448,"raw_usage":{"total_tokens":2294,"prompt_tokens":1011,"completion_tokens":1283,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":1207}},"tokens_in":627,"tokens_out":1283,"duration_ms":8378,"temperature":1.0,"reasoning_tokens":1207,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:41:32.936236+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On real hourly data from a distribution grid that already has PV and battery penetration comparable to S1, run the same LSTM and CNN-LSTM day-ahead forecasts; if their RMSE is not higher than on low-PV current-state data (or if the ML models no longer beat the day-before benchmark), the paper's degradation claim is falsified.","supporting_citations":[{"cited_title":"A digital twin of a local energy system based on real smart meter data,","cited_arxiv_id":null,"evidence_quote":"This reference supplies the digital twin of the local energy system that generates the simulated future demand profiles used to train and test all forecasters."},{"cited_title":"Approval of the scenario framework 2023-2037/2045,","cited_arxiv_id":null,"evidence_quote":"This reference provides Germany's national 2037 PV and battery expansion targets, which the paper regionalizes to build scenario S1."},{"cited_title":"Predicting residential energy consumption using CNN-LSTM neural networks,","cited_arxiv_id":null,"evidence_quote":"This reference gives the CNN-LSTM encoder-decoder architecture that yields the best RMSE in the current state and in scenario S2."},{"cited_title":"The development of stationary battery storage systems in germany – a market review,","cited_arxiv_id":null,"evidence_quote":"This reference justifies the battery size (7 kW / 10.5 kWh) and the self-consumption control rule used in the digital-twin simulation."},{"cited_title":"Review of low voltage load forecasting: Methods, ap- plications, and recommendations,","cited_arxiv_id":null,"evidence_quote":"This reference supplies the review of low-voltage load forecasting methods and the standard choice of RMSE and MAPE metrics, as well as LSTM as a state-of-the-art architecture."},{"cited_title":"Multi-step time series forecasting of electric load using machine learning models,","cited_arxiv_id":null,"evidence_quote":"This reference provides the vanilla LSTM setup that the paper adapts as one of its two ML forecasters."},{"cited_title":"Short-term residential load forecasting based on lstm recurrent neural network,","cited_arxiv_id":null,"evidence_quote":"This reference offers the published LSTM MAPE benchmark against which the paper compares its own error levels."}],"review_version":1}