Pith. sign in

REVIEW 4 major objections 5 minor 17 references

A generative AI model trained only on pre-market data can forecast the Irish grid's largest single infeed/outfeed up to 38 hours ahead with accuracy nearly matching the operational 8-hour rolling forecast.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:58 UTC pith:XDYSEY7G

load-bearing objection A genuinely new operational target for day-ahead forecasting, with a solid real benchmark comparison, but the feature-availability handling and the 15% cost-saving figure need much more transparency before the headline accuracy claim is bulletproof. the 4 major comments →

arxiv 2607.15900 v1 pith:XDYSEY7G submitted 2026-07-17 eess.SY cs.SY

Day-Ahead Forecasting of Largest Single Infeed/Outfeed on the Irish Power Grid: A Generative Artificial Intelligence Approach

classification eess.SY cs.SY
keywords largest single infeedlargest single outfeedreserve dimensioningday-ahead forecastinggenerative AItransformer time seriesIrish power systemprobabilistic forecast
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper is trying to establish that a generative-AI forecasting model, fed only information available before 9am on the day before operation, can predict the largest single infeed (LSI) and largest single outfeed (LSO) of the Irish power system almost as accurately as the operational 8-hour rolling forecast that uses full market data. The reported result is a mean absolute percentage error only 1.1 percentage points worse than that benchmark, despite a longer lead time and far less information. The motivation is reserve dimensioning: reserves are sized against the largest single contingency, so better day-ahead forecasts of LSI/LSO could let system operators procure reserve capacity more efficiently. The paper estimates that using the model's P95 predictions instead of a flat-line reserve could cut reserve procurement costs by around 15%. A sympathetic reader would care because this is a real-world grid application where forecast lead time, not just accuracy, determines operational value.

Core claim

The central claim is that a Transformer-style generative model, trained on pre-market gate closure data, produces day-ahead LSI and LSO forecasts whose MAPE is only 1.1% higher than the current operational 8-hour unit-commitment benchmark, while forecasting up to 38 hours ahead. On the test set, the proposed model achieves an LSI MAPE of 9.8% and MAE of 44.5 MW, compared with 19.9% MAPE for a schedule-based day-ahead baseline and 8.7% MAPE for the rolling 8-hour model. The authors argue this closes most of the accuracy gap that previously forced reserve dimensioning to rely on near-real-time information, and they estimate that early, accurate forecasts of interconnector flows, wind generatio

What carries the argument

The load-bearing machinery is the reference-incident identity RI = LSI (or LSO) + consequential losses, which ties a single forecast quantity directly to reserve volume requirements, combined with a Transformer-style sequence model that produces 48 half-hourly forecasts of large-generator outputs and interconnector flows along with percentile bands. The Transformer encodes each signal as a temporal sequence and jointly models spatial dependencies across multiple inputs, allowing it to capture relationships across time horizons and support what-if analyses of wind forecasts or market price differentials. The attention mechanism is what lets the model identify interconnector flows, wind genera

Load-bearing premise

The result stands only if the day-ahead model genuinely had no access to post-gate-closure information at inference; the paper states this but never explains how the outturn-only training features (bidding/offer behaviour, interconnector flows, top-10 generator output) are masked or imputed, and the evaluation window is short enough—five weeks by the authors' own description—that the 1.1-point gap may not generalize.

What would settle it

A reader could settle the claim by re-running the test window with all outturn-only features replaced by their forecast-only counterparts at inference and measuring the LSI MAPE gap against the 8-hour benchmark; if the gap jumps materially, the model relied on future information. A complementary check is to extend the evaluation beyond the short window the authors flag, including cold spells, interconnector outages, and high-renewable days, and see whether the 1.1-point gap holds out of sample.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Reserve dimensioning can move from near-real-time to a day-ahead basis, giving system operators a full day of extra planning time before the operating day.
  • Using the model's P95 percentile predictions rather than a flat-line reserve could reduce reserve procurement costs by about 15%, if the sensitivity estimate holds in actual auction conditions.
  • The day-ahead LSI MAPE drops from 19.9% with a schedule-based baseline to 9.8%, roughly halving the error of the previous day-ahead approach.
  • The approach is extensible to other TSO planning tasks beyond reserve sizing, such as operational planning and real-time management, as power systems become more dynamic.
  • The five-week development cycle suggests that AI-based forecasting can be deployed quickly when the platform allows iterative collaboration with system experts.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the accuracy gap survives a clean test that removes all outturn-only features at inference, the practical consequence is that day-ahead reserve auctions could size reserves from probabilistic forecast bands rather than conservative flat-line assumptions, and the cost saving could be measured directly once the day-ahead auction arrangement goes live.
  • The approach is likely to transfer to other low-inertia island power systems where a single HVDC interconnector or large unit is a significant share of demand, though that is an extrapolation beyond the paper's evidence.
  • The 15% cost-saving figure is only a sensitivity estimate; actual savings depend on how P95 prediction errors correlate across trading periods and on auction price formation, neither of which the paper models.
  • A testable extension would be to compare the Transformer against a simple parametric benchmark using just residual load, wind, and interconnector price differentials; if that benchmark approaches 9.8% MAPE, the advantage of the deep learning architecture is smaller than implied.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a Transformer-based generative AI platform (GridZero.ai) to forecast, at day-ahead stage from 9am D-1, the largest single infeed (LSI) and largest single outfeed (LSO) on the Irish power system, for use in reserve dimensioning under the future DASSA auction. The model is trained on data available before market gate closure and is compared with a schedule-based D-1 baseline and a rolling 8-hour unit-commitment benchmark. Reported LSI MAPE is 9.8% versus 8.7% for the rolling benchmark, which the paper summarizes as a 1.1 percentage-point gap, and it also claims about 15% potential reserve-cost savings from using P95 predictions. The paper is an application case study rather than a methodological advance, with a short test window and no code/data release.

Significance. If the central near-parity claim holds, the result is practically significant: a day-ahead LSI forecast from pre-market data would enable earlier and potentially cheaper reserve dimensioning than the current operational 8-hour benchmark. The paper has the strength of being tested against an actual operational rolling unit-commitment benchmark over the same period, and it includes a feature-contribution analysis and explicit temporal train/validation/test splits. However, the significance is currently limited by three load-bearing gaps: (i) the training features include outturn information that is not available at the 9am D-1 inference time, with no masking protocol specified; (ii) the LSO results are substantially worse than LSI but are not reported with MAPE; and (iii) the claimed 15% reserve-cost saving is asserted without a derivation or an independent cost model. The paper is plausible and within the journal's scope, but the evidence as presented is not yet sufficient to support the headline economic and accuracy claims.

major comments (4)
  1. [Section III-A and III-B] Train/serve feature mismatch. The Inputs list in Section III-A includes outturn features: generator bidding and offer behavior, interconnector exchanges, and output of the ten largest generators. These are not observable at 9am D-1. The text only says 'the model accounts for the fact that only a subset of inputs is available at prediction time' without specifying masking, imputation, or architectural gating. Section III-B then claims the model is 'trained on pre-market gate closure system and forecast data available before 9 am D-1' with 'strict temporal causality.' These statements are inconsistent unless the training pipeline explicitly prevents conditioning on outturn features. If those features are present in training and absent/zero-filled at inference, the reported 9.8% LSI MAPE could reflect shortcuts that disappear in deployment. This is load-bearing for the central near-parity c
  2. [Section III-B and Table II] The headline comparison is LSI-only. Table II reports LSO MAE for the GridZero.ai model as 12.7 MW versus 6.1 MW for the rolling 8-hour model — more than double — and no LSO MAPE is given. The abstract and Section III-B state that accuracy is 'only 1.1% higher' without noting that this refers to LSI MAPE only. Given that LSO is a key reserve dimensioning input (positive/negative reference incident), the LSO performance must be reported with the same metrics, and the near-parity claim should be qualified to LSI. In addition, the test window is described as 'five-week' in the Conclusions, but Table II says June 6–August 14, 2024, which is about ten weeks; this discrepancy should be corrected and the sensitivity of the MAPE gap to the chosen window discussed.
  3. [Section I-C and Section III-B] The 15% reserve-cost saving estimate is unsupported. Section III-B states that 'sensitivity tests suggest' this figure when using P95 predictions versus 'a standard flat-line reserve procurement,' but no cost model, reserve procurement formulation, sensitivity analysis, or derivation is provided. This estimate appears in the Contributions list (Section I-C) and in the abstract, so it is a substantive claim, not a side remark. The authors should either provide a transparent calculation with assumptions, or remove the quantitative cost-saving claim from the abstract and contributions until properly supported.
  4. [Section III-B] No uncertainty quantification on the reported errors. The comparison of MAPE 9.8% vs 8.7% is presented as a point estimate over a single test period, with no confidence intervals, significance test, or day-by-day distribution. Given that the test window is short and the benchmark is operational, the authors should report the variability of the MAPE/MAE difference (e.g., by trading period or day) to show that the 1.1 percentage-point gap is not driven by a few anomalous days.
minor comments (5)
  1. [Section II-B] The term 'generative AI' is never defined. The model is a Transformer for multi-step time-series prediction; it is not shown to be generative in the sense of sampling from a learned distribution. Clarify what 'generative' means in this context.
  2. [Section III-A] Figure 3 reports Pearson correlations without numerical values, error bars, or significance levels. The feature-contribution discussion is qualitative; include the coefficients or a table for reproducibility.
  3. [Section IV] The Conclusions acknowledge a 'five-week test window' but the case study dates in Table II span approximately ten weeks. Please reconcile the dates and the stated evaluation length.
  4. [Section II-A] The phrase 'up to 38 hours ahead' is clear from the 9am D-1 prediction to 11pm D-day, but it would help to restate the forecast horizon explicitly in the abstract or in the methodology to avoid confusion with the 24-hour output window.
  5. [General] The paper does not mention whether code, anonymized data, or model weights will be made available. Given the train/serve mismatch concern above, even a minimal data-availability statement would improve reproducibility and trust.

Circularity Check

0 steps flagged

No significant circularity: the headline accuracy comparison is against an external operational benchmark; residual train/serve feature questions are leakage concerns, not definitional circularity.

full rationale

The paper's central accuracy claim (GridZero.ai D-1 9am LSI MAPE 9.8% vs rolling 8-hour benchmark 8.7%, Table II) is a comparison against an external security-constrained unit-commitment model, not a quantity derived from the model's own fitted values. LSI/LSO are defined in Eqs. (1)-(2) via the reference incident and are computed from predicted generator outputs/interconnector flows; no equation is fitted to the target and then reported as a prediction. The 15% reserve-cost reduction is a sensitivity estimate from the model's P95 outputs compared with flat-line procurement; it is self-referential in provenance but not a fitted-input-equals-prediction reduction. The self-citations [1]-[3], [16] supply system context and DASSA background; none is load-bearing for the forecast result. The unresolved issue that 'outturn features' (including output of the ten largest generators) are listed as inputs while the model is claimed to use only pre-9am D-1 data with 'strict temporal causality' (Section III) is a potential information-leakage/correctness problem, not circularity, because the paper nowhere defines the prediction as equal to a training input by construction. The model's masking mechanism is also underspecified in the sentence 'the model accounts for the fact that only a subset of inputs is available at prediction time.' The paper itself acknowledges the short test window in Section IV ('This evaluation is based on a five-week test window...') and defers additional baselines and out-of-distribution analyses to future work, further confirming the claims are empirical and not definitional.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

No new physical or mathematical entities are introduced; the GridZero.ai platform is a commercial system, not a postulated entity. The central claim rests on standard Transformer assumptions, the operational reference-incident definition, and an unverified leakage-free train/serve split.

free parameters (3)
  • Transformer hyperparameters (architecture and training choices) = not disclosed
    The central model is described only as a 'Transformer-style, sequence-based multi-input architecture'; no layer counts, heads, learning rates, or regularization are given, so performance may depend on undocumented tuning.
  • P95 prediction interval used for reserve sizing = P95 (95th percentile) from model output
    The 15% cost-saving claim uses P95 predictions as the reserve requirement; the calibration of these intervals is not validated against actual exceedances.
  • Reserve cost reduction estimate = ~15%
    Appears from 'sensitivity tests' in Section III-B but no cost model, flat-line baseline, or sensitivity inputs are shown; the number is effectively a hand-stated result.
axioms (5)
  • domain assumption Reference incident is LSI/LSO plus a fixed consequential-loss term, and reserve volumes scale with it.
    Section II-A equations (1)-(2) and Table I; this is the operational premise that makes the forecast target important.
  • domain assumption A Transformer-style sequence model can represent the spatiotemporal generation/interconnector dynamics from the chosen features.
    Section II-B; no comparison with non-transformer ML baselines is provided, so the modeling assumption is untested against simpler alternatives.
  • domain assumption Temporal 80/10/10 split and 'strict temporal causality' guarantee no information leakage, and missing inputs at inference are handled correctly.
    Section III-A/B; the model is trained with outturn features not available at 9am D-1; no masking or leakage checks are described.
  • ad hoc to paper The rolling 8-hour model is an upper bound on achievable forecast accuracy.
    Section III-B states this benchmark 'represents an upper bound on forecast accuracy,' but it is an operational solver, not a theoretical ceiling; treating it as upper bound is an assumption.
  • domain assumption Consequential loss factor is set as a fixed percentage of generation/demand.
    Section II-A states C_loss will be fixed by TSOs (e.g., 50% of large energy user demand); the model does not estimate it.

pith-pipeline@v1.3.0-alltime-deepseek · 6558 in / 14083 out tokens · 143850 ms · 2026-08-01T21:58:57.812959+00:00 · methodology

0 comments
read the original abstract

This paper presents a generative artificial intelligence (Gen AI) approach for forecasting, at a day-ahead stage, the largest single infeed (LSI) and largest single outfeed (LSO) on the Irish power system to assist in reserve dimensioning. Developed collaboratively between EirGrid, the electric transmission system operator (TSO) for Ireland, and GridZero.ai using the GridZero.ai platform, the system delivers accurate forecasts up to 38 hours ahead of real-time using limited data available before the day-ahead and intra-day energy market gate closure timings. Initial performance demonstrates an accuracy with a mean absolute percentage error (MAPE) that is only 1.1\% higher than the results possible using full market data (8-hours ahead). Thus, if this approach is integrated into operational systems and such high levels of accuracy are maintained, reserve procurement costs could be significantly reduced. The results also demonstrate the practicality and extensibility of AI-powered resource planning for TSOs.

Figures

Figures reproduced from arXiv: 2607.15900 by Amir Moshari, Bryan Murray, Chotiya Mahittigul, Colm Gaffney, Eoin Kennedy, Manuel Hurtado, Michael Walsh, Mo Cloonan, Ritesh Madan, Simon Tweed, Taulant Kerci, Zhi Li.

Figure 1
Figure 1. Figure 1: Illustration of the variation of LSI and frequent LSI unit changes [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The GridZero.ai platform consists of three closely connected layers: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Pearson correlation coefficients between the LSI value and various [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Actual vs. predicted LSI values for the three models—(top) GridZero.ai [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Actual vs. predicted LSO values for the three models—(top) [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

17 extracted references · 1 linked inside Pith

  1. [1]

    Analysis of wind energy curtailment in the Ireland and Northern Ireland power systems,

    M. Hurtado, T. K ¨erc ¸i, S. Tweed, E. Kennedy, N. Kamaluddin, and F. Milano, “Analysis of wind energy curtailment in the Ireland and Northern Ireland power systems,” in2023 IEEE Power & Energy Society General Meeting (PESGM), 2023, pp. 1–5

  2. [2]

    Stability assessment of low-inertia power systems: A system operator perspective,

    M. Hurtado, M. Jafarian, T. K ¨erc ¸i, S. Tweed, M. V . Escudero, E. Kennedy, and F. Milano, “Stability assessment of low-inertia power systems: A system operator perspective,” in2024 IEEE Power & Energy Society General Meeting (PESGM), 2024, pp. 1–5

  3. [3]

    A comprehensive approach to evaluate fre- quency control strength of power systems,

    T. K ¨erc ¸i and F. Milano, “A comprehensive approach to evaluate fre- quency control strength of power systems,”IEEE Open Access Journal of Power and Energy, vol. 13, pp. 39–50, 2026

  4. [4]

    Day-ahead system services auction (DASSA) design recommendations paper v1.0,

    EirGrid & SONI, “Day-ahead system services auction (DASSA) design recommendations paper v1.0,” 2024. [Online]. Available: https://consult.eirgrid.ie/

  5. [5]

    Day-ahead system services auction (DASSA) volume forecasting methodology consultation paper v1.0,

    ——, “Day-ahead system services auction (DASSA) volume forecasting methodology consultation paper v1.0,” 2024. [Online]. Available: https://consult.eirgrid.ie/

  6. [6]

    Dynamic dimensioning approach for operating reserves: Proof of concept in Belgium,

    K. De V os, N. Stevens, O. Devolder, A. Papavasiliou, B. Hebb, and J. Matthys-Donnadieu, “Dynamic dimensioning approach for operating reserves: Proof of concept in Belgium,”Energy Policy, vol. 124, pp. 272–285, 2019

  7. [7]

    Power systems optimization under uncertainty: A review of methods and applications,

    L. A. Roald, D. Pozo, A. Papavasiliou, D. K. Molzahn, J. Kazempour, and A. Conejo, “Power systems optimization under uncertainty: A review of methods and applications,”Electric Power Systems Research, vol. 214, p. 108725, 2023

  8. [8]

    Operating dynamic reserve dimensioning using probabilistic forecasts,

    N. Costilla-Enriquez, M. A. Ortega-Vazquez, A. Tuohy, A. Motley, and R. Webb, “Operating dynamic reserve dimensioning using probabilistic forecasts,”IEEE Transactions on Power Systems, vol. 38, no. 1, pp. 603–616, 2023

  9. [9]

    Assessing dynamic reserves vs. stochastic optimization for effective integration of operating probabilistic forecasts,

    Q. Wang, M. Ortega-Vazquez, A. Tuohy, E. Ela, M. Bello, D. Kirk- Davidoff, W. B. Hobbs, and V . Kumar, “Assessing dynamic reserves vs. stochastic optimization for effective integration of operating probabilistic forecasts,”IEEE Transactions on Sustainable Energy, vol. 16, no. 3, pp. 2132–2143, 2025

  10. [10]

    Assessing potential benefits of dynamic frequency restoration reserve dimensioning in multi-area power systems,

    A. Khodadadi, H. Nordstr ¨om, R. Eriksson, and L. S ¨oder, “Assessing potential benefits of dynamic frequency restoration reserve dimensioning in multi-area power systems,”Electric Power Systems Research, vol. 247, p. 111807, 2025

  11. [11]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”Advances in neural information processing systems, vol. 30, 2017

  12. [12]

    D. Rothman,Transformers for Natural Language Processing: Build, train, and fine-tune deep neural network architectures for NLP with Python, Hugging Face, and OpenAI’s GPT-3, ChatGPT, and GPT-4. Packt Publishing Ltd, 2022

  13. [13]

    Trans- formers in time series: A survey,

    Q. Wen, T. Zhou, C. Zhang, W. Chen, Z. Ma, J. Yan, and L. Sun, “Trans- formers in time series: A survey,”arXiv preprint arXiv:2202.07125, 2022

  14. [14]

    Adaptive temporal transformer method for short-term wind power forecasting considering shift in time series distribution,

    D. Li, Y . Hu, S. Miao, Z. Fang, Y . Liang, and S. He, “Adaptive temporal transformer method for short-term wind power forecasting considering shift in time series distribution,”AIP Advances, vol. 14, no. 2, 2024

  15. [15]

    Fourcastnet: Accelerating global high-resolution weather forecasting using adaptive fourier neural operators,

    T. Kurth, S. Subramanian, P. Harrington, J. Pathak, M. Mardani, D. Hall, A. Miele, K. Kashinath, and A. Anandkumar, “Fourcastnet: Accelerating global high-resolution weather forecasting using adaptive fourier neural operators,” inProceedings of the platform for advanced scientific com- puting conference, 2023, pp. 1–11

  16. [16]

    New reserve services for the Ireland and Northern Ireland power systems,

    T. K ¨erc ¸i, M. Hurtado, B. Franken, M. Cloonan, D. McGowan, S. Tweed, and E. Kennedy, “New reserve services for the Ireland and Northern Ireland power systems,” in2025 IEEE Power & Energy Society General Meeting (PESGM), 2025, pp. 1–5

  17. [17]

    Establishing a guideline on electricity transmission system operation,

    European Commission, “Establishing a guideline on electricity transmission system operation,” 2017. [Online]. Available: https: //eur-lex.europa.eu/