REVIEW 3 major objections 2 minor 3 cited by
Towards Generalization of Graph Neural Networks for AC Optimal Power Flow
T0 review · 3 major / 2 minor · reviewed 2026-05-18 · grok-4.3
Pith's one-line read A hybrid graph neural network achieves less than 1% optimality gap for AC optimal power flow across grid sizes from 14 to 2000 buses and generalizes to N-1 contingencies.
desk verdict The hybrid HH-MPNN gets under 1% gap on default topologies up to 2000 buses and under 3% on some zero-shot N-1 cases, but the benchmarks' coverage of real contingencies is the open question. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Hybrid Heterogeneous Message Passing Neural Network (HH-MPNN) that combines heterogeneous GNN for component-specific modeling with transformer-based global attention and physics-informed positional encodings.
What would settle it
Running the model on a grid topology or contingency set substantially different from the training distribution and observing optimality gaps consistently above 5% would indicate the generalization claims do not hold.
Extended reading notes
Core claim
The HH-MPNN architecture integrates a heterogeneous graph neural network with a scalable transformer and physics-informed positional encodings to explicitly model distinct power system components for local features and employ global attention for long-range dependencies, resulting in less than 1% optimality gap on default topologies from 14 to 2000 buses and less than 3% gap in zero-shot N-1 generalization on test cases after training only on default topologies.
Load-bearing premise
The selected benchmarks and augmentation strategies capture enough of the variety in actual power grid topologies and critical contingencies that the observed generalization holds for unseen real-world cases.
Editorial extensions
If this is right
- Models can be trained on small grids and applied to much larger ones with improved performance.
- Targeted data augmentation allows robust handling of high-impact contingencies without simulating all possible cases.
- Computational speedups reach up to 5000 times compared to interior point methods for practical deployment.
- Topology flexibility reduces the need for retraining when grid configurations change.
Reading between the lines
- The same architecture principles could apply to other network flow optimization tasks beyond power systems.
- Real-world validation on proprietary utility data would be needed to confirm transfer beyond the public benchmarks used.
- Pre-training strategies might reduce data requirements for new grid deployments.
- Integration with online monitoring systems could enable adaptive control in dynamic grids.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a Hybrid Heterogeneous Message Passing Neural Network (HH-MPNN) that combines a heterogeneous GNN, a scalable transformer, and physics-informed positional encodings to solve AC Optimal Power Flow (ACOPF). It reports empirical results on PGLearn and GridFM-DataKit benchmarks showing <1% optimality gap on default (N-0) topologies for grids ranging from 14 to 2000 buses, zero-shot N-1 generalization with <3% gap on several test cases when trained only on default topologies, further gains from targeted training-data augmentation for high-impact contingencies, and improved large-grid performance via pre-training on small grids, together with speedups up to 5000× versus interior-point solvers.
Significance. If the reported zero-shot and size generalization results hold under broader validation, the work would advance practical ML surrogates for real-time ACOPF by addressing topology flexibility and scalability, two persistent barriers in prior GNN-based approaches. The explicit modeling of heterogeneous power-system components and the demonstration that exhaustive N-1 simulation may not be required are potentially useful contributions, provided the benchmarks adequately sample real-world topology and contingency distributions.
major comments (3)
- [§4.3] §4.3 (Targeted Augmentation for N-1): the post-hoc selection of high-impact contingencies for augmentation introduces a potential selection bias that is not quantified; the central zero-shot generalization claim (abstract and §5.2) therefore rests on an unverified assumption that the chosen augmentation set is representative of the full distribution of critical N-1 cases.
- [§5.1] §5.1 and Table 2: optimality-gap results are reported without error bars, standard deviations across random seeds, or exact train/validation/test split ratios; this omission makes it impossible to assess statistical reliability of the <1% gap claim across grid sizes 14–2000 buses.
- [§5.3] §5.3 (Size Generalization): the pre-training experiment on small grids improving large-grid performance lacks an ablation that isolates the contribution of the physics-informed positional encodings versus the heterogeneous message-passing layers; without this, the source of the reported size-transfer benefit remains unclear.
minor comments (2)
- [Figure 1] Notation for the heterogeneous node/edge types in the MPNN is introduced without an explicit legend in Figure 1; adding a small table would improve readability.
- [Abstract] The abstract states “several test cases” for the <3% N-1 result; the main text should list the exact contingency indices and grid identifiers used.
Simulated Author's Rebuttal
We thank the referee for their constructive comments on our manuscript. We provide point-by-point responses to the major comments below, indicating the revisions we plan to make.
read point-by-point responses
-
Referee: [§4.3] §4.3 (Targeted Augmentation for N-1): the post-hoc selection of high-impact contingencies for augmentation introduces a potential selection bias that is not quantified; the central zero-shot generalization claim (abstract and §5.2) therefore rests on an unverified assumption that the chosen augmentation set is representative of the full distribution of critical N-1 cases.
Authors: We clarify that the zero-shot N-1 generalization results reported in the abstract and §5.2 are obtained by training exclusively on default (N-0) topologies and evaluating directly on N-1 contingency cases without any data augmentation. These results do not depend on the targeted augmentation strategy. The targeted augmentation in §4.3 is an additional technique to achieve robust performance specifically on high-impact contingencies, demonstrating that exhaustive N-1 simulation is not required. We acknowledge that the selection of high-impact contingencies could introduce bias and will revise the manuscript to explicitly describe the selection criterion (e.g., based on impact on objective value or constraint violations) and discuss its representativeness to the extent possible with the available data. revision: partial
-
Referee: [§5.1] §5.1 and Table 2: optimality-gap results are reported without error bars, standard deviations across random seeds, or exact train/validation/test split ratios; this omission makes it impossible to assess statistical reliability of the <1% gap claim across grid sizes 14–2000 buses.
Authors: We agree that including measures of statistical variability would improve the assessment of our results. In the revised manuscript, we will report standard deviations computed over multiple random seeds for the optimality gaps in Table 2 and specify the exact proportions used for train, validation, and test splits in each experiment. revision: yes
-
Referee: [§5.3] §5.3 (Size Generalization): the pre-training experiment on small grids improving large-grid performance lacks an ablation that isolates the contribution of the physics-informed positional encodings versus the heterogeneous message-passing layers; without this, the source of the reported size-transfer benefit remains unclear.
Authors: We appreciate this suggestion for clarifying the contributions of different components. We will add an ablation study in the revised §5.3 that compares variants with and without the physics-informed positional encodings, as well as with and without the heterogeneous message-passing layers, to isolate their individual impacts on the size generalization performance. revision: yes
Circularity Check
No circularity: empirical claims rest on external benchmarks
full rationale
The paper presents an HH-MPNN architecture for ACOPF and reports performance metrics (optimality gaps, zero-shot N-1 generalization, size scaling) evaluated on independent external datasets (PGLearn, GridFM-DataKit). No derivation chain, equations, or fitted parameters are described that reduce by construction to self-defined quantities, self-citations, or renamed inputs. Claims are statistically falsifiable against held-out test cases and do not invoke uniqueness theorems or ansatzes from prior self-work as load-bearing justification. This is a standard empirical ML evaluation paper with no detectable circularity patterns.
Assumptions & free parameters
free parameters (1)
- Neural network hyperparameters and training schedule
assumptions (1)
- domain assumption Power system components can be represented as distinct node and edge types in a heterogeneous graph whose local structure encodes the AC power flow equations sufficiently for learning.
Cite this review
Pith. "Pith review of Towards Generalization of Graph Neural Networks for AC Optimal Power Flow." pith.science (2026). https://pith.science/paper/2510.06860
@misc{pith2026251006860,
author = {Pith},
title = {Pith review of: Towards Generalization of Graph Neural Networks for AC Optimal Power Flow},
year = {2026},
howpublished = {\url{https://pith.science/paper/2510.06860}},
note = {Machine review of arXiv:2510.06860}
}
read the original abstract
AC Optimal Power Flow (ACOPF) is computationally intensive for large-scale grids, often requiring prohibitive solution times with conventional solvers. Machine learning offers significant speedups, but existing models struggle with scalability and topology flexibility. To address these challenges, we propose a Hybrid Heterogeneous Message Passing Neural Network (HH-MPNN) that integrates a heterogeneous graph neural network (GNN) with a scalable transformer and physics-informed positional encodings. Our architecture explicitly models distinct power system components to capture local features while using global attention for long-range dependencies. Evaluated on diverse benchmarks, including PGLearn and GridFM-DataKit datasets, HH-MPNN achieves less than 1% optimality gap on default topologies across grid sizes from 14 to 2,000 buses. For N-1 contingencies, our approach demonstrates zero-shot N-1 generalization with less than 3% optimality gap on several test cases despite training only on default topologies. We further develop an approach that ensures robust N-1 generalization to high-impact contingencies through targeted augmentation of the training data, showing that exhaustive simulation is unnecessary for topologically flexible models. Finally, size generalization experiments demonstrate that pre-training on small grids significantly improves performance on large-scale systems. Achieving computational speedups of up to 5,000 times compared to interior point solvers, these results advance practical, generalizable machine learning for real-time power system operations.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 3 Pith papers
-
GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
A unified neural solver, GENCO, handles power flow, optimal power flow, and state estimation in one architecture, with large speedups over classical AC solvers on large grids.
-
Towards Systematic Generalization for Power Grid Optimization Problems
A shared graph neural network framework jointly solves ACOPF and SCUC problems using physics constraints and shows improved generalization to unseen grid topologies.
-
Learning to Route Electric Trucks Under Operational Uncertainty
A reinforcement learning framework formulated as an event-driven semi-Markov decision process with graph states and action masking outperforms heuristic and optimization baselines for stochastic electric truck routing...
Reference graph
Works this paper leans on
-
[1]
History of Optimal Power Flow and Formulations,
M. B. Cain, R. P. O’Neill, and A. Castillo, “History of Optimal Power Flow and Formulations,” 2012
work page 2012
-
[2]
Review of Ma- chine Learning Techniques for Optimal Power Flow,
H. Khaloie, M. Dolanyi, J.-F. Toubeau, and F. Vall ´ee, “Review of Ma- chine Learning Techniques for Optimal Power Flow,”SSRN Electronic Journal, 2024
work page 2024
-
[3]
Operating in the Fog: Security Management Under Uncertainty,
P. Panciatici, G. Bareux, and L. Wehenkel, “Operating in the Fog: Security Management Under Uncertainty,”IEEE Power and Energy Magazine, vol. 10, no. 5, pp. 40–49, Sep. 2012
work page 2012
-
[4]
Foundation models for the electric power grid,
H. F. Hamann, B. Gjorgiev, T. Brunschwiler, L. S. Martins, A. Puech, A. Varbella, J. Weiss, J. Bernabe-Moreno, A. B. Mass ´e, S. L. Choi, I. Foster, B.-M. Hodge, R. Jain, K. Kim, V . Mai, F. Mirall`es, M. De Mon- tigny, O. Ramos-Lea ˜nos, H. Supr ˆeme, L. Xie, E.-N. S. Youssef, A. Zinflou, A. Belyi, R. J. Bessa, B. P. Bhattarai, J. Schmude, and S. Sobolev...
work page 2024
-
[5]
S. Shin, M. Anitescu, and F. Pacaud, “Accelerating optimal power flow with GPUs: SIMD abstraction of nonlinear programs and condensed- space interior-point methods,”Electric Power Systems Research, vol. 236, p. 110651, Nov. 2024
work page 2024
-
[6]
DeepOPF: A Feasibility- Optimized Deep Neural Network Approach for AC Optimal Power Flow Problems,
X. Pan, M. Chen, T. Zhao, and S. H. Low, “DeepOPF: A Feasibility- Optimized Deep Neural Network Approach for AC Optimal Power Flow Problems,”IEEE Systems Journal, vol. 17, no. 1, pp. 673–683, Mar. 2023
work page 2023
-
[7]
Z. Yang, H. Zhong, A. Bose, T. Zheng, Q. Xia, and C. Kang, “A Linearized OPF Model With Reactive Power and V oltage Magnitude: A Pathway to Improve the MW-Only DC OPF,”IEEE Transactions on Power Systems, vol. 33, no. 2, pp. 1734–1745, Mar. 2018, publisher: Institute of Electrical and Electronics Engineers (IEEE)
work page 2018
-
[8]
Implementation of a Large-Scale Optimal Power Flow Solver Based on Semidefinite Programming,
D. K. Molzahn, J. T. Holzer, B. C. Lesieutre, and C. L. DeMarco, “Implementation of a Large-Scale Optimal Power Flow Solver Based on Semidefinite Programming,”IEEE Transactions on Power Systems, vol. 28, no. 4, pp. 3987–3998, Nov. 2013, publisher: Institute of Electrical and Electronics Engineers (IEEE)
work page 2013
Show all 34 references
-
[9]
High- Fidelity Machine Learning Approximations of Large-Scale Optimal Power Flow,
M. Chatzos, F. Fioretto, T. W. K. Mak, and P. V . Hentenryck, “High- Fidelity Machine Learning Approximations of Large-Scale Optimal Power Flow,” Jun. 2020, arXiv:2006.16356 [eess]
2020
-
[10]
Solutions of DC OPF are Never AC Feasible,
K. Baker, “Solutions of DC OPF are Never AC Feasible,” inProceedings of the Twelfth ACM International Conference on Future Energy Systems. Virtual Event Italy: ACM, Jun. 2021, pp. 264–268
2021
-
[11]
Learning Optimal Solutions for Extremely Fast AC Optimal Power Flow,
A. S. Zamzam and K. Baker, “Learning Optimal Solutions for Extremely Fast AC Optimal Power Flow,” in2020 IEEE International Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm). Tempe, AZ, USA: IEEE, Nov. 2020, pp. 1–6
2020
-
[12]
Compact Optimization Learning for AC Optimal Power Flow,
S. Park, W. Chen, T. W. Mak, and P. Van Hentenryck, “Compact Optimization Learning for AC Optimal Power Flow,”IEEE Transactions on Power Systems, vol. 39, no. 2, pp. 4350–4359, Mar. 2024
2024
-
[13]
Unsupervised Learning for Solving AC Optimal Power Flows: Design, Analysis, and Experiment,
W. Huang, M. Chen, and S. H. Low, “Unsupervised Learning for Solving AC Optimal Power Flows: Design, Analysis, and Experiment,”IEEE Transactions on Power Systems, vol. 39, no. 6, pp. 7102–7114, Nov. 2024
2024
-
[14]
Unsupervised Optimal Power Flow Using Graph Neural Networks,
D. Owerko, F. Gama, and A. Ribeiro, “Unsupervised Optimal Power Flow Using Graph Neural Networks,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Korea, Republic of: IEEE, Apr. 2024, pp. 6885– 6889
2024
-
[15]
Learning Warm-Start Points for AC Optimal Power Flow,
K. Baker, “Learning Warm-Start Points for AC Optimal Power Flow,” May 2019, arXiv:1905.08860 [math]
2019 arXiv
-
[16]
Data-driven Probabilistic Constraint Elimi- nation for Accelerated Optimal Power Flow,
C. Crozier and K. Baker, “Data-driven Probabilistic Constraint Elimi- nation for Accelerated Optimal Power Flow,” in2022 IEEE Power & Energy Society General Meeting (PESGM). Denver, CO, USA: IEEE, Jul. 2022, pp. 1–5
2022
-
[17]
Spatial Network Decomposition for Fast and Scalable AC-OPF Learning,
M. Chatzos, T. W. K. Mak, and P. V . Hentenryck, “Spatial Network Decomposition for Fast and Scalable AC-OPF Learning,”IEEE Trans- actions on Power Systems, vol. 37, no. 4, pp. 2601–2612, Jul. 2022
2022
-
[18]
Leveraging Power Grid Topology in Machine Learning Assisted Optimal Power Flow,
T. Falconer and L. Mones, “Leveraging Power Grid Topology in Machine Learning Assisted Optimal Power Flow,”IEEE Transactions on Power Systems, vol. 38, no. 3, pp. 2234–2246, May 2023
2023
-
[19]
ConvOPF-DOP: A Data-Driven Method for Solving AC-OPF Based on CNN Considering Different Operation Patterns,
Y . Jia, X. Bai, L. Zheng, Z. Weng, and Y . Li, “ConvOPF-DOP: A Data-Driven Method for Solving AC-OPF Based on CNN Considering Different Operation Patterns,”IEEE Transactions on Power Systems, vol. 38, no. 1, pp. 853–860, Jan. 2023
2023
-
[20]
DeepOPF-FT: One Deep Neural Network for Multiple AC-OPF Problems With Flexible Topology,
M. Zhou, M. Chen, and S. H. Low, “DeepOPF-FT: One Deep Neural Network for Multiple AC-OPF Problems With Flexible Topology,”IEEE Transactions on Power Systems, vol. 38, no. 1, pp. 964–967, Jan. 2023
2023
-
[21]
Optimal Power Flow Using Graph Neural Networks,
D. Owerko, F. Gama, and A. Ribeiro, “Optimal Power Flow Using Graph Neural Networks,” inICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Barcelona, Spain: IEEE, May 2020, pp. 5930–5934
2020
-
[22]
A Physics-Guided Graph Con- volution Neural Network for Optimal Power Flow,
M. Gao, J. Yu, Z. Yang, and J. Zhao, “A Physics-Guided Graph Con- volution Neural Network for Optimal Power Flow,”IEEE Transactions on Power Systems, vol. 39, no. 1, pp. 380–390, Jan. 2024
2024
-
[23]
Topology-Aware Graph Neural Networks for Learning Feasible and Adaptive AC-OPF Solutions,
S. Liu, C. Wu, and H. Zhu, “Topology-Aware Graph Neural Networks for Learning Feasible and Adaptive AC-OPF Solutions,”IEEE Transac- tions on Power Systems, vol. 38, no. 6, pp. 5660–5670, Nov. 2023
2023
-
[24]
CANOS: A Fast and Scalable Neural AC-OPF Solver Robust To N-1 Perturbations,
L. Piloto, S. Liguori, S. Madjiheurem, M. Zgubic, S. Lovett, H. Tom- linson, S. Elster, C. Apps, and S. Witherspoon, “CANOS: A Fast and Scalable Neural AC-OPF Solver Robust To N-1 Perturbations,” Mar. 2024, arXiv:2403.17660 [cs]
2024
-
[25]
Recipe for a General, Powerful, Scalable Graph Trans- former,
L. Ramp ´aˇsek, M. Galkin, V . P. Dwivedi, A. T. Luu, G. Wolf, and D. Beaini, “Recipe for a General, Powerful, Scalable Graph Trans- former,” Jan. 2023, arXiv:2205.12454 [cs]
2023
-
[26]
Interaction Networks for Learning about Objects, Relations and Physics,
P. W. Battaglia, R. Pascanu, M. Lai, D. Rezende, and K. Kavukcuoglu, “Interaction Networks for Learning about Objects, Relations and Physics,” Dec. 2016, arXiv:1612.00222 [cs]
2016 arXiv
-
[27]
Rethinking Attention with Performers,
K. Choromanski, V . Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, D. Belanger, L. Colwell, and A. Weller, “Rethinking Attention with Performers,” Nov. 2022, arXiv:2009.14794 [cs]
2022 arXiv
-
[28]
A Topological Investigation of Power Flow,
H. Cetinay, F. A. Kuipers, and P. Van Mieghem, “A Topological Investigation of Power Flow,”IEEE Systems Journal, vol. 12, no. 3, pp. 2524–2532, Sep. 2018
2018
-
[29]
The impact of the topology on cascading failures in a power grid model,
Y . Koc ¸, M. Warnier, P. V . Mieghem, R. E. Kooij, and F. M. Brazier, “The impact of the topology on cascading failures in a power grid model,” Physica A: Statistical Mechanics and its Applications, vol. 402, pp. 169– 179, May 2014
2014
-
[30]
OPFData: Large-scale datasets for AC optimal power flow with topological perturbations
S. Lovett, M. Zgubi ˇc, S. Liguori, S. Madjiheurem, H. Tomlinson, S. Elster, C. Apps, S. Witherspoon, and L. Piloto, “OPFData: Large-scale datasets for AC optimal power flow with topological perturbations.”
-
[31]
The Power Grid Library for Benchmarking AC Optimal Power Flow Algorithms,
S. Babaeinejadsarookolaee, A. Birchfield, R. D. Christie, C. Coffrin, C. DeMarco, R. Diao, M. Ferris, S. Fliscounakis, S. Greene, R. Huang, C. Josz, R. Korab, B. Lesieutre, J. Maeght, T. W. K. Mak, D. K. Molzahn, T. J. Overbye, P. Panciatici, B. Park, J. Snodgrass, A. Tbaileh,...
2021
-
[32]
Exploring extrapolation of machine learning models for power system time domain simulation,
O. Arowolo, J. Stiasny, and J. Cremer, “Exploring extrapolation of machine learning models for power system time domain simulation,” Sustainable Energy, Grids and Networks, vol. 43, p. 101908, Sep. 2025
2025
-
[33]
From Local Structures to Size Generalization in Graph Neural Networks,
G. Yehudai, E. Fetaya, E. Meirom, G. Chechik, and H. Maron, “From Local Structures to Size Generalization in Graph Neural Networks,” Jul. 2021, arXiv:2010.08853 [cs]
2021
-
[34]
Predicting AC Optimal Power Flows: Combining Deep Learning and Lagrangian Dual Methods,
F. Fioretto, T. W. K. Mak, and P. V . Hentenryck, “Predicting AC Optimal Power Flows: Combining Deep Learning and Lagrangian Dual Methods,” Dec. 2019, arXiv:1909.10461 [eess]. APPENDIXA PROBLEM FORMULATION For a power system withNbuses,Ebranches,Ggenerators, Lloads andSshunts,...
2019
Reviewed May 18, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.