REVIEW 4 major objections 5 minor 1 cited by
AI-Driven Stabilization in Power Grids through Controlling Line Admittances
T0 review · 4 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read A single reinforcement-learning controller, applied once after a transmission-line fault, reduces average frequency fluctuations in a real power-grid model by about half.
desk verdict The genuinely new thing here is the policy-derived placement ranking that beats PTDF; the 53% reduction figure is a proof-of-concept from an idealized actuator in a simplified model, not a field-ready number. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is a graph-neural-network policy that maps the grid state—inertias, dampings, powers, phases, angular velocities, fault-line indicator, and admittance matrix—to two outputs per line: a Bernoulli control decision c_ij and a Gaussian adjustment exponent q_ij. The applied action rescales admittance as Ŷ_ij = 2^{c_ij q_ij χ_ij} Y_ij, where χ_ij marks whether the line has a regulator, allowing both increases and decreases. Trained with a clipped policy-gradient update to maximize a reward that balances the reduction of inertia-weighted frequency fluctuation ΔΞ(ℓ) against a penalty on the number of controlled lines, this single-step policy selects both which lines to act on and h
What would settle it
A hardware-in-the-loop or electromagnetic-transient simulation of the same UK grid with realistic FACTS constraints—say, ±30% admittance limits, 50 ms response delay, and discrete tap steps—would settle the claim. If the same trained policy, applied at the delayed time with bounded actions, yields an average fluctuation reduction well below 53% or no longer reaches near-minimal fluctuation with the top five S(2) regulators, the paper's central quantitative conclusions would be refuted.
Extended reading notes
Core claim
The paper's central claim is that a single reinforcement-learning agent (the Adaptive Admittance Controller, AAC) can, in one immediate action, rescale transmission-line admittances after a line fault and thereby cut the average inertia-weighted frequency fluctuation by about 53% on the reduced UK grid and about 55% on a synthetic homogeneous grid. The agent intervenes selectively—strongly damping high-impact faults and barely touching minor ones. Its action statistics yield three placement rankings; the best (S(2), based on the ratio of adjusted to original admittance) reaches minimal fluctuation with 35 of 114 regulators, and five top-ranked regulators give near-optimal cost-effective stab
Load-bearing premise
The load-bearing premise is that line admittance can be changed instantly, without bound, and essentially for free (apart from a small penalty on the number of controlled lines) at the moment a fault is detected; real flexible-AC-transmission devices have limited compensation ranges, response delays, and operating costs, so the 53% reduction and the 'five regulators suffice' conclusion depend on this idealization.
Editorial extensions
If this is right
- On the reduced UK grid, AAC reduces the average inertia-weighted frequency fluctuation by about 53% over all 105 single-line fault scenarios; on the homogeneous SHK grid it achieves about 55%.
- Using the S(2) ranking, the minimum average fluctuation is reached with 35 of 114 regulators on the UK grid and 74 of 115 on the SHK grid; the top five S(2) lines already provide near-optimal, cost-effective stabilization.
- The AAC-derived placement rankings outperform the traditional PTDF-based placement heuristic in terms of achievable fluctuation reduction.
- The controller intervenes selectively: it strongly damps high-impact faults, leaves low-impact faults alone, and only one minor scenario shows a slight worsening.
- The reduction in transient frequency fluctuation and the restoration of steady-state power flow are strongly correlated (ρ≈0.83 on the UK grid, 0.81 on SHK), indicating simultaneous stabilization of both regimes.
Reading between the lines
- An extension not tested in the paper: training the same controller under actuator constraints—bounded admittance range, finite response time, and switching costs—would show whether the ~53% reduction and the top-five placement result survive realistic device limitations.
- Because the action only rescales line coupling strengths in a Kuramoto-like swing network, the same framework could be applied proactively (before a fault) or to other flow networks where a few link-weight changes could suppress cascades, such as road or data networks.
- The finding that S(2) (typical control effort) beats both intervention frequency and PTDF suggests a testable hypothesis: the most effective regulator locations are those where the grid's response is most sensitive to admittance changes, a quantity that could be computed directly from the swing-equation Jacobian.
- The near-optimality of five regulators hints at an underlying low-dimensional structure in which only a handful of lines control the slowest, most vulnerable modes; analyzing the network's dominant eigenmodes could connect this empirical ranking to modal controllability.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Adaptive Admittance Controller (AAC), a PPO-trained graph neural network that, in a single step after a line outage, outputs admittance multipliers for all controlled lines according to Eq. (5). On a reduced 54-bus model of the UK grid (105 fault cases) and on an SHK synthetic grid, the authors report that AAC reduces the inertia-weighted frequency-fluctuation measure Xi by about 53% and 55%, respectively. They then derive three line-ranking metrics S(1)-S(3) from the trained policy, report that S(2) reaches the minimum average fluctuation with about 35 regulators, and propose that the top five S(2) lines are a cost-effective regulator set. A phase-space analysis is used to claim that the controller simultaneously restores pre-fault power flows. The central quantitative claims are the 53%/55% reductions and the sufficiency of the five-regulator set.
Significance. If established, the results would be a useful contribution to the growing literature on RL-based grid control: a single framework that outputs both regulator placement and control settings, with a control-aware ranking that outperforms the PTDF heuristic. The paper is also careful to state its simulation equations, the single-step episode design, and the reward function. However, the paper does not yet provide independent evidence that the main numbers are robust. All headline figures are single-run outcomes, the evaluation faults overlap with the training faults, the actuation model in Eq. (5) is unbounded and instantaneous, and the PPO loss in Eq. (7) is written in a non-standard form. These are fixable, but they are load-bearing for the abstract's claims of real-time stabilization and cost-effective placement.
major comments (4)
- [Sec. V, Eq. (5); Sec. II.A] The control action is idealized. In Eq. (5), \hat{Y}_{ij}=2^{\delta y_{ij}}Y_{ij} with \delta y_{ij}=c_{ij}q_{ij}\chi_{ij}; q_{ij} is sampled from N(\mu_{ij},\sigma_{ij}) during training and set to \mu_{ij} during evaluation, and no bound is placed on \mu_{ij} or \sigma_{ij}. The admittance can therefore be rescaled by arbitrary powers of two instantaneously after the fault. Real TCSC/FACTS devices have bounded compensation ranges, finite switching times, and communication/control latency. The Discussion flags only the simplified UK model and the single-fault restriction, not this actuator model. Because the 53%/55% reductions and the 'five regulators suffice' conclusion are obtained under this unbounded instantaneous actuation, the practical interpretation of the results is not yet supported. I ask for experiments with bounded actions (e.g., |\delta y| \le \delta_max), finite actuation
- [Sec. V (reward); Sec. II.A and Fig. 1(b)] The headline 53% is not an independent out-of-sample result. The reward is R(\ell)=\Delta\Xi(\ell)/(1+0.05\sum c), with \Delta\Xi defined from Eq. (2), and the same \Delta\Xi is the quantity averaged in Fig. 1(b) to obtain the 53%. The policy is trained on all 105 single-line faults that are then used for evaluation (Sec. V: 'L(\ell) is averaged on all \ell-th single-line failures'; Sec. II.A: 'we consider all single-line failure scenarios'). No train/test split, no seed count, and no confidence intervals are reported; Figs. 1 and 3 show single curves. At minimum, please report multi-seed mean \pm standard deviation and a held-out or out-of-distribution evaluation (e.g., cross-validated over fault scenarios, or on SHK parameter perturbations). Without this, the numbers in the abstract cannot be distinguished from overfitting to the training objective.
- [Sec. V, Eq. (7)] The PPO surrogate loss is written with max, not the standard min: L(\ell) = - max[ (\pi/\pi_old)R, clip(\pi/\pi_old,1-\epsilon,1+\epsilon)R ]. In the standard clipped surrogate the outer operation is min, which is deliberately pessimistic and prevents the policy from exploiting the unclipped ratio when the ratio is outside the clipping range. With the max operation as written, positive rewards are not clipped for ratios above 1+\epsilon and the update has an optimistic bias. If this is a typographical error, it must be corrected; if the authors intentionally use max, the choice needs a derivation. Since no code is included, the reader cannot tell which objective was actually optimized.
- [Sec. II.B, Fig. 3; Sec. V (reward)] The placement conclusions depend on hand-tuned reward coefficients without sensitivity analysis. The penalty coefficient 0.05 and the no-effect reward Rc are described as empirically chosen, and Fig. 3 shows the resulting average \Xi versus N_regulator for a single training run. The statement that the minimum is reached with 35 regulators, and that five regulators are a 'sweet spot', is therefore specific to one choice of these hyperparameters and one random seed. Moreover, the Discussion's claim of 'cutting implementation costs by more than 95%' equates a reduction in regulator count with cost reduction, which is not justified for FACTS installations. Please add sensitivity analysis to (0.05, Rc), a random-placement baseline, and error bars for each N_regulator curve; additionally, either rephrase or support the cost claim.
minor comments (5)
- [Abstract and Sec. IV] The abstract says 'real UK power grid' and 'real time'; the model is the reduced Pagnier-Jacquod UK grid, and control is a one-shot action after a fault. Please qualify the wording to avoid overstatement.
- [Figs. 4 and 5] Typos: 'Uncontrlled' in Fig. 4 (and Fig. 9 in the SI) should be 'Uncontrolled'; 'venctor' in Fig. 5 should be 'vector'; 'F ACTS' in the Introduction should be 'FACTS'.
- [General] No data or code availability statement is provided. For a paper whose results depend on a trained stochastic policy, adding code, seeds, and trained-model details is important for reproducibility.
- [Eq. (2)] The displayed formula has a typesetting issue ('1P i mi'); please correct so that the inertia-weighted variance is unambiguous.
- [SI Table I] The SHK network's clustering coefficient (0.4070) differs noticeably from the UK value (0.5025); 'closely match' should be quantified or softened.
Circularity Check
Headline 53%/55% reduction is the RL training objective evaluated on the training scenarios; no independent held-out prediction.
-
fitted input called prediction
[Sec. II.A (Fig. 1b, 'On average...') and Sec. V (Reward design and PPO loss)]
"On average, it decreases Ξ(ℓ) by approximately 53%. ... R(ℓ) = ΔΞ(ℓ)/(1 + 0.05 Σ_{(i,j)} c_{ij}) if ΔΞ(ℓ) ≠ 0, Rc otherwise. ... To compute the decrease in frequency fluctuation ΔΞ(ℓ), we set T = 2 seconds for rapid policy updates during training, while T = 10 seconds is used in evaluation ... L(ℓ) is averaged on all ℓ-th single-line failures, and the policy network is updated through gradient ascent to maximize expected reward."
The headline 53% (and SI 55%) is not a prediction from an independent model: it is the average reduction in Ξ on the same single-line-fault scenarios used to train the policy, and the reward is a monotone function of that same ΔΞ. The PPO loss is averaged over all ℓ-th failures and maximizes expected reward, so the direction and rough magnitude of the reported reduction are forced by the training objective. The only deviations from exact identity are the shorter training window (T=2 s) vs. evaluation window (T=10 s) and the sparsity denominator, but no held-out or out-of-sample fault split is reported.
full rationale
The paper's internal simulation is coherent, and most of the claimed results (ranking comparisons, phase-space correlations, placement curves) are empirical and self-contained rather than derived by definition or by a self-citation chain. The only self-citation, ref. [36] for the definition of Ξ, is not load-bearing; Ξ is a standard inertia-weighted variance measure. No uniqueness theorem or ansatz is imported from the authors' prior work. The unbounded instantaneous admittance actuation in Eq. (5) is a serious modeling idealization, but that is a correctness/realism risk, not circularity. The one genuine circularity concern is that the central quantitative claim—the 53%/55% reduction—is the training objective evaluated on the training distribution: the reward is built from ΔΞ and the policy is trained and evaluated on the same fault set. This is partial circularity because a policy could in principle fail to reduce Ξ, and the training/evaluation windows differ, so the reduction is not a mathematical identity; nevertheless, the reported figure is an in-sample measure of the fitted objective rather than an independent prediction. Score 5 reflects this partial reduction-by-construction while recognizing the substantial independent empirical content elsewhere in the paper.
Assumptions & free parameters
free parameters (6)
- reward control-penalty coefficient =
0.05
- no-effect reward R_c =
10^-5 (UK grid; chosen again for SHK)
- evaluation window T =
10 s evaluation, 2 s training
- SHK generator parameters =
p=0.4, q=0.9, r=0.1, s=0.2
- PPO clipping epsilon =
0.1
- 'sweet spot' regulator count =
5
assumptions (7)
- domain assumption Post-fault grid dynamics are governed by the swing equation (Eq. 1) with constant bus voltages.
- domain assumption All bus voltages equal 1 in per-unit, so coupling Kij equals admittance Yij.
- domain assumption A fault is the instantaneous removal of one line (Yuv = 0); the grid must remain connected, restricting scenarios to 105 of 114 lines.
- domain assumption Ξ(ℓ) (Eq. 2) is the operative stability measure: reducing inertia-weighted frequency variance over 10 s equals stabilizing.
- domain assumption Iij = |sin(θi − θj)| is the fraction of line capacity used, so power-flow restoration is measured as return of phase differences to pre-fault values.
- ad hoc to paper The reward R(ℓ) = ΔΞ/(1 + 0.05 Σc) or R_c is the correct scalar proxy for control value, with a one-shot action structure.
- ad hoc to paper Admittance actions in Eq. (5) are instantly and boundlessly realizable with no device dynamics.
Cite this review
Pith. "Pith review of AI-Driven Stabilization in Power Grids through Controlling Line Admittances." pith.science (2026). https://pith.science/paper/XEW7M62E
@misc{pith2026260102114,
author = {Pith},
title = {Pith review of: AI-Driven Stabilization in Power Grids through Controlling Line Admittances},
year = {2026},
howpublished = {\url{https://pith.science/paper/XEW7M62E}},
note = {Machine review of arXiv:2601.02114}
}
read the original abstract
The global transition from traditional power plants to renewable energy sources introduces new challenges in grid stability, primarily because inverter-based technologies provide insufficient inertia. To address this, we introduce an artificial intelligence algorithm that autonomously stabilizes power grids by adaptively tuning admittance regulators in response to disturbances. This Adaptive Admittance Controller (AAC) algorithm not only stabilizes the system in real time but also identifies the best regulator locations, thereby unifying grid planning and real time control within a single framework. When tested on a real UK power grid, the AAC markedly reduces frequency deviations and rapidly restores nominal operation. In addition, the algorithm isolates a small number of key regulators and intervenes only on these, lowering both system complexity and cost. The AAC algorithm further reduces the nonlinearity effect, quickly stabilizing the frequency and power flow. This intelligent control scheme enables power grids to reliably return to stable operating conditions under a broad spectrum of fault scenarios. The proposed framework can also be used to mitigate cascading failures by adaptively controlling critical links in a variety of networked infrastructures, such as cascades of traffic congestion on road networks or fuse failures in energy-saving systems.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
Paths to synchronization in the Kuramoto model with inertia
In the inertial Kuramoto model, Gaussian frequencies produce hierarchical central entrainment while uniform frequencies produce homogeneous peripheral mergers, yielding smooth versus stepwise order-parameter growth.
Reference graph
Works this paper leans on
-
[1]
Obama, Presidential policy directive 21–critical infras- tructure security and resilience (2013)
B. Obama, Presidential policy directive 21–critical infras- tructure security and resilience (2013)
2013
-
[2]
S. M. Rinaldi, J. P. Peerenboom, and T. K. Kelly, Identi- 8 fying, understanding, and analyzing critical infrastructure interdependencies, IEEE control systems magazine 21, 11 (2001)
2001
-
[3]
Shaukat, S
N. Shaukat, S. Ali, C. Mehmood, B. Khan, M. Jawad, U. Farid, Z. Ullah, S. Anwar, and M. Majid, A survey on consumers empowerment, communication technologies, and renewable generation penetration within smart grid, Renewable and Sustainable Energy Reviews 81, 1453 (2018)
2018
-
[4]
Gulraiz, S
A. Gulraiz, S. Sajjad Haider Zaidi, and B. Moham- mad Khan, Advancing energy integration: renewable sources, ancillary services, and stability, PloS one 20, e0324812 (2025)
2025
-
[5]
Crivellaro, A
A. Crivellaro, A. Tayyebi, C. Gavriluta, D. Groß, A. Anta, F. Kupzog, and F. D¨ orfler, Beyond low-inertia systems: Massive integration of grid-forming power converters in transmission grids, in 2020 IEEE power & energy society general meeting (PESGM)(IEEE, 2020) pp. 1–5
2020
-
[6]
Khalid, Smart grids and renewable energy systems: Perspectives and grid integration challenges, Energy Strat- egy Reviews 51, 101299 (2024)
M. Khalid, Smart grids and renewable energy systems: Perspectives and grid integration challenges, Energy Strat- egy Reviews 51, 101299 (2024)
2024
-
[7]
K. Y. Yap, C. R. Sarimuthu, and J. M.-Y. Lim, Virtual inertia-based inverters for mitigating frequency instability in grid-connected renewable energy system: A review, Applied Sciences 9, 5300 (2019)
2019
-
[8]
Kerdphol, F
T. Kerdphol, F. S. Rahman, and Y. Mitani, Virtual iner- tia control application to enhance frequency stability of interconnected power systems with high renewable energy penetration, Energies 11, 981 (2018)
2018
Show all 48 references
-
[9]
Smith, O
O. Smith, O. Cattell, E. Farcot, R. D. O’Dea, and K. I. Hopcraft, The effect of renewable energy incorporation on power grid stability and resilience, Science advances 8, eabj6734 (2022)
2022
-
[10]
Bialek, What does the gb power outage on 9 august 2019 tell us about the current state of decarbonised power systems?, Energy Policy 146, 111821 (2020)
J. Bialek, What does the gb power outage on 9 august 2019 tell us about the current state of decarbonised power systems?, Energy Policy 146, 111821 (2020)
2019
-
[11]
Zhang, H
G. Zhang, H. Zhong, Z. Tan, T. Cheng, Q. Xia, and C. Kang, Texas electric power crisis of 2021 warns of a new blackout mechanism, CSEE journal of Power and Energy Systems 8, 1 (2022)
2021
-
[12]
N. M. Flores, H. McBrien, V. Do, M. V. Kiang, J. Schlegelmilch, and J. A. Casey, The 2021 texas power crisis: distribution, duration, and disparities, Journal of exposure science & environmental epidemiology 33, 21 (2023)
2021
-
[13]
Sharma, A
N. Sharma, A. Acharya, I. Jacob, S. Yamujala, V. Gupta, and R. Bhakar, Major blackouts of the decade: Underlying causes, recommendations and arising challenges, in 2021 9th IEEE International Conference on Power Systems (ICPS) (IEEE, 2021) pp. 1–6
2021
-
[14]
M. A. Raza, K. L. Khatri, A. Hussain, M. H. A. Khan, A. Shah, and H. Taj, Analysis and proposed remedies for power system blackouts around the globe, Engineering Proceedings 20, 5 (2022)
2022
-
[15]
St¨ urmer, A
J. St¨ urmer, A. Plietzsch, T. Vogt, F. Hellmann, J. Kurths, C. Otto, K. Frieler, and M. Anvari, Increasing the re- silience of the texas power grid against extreme storms by hardening critical lines, Nature Energy 9, 526 (2024)
2024
-
[16]
Pagnier and P
L. Pagnier and P. Jacquod, Inertia location and slow network modes determine disturbance propagation in large-scale power grids, PloS one 14, e0213550 (2019)
2019
-
[17]
Fern´ andez-Guillam´ on, E
A. Fern´ andez-Guillam´ on, E. Muljadi, and A. Molina- Garc ´ ıa, Frequency control studies: A review of power sys- tem, conventional and renewable generation unit model- ing, Electric Power Systems Research 211, 108191 (2022)
2022
-
[18]
Sch¨ afer, D
B. Sch¨ afer, D. Witthaut, M. Timme, and V. Latora, Dy- namically induced cascading failures in power grids, Na- ture communications 9, 1975 (2018)
1975
-
[19]
Pahwa, C
S. Pahwa, C. Scoglio, and A. Scala, Abruptness of cascade failures in power grids, Scientific reports 4, 3694 (2014)
2014
-
[20]
Zhang and O
Y. Zhang and O. Ya˘ gan, Optimizing the robustness of elec- trical power systems against cascading failures, Scientific reports 6, 27625 (2016)
2016
-
[21]
Y. Dai, R. Preece, and M. Panteli, Risk assessment of cascading failures in power systems with increasing wind penetration, Electric Power Systems Research211, 108392 (2022)
2022
-
[22]
Alexander, Oscillatory solutions of a model system of nonlinear swing equations, International Journal of Electrical Power & Energy Systems 8, 130 (1986)
J. Alexander, Oscillatory solutions of a model system of nonlinear swing equations, International Journal of Electrical Power & Energy Systems 8, 130 (1986)
1986
-
[23]
Q. Qiu, R. Ma, J. Kurths, and M. Zhan, Swing equation in power systems: Approximate analytical solution and bifurcation curve estimate, Chaos: An Interdisciplinary Journal of Nonlinear Science 30 (2020)
2020
-
[24]
´Odor, I
G. ´Odor, I. Papp, K. Benedek, and B. Hartmann, Improv- ing power-grid systems via topological changes or how self-organized criticality can help power grids, Physical Review Research 6, 013194 (2024)
2024
-
[25]
Sch¨ afer, T
B. Sch¨ afer, T. Pesch, D. Manik, J. Gollenstede, G. Lin, H.-P. Beck, D. Witthaut, and M. Timme, Understanding braess’ paradox in power grids, Nature Communications 13, 5396 (2022)
2022
-
[26]
N. L. Cain and H. T. Nelson, What drives opposition to high-voltage transmission lines?, Land use policy 33, 204 (2013)
2013
-
[27]
Cohen, K
J. Cohen, K. Moeltner, J. Reichl, and M. Schmidthaler, An empirical analysis of local opposition to new transmission lines across the eu-27, The Energy Journal 37, 59 (2016)
2016
-
[28]
Lee, S.-H
H.-J. Lee, S.-H. Kim, K. Hur, J.-S. Choi, H.-J. Oh, B.-J. Lee, G. Jang, and J. H. Chow, Integrating tcsc to enhance transmission capability and security: Feasibility studies for korean electric power system, in 2016 IEEE Power and Energy Society General Meeting (PESGM)(IEEE,
2016
-
[29]
Azimi and G
Z. Azimi and G. Shahgholian, Power system transient stability enhancement with tcsc controller using genetic algorithm optimization, International Journal of Natural and Engineering Sciences 10, 09–15 (2019)
2019
-
[30]
I. E. Nkan, E. E. Okpo, and A. B. Inyang, Enhancement of power systems transient stability with tcsc: a case study of the nigerian 330 kv, 48-bus network, International Journal of Multidisciplinary Research and Analysis 6, 4828 (2023)
2023
-
[31]
Silver, A
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al., Mastering the game of go with deep neural networks and tree search, nature 529, 484 (2016)
2016
-
[32]
Mirhoseini, A
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, A. Nova, et al., A graph placement methodology for fast chip design, Nature 594, 207 (2021)
2021
-
[33]
Degrave, F
J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de Las Casas, et al., Magnetic control of tokamak plasmas through deep reinforcement learning, Nature 602, 414 (2022)
2022
-
[34]
Jo, J.-Y
S. Jo, J.-Y. Oh, Y. T. Yoon, and Y. G. Jin, Self-healing 9 radial distribution network reconfiguration based on deep reinforcement learning, Results in Engineering 22, 102026 (2024)
2024
-
[35]
R. A. Jacob, S. Paul, S. Chowdhury, Y. R. Gel, and J. Zhang, Real-time outage management in active distri- bution networks using reinforcement learning over graphs, Nature Communications 15, 4766 (2024)
2024
-
[36]
Y. Lee, H. Choi, L. Pagnier, C. H. Kim, J. Lee, B. Jhun, H. Kim, J. Kurths, and B. Kahng, Reinforcement learning optimizes power dispatch in decentralized power grid, Chaos, Solitons & Fractals 186, 115293 (2024)
2024
-
[37]
Huang, W
R. Huang, W. Gao, R. Fan, and Q. Huang, Damping inter- area oscillation using reinforcement learning controlled tcsc, IET Generation, Transmission & Distribution 16, 2265 (2022)
2022
-
[38]
Ernst, M
D. Ernst, M. Glavic, and L. Wehenkel, Power systems stability control: reinforcement learning framework, IEEE transactions on power systems 19, 427 (2004)
2004
-
[39]
Schultz, J
P. Schultz, J. Heitzig, and J. Kurths, A random growth model for power grids and other spatially embedded in- frastructure networks, The European Physical Journal Special Topics 223, 2593 (2014)
2014
-
[40]
A. J. Wood, B. F. Wollenberg, and G. B. Shebl´ e,Power generation, operation, and control(John wiley & sons, 2013)
2013
-
[41]
Habur and D
K. Habur and D. O’Leary, Facts-flexible alternating cur- rent transmission systems: for cost effective and reliable transmission of electrical energy, Siemens-World Bank document–Final Draft Report, Erlangen 46, 244 (2004)
2004
-
[42]
Longoria, M
G. Longoria, M. Lynch, N. Farrell, and J. A. Curtis, The impact of planning and regulatory delays for major energy infrastructure, Tech. Rep. (ESRI Working Paper, 2022)
2022
-
[43]
Defferrard, X
M. Defferrard, X. Bresson, and P. Vandergheynst, Con- volutional neural networks on graphs with fast localized spectral filtering, Advances in neural information process- ing systems 29 (2016)
2016
-
[44]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347 (2017)
2017 arXiv
-
[45]
Moalla, A
S. Moalla, A. Miele, D. Pyatko, R. Pascanu, and C. Gul- cehre, No representation, no trust: connecting representa- tion, collapse, and trust issues in ppo, Advances in Neural Information Processing Systems 37, 69652 (2024)
2024
-
[46]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al., Deepseekmath: Pushing the limits of mathematical reasoning in open language models, arXiv preprint arXiv:2402.03300 (2024)
2024 arXiv
-
[47]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, Direct preference optimization: Your language model is secretly a reward model, Advances in neural information processing systems 36, 53728 (2023)
2023
-
[48]
Zhang and L
Q. Zhang and L. Ying, Zeroth-order policy gradient for reinforcement learning from human feedback without re- ward inference, arXiv preprint arXiv:2409.17401 (2024). SUPPLEMENT AR Y INFORMA TION A. T opological and Dynamical Characteristics of the SHK Network To verify that th...
2024 arXiv
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.